CaseIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #10
Describe how you would run an internal AI literacy session for your leadership team.
SPARK the session built around two Deskmend tickets from the same week, one clean, one confidently wrong
Farlow Systems runs its own employee IT helpdesk on Deskmend, an AI tool that reads a support ticket and tries to close it without a person touching it. Yulia Wenzel is the AI PM who owns Deskmend's roadmap. Obadiah Ottway is Farlow's CEO. Kessandra Fossett is the CFO who sets the IT support budget every year.
The direct answer
Run one short session, about forty five minutes, with the whole leadership team in the room together. Build it around two real Deskmend tickets from the same week: one it resolved cleanly on its own, and one it closed as "Resolved" after only fixing half the problem. Skip the model internals and skip any ask for more budget. The point is one shared question leadership can carry into every future decision: is this the kind of ticket Deskmend is actually good at, or the kind it only sounds confident about.
Do this, in order
Build the session around two real tickets, never slides about AI in general.Why: this is the actual mechanism that gives leadership something concrete to calibrate against, instead of a mood about AI.
Put the whole leadership team in the room at once, no proxies, no recap later.Why: a shared question only works if everyone heard it form together. One person relaying it secondhand loses the calibration.
Pick one clean win and one confidently wrong ticket from the same week.Why: same week proves this is how Deskmend behaves right now, not an old bug or a cherry picked outlier from months back.
Cut anything that explains the classifier or asks for headcount or budget.Why: either one kills the room, a technical tangent loses them, a budget ask turns the session into a pitch.
Judge it by the next real decision, not by nodding heads in the room.Why: the only real test is whether leadership asks a sharper question the next time an AI number lands on their desk.
Never make Yun Ridgewell's support engineers sit through this.Why: they already calibrate their trust in Deskmend by hand, every shift. This session is for the people who never open a ticket at all.
01How to answer this, stage by stage
Nobody is grading whether you can say "leadership needs to understand AI better." They are grading whether you can name the actual room, the actual two tickets, and the actual line between a session that calibrates judgment and one that either bores a CEO or flatters him.
01
Name the exact meeting, not the general idea
Say it like this
"I'll ground this in one team. Farlow Systems runs its own IT helpdesk on Deskmend. I'm the AI PM, Yulia Wenzel. Obadiah Ottway is the CEO. Kessandra Fossett is the CFO who owns the support budget."
Why this works
Stops the answer sliding into a generic lecture on "leadership education," which is a different, weaker question.
02
Announce the shape before the content
Say it like this
"I'll run this as SPARK. Situation, how leadership actually forms opinions about Deskmend today, with nothing designed for it. Payoff, the one habit I want the session to build. Anchor, the actual forty five minutes. Risk, what breaks if I get the format wrong, in either direction. Keep out, what this session will never turn into."
Why this works
Two seconds of structure tells the interviewer you have a plan, not just a feeling that "leadership should know more."
03
Say what the question is actually testing
Say it like this
"This isn't really 'how do you run a good workshop.' It's whether Obadiah and Kessandra can tell the difference between a ticket Deskmend is genuinely good at and one it just sounds confident about, before that gap costs the company a real decision."
Why this works
Separates a real design answer from a generic training plan that happens to be about an AI product.
04
Open with the two tickets, not a slide
Say it like this
"I bring two real tickets from the same week. One: someone's locked out of their laptop, Deskmend verifies them and resets the password in under ninety seconds. Two: someone's VPN drops and their badge access breaks too, Deskmend fixes the VPN, closes the ticket as Resolved, and never touches the badge problem at all."
Why this works
This is the anchor, the actual concrete decision, not a description of "showing some examples."
05
Name both ways the room can go wrong
Say it like this
"Get this too technical, start explaining how the ticket classifier scores intent, and Obadiah and Kessandra check out inside ten minutes. Nobody talked down to a room ever thanked you for it. Get it too soft, show only the password win and skip the badge ticket, and I've handed them the exact same rosy picture the vendor's own benchmark already gave them."
Why this works
Naming both failure directions, not just one, shows you're not about to overcorrect into the opposite mistake.
06
Ground it in the real near miss
Say it like this
"Here's why this mattered enough to build. The same week a rival's ticket bot closed a security breach report as a routine password reset, and it made the tech press, Obadiah forwarded a different vendor's number, ninety percent full closure, and floated cutting our support team in half next quarter. Kessandra, spooked by the other story, wanted to freeze Deskmend entirely. Neither of them had ever seen one real Deskmend ticket."
Why this works
One real week carries more weight than any abstract claim about "leadership needs more AI awareness."
07
Say the trade off out loud
Say it like this
"The fix coming out of this session costs something real. Routing every multi-issue ticket to a person instead of letting Deskmend auto-close it adds a few minutes of staff time to roughly one ticket in ten. That's the price for not silently dropping someone's badge access for a week."
Why this works
Shows this is a real decision with a real cost, not a wish that speed and accuracy were both free.
08
Draw the keep out line, then close in one breath
Say it like this
"I'd never let this turn into a pitch for a bigger Deskmend budget, and I'd never open the classifier's scoring logic in that room, that's a different meeting for a different audience. So: one session, two real tickets from the same week, no ask, no architecture, and one question leadership carries into every AI decision after this: which kind of ticket is this."
Why this works
Shows judgment about where the idea stops, then restates the whole decision in one breath.
02Let's learn
For three years running, Farlow's IT support headcount did not move by a single person, up or down.
This is the gap the session exists to close. Not a lack of interest in the room. A lack of anything real to look at.
Deskmend reads an employee's IT ticket and tries to close it without a person touching it. Before Farlow built it, a nine person support team led by Yun Ridgewell handled roughly 2,300 tickets a month by hand, about twelve minutes each on average, split across VPN failures, password resets, printer trouble, and requests that never fit neatly into one box.
Knowledge spark: what is a multi issue ticket?
One ticket describing more than one problem at once, say a broken VPN and a broken badge reader in the same message. A model has to notice it's actually two separate jobs, not one, before it can safely say the ticket is done.
Fourteen months in, Deskmend closed 71 percent of all tickets without a person touching them, and that single number sat on the same slide in the monthly ops review the entire time. Password reset tickets, a third of total volume, told one story: Deskmend got the right answer 97 percent of the time, and handle time fell from twelve minutes to under ninety seconds. Multi issue tickets, about one in eleven, told a different one. Deskmend fully and correctly fixed both parts only 22 percent of the time. But it marked the ticket "Resolved," closed and done, 61 percent of the time regardless, because closing only needed one detected problem to clear its own cut off point.
Deskmend's outcomes, single issue vs multi issue tickets
Genuinely doneActually correctMarked done, often wrong
The gap between the middle bar and the right one is the whole problem. Only 22 percent of multi issue tickets were actually fixed in full, but 61 percent got closed as though they were.
Closed and wrong looks exactly like closed and right, right up until someone needs their badge back.
Neither message was about Deskmend. Both of them shaped what Farlow's leadership believed about it anyway.
The Tuesday that forced the issue, a company using a rival product called Wybrant Resolve had its ticket bot close an employee's security breach report as a routine password reset, and the story made the tech press by lunch. That same morning, Obadiah had forwarded a different vendor's press release to the leadership channel, a claim of ninety percent full ticket closure with zero human touch, and asked whether Farlow could match it by cutting Tier 1 support in half next quarter. Kessandra, reading the Wybrant story an hour later, replied that she wanted Deskmend's rollout frozen until someone could prove it was safe. Neither message mentioned a single real Deskmend ticket, because neither of them had ever seen one.
The whole answer sits in this one picture. Not a claim about Deskmend in general. Two tickets, the same week, side by side.
What Yulia built instead of answering either message directly: a forty five minute session, both of them in the room together, no slides about AI in general. She opened with the password reset ticket, Deskmend's clean win, then the VPN and badge ticket, where it fixed one problem and silently dropped the other while still marking the whole thing Resolved. Same model, same week, two very different tickets.
Same ticket type, same model. What changes the outcome is only whether anyone in the room has ever seen one.
What it costs at its worst: if Obadiah had cut Tier 1 support in half on the strength of a competitor's headline number, Farlow would have gutted the exact team that catches a dropped badge ticket before it becomes a locked door on a Monday morning. If Kessandra had frozen the rollout instead, Farlow would have thrown away the 97 percent of tickets Deskmend genuinely handles well, to guard against a failure mode a five minute look at the real data could have named and fixed.
The choice I would take back
A year earlier, when Deskmend first launched, Farlow decided the monthly ops review would carry one line: "Auto-resolve rate: 71 percent." That was a fine amount of detail when Deskmend was a minor line item nobody's budget depended on. It stopped being fine the moment real decisions, headcount, rollout speed, started resting on what leadership imagined was sitting behind that one number. Nobody chose to hide the multi issue gap on purpose. Nobody ever asked whether one number was still enough, either.
What I would leave alone: Yun Ridgewell's support engineers, who triage Deskmend's escalations every shift, don't need this session at all. They already know exactly which ticket types to double check, because they've watched the tool get it wrong in person. This session exists for the two people in the building who never open a ticket.
The lesson: a single number on a slide isn't knowledge, it's a placeholder for knowledge. Leadership needed to see two real tickets before that 71 percent meant anything true to them at all.
03Now here is the same thing as a story
The short version above is what you'd actually say in the room. Read this one for why the fix had to be a session built on real tickets, not a promise to "brief leadership more often."
Every Monday morning, Kessandra Fossett opens the same spreadsheet and leaves the support line exactly where it was the week before.
Yulia Wenzel has built support tooling for eight years, and she can usually tell within a week whether a rollout is actually working or just quiet because nobody's looked hard enough yet. When Deskmend went live at Farlow, she built the eval set herself, tracked it by ticket category from day one, and trusted the number climbing on her own dashboard, because she was the one watching it climb.
For the first ten months, that felt like enough. Yulia sent a short written update to the leadership channel each month: a graph going up, a line or two of context. Obadiah skimmed it. Kessandra skimmed it less. By month six, the updates had shrunk to one line: "Deskmend continues to perform well." By month nine, nobody replied to them at all. Nobody had stopped trusting Deskmend. They had simply stopped needing to think about it, because the number in the update never seemed to change.
Then came the Tuesday two messages landed in the same channel within an hour of each other, and neither one came from Yulia.
Yulia drew this line early, and held it on purpose. Not because leadership couldn't handle more. Because more wasn't the problem.
Obadiah forwarded a competitor's press release, a claim of ninety percent full closure with zero human involvement, and asked, half rhetorically, whether Farlow could halve its Tier 1 team by next quarter to match it. An hour later, Kessandra replied to the whole channel with a link to the Wybrant Resolve story, a rival's ticket bot that had just closed a security breach report as a routine password reset, and said she wanted Deskmend's rollout paused until someone could show her it wasn't doing the same thing.
We didn't lose an afternoon to a badge access ticket. We lost the only two people in the building who could still ask a real question about Deskmend.
Yulia's first instinct was the tempting one: reply to both messages with the real number, 71 percent, and a paragraph explaining why the comparison to Wybrant wasn't fair. She drafted it, then deleted it. A written reply to a written panic would just become a third message in a thread neither of them was reading carefully to begin with.
She thought back to the meeting, more than a year earlier, where the one line ops review update had been decided. It hadn't been a lazy choice. Deskmend was brand new then, a small pilot, and one honest number felt like plenty of transparency for something that small. Nobody in that room imagined a year later that number would be standing in for the entire company's understanding of what the tool could and couldn't do.
What she built instead: a forty five minute session, two days later, both Obadiah and Kessandra in the same room, no deck about AI in general. She opened with a password reset ticket from the week before, an employee locked out at 7am, verified and reset in eighty three seconds. Then she opened the VPN and badge ticket from the same week: VPN fixed cleanly, badge access never touched, ticket marked Resolved anyway. Same model. Same week. Two very different outcomes.
She also considered asking Deskmend's engineering lead to walk the room through how the ticket classifier actually scores intent, confidence numbers, cut off points, the works. She killed that idea before the invite went out. Neither Obadiah nor Kessandra needed to know how the number got computed. They needed to know what it meant for the next ticket that looked like the badge one.
Same headcount conversation Farlow was always going to have. What changed is only what anyone in the room knew before it started.
Six weeks later, Obadiah opened the actual quarterly headcount meeting. He didn't mention the ninety percent claim once. Instead he asked Yulia straight: "What's our real rate on tickets shaped like that badge one?" Kessandra, for the first time in three years, added a specific new line to the support budget instead of leaving it untouched: two review hours a week, paid time for a person to check every flagged multi issue ticket before it closed. The headcount cut wasn't promised that day. It also wasn't ruled out for good. It just stopped being a decision made off a stranger's number.
One design handed leadership a single slide and asked them to trust it. The other handed them two real tickets and asked them to look. Only one of those two designs survives a Tuesday it didn't see coming.
What I'd tell myself, reading Obadiah's message that first Tuesday: the mistake wasn't that leadership reacted to outside news. It was that we'd never given them anything of our own to react to instead.
04SPARK, the five checks behind forty minutes with the leadership team
Not a script for sounding informative. SPARK is what forces you to name the one concrete session, and prove it survives both ways a literacy session actually breaks.
SSituation. How does leadership actually form its picture of Deskmend today?
Without a designed session, Obadiah picks up his mental model from vendor press releases and competitor headlines, and Kessandra avoids the topic entirely in her own budget planning. Neither extreme is grounded in a single real Deskmend ticket. Both are decisions waiting to happen with no real information behind them.
One CEO, one CFO, one real product. Never a segment called "executive AI readiness."
PPayoff. What habit do I want this to build?
I want leadership to carry one question into every future AI related decision: is this the kind of thing Deskmend is actually good at, or the kind it only sounds confident about. The habit is the product. Fewer panicked headcount swings and budget freezes are downstream of that habit, not the goal itself.
Name the question they'll ask next time, not the meeting's length. That's the payoff.
AAnchor. The one decision everything else hangs on.
A forty five minute session, both leaders in the room, built around two real Deskmend tickets from the same week: a password reset it resolved cleanly, and a VPN plus badge ticket it closed as Resolved after only fixing half of it. Yulia also considered a technical walkthrough of the classifier's confidence scoring from Deskmend's engineering lead. She rejected it: neither Obadiah nor Kessandra needed to know how the number gets computed, only what it means for the next ticket shaped like the badge one.
Concrete enough to argue with. This is the answer to the question.
RRisk. What breaks the first time it's wrong?
Run it too technical, open the classifier's internals, and leadership checks out inside ten minutes, then goes right back to reacting to whatever headline lands next. Run it too soft, show only the clean win, and the session just reinforces the same rosy picture a vendor's own marketing slide already gave them. Yulia accepts a real cost here: routing every multi issue ticket to a person instead of letting Deskmend auto close it adds staff minutes to roughly one ticket in ten, forever, because the alternative is a badge request silently dropped for a week.
Not "leadership trusts it more." What each of them actually does in the very next AI related decision.
Same ticket, same model, same leadership team. What changes the outcome is only whether the session ever happened.
KKeep out. What I deliberately will not build into this.
No ask for a bigger Deskmend budget, no roadmap pitch riding on the session's goodwill. No walkthrough of the classifier's scoring logic, that belongs in a different meeting with a different audience. No company wide roadshow either, this session is sized for two people who make company level calls, not for general awareness.
Shows judgment instead of a wish to cover everything. Ties straight back to Risk: the wrong kind of depth and the wrong kind of ask are both real ways to lose the room.
The recap, one line per letter: situation is a leadership team forming its picture of Deskmend from outside noise alone, payoff is one shared, calibrated question carried into every future decision, anchor is the forty five minutes built on two real tickets from the same week, risk is either extreme costing months of bad decisions, and keep out draws the line at a budget ask and a technical deep dive, never at the tickets themselves.
05And if you want to be sure it really works, try it somewhere else
Same five letters, a county building department instead of an IT helpdesk, and this time the blind spot isn't a badge reader. It's a mixed use permit that touches two code categories at once.
Doskoch Civic Systems builds Permitto, an AI tool that reads a construction permit application for Pellew County and auto approves the straightforward ones before a person ever opens the file. Ozella Menlove owns Permitto's roadmap. Pellew County's leadership formed its own opinion of Permitto the same ungrounded way Farlow's did: the County Administrator wanted to publicize a faster average approval time after reading about a neighboring county's rollout, while the Finance Director had quietly left Permitto out of every budget conversation for two years, treating it as something too technical to plan around.
Mapped onto SPARK: the situation is leadership forming a picture of Permitto from a neighboring county's press coverage and their own avoidance, never from a real Pellew County permit. The payoff is one shared question carried into every future permitting decision: is this a single code category permit, or a mixed use one. The anchor is the same structure, a short session built on two real permits from the same week, a simple fence permit Permitto approved correctly, and a garage to office conversion permit it approved without ever flagging the required electrical inspection. The risk runs the same both ways: too technical and the room glazes over at zoning code citations, too soft and it just repeats the "permits now move faster" line that started the confusion. Keep out draws the same line: no pitch to expand Permitto's contract, no walkthrough of its document parsing model.
Pellew County, mixed use permits approved without a flagged second review, week by week
Before the routing fixAfter the routing fix
Flat around a third of mixed use permits for five straight weeks, the same weeks Pellew County's leadership was arguing about a headline instead of a real file. Once the session led to routing mixed use permits to a person, the rate dropped by more than three quarters within two weeks.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, the two tickets and the one question they leave with.
Cost: no time on the leadership calendar for a standing session at all. Run it once, as a single fixed session tied to the next real decision on the table, rather than a recurring meeting, the calibration still lands, it just isn't refreshed automatically.
The model got better, for real: say Deskmend's multi issue accuracy climbs to 80 percent. The session still matters, because leadership still needs the vocabulary to ask "which kind of ticket is this," even when the honest answer is "the safer kind now." A better model doesn't teach anyone how to ask a sharper question on its own.
Where people run it wrong.
They let an engineer take over the room and explain how confidence scoring actually works, and the session quietly becomes the technical tutorial it was built to avoid.
They pick two tickets from months apart instead of the same week, and leadership assumes the bad one has since been fixed, when nothing about it has changed.
They end the session with a slide asking for more AI budget, and every calibrated question from the last forty minutes gets read backward as a sales pitch.
How to use it live. Before answering a "how would you run a literacy session" question cold, ask yourself one thing: what two real, current outputs would you actually put in front of the room, one good and one wrong. Naming those two things, not a value word like "understanding" or "alignment," is usually exactly what the question is listening for.
06Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question about designing a session, not reacting to a single failure?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It runs forward from how leadership forms its picture of the product today, instead of working backward from one incident.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Yulia Wenzel, the AI PM who owns Deskmend's roadmap at Farlow Systems. Obadiah Ottway is the CEO. Kessandra Fossett is the CFO who sets the IT support budget.
3 · THE PAYOFF
What habit does the session exist to build?
Tap to flip
ANSWER
Leadership carries one shared question into every future AI related decision: is this the kind of ticket Deskmend is actually good at, or the kind it only sounds confident about.
4 · THE ANCHOR
What's the one concrete decision in this answer?
Tap to flip
ANSWER
A forty five minute session, both leaders in the room, built on two real tickets from the same week, one clean win and one Deskmend closed as Resolved after only half fixing it.
5 · THE OLD DECISION
What old decision would Yulia take back?
Tap to flip
ANSWER
Summarizing Deskmend for leadership as one aggregate number on one slide, for a full year. It was fine when the tool was a small pilot. It stopped being fine once real headcount and budget decisions started resting on it.
6 · THE NUMBER
Fill in the blank: Deskmend marked a multi issue ticket "Resolved," whether or not both problems were fixed, ___ percent of the time.
Tap to flip
ANSWER
61 percent. Only 22 percent of multi issue tickets were actually fixed in full, but the ticket closed as done in nearly three times that many cases.
7 · THE RISK, SURVIVED
What breaks if the session goes wrong in either direction, and how does the anchor survive it?
Tap to flip
ANSWER
Too technical and leadership checks out, then reacts to headlines again. Too soft and it repeats the same rosy picture a vendor's slide already gave them. It survives because the two tickets stay concrete and current, with no architecture and no budget ask attached.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs SPARK again on a different product. Which one, and what's the equivalent anchor?
Tap to flip
ANSWER
Permitto, Doskoch Civic Systems' tool for approving Pellew County construction permits. The equivalent anchor is the same short session, built on a correctly approved fence permit and a mixed use permit that skipped a required electrical review.
07Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: for a full year, Deskmend's monthly leadership update carried exactly one number, "Auto-resolve rate: ___ percent," with no example ticket ever shown.
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
71 percent. That number was true and, by itself, told leadership almost nothing about where Deskmend actually struggled.
Multiple choice
2. Why did Deskmend's high overall auto-resolve number never catch the badge access problem?
A. Because Deskmend had never been trained on badge access tickets at all.
B. Because the one blended number hid a much lower success rate on the smaller, harder category of multi issue tickets.
C. Because the ops review update was only sent once a year.
D. Because Yun Ridgewell's team was hiding failed tickets from leadership on purpose.
Show hint
Compare the 71 percent headline number to the 22 percent and 61 percent figures in the bar chart.
Show answer
B. The blended 71 percent mixed a strong password reset category with a weak multi issue one, and the average looked healthy the entire time the gap existed.
True or false
3. True or false: having Deskmend's engineering lead walk leadership through the classifier's confidence scoring would have made this session more convincing.
True
False
Show hint
Look at what Yulia considered and rejected in the Anchor step block.
Show answer
False. Yulia considered this and rejected it. Leadership didn't need to know how the confidence number gets computed, only what it means for the next ticket shaped like the badge one. A technical walkthrough risked losing the room entirely.
Short answer, name the reversal
4. What old decision would Yulia take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Summarizing Deskmend for leadership as one aggregate number on one slide, with no real example ever shown. It made sense when Deskmend was a small, new pilot and one honest number felt like plenty of transparency. It stopped making sense once real headcount and budget decisions started resting on what leadership imagined was behind that number.
Short answer, apply it yourself
5. Think of an AI product you use or have worked on. What's one real, current output, a genuine win and a genuine miss, you could put in front of a leadership team to replace a vague claim about how well it works?
Show hint
Look for two outputs from roughly the same week, not one dramatic failure pulled from months ago.
Show answer
Model answer: A meeting notes tool might show one meeting where its summary was used verbatim, and another from the same week where it confidently summarized a decision that was actually still under debate, sending it out as though it were final.
Short answer, work the number
6. If the multi issue ticket share doubled from 9 percent of total volume to 18 percent, with the same 61 percent wrongly-closed rate, would the fix's staff time cost roughly double too? Why or why not?
Show hint
The guardrail routes based on how many tickets fall into the multi issue category, not on the total ticket volume.
Show answer
Model answer: Roughly yes. The guardrail flags a fixed share of multi issue tickets for human review, so doubling that category's volume roughly doubles the number of tickets needing the extra minutes, even though the total ticket count and the password reset category haven't changed at all.
Before you close the answer
Why this works
Tests whether you can turn "leadership needs AI literacy" into an actual, repeatable session, two real tickets, one question, instead of a training plan. And whether you know exactly where the line sits between calibrating judgment and either boring a CEO with architecture or handing him a sales pitch.
Follow-up traps
"Isn't forty five minutes too short to actually teach anyone anything?" Response: it isn't trying to teach the model, only to install one repeatable question. That length is right for a question, wrong for a curriculum.
"What if a leader still overreacts to the wrong ticket after this session?" Response: that's still real progress, they're now reacting to something Deskmend actually did instead of a stranger's headline. The goal is calibration, not a guarantee of perfect judgment.
If pressed
The guardrail that came out of this session routes a ticket to a person whenever Deskmend detects a second, lower confidence intent still present in the same message, rather than auto closing the moment any single intent clears its own cut off point. That's a probability call, not a fixed rule, so it still gets reviewed against fresh tickets every quarter to make sure the cut off hasn't quietly drifted.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.