CaseIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #13
How do you communicate a quality bar to people who have never seen an eval report?
SPARK the quality card built from one real invoice fraud email, and the two calls that failed before it
Kessley Mail sells hosted business email with a built in phishing and spam screen called Talonguard. Ottavine Thorbecke is the AI PM who owns Talonguard's quality bar. Brumley Grimwade runs Customer Success at Kessley Mail and has to answer for that quality bar on the phone. Wickstead Hardware is a fourteen store customer that nearly lost real money to a phishing email Talonguard let through.
The direct answer
Don't send the raw eval report, and don't just say "it's working well" either. Build one plain language quality card: how often Talonguard catches a phishing email, how often it wrongly flags a real one, plus one real named example of the kind it still misses. Give that card to anyone who has to answer for Talonguard's quality out loud, and rebuild it every quarter so it never goes stale.
Do this, in order
Build one plain language card pairing the catch rate and the false flag rate with a real named miss.Why: this is the actual translation, not a summary of the translation.
Never hand a non technical person the raw eval report.Why: precision, recall, and a threshold sweep chart lose the room before the real conversation starts.
Never leave the number out either, no "it's working well."Why: a claim with nothing behind it gives nobody a sharp question to ask back.
Name where the misses cluster, not just the average.Why: a flat catch rate lets people assume every miss is small and rare, when the real ones concentrate hard on one weak spot.
Rebuild the card every quarter against fresh cases.Why: phishing senders adapt into whatever Talonguard is currently worst at catching, so a one time claim goes stale fast.
Keep architecture and any upsell pitch out of the card completely.Why: a technical deep dive or a sales ask both derail the one conversation the card exists to have.
01How to answer this, stage by stage
Nobody is grading whether you can define precision and recall. They are grading whether you can turn a probability into something a person with no model background can actually act on, and defend the exact words you'd use.
01
Ground it in one real account, not "stakeholders" in general
Say it like this
"I'll ground this in one team. Kessley Mail sells hosted email with a phishing screen called Talonguard. I'm the AI PM, Ottavine Thorbecke. Brumley Grimwade runs Customer Success and has to answer for Talonguard's quality on live calls."
Why this works
Stops the answer sliding into "communicate clearly," which is a much weaker, generic question.
02
Say the shape out loud before the content
Say it like this
"I'll run this as SPARK. Situation, how people talk about Talonguard's quality today with nothing designed for it. Payoff, the one habit I want that to build. Anchor, the actual card. Risk, what breaks if the card is wrong in either direction. Keep out, what the card will never turn into."
Why this works
Two seconds of structure tells the interviewer you have a plan, not a vague feeling that "people should understand it better."
03
Say what the question is actually testing
Say it like this
"This isn't really 'how do you explain a metric.' It's whether Brumley can turn a probability, Talonguard is right most of the time but not all of the time, into something Wickstead Hardware's ops director can actually weigh, before the next fraud email decides it for her."
Why this works
Separates a real answer about model behavior from generic advice about clear writing.
04
Give the plain language numbers, straight out
Say it like this
"I'd tell her straight. Talonguard catches about nineteen of every twenty phishing emails that try to get through. About one in five hundred real customer emails gets wrongly flagged and held for someone to check. Those are the two numbers that actually matter to her."
Why this works
This is the anchor, the actual concrete decision, not a description of "showing some numbers."
05
Name the one real miss, specifically
Say it like this
"And here's what it's still weakest against: a first contact sender pretending to be a vendor you already trust, no message history to compare it to. That's exactly the invoice that got past us at Wickstead. I did look at just putting the full eval numbers on a live dashboard customers could check any time, and I killed that, it's the same unreadable sheet, just on a webpage instead of an email."
Why this works
Names a specific caught/missed example instead of a category label, and names the rejected alternative so the anchor reads as a real decision.
06
Say both ways this breaks
Say it like this
"Flatten this to one number and people assume every miss is small and rare, when they're not, almost half of Talonguard's real misses come from that one narrow sender type. Go the other way, hand someone the threshold curve, and I've lost the room before the real conversation even starts."
Why this works
Naming both failure directions, not just one, shows you won't overcorrect into the opposite mistake.
07
State the cost you're accepting
Say it like this
"Fixing the weak spot costs something real. Routing every first contact vendor claim to a person before it clears adds a few hours' delay to about one in forty genuine new vendor invoices. That's the price for catching the kind of email that hit Wickstead."
Why this works
Shows this is a real decision with a real cost, not a wish that catch rate and speed were both free.
08
Draw the keep out line, then close in one breath
Say it like this
"This card will never turn into a walkthrough of Talonguard's threshold curve, and it will never turn into a pitch for a bigger security add on. So: one page, two real numbers, one real miss, refreshed every quarter. That's the whole habit I want every call at Kessley Mail to run on."
Why this works
Shows judgment about where the idea stops, then restates the whole decision in one breath.
02Let's learn
What do you say to someone who has never once looked at an eval report? That's really the whole question, and the honest answer is: not the report, and not a vibe either.
This is the gap the card exists to close. Not a lack of care from Brumley's team. A lack of anything real to hand a worried customer.
Talonguard reads every email that reaches a Kessley Mail inbox and decides, in a fraction of a second, whether it's safe, spam, or a phishing attempt. Before Kessley Mail built it, a business like Wickstead Hardware relied on its own spam folder and a wary eye, and still lost the occasional afternoon to a fake shipping notice or a bad link. Talonguard screens roughly 40,000 emails a day for Kessley Mail's customers, and about 1 in 900 of those is a genuine phishing attempt.
Knowledge spark: what is a false flag?
A real, safe email that Talonguard wrongly holds back and marks as risky. It costs someone a delayed invoice or an annoyed reply. It's the price of catching more of the real threats, and it's never zero.
Over the last quarter, checked case by case against a set of real, confirmed emails, Talonguard caught 95 percent of confirmed phishing attempts, about 19 of every 20. It wrongly held about 1 in 500 real customer emails for review. Those two numbers on their own sound close to finished. They aren't, because they hide exactly where the misses live.
Share of phishing attempts vs. share of Talonguard's misses, by sender type
Share of confirmed phishing attemptsShare of Talonguard's misses
First contact vendor impersonation is rare, just 4 percent of attempts, but it's behind nearly half of everything Talonguard missed last quarter. A flat 95 percent catch rate never shows this on its own.
A flat 95 percent doesn't lie. It just doesn't say where the other 5 percent lives.
A single badge like this one looks calm because it's built to look calm. The real failures underneath it were never spread evenly.
What it costs at its worst: hand a worried customer a single reassuring number, and the very first bad email that reaches them lands as proof the number was a lie, not proof that a rare, hard category slipped through. Hand them the raw eval sheet instead, and most people can't read it well enough to know whether they should be relieved or terrified, so they assume the worst by default.
The choice I would take back
About a year before this, when Ottavine's team first put Talonguard's numbers in front of customers, they built one line for Customer Success to quote: "Trust score: 97 out of 100." It read clean and confident. That was fine when nobody's account had been tested by a real fraud attempt yet. It stopped being fine the moment a real invoice fraud email needed someone to say something true and specific about it, and there was nothing behind the badge to say.
What I would leave alone: Talonguard's engineers already read the full eval report every release, threshold sweep and all, and they should keep doing exactly that. The card isn't for them. It's only for the people who will never open that report, on purpose or otherwise.
The lesson: a number on its own isn't understanding, it's a placeholder for understanding. Brumley needed a real example and a real second number before "95 percent" meant anything he could act on.
03Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one for why the fix had to be a one page card, not a promise to "communicate more clearly."
Every Thursday afternoon, Brumley Grimwade clears his calendar for renewal calls. He's run Customer Success at Kessley Mail for five years, long enough to know that a calm voice and a real answer save an account faster than any discount ever has.
Talonguard had been quiet in his calls for months. Nobody at Wickstead Hardware, a fourteen store chain that had used Kessley Mail's inboxes since 2019, had ever asked him a hard question about it. He didn't bring it up either. Some Thursdays he didn't think about Talonguard at all.
Then, on a Tuesday, an email reached Salvestro Corcoran, Wickstead's accounts payable clerk. It looked exactly like an invoice from Wickstead's regular lumber supplier: same logo, same layout, one line changed, a new routing number tucked into the wire details. Salvestro did what anyone would have. The logo matched. The format matched. Nobody double checks a routing number by habit. Talonguard let it through, and he wired $22,650 before the mismatch turned up on the next month's statement.
Two honest attempts, two failed calls. Neither one gave Wickstead's ops director anything she could actually use.
Wickstead's ops director called Brumley that same afternoon. Her question was simple: "How good is this thing, actually?" Brumley did the thing that felt safest. He forwarded her Ottavine's eval sheet: precision 0.951, recall 0.949, F1 0.950, AUC 0.978, a threshold sweep chart attached underneath. She called back an hour later angrier, not calmer. She hadn't understood a single row of it, and now assumed nobody at Kessley Mail knew whether Talonguard worked either.
On the next call, three days later, Brumley tried the opposite. "It's working really well," he said. "This was a rare thing." No number behind it, nothing she could check against anything else. She didn't believe that one either. By Friday she'd drafted a note to her own team recommending they start evaluating a different vendor.
We didn't lose Wickstead Hardware to one fraud email. We lost them to two calls where nobody could tell them anything they could actually use.
I want to say the problem was that Talonguard missed one email. It did miss one, the kind it's genuinely worst at: a first contact sender pretending to be a known vendor, no history to compare it to. But that's not really the story. Wickstead's ops director never had a real number in her head about Talonguard, not once in five years. She had a mood, built from a smooth onboarding call and two quiet years. One bad Tuesday, and there was nothing underneath the mood to catch her.
So here's the decision I would take back. A year earlier, Ottavine's team built a single line for Customer Success to quote: "Trust score: 97 out of 100." A single number was the easiest thing to hand a salesperson, and nobody built it to mislead anyone. It just flattened every kind of mistake Talonguard makes into something that looked evenly spread, when the real failures were nothing like even.
I would put a real quality card back in its place instead. Not a bigger number, a different shape of information: catches about 19 of every 20 phishing emails, wrongly holds about 1 in 500 real ones, and here's the exact kind it still misses, the kind that reached Salvestro. Replay that second call with the card in Brumley's hand: he opens with the real numbers, names the one real miss, and explains the fix already in motion. That call runs about ten minutes, not forty. Wickstead renews for two more years instead of drafting a cancellation that Friday.
And the part I'd want to tell myself, if I could go back: we thought we were protecting customers from a confusing spreadsheet. We were really just leaving them with nothing real to hold onto instead.
04SPARK, the five beats behind the one page card
Not a script for sounding reassuring. SPARK is what forces you to name the one concrete card, and prove it survives both ways a quality bar actually breaks.
SSituation. How does anyone form a picture of Talonguard's quality today?
Without a designed way to translate it, Brumley's team either hands over the raw eval sheet, which a non technical customer can't read, or says "it's working well" with no number behind it, which gives nobody a real question to ask. Neither one lets a real quality conversation happen.
One customer, one call, one real fraud email. Never a segment called "non technical stakeholders."
PPayoff. What habit do I want this to build?
Anyone talking about Talonguard, Brumley included, can state the real catch rate and the real false flag rate in one breath, tied to a real example, and use it to ask a sharper question instead of trusting a mood. The habit is the product. Calmer renewal calls are downstream of that habit, not the goal itself.
Name the question they can now ask, not the card's length. That's the payoff.
AAnchor. The one decision everything else hangs on.
A one page quality card: catches about 19 of 20 phishing emails, wrongly holds about 1 in 500 real ones, plus one real named miss, first contact vendor impersonation, the exact shape of the email that reached Wickstead. Ottavine also considered building a live dashboard exposing the full precision, recall, and threshold numbers for customers to check any time. She rejected it: it's the same unreadable eval sheet, just moved onto a webpage instead of an email attachment.
Concrete enough to argue with. This is the answer to the question.
This is the whole anchor in one picture. Not a claim about Talonguard in general. Two real numbers and one real miss, on a single page.
RRisk. What breaks the first time the card is wrong?
Flatten it to just "19 of 20" with no context, and people assume the misses spread evenly, when almost half of them cluster on one narrow sender type. Open with the full threshold curve instead, and the room checks out before the real conversation starts. Ottavine accepts a real cost here: routing every first contact vendor claim to a person instead of clearing it automatically adds a few hours' delay to about one in forty genuine new vendor invoices, forever, because the alternative is another Salvestro.
Not "the customer trusts it more." What Brumley actually says on the very next call about a miss.
Same fraud email, same Brumley. What changes the outcome is only whether the card exists before the call happens.
KKeep out. What I deliberately will not build into this.
No walkthrough of Talonguard's confusion matrix or threshold sweep, that belongs in a different room with a different audience. No pitch for a bigger security add on riding on the card's goodwill either. And no promise that the number will ever hit 100, a probabilistic model earns trust by being honest about its edges, not by pretending it doesn't have any.
Ties straight back to Risk: the wrong kind of depth and the wrong kind of ask are both real ways to lose the room.
Ottavine drew this line early and held it on purpose. Not because customers can't handle detail. Because detail wasn't the problem.
The recap, one line per letter: situation is a customer success team with no designed way to translate a number, payoff is one shared, calibrated question carried into every hard call, anchor is the one page card built from two real numbers and one real miss, risk is either extreme costing a real account, and keep out draws the line at a threshold walkthrough and a sales pitch, never at the numbers themselves.
05And if you want to be sure it really works, try it somewhere else
Same five letters, a municipal water utility instead of an email screen, and this time the blind spot isn't a vendor's invoice. It's a leak nobody can see from the street.
Cinnabar Penrhyn owns the quality bar for Leaklight, an AI tool Yewcroft Regional Water uses to flag meter readings that look like a hidden leak or a tampered meter, so a crew truck gets sent out before a street floods or a bill goes wrong. Yewcroft's council formed its picture of Leaklight the same ungrounded way Kessley Mail's customers did: a neighboring utility's press release claimed it "catches nearly every leak," while a resident's angry call after a truck showed up at a house with nothing wrong left the opposite impression. Hallward Osbaldeston, the council member who owns the water budget, had never seen one real Leaklight flag before either story reached him.
Mapped onto SPARK: situation is a council forming its picture of Leaklight from a rival utility's marketing and one angry resident, never from a real flag. Payoff is one shared question the council carries into every future budget vote: is this the kind of anomaly Leaklight is actually good at, or the kind it only sounds confident about. Anchor is the same shaped card: Leaklight catches about 9 of every 10 real leaks worth a crew visit, and sends a crew to about 1 in 200 homes with nothing wrong, paired with the one thing it still misses most, a slow leak inside a house wall, where the meter barely moves for weeks before a resident notices water damage on their own. Risk runs the same both directions: call it "90 percent, basically solved" and the council assumes every miss is minor, when the in wall leaks are exactly the expensive kind. Keep out draws the same line: no pitch for a bigger Leaklight contract, no walkthrough of the anomaly model's feature weights.
Yewcroft Regional Water: unnecessary crew visits per week, before and after the council session
Before the routing fixAfter the routing fix
Flat around 14 to 16 unnecessary visits a week for six straight weeks, the same weeks the council was arguing about a headline instead of a real flag. Once the session led to routing ambiguous readings to a person before dispatch, unnecessary visits fell by more than two thirds within two weeks, while genuine in wall leaks got caught earlier instead of later.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, the two numbers and the one named miss.
Cost: no budget for a standing quarterly rebuild. Rebuild the card once, tied to the next real renewal or vote on the calendar, rather than a fixed schedule, the calibration still lands, it just isn't refreshed automatically.
The model got better, for real: say Talonguard's first contact catch rate climbs to 80 percent. The card still matters, because Brumley still needs the words to say "which kind of miss is this," even when the honest answer is "a rarer one now." A better model doesn't teach anyone how to ask a sharper question on its own.
Where people run it wrong.
They let an engineer take over the call and start explaining how the classifier's threshold actually works, and the card quietly becomes the technical walkthrough it was built to avoid.
They quote the card's numbers from launch day and never touch them again, so a customer hears "97 out of 100" eighteen months after the real number moved.
They tack a renewal pitch onto the end of the card, and every honest number in it gets read backward as a sales tactic.
How to use it live. Before answering a "how would you explain the quality bar" question cold, ask yourself one thing: what's the one real, current example, a genuine catch and a genuine miss, you'd actually put in front of the room. Naming that example, not a value word like "transparency" or "trust," is usually exactly what the question is listening for.
06Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
Which framework fits a question about designing a way to say something clearly, not reacting to one bad email?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It works forward from how people already form their picture of quality, instead of backward from a single failure.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Ottavine Thorbecke, the AI PM who owns Talonguard's quality bar at Kessley Mail. Brumley Grimwade is the VP of Customer Success who has to answer for it out loud. Wickstead Hardware is the customer who nearly lost money to a phishing email.
3 · THE PAYOFF
What habit does the quality card exist to build?
Tap to flip
ANSWER
Anyone talking about Talonguard, on any call, can say the real catch rate and the real false flag rate in one breath, tied to a real example, and use it to ask a sharper question instead of trusting a mood.
4 · THE ANCHOR
What's the one concrete thing in this answer?
Tap to flip
ANSWER
A one page quality card: catches about 19 of 20 phishing emails, wrongly holds about 1 in 500 real ones, plus one real named miss, rebuilt every quarter against fresh data.
5 · THE OLD DECISION
What old decision would Ottavine take back?
Tap to flip
ANSWER
Quoting customers a single "Trust score: 97 out of 100" badge for about a year. It was the easiest thing to hand a salesperson early on. It stopped being safe once one narrow sender type started carrying almost half of Talonguard's real misses.
6 · THE NUMBER
Fill in the blank: first contact vendor impersonation made up only ___ percent of confirmed phishing attempts last quarter, but accounted for ___ percent of everything Talonguard missed.
Tap to flip
ANSWER
4 percent of attempts, 46 percent of the misses. That gap is the whole reason a flat catch rate hides more than it shows.
7 · THE RISK, SURVIVED
What breaks if the card goes wrong in either direction, and how does the anchor survive it?
Tap to flip
ANSWER
Too simple, just "19 of 20," and people assume the misses spread evenly, when they cluster hard on one sender type. Too technical, the full eval report, and the room checks out before the conversation starts. The card survives both because it names the plain numbers and the one real miss, nothing more.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs SPARK again on a different product. Which one, and what's the equivalent anchor?
Tap to flip
ANSWER
Leaklight, Yewcroft Regional Water's tool for flagging possible leaks. The equivalent anchor is the same shaped card: catches about 9 of 10 real leaks, wrongly sends a crew to about 1 in 200 homes, and names the slow in wall leak as the kind it still misses most.
07Check yourself Score: 0 / 0
True or false
1. True or false: forwarding Wickstead Hardware's ops director the full eval report, precision, recall, F1, AUC, was the safest way to answer her question.
True
False
Show hint
Look at what happened on Brumley's first call, in the story.
Show answer
False. She couldn't read the report, and it made her more scared, not less. A raw eval report is not the same thing as a plain language quality bar.
Multiple choice
2. Why did Talonguard's high overall catch rate, 19 of 20, fail to protect Wickstead Hardware from the invoice fraud email?
A. Talonguard had never been trained on invoice style emails at all.
B. The one blended catch rate hid a much lower catch rate on the smaller, harder category of first contact vendor impersonation.
C. Brumley never enabled Talonguard for Wickstead's account.
D. Salvestro ignored a warning Talonguard had already shown him.
Show hint
Compare the 95 percent headline number to the 4 percent and 46 percent figures in the bar chart.
Show answer
B. The blended catch rate mixed a strong result on common phishing with a much weaker one on first contact vendor impersonation, and the average looked healthy the whole time that gap existed.
Fill in the blank
3. Fill in the blank: for about a year, Kessley Mail's customer success team quoted one number to customers, "Trust score: ___ out of 100," with no real example ever shown.
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
97. That number was true and, on its own, told nobody anything about where Talonguard actually struggled.
Short answer, name the reversal
4. What old decision would Ottavine take back, and why did it make sense when it was first made?
Show hint
Look at the key point box in Let's learn and the same beat in the story section.
Show answer
Model answer: Quoting customers a single "Trust score: 97 out of 100" badge, with no real example behind it. It made sense when Talonguard was new and no customer's account had been tested by a real fraud attempt yet. It stopped making sense once one narrow sender category started carrying real financial cost.
Short answer, apply it yourself
5. Think of an AI product you use or have worked on. What's one real, plain language quality bar, tied to something you already care about, you could give someone instead of a raw accuracy number?
Show hint
Pair a plain language rate with one real, named example of what it still gets wrong.
Show answer
Model answer: A resume screening tool could say "shortlists about 8 of 10 qualified candidates, and about 1 in 30 shortlisted resumes turns out to be a poor fit" instead of quoting a precision score of 0.91.
Short answer, work the number
6. If first contact vendor impersonation doubled from 4 percent to 8 percent of all confirmed phishing attempts, while still causing 46 percent of the misses, would tightening the review threshold for that category alone still be the right fix? Why or why not?
Show hint
Think about how much volume the guardrail now has to handle, not just the share of misses.
Show answer
Model answer: Yes, and it becomes more urgent. The category causing the concentrated misses now also makes up a bigger share of total attempts, so the same guardrail has to review more volume, not just the same volume, which means the manual review step needs more capacity, not a different fix.
Before you close the answer
Why this works
Tests whether you can turn a probabilistic number into something a non technical person can actually act on, not just whether you can define precision and recall. It's also testing whether you know a single blended number can hide a concentrated, expensive failure.
Follow-up traps
"Isn't a plain language card just as much of an oversimplification as the trust score badge was?" Response: no, because it names the real miss category and gets rebuilt every quarter. A badge with no example and no update cycle is what actually oversimplifies.
"What if the customer still doesn't trust the number right after the fraud email?" Response: that's expected once, and fine. The card's job is giving them something specific to check the next email against, not making them feel good about this one.
If pressed
Talonguard checks whether it has seen three or more prior messages from that exact sender domain before it treats a vendor claim as lower risk. Anything under three prior messages gets the stricter routing threshold, which is why a real long time supplier's invoice almost never gets held, only a brand new one does.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.