CaseIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #18
Your feature works well for ninety percent of users and badly for a vocal ten percent. Communicate that.
TRACE · a vocal ten percent, tested on Glyphcast's alt text for Quorum Weekly
Glyphcast writes a description of every image a partner newsroom publishes, read aloud by a screen reader for blind and low vision readers. Neelima Tarrant owns that product. Nine out of ten readers trust it. This is the week she found out what the tenth one actually needed to hear.
The direct answer
Do not tell leadership "it's just ten percent, ignore it," and do not tell them "pull the whole feature." Check whether the unhappy ten percent share a real, checkable trait first. If they do, name it specifically and name the fix, to leadership and to the affected readers both. If they genuinely do not, say that plainly too. Only one of those is a guess dressed up as an answer.
Do this, in order
Find out whether the ten percent share a real trait before you say anything to anyone.Why: calling them noise and pulling the feature are both guesses about the cause. Only a real check tells you which guess, if either, is true.
Slice the complaints by every real dimension you have, not just the loudest one.Why: a ten percent average can hide one group that is almost entirely unhappy sitting next to a group that never had a problem.
Check the actual output against the actual image, not whether the writing sounds fluent.Why: a wrong description of a chart reads exactly as calm and clean as a right one. Fluent is not the same thing as true.
If a real pattern shows up, name it out loud, specifically, to leadership and to the readers it hits.Why: naming the true story is what tells the ten percent you actually understood the problem, not just heard how loud they were.
Leave the parts that are already working alone.Why: slowing down or hedging the ninety percent who are fine only punishes readers who never had a problem to begin with.
If the check turns up no pattern at all, say that honestly instead of inventing one.Why: a made-up explanation you don't actually have is its own kind of dishonesty, even if it sounds tidier.
How to answer this, stage by stage
Nobody is grading whether you know screen readers. They're grading whether you'll find the real shape of the ten percent before you decide how to talk about them.
1
Scope it to one product, one reader
Say it like this
"Let me ground this in one real case. Glyphcast is a plug-in newsrooms install once. Every image they post gets a written description a screen reader can read out loud, so a blind or low vision reader knows what a photo or a chart actually shows, not just that one exists. Nine out of ten readers rate it accurate enough to trust. I want to talk about the tenth."
Why this works
One real product with a real reader on the other end stops the answer from turning into a lecture about handling unhappy customers in general.
2
Reframe what's actually being asked
Say it like this
"The real question isn't how to word a message to calm ten percent of people down. It's whether that ten percent is one thing or ten different things wearing the same number. Those need completely different messages, and on day one they look identical on a dashboard."
Why this works
Naming the reframe up front shows the interviewer you're not about to write a press release before you've checked what's true.
3
Say your structure out loud
Say it like this
"I'd run this as TRACE. Timeline: when the complaints actually started building, not just when someone noticed. Recut: slice the unhappy ten percent every way I can. Assume nothing: don't default to 'they're just picky' or 'something's badly broken.' Cause candidates: three honest guesses at what's really going on. Evidence test: the one check that tells me which guess is real."
Why this works
Two seconds of structure tells the interviewer this is a method, not a reflex dressed up as one.
4
Give the decision, committed
Say it like this
"So here's what I'd actually do. I would not tell leadership to ignore the ten percent, and I would not tell them to pull the feature. I'd pull a real sample of what Glyphcast wrote, check it against what the images actually show, and slice the results by everything I can, image type, publication, device, screen reader software. Whatever that turns up is what I communicate. Not before."
Why this works
This is the direct answer, said plainly, before a single detail of the story shows up.
5
Recut before you assume anything
Say it like this
"Before I check a single number, I want to name three honest guesses. One, it's random. Some slice of readers is just unlucky, spread evenly, and it's not really about any one kind of image. Two, it's really about one publication's audience, for reasons that have nothing to do with what Glyphcast wrote. Three, it's a genuine, identifiable weakness in the model on one kind of image, and it shows up wherever that kind of image shows up. All three would look like the exact same ten percent on a dashboard."
Why this works
Naming all three before checking anything is what keeps the diagnosis honest instead of a foregone conclusion.
6
Run the evidence test, numbers first
Say it like this
"Here's what I actually found. I pulled four hundred Glyphcast descriptions across six partner newsrooms and had a sighted reviewer check each one against the real image, not against whether the sentence read cleanly. Photos came back ninety six percent accurate. Illustrations, ninety three. Charts and data visualizations, sixty one. That gap showed up at every single newsroom that published charts, not just the one where the complaints were loudest."
Why this works
A real number, checked against the real image, is what makes the evidence test checkable instead of a guess about which story sounds more convincing.
7
Name the trade-off, then close
Say it like this
"One thing worth saying straight: the honest fix is slower for charts specifically. Routing them to a verified pipeline can take up to a day instead of a second. I'd take that trade every time, on that slice only, because handing a blind reader a wrong number about a vaccine chart, stated with total confidence, is worse than making them wait. So: find out what the ten percent actually share before you say a word, then say exactly that."
Why this works
Naming the real cost keeps the close honest instead of a tidy resolution nobody would believe.
Let's learn
Glyphcast is a plug-in a newsroom installs once. Every time it posts a new image, Glyphcast writes a description, so a screen reader can read it out loud and a blind or low vision reader knows what the image actually shows.
The second box is the only place in this whole chain where the picture actually gets turned into words. Everything after it just carries whatever that box decided.
Before Glyphcast, only about twelve percent of images across these newsrooms ever got a real description at all. A person writing an accurate one for a data chart took about six minutes, and nobody had six minutes for the fortieth chart of the day. Most images just went undescribed. With Glyphcast, close to every image gets a description, written in under a second.
Knowledge spark: why would a model do worse on charts specifically?
A model only gets good at what it has seen a lot of. Glyphcast was trained mostly on photos, over seventy percent of its examples. Charts and data visualizations made up about four percent. Ask it to describe a photo and it has seen a million like it. Ask it to describe a chart and it has barely seen the shape of the problem, so it guesses, confidently, in the same calm voice either way.
Six weeks after launch, a survey across every partner newsroom came back strong: ninety percent of readers said Glyphcast's descriptions were accurate enough to trust. Here's the turn. The other ten percent were not just a little less happy. Their real problem was that Glyphcast handed them a wrong fact, stated as if it were certain, about the one kind of image they had no other way to check.
We didn't just get one chart wrong. We handed her a wrong fact and asked her to trust it.
What it costs at its worst: leadership hears "ten percent" and does nothing, so a small, real group of readers keeps getting fed wrong numbers about things like vaccine efficacy and public health data, quietly, indefinitely. Or leadership panics and pulls the whole alt text feature company-wide, and every newsroom drops straight back to twelve percent image coverage, taking accurate, instant descriptions away from the ninety percent to react to a problem that only ever lived in one kind of image.
The choice I would take back
Glyphcast's feedback button only ever logged one thing: a thumbs up or down per reader, added into one aggregate score. It never recorded which image, or what kind of image, a vote was about. That was fine when every partner published roughly the same mix of photos and charts. It stopped being fine the day one partner started publishing far more charts than anyone else, because the model's real weakness had nowhere to show up except as a slightly duller version of one company-wide number.
What I would leave alone: Glyphcast's plain photo and illustration descriptions, verified accurate on ninety three to ninety six percent of outputs, written in under a second. Slowing those down with an extra review, just because charts turned out to be unreliable, would only cost the ninety percent who were never the problem.
The lesson: an average hides a cliff. The only way to find the cliff is to ask what's actually different about the people standing on it, not how loud they happen to be.
Now here is the same thing as a story
Say the short version out loud in an interview. Read this one when you want to feel exactly how one wrong chart, repeated once out loud in a meeting, turns into a phone call nine weeks later.
Every weekday morning at half past six, Suhaila Naidoo opens Quorum Weekly before she opens anything else.
Before Glyphcast, a chart she needed for work meant a phone call to the newsroom's research desk, and however long they took to answer.
She's read data for a public health nonprofit for nine years, blind since birth, and she moves through her screen reader at a speed most sighted colleagues can't follow by ear. Long before Glyphcast, she'd built her own workaround for Quorum Weekly's charts: call the research desk, ask them to describe the numbers, wait. Sometimes she had an answer in an hour. Sometimes two days, by which point the newsletter had moved on to the next issue.
Glyphcast arrived at Quorum Weekly in April. For the first two months, it was the best change to her mornings in years. She could read a chart the same day it published, out loud, with nobody standing between her and the numbers. She checked the first dozen or so against a sighted colleague, just to see. They all held up. So she stopped checking.
She stopped double-checking the simple bar charts first, because they always matched. Then the line charts. By August she was reading anything Glyphcast described and repeating it as fact, the same way she'd repeat a number from her own spreadsheet.
In week nine, in a funding meeting at her nonprofit, she cited a chart from that week's Quorum Weekly: vaccine efficacy in a certain age group, she said, had declined toward the end of the study period, according to what Glyphcast had read her. A colleague, checking the same report afterward for the meeting notes, came back to her quietly the next morning. The chart hadn't shown a decline. Efficacy held steady, near the top of its range, for the whole period. Glyphcast had described it backward.
Neither reaction actually asks what the ten percent have in common. That's the part TRACE is for.
Suhaila didn't say anything publicly that week. She went back through two months of Glyphcast descriptions instead, chart by chart, cross-checking each one with the same colleague. She found six that were flatly wrong, all from Quorum Weekly, all charts. She wrote it up in detail on an accessibility forum she'd used for years: what Glyphcast said, what the chart actually showed, side by side, six times.
Suhaila didn't lose an image. She lost a fact she'd already repeated out loud, in a room, to people making a decision.
A Glyphcast support lead forwarded the post to Neelima Tarrant with one line: "You're not still calling this ten percent noise, are you?" By then the numbers already told a story if anyone had looked. Support tickets citing a wrong or misleading description had been climbing for weeks, not spiking on one bad day.
Nobody could point to one bad day. The number just kept climbing, quietly, for ten weeks before anyone treated it as a pattern instead of background noise.
Support tickets citing a wrong or misleading description, by week
Weekly ticket countAggregate survey, 90% positiveDetailed public thread
The ninety percent survey and the climbing ticket count happened in the same week. The aggregate number was true. It just wasn't the whole story.
Neelima didn't guess who was right. She pulled a real sample: four hundred Glyphcast descriptions across all six partner newsrooms, checked one by one by a sighted reviewer against the actual image, not against whether the sentence sounded reasonable.
Four of the five slices came back flat. The fifth one, image category, lit up everywhere charts appeared, not just at Quorum Weekly.
Verified-accurate rate, by image category, across all six partner newsrooms
Holding up across every partnerThe one slice that cratered
The average across all four categories still lands around ninety percent. The reader who only ever sees charts never experiences the average.
The old decision came back to Neelima the way you remember a meeting rather than a policy. A year earlier, a small team had debated whether to build image type tags into the feedback button. Someone said the team was small enough to just watch the one aggregate score and notice if it dropped. Nobody built the tags. Nobody had to, for a year, until one partner started publishing far more charts than anyone else and the model's real weakness had nowhere left to hide except inside an average.
Neelima called Suhaila first, not to explain the number away, but to ask exactly which charts had been wrong. Suhaila sent her thread again, six examples, chart and description side by side. Every one of them was a chart. None of them was a photo.
The loudest, most specific complaint came from the reader who read the most charts. That correlation is the whole diagnosis in one picture.
She called Aldena Choi, Quorum Weekly's editor, the same afternoon. Quorum published roughly forty percent of its images as charts and data visualizations, against under five percent for the average partner. That single fact, not anything about Quorum's readers being harder to please, was the real reason its readers had been carrying almost all of Glyphcast's chart problem alone.
TRACE, so a vocal ten percent gets a real answer instead of a guess
Not a way to decide whether Suhaila was right to be upset. TRACE is what stops "she's just one loud reader" and "the whole feature is broken" from getting communicated with equal confidence, when only one of them was ever actually checked.
TTimeline. Lay out when the pattern actually started, not when someone noticed.
Glyphcast launched in April. The six-week survey in May showed ninety percent satisfaction, a true number. Ticket volume had already begun climbing before that survey ran, and kept climbing for four more weeks before Suhaila's thread in week ten made it impossible to read as background noise.
The gap that mattered wasn't the thread. It was the four weeks between the survey and the thread, where the real pattern was already visible and nobody had looked yet.
RRecut. Slice the ten percent every real way you have.
Neelima sliced complaints by screen reader software, device, region, which publication, and image category. Four of those five came back flat, no signal. The fifth, image category, showed the same pattern everywhere charts appeared, at every partner, not only at Quorum Weekly.
A ten percent average that looks the same across every publication except one is not evidence about that one publication. It's evidence about whatever that publication happens to publish more of.
AAssume nothing. Neither easy story gets the benefit of the doubt.
Neelima didn't assume the ten percent were simply pickier readers, because a minority complaining loudly is not proof they're wrong. She also didn't assume Quorum's readers were right just because they were specific and organized. A specific, well-written complaint and a genuinely random unlucky streak can both look identical from across a support queue.
Both easy guesses are cheap to reach for and expensive to be wrong about. Neither one gets to stand in for a real check.
CCause candidates. Three honest guesses, checked against real evidence.
One, it's random: an unlucky slice of readers, no real pattern by image type. Two, it's about Quorum specifically, something about its audience, unrelated to what Glyphcast actually wrote. Three, it's a genuine, identifiable weakness in Glyphcast's model on charts and data visualizations, showing up wherever that kind of image appears, and Quorum just happens to publish the most of them.
Only the third one survived contact with the four hundred sample. It wasn't a hunch. It was the only explanation the evidence actually supported.
EEvidence test. Check the output against the real image, not against how it reads.
Pull a real sample across every image category and every partner, and have someone who can see the image check whether the description is actually true, not whether it sounds fluent. Wrong-but-confident reads identical to right on the page. Only checking it against the real picture tells the two apart.
This is the strongest move in the whole framework. It's checkable against a real image, not a guess about which explanation is more convenient to believe.
Both readings of the same number are honest until you check. Only one of them survived the check here.
The fix isn't a blanket warning on every image. It's one new fork, added at exactly the point where the model's own weakness actually lives.
Three things worth saying plainly, since interviewers push here. Neelima's team first floated a faster option: put a short disclaimer on every single description, "written by AI, may be wrong," and call it honesty. She rejected it, because it would have added a sentence a screen reader reads aloud on every one of maybe thirty images in an issue, forever, for the ninety percent who were never the problem, and it still wouldn't tell the ten percent which images to actually worry about. The AI-specific failure worth naming by name: Glyphcast's confident wrongness on out-of-distribution images, charts and data visualizations made up about four percent of its training examples against over seventy percent photos, so it answers a chart with the same calm certainty it uses on a photo, with far less reason to. The guardrail is the new fork itself: chart-type images route to a verified pipeline, checked against a labeled sample continuously, not just an aggregate satisfaction score. And the trade-off, accepted on purpose: charts take up to a day instead of a second, in exchange for a reader never again hearing a wrong number about something like vaccine data stated as settled fact.
And if you want to be sure it really works, try it somewhere else
Same five letters, a grain co-op instead of a newsroom, and this time the honest finding runs the other way: no pattern at all, just a genuine toss-up the model itself already knew about.
Bushelworks scans a camera sample of every truckload of grain that pulls up to a co-op elevator and grades it for moisture, foreign material, and damage, a grade that sets the price paid to the farmer that load. Godfrida Larrabee runs product there, and hit her own version of Neelima's week during her first harvest season: ninety percent of loads match what a human grader would call. The other ten percent get disputed, loudly, sometimes right at the scale house window with three trucks waiting behind.
The same two honest guesses apply here. This time the check pointed the other direction entirely.
Godfrida ran the same recut Neelima did: crop variety, moisture sensor, time of day, driver, truck type. Every one of them came back flat. Disputes weren't concentrated in one crop, one sensor, one shift, or one driver's routes. What they were concentrated in was Bushelworks' own confidence number: nearly every disputed load sat inside the same narrow band, the range where the model itself was already least sure.
Disputed grain loads, by the model's own confidence at the time of grading
Accepted loadDisputed load
Height on the page is only spread, so the dots don't overlap. The one thing that matters is left to right: every disputed load sat inside the same narrow band, no matter whose truck it was.
Godfrida's team briefly considered publishing a table of "what disputed loads have in common" anyway, because farmers wanted an explanation and a real pattern is easier to sell than genuine uncertainty. She rejected it. Inventing a tidy cause that the evidence didn't support would have looked better for one meeting and cost more trust the first time a farmer noticed the pattern didn't actually hold.
The decision Godfrida would take back
Bushelworks showed farmers a single grade and a single price, never the model's own confidence at the time it graded. A number that told the truth, "this one's close, we're not fully sure," was sitting inside the system the whole time. It just never got shown to the one person it would have mattered to.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: don't dismiss or panic, check whether the loud ten percent share a real trait before you say anything.
Cost: no time to run a real sample. Ask one question instead: does this show up wherever the same content type appears, or only in one place?
The model got better, for real: say Glyphcast's chart accuracy climbs to ninety percent next year. The evidence test still runs exactly the same way. It just moves what the honest answer turns out to be, not whether you need to check.
Where people run it wrong.
They pick a story, noise or crisis, before they've checked anything, because both stories are easier to say out loud than "I don't know yet."
They slice the data by whoever complained loudest instead of by every real dimension, and mistake a loud subgroup for the whole pattern.
They find a real pattern and then hedge everything anyway, punishing the readers who were never affected along with the ones who were.
How to use it live. When an interviewer throws this at you cold, buy two seconds by asking one thing back: "do we know yet if the ten percent share anything in common, or is that still open?" That question alone usually is exactly what a question shaped like this one is listening for.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits figuring out why a loud ten percent is unhappy?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. Built to find the real cause before you decide how to talk about it.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Neelima Tarrant, who owns Glyphcast's image description product, and Suhaila Naidoo, a blind data analyst and six-year screen reader user who read Quorum Weekly's charts through it.
3 · THE TIMELINE
What happened by week six, and what happened by week ten?
Tap to flip
ANSWER
Week six: an aggregate survey showed ninety percent of readers found Glyphcast accurate enough to trust. Week ten: Suhaila posted a detailed catalogue of six wrong chart descriptions from one newsletter.
4 · THE RECUT
Name the three cause candidates Neelima had to rule between.
Tap to flip
ANSWER
Random, unlucky readers with no real pattern. Something specific to Quorum Weekly's audience, unrelated to the images. Or a genuine weakness in Glyphcast on charts and data visualizations, showing up everywhere they appear. Only the third one held up.
5 · THE OLD DECISION
What decision would Neelima take back?
Tap to flip
ANSWER
Glyphcast's feedback button logged one aggregate score per reader, never which image or image type a vote was about. Fine with a similar mix of images across partners. It stopped being fine once one partner published far more charts than anyone else.
6 · THE NUMBER
Fill in the blank: Glyphcast's descriptions were verified accurate on ___ percent of charts, against ___ percent or higher on everything else. Charts made up about ___ percent of its training data.
Tap to flip
ANSWER
61 percent. 90 percent. 4 percent. That last number is the actual reason charts, specifically, needed a different pipeline.
7 · THE EVIDENCE TEST
What's the one check that told Neelima this was real, not noise?
Tap to flip
ANSWER
Pull a real sample of descriptions across every image category and every partner, and have a sighted reviewer check each one against the actual image, not against whether the sentence reads smoothly.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what did the recut find that time?
Tap to flip
ANSWER
Bushelworks, a grain-grading tool for a co-op elevator, run by Godfrida Larrabee. That time the honest finding was the opposite: no pattern by farmer, crop, or sensor, just genuine uncertainty near the model's own confidence line.
Check yourself Score: 0 / 0
True or false
1. True or false: the fact that a vocal ten percent complained proves Glyphcast's ninety percent satisfaction number was fake.
True
False
Show hint
Look at the chart note under the ticket volume line chart.
Show answer
False. The aggregate number was true. It just wasn't the whole story. A real average and a hidden cliff underneath it can both be true at the same time.
Multiple choice
2. Which of the three cause candidates actually explained what Neelima found?
A. A software bug in the screen reader itself.
B. Quorum Weekly's readers being unusually demanding, regardless of what the images were.
C. Glyphcast's model doing genuinely worse on charts and data visualizations, wherever they appeared.
D. Random bad luck, spread evenly across every kind of image.
Show hint
Check the verified-accurate rate by image category, across all six partners, not just Quorum.
Show answer
C. The chart accuracy gap showed up at every partner that published charts, not only at Quorum Weekly. That's what ruled out both the "just this newsroom" guess and the "just random" guess.
Fill in the blank
3. Glyphcast's descriptions were verified accurate on ___ percent of charts and data visualizations, and ___ percent or higher on everything else Neelima checked.
Show hint
Look at the bar chart under "Let's learn" and "Now here is the same thing as a story."
Show answer
61. 90. Photos, illustrations, and screenshots all landed between 90 and 96 percent. Charts alone dropped to 61.
Short answer, apply it yourself
4. Think of a product you use where a small group of users seem to complain much more than everyone else. What's one question from this answer you'd ask before deciding they're "just loud"?
Show hint
Think about what they might share, not how they might be different as people.
Show answer
Model answer: Does what they're complaining about share a specific trait the rest of us don't have, a device, a use case, a kind of content, rather than assuming they're simply more sensitive or more vocal by nature.
Short answer, where it wouldn't matter
5. Name a part of Glyphcast Neelima should NOT slow down or hedge just because charts turned out to be unreliable. Why not?
Show hint
Look at "What I would leave alone" in Let's learn.
Show answer
Model answer: Plain photo and illustration descriptions, verified accurate on 93 to 96 percent of outputs, near instant. Adding a review step there would only cost the ninety percent who were already being served well, for a problem that lives somewhere else.
Short answer, the number question
6. If Glyphcast's next model update genuinely lifts chart accuracy from 61 percent to 90 percent, can Neelima stop tracking accuracy by image category and go back to watching one aggregate number? Why or why not?
Show hint
Think about how the 61 percent problem got found in the first place.
Show answer
Model answer: No. The 61 percent was hiding inside one aggregate number for ten weeks before anyone sliced it out. Dropping that tracking just because one category improved risks the exact same blind spot opening up on some other image type later, with nobody watching for it.
Before you close the answer
Why this works
Tests whether you'll dig for a real, checkable pattern before deciding how to talk about a vocal minority, instead of reaching for whichever story, noise or crisis, is easier to say out loud. Most candidates skip straight to a message and never check if there's a real thing to communicate.
Follow-up traps
"Isn't slicing the data by chart-heavy publishers just cherry-picking a story that fits the loudest complaint?" Response: no, because the check ran across every category and every partner, not just Quorum, and the same chart weakness showed up wherever charts appeared, not only where the loudest complaint happened to come from.
"What if leadership wants an announcement before you've confirmed anything?" Response: say what's actually known today, the aggregate is strong, one content type is weak, and commit to a date for the plan, not a date for a fix that hasn't been verified. A wrong promise to blind readers about a fix is exactly the kind of confident wrongness this whole answer is trying to avoid.
If pressed
The four hundred output sample was checked by sighted staff reading full resolution images, not the compressed thumbnails Glyphcast's model actually sees when it writes a description in production. The real gap on charts at production image quality may run worse than 61 percent, a real limit inside the evidence test itself, not yet re-checked.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.