CaseIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #12

A stakeholder cherry-picks one bad output as proof the feature is broken. How do you respond?

GUARD · what one photo proves, and what it doesn't

Latent is Nitrate Studio's tool for restoring and enhancing old or damaged photographs. Brennus Kestner, the product manager who owns Latent's restoration model, tracks how it performs by damage type most mornings before anything else. Tabitha Castillon, Nitrate Studio's VP of Customer Success, found one badly warped photo and posted it straight into the leadership channel. Fintry Larsholm just wanted her grandfather Horace's eye to look like an eye again before Saturday.

The direct answer
Pull the exact photo and find out what actually went wrong: which kind of damage, and whether it's a pattern Latent already tracks. Check that pattern's real fail rate in the eval set before deciding the complaint proves the feature is broken, or deciding it proves nothing at all. Then answer with both things together: the fix for this one photo, and the real number for how often this happens, not a verdict built on a single image.
Do this, in order
  1. Pull the exact photo and name the damage subtype before reacting either way.Why: "the feature is broken" and "this is nothing" are both guesses until you know if it's a known, tracked pattern.
  2. Check that subtype's real fail rate in the eval set, not the headline pass rate.Why: a 96 percent number can sit on top of a subtype failing 15 percent of the time and nobody would see it.
  3. Fix the specific photo on its own, the same day if you can.Why: whatever the aggregate says, Horace's photo still has to come out right for Fintry by Saturday.
  4. Answer the stakeholder with both numbers together, the fix and the real rate.Why: either number alone reads as a dismissal or a panic. Together, they read as an answer.
  5. Log the ticket against the damage taxonomy so it counts, every time, not just this once.Why: a team that only checks the real rate when an executive is watching never actually fixes the loudest-voice problem.
  6. Leave the wider rollout running for the 97.5 percent of photos this subtype doesn't touch.Why: pausing the whole feature over one rare, already-tracked pattern costs every unaffected customer, for nothing.

How to answer this, stage by stage

Nobody is grading whether you think this customer is right. They're grading whether you can turn "one bad photo" into a specific, checkable decision instead of a gut call in either direction.

1
Scope it to one product and one photo
Say it like this
"Let me make this concrete. Say a company builds an AI tool that restores old, damaged photos. One customer's grandfather's photo comes back with a warped eye. A VP sees it and posts it to the whole leadership channel as proof the feature's broken. That's the scenario I'll run, because 'a stakeholder cherry-picks one bad output' means nothing until there's an actual photo and an actual Slack message."
Why this works
Keeps the answer from turning into a lecture about trusting anecdotes in general.
2
Say your structure out loud
Say it like this
"I'll run this as GUARD. Groups: who's actually carrying risk in each direction. Unequal: what gets under-fixed if the loudest complaint always wins the sprint. Ability to contest: can the team check this fast, or is trusting the VP the only move available. Reduce: the actual fix. Detect: how you'd know it's being handled right."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion about which side to believe.
3
Reframe what the question is actually testing
Say it like this
"This isn't really 'is the photo fake or real.' It's real, and it's genuinely bad. The real question is whether it's representative: is this a known, elevated-risk pattern, or a rare one-off dressed up as a trend because one senior person happened to see it."
Why this works
Separates this from a generic "handle an upset customer" answer and locates the real judgment: is one output a fair sample of the model's real behavior on this kind of input.
4
Give the one decision
Say it like this
"Here's what I'd actually do: pull the photo, run it against Latent's damage taxonomy, find out this is a 'crease crosses the eye' case, and check how often that specific pattern fails in the eval set. Then I answer with both numbers, the fix for this photo and the real rate, not one or the other."
Why this works
This matches the direct answer word for word. If it doesn't, the interviewer notices before you do.
5
Prove it with the compressed failure, both directions
Say it like this
"Here's what happens if you dismiss it: the crease-over-eye pattern really does fail 15 percent of the time, and nobody with the power to fix it took the complaint seriously. Here's what happens if you panic instead: you pause auto-enhance for every customer over a pattern that hits 2.5 percent of photos, while the real bigger problem, a skin-tone shift on sepia photos, keeps failing quietly because nobody senior has personally seen it yet."
Why this works
Shows both failure modes are real, which is the entire tension the question is testing.
6
Close on what you'd watch and what you'd leave alone
Say it like this
"I'd track what share of complaints get logged against a known subtype and checked for real frequency, every time, not just when someone senior is watching. And I'd leave the wider rollout alone. It's not the part that's actually broken."
Why this works
Ends on something the interviewer could go verify, and shows judgment about scope instead of blanket caution.

Let's learn

Latent is an app that takes a photo of an old, torn, faded, or water-stained picture and hands back a restored, sharpened, sometimes colorized version, usually inside a minute.

Hand sketched comparison diagram titled Whose word moves the roadmap right now. Left, a person icon labeled Tabitha, caption one post in the exec channel, the roadmap moves same day. Right, a person icon labeled Brennus, caption a printed eval sheet, no channel built yet to make it count.
Neither person here did anything wrong. One of them just had a faster way to be heard than the other.

Nitrate Studio checks Latent's restorations the way a lab checks film: a golden eval set of 2,400 damaged photos, each one graded against the original by a professional retoucher on a fixed rubric. Overall, Latent passes 96 percent of them. That's the number on the company's homepage, and it's a real, checked number.

Knowledge spark: what's a damage subtype? Not every damaged photo breaks the same way. A crease, a water stain, a torn corner, a faded sepia tone: Latent's eval set tags each photo by which kind of damage it has, so the team can check accuracy for each kind separately, not just as one blended number.

Inside that 96 percent, the subtypes don't all perform the same. Sixty of the 2,400 eval photos have a crease or fold line running directly across an eye. On those, Latent gets it wrong 15 percent of the time: the model fills in a plausible-looking eye where the crease destroyed the real one, and sometimes that guess comes out warped, doubled, or slightly misplaced. A much bigger slice, 745 photos, are sepia-toned originals where the model has to guess a real skin color from a monochrome tint with no ground truth to check against. There, it's wrong 5.5 percent of the time, a lower rate, but on so much more volume that it accounts for 41 of the eval set's 96 total failures, against 9 for crease-over-eye.

Hand sketched quadrant diagram titled Which damage subtype actually deserves the sprint. X axis, how many photos hit this subtype, from rare to common. Y axis, how often it fails within its own subtype, from low to high. Crease over an eye sits rare and high failing. Sepia skin tone shift sits common and moderately failing. Torn corner and ordinary faded photo sit low failing.
Crease-over-eye fails more often when it happens. Sepia shift happens so much more often that it causes more real failures overall.

Here's the turn. On a Tuesday morning, Fintry Larsholm uploaded a water-and-crease-damaged photo of her late grandfather Horace for a family reunion slideshow due that Saturday. The crease crossed his left eye. Latent's restoration came back with the eye rendered strange, a doubled eyelid, an iris sitting slightly off-center. Fintry emailed support, upset, the original and the bad output side by side. A support rep looped in Tabitha Castillon, who runs Customer Success, because of the deadline. Tabitha didn't check the eval set. She posted the two photos straight into the #leadership channel at 8:50 that morning: "This is what we're shipping. This is broken. Pull auto-enhance before it wrecks someone else's weekend."

One photo is real evidence. It is not, on its own, a verdict on 2,400 others.

At its worst, this goes two ways. If Brennus waves the photo off as noise, the crease-over-eye pattern keeps failing 15 percent of the time on exactly the photos where it matters most, a face, and nobody with the power to fix it ever logs a reason to. If the team instead pauses auto-enhance for every customer that Friday, on the strength of one photo, the 97.5 percent of restorations that were never at risk get delayed too, while the sepia shift problem, the one actually responsible for 43 percent of all real failures, still gets no attention at all, because it never produces a photo dramatic enough for anyone to screenshot.

The choice I would take back Before this, a bad-output complaint had two paths: get closed as a one-off, or get escalated based on how upset or how senior the person raising it was. Nothing connected an individual ticket to Latent's own damage taxonomy, so nobody could answer "how often does this actually happen" in the moment it mattered. That was fine when complaints were rare. It stopped being fine the day one of them reached a VP with a Slack channel and a deadline.
What I would leave alone The wider Friday rollout of auto-enhance to all users never needed to pause. It was never built on the crease-over-eye subtype specifically, and that subtype touches only 2.5 percent of photos coming through.

The lesson: a single bad output is never nothing, and it's never proof either. It's a data point that needs a distribution to sit inside before anyone can say what it means. The team's job isn't deciding whether to believe the photo. It's building a fast way to find out.

Now here is the same thing as a story

Stage five above compresses this into four sentences. Here's the rest of that Tuesday, the part a stand-up answer skips.

Brennus Kestner has owned Latent's restoration model for two years, long enough to know its weak spots better than its marketing page does. Most mornings, before his coffee's finished, he pulls up the eval dashboard, sorted by damage subtype, just to see if anything's drifted since yesterday.

Fintry Larsholm found Latent three weeks earlier, looking for a way to fix the only photo she had of her grandfather Horace before he died, a water-stained print with a deep crease folded straight across his left eye for sixty years in a shoebox. She uploaded it on a Tuesday morning, five days before a family reunion slideshow, and watched the restoration come back in under a minute. It was close. The eye wasn't. Where the crease had been, Latent had filled in something that read, on a phone screen, like a second eyelid layered over the first.

She wrote to support at 8:20 that morning, original and result side by side, polite but clearly rattled: "Is this what my grandfather's photo is going to look like at his own reunion?"

The rep on shift, new enough to be unsure what counted as urgent, forwarded it to Tabitha Castillon with a note: "family reunion Saturday, thought you'd want to see this." Tabitha did want to see it. What she didn't have, standing in the support queue at 8:47 with a screenshot in hand, was any way to ask how often this actually happens. So she asked the only question she could answer on her own: does this look bad. It did. Three minutes later it was in #leadership, tagged to the CEO, with a demand to pull the rollout scheduled for Friday.

Brennus saw the post at 9:05. He didn't argue with the photo. It was real, and it was genuinely bad. What he did instead was open the eval dashboard and filter for the damage tag that matched: crease crosses eye region. Sixty photos in the golden set carried that tag. Nine of them failed. Fifteen percent, against a 96 percent number the whole company was used to trusting as the whole story.

We hadn't lost the photo's meaning. We'd just never built a way to look it up.

He also pulled the wider breakdown, the one nobody had asked for that morning. The sepia skin-tone-shift subtype, 745 photos, failed 5.5 percent of the time, a lower rate, but on so much more volume that it was responsible for 41 of the eval set's 96 total failures, against 9 for the pattern in Tabitha's screenshot. Nobody had ever posted a sepia photo to #leadership. It doesn't look grotesque. It just looks a little off, the kind of thing a customer mentions once in a support ticket and nobody escalates.

Hand sketched labeled parts diagram titled What one bad restoration could actually mean. A central document icon labeled Horace's photo, with four callouts around it: a known tracked subtype, a brand new failure, one off noise rare, stale build already fixed.
Before Brennus checked, the photo could have meant any of these four things. Only the eval set could say which one it actually was.

By 9:40 he had a reply ready for the leadership channel, and it carried two things together on purpose. First: Horace's photo was being reprocessed by hand right now, and Fintry would have a corrected version before lunch. Second: the crease-over-eye pattern was a known, tracked subtype, currently failing 15 percent of the time on a slice that touches 2.5 percent of all photos, and it was already the top item in next sprint's queue, not because of the Slack post, but because the rate had already crossed the team's own threshold two weeks earlier.

Fintry had her grandfather's photo back, eye repaired, by 12:15. The Friday rollout shipped as planned. And the sepia subtype, the one nobody had ever screenshotted, moved to the top of the actual priority list the following sprint, on the strength of the number, not a photo.

Hand sketched comparison diagram titled One ticket, two ways to triage it. Left, a question mark icon labeled Before, caption gut call, ignore it or forward it up, no number attached. Right, a document icon labeled After, caption logged to the subtype, checked against the real rate.
Same ticket, same photo. The only thing that changed was whether a number got attached to it before anyone reacted.

The decision Brennus would take back sits further back than that Tuesday. The team had tried a version of "just pause it" once before, three separate times last quarter, in fact, each one triggered by a single dramatic screenshot that turned out, on closer look, to be a known, already-mitigated rare case. One of those pauses cost a scheduled colorization launch two full weeks. All three taught support to escalate everything, just in case, which made the noise worse, not better.

So the real fix wasn't a rule about this one photo. It was a rule about every future one: any ticket with a bad output now gets logged against the damage taxonomy as a required field, and any reply to whoever raised it has to carry both numbers, the specific fix and the real rate, every time, not only on the mornings a VP happens to be watching.

What I'd tell myself, standing in Brennus's spot at 9:05 that morning: the photo was never the thing to argue with. The argument was always going to be about whether anyone had a fast way to check it against everything else.

GUARD, run against one Slack post and one photo

This was never really about whether Tabitha was right to be alarmed. GUARD is for naming the real risk sitting in both directions at once, and building a fix specific enough to survive the next photo too.

GGroups. The real risk sitting on each side.
Dismiss the photo outright, and the risk lands on customers hitting a genuine, elevated-risk pattern the eval set already knows about: 15 percent of crease-over-eye restorations, a real failure the team could catch and isn't. Overreact instead, pausing the whole feature on one image, and the risk lands on every customer waiting on the other 97.5 percent of photos, plus the team itself, which just taught engineering that the roadmap moves on volume of complaint, not severity.
Neither direction is free. Naming both, before reacting to just the loudest one, is the whole first move.
UUnequal. What gets under-fixed when the loudest voice always wins the sprint.
The sepia skin-tone-shift subtype causes 41 of the eval set's 96 total failures, more than four times the crease-over-eye pattern's 9. It has never been the subject of a leadership Slack post, because a slightly wrong skin tone doesn't look alarming the way a warped eye does. If the team only ever fixes what gets screenshotted, the biggest real problem stays quietly under-addressed while a smaller, more visually dramatic one eats every sprint.
Attention allocated by who complains loudest is not the same thing as attention allocated by actual harm. The gap between the two is exactly where the sepia problem was hiding.
AAbility to contest. Can the team check this fast, or is trusting the loudest voice the only option.
Before the fix, Brennus could check the eval set himself in twenty minutes, but nothing required anyone to. A less senior PM without his habits, or a team without a damage taxonomy at all, would have had exactly one input to work from: how upset, and how senior, was the person who complained. That's not a real check. That's deferring to rank.
This is GUARD's sharpest question for this exact scenario: not "is the anecdote true," but "does anyone have a fast, real way to test whether it's representative."
Hand sketched flow diagram titled Where a real check should sit, and doesn't. Five connected boxes reading Ticket lands, No check run, Closed ignored, Escalated loud, Roadmap shifts, with the second box emphasized in red.
Four of these five steps happen whether or not anyone built a real check. Only the missing one has to be built on purpose.
RReduce. The actual fix, not a policy memo.
Every ticket with a bad output gets logged against Latent's damage taxonomy, a required field, using the same classifier that already tags eval-set photos. Any reply to the person who raised it has to carry both the specific fix and the real subtype rate, together, every time. The crease-over-eye subtype, specifically, now routes to mandatory human review before delivery, adding roughly six minutes and forty cents per photo, on the 2.5 percent of volume where that's worth it.
The alternative worth naming and rejecting: auto-pause the feature the instant any complaint reaches an executive channel. The team tried an informal version of exactly that last quarter, three times, on three anecdotes that turned out to be known rare cases, costing a colorization launch two weeks and teaching support to escalate everything, which made the real signal harder to find, not easier.
DDetect. How you'd know it's being handled right, not just quiet.
Track the share of incoming complaints that get logged against a known subtype and checked for real frequency, week over week. Before the fix, that share sat around 11 percent, whatever a support rep happened to remember to do. After it shipped, it climbed to 95 percent within three weeks, because logging stopped being optional.
The failure worth naming plainly: if next sprint's priority order can be predicted by asking whose Slack message got the most replies this week, instead of by a stable severity-times-frequency ranking, it's being handled wrong, no matter how calm the quarter looks.
Fail rate by damage subtype, Latent's 2,400-photo eval set
15% 7.5% 0 4% Aggregate, all 2,400 15% Crease over an eye, 60 5.5% Sepia tone shift, 745
Whole eval setCrease over an eyeSepia skin tone shift
Crease-over-eye fails more often per photo. Sepia shift happens on so much more volume that it causes more than four times as many real failures, 41 against 9.
Share of complaints logged against a known subtype, by week
100% 50% 0 fix ships Wk 1 Wk 2 Wk 3 Wk 4, ticket Wk 5 Wk 6 Wk 7 Wk 8
Share of complaints logged against a known subtype, per week
Before the fix, logging happened only when someone remembered, around 11 percent of the time. Once it became a required field, it climbed to 95 percent inside three weeks.
The trade-off, said out loud Routing crease-over-eye photos to mandatory human review adds about six minutes and forty cents per photo, on the 2.5 percent of volume where that subtype shows up. Nitrate Studio took that trade on purpose, for the narrow slice carrying real, elevated risk, instead of either the free-but-risky option of auto-delivering every photo, or the slow-and-expensive option of reviewing all 2,400 monthly restorations by hand, which would have erased the instant turnaround that makes Latent worth using at all.

And if you want to be sure it really works, try it somewhere else

Same five letters, a veterinary diagnostic tool instead of a photo app, and this time the cherry-picked photo is a dog's harmless fatty lump.

Lucent, built by Aldergate Veterinary Group, reads an intake photo of a skin lesion on a pet and suggests a likely category to a vet tech before the vet's own exam. Umbriel Achterberg owns Lucent's diagnostic model the way Brennus owns Latent's restoration model.

Hand sketched icon list diagram titled What Lucent's fix actually does. Four numbered rows: flags the lighting subtype before delivery, logs the ticket against a known subtype, checks the real rate not just the average, keeps the wider rollout live and running.
Same four moves as Latent's fix, aimed at a different subtype entirely: bad exam-room lighting instead of a crease over an eye.

Marlowe Redshaw, who manages one of Aldergate's busiest clinics, posted a screenshot in the regional channel: a dog's harmless fatty lump, flagged by Lucent as "likely malignant, urgent." "Turn this thing off before we scare another client into a biopsy they don't need." Lucent's eval set, 1,800 graded lesion photos, passes 94 percent overall. Inside that number, 150 photos were shot under the clinic's older yellow-tinted exam lights, and on those, the false "urgent" flag rate is 12 percent, four times the 3 percent rate everywhere else.

Same rank, mapped onto Lucent: route photos shot under that lighting subtype to a mandatory second look from the vet before an "urgent" label ever reaches a client, and answer Marlowe with both numbers, the fix for this one dog's photo and the real 12 percent rate, instead of turning the feature off. Leave the other clinics' rollout running; the newer clinics don't have that lighting at all.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: pull the subtype's real rate before reacting to the photo at all. A verdict built on one image is free to say and expensive to be wrong about.
Cost: no budget to build a full damage-taxonomy pipeline this quarter. Start with the one field that matters: require every escalation reply to name a known subtype and its real rate, even if it's still a spreadsheet someone checks by hand.
The model got better, for real: say the crease-over-eye rate drops to 3 percent next quarter after a training fix. That's a reason to widen the claim and pull the mandatory review back, on purpose, with a new eval run backing it up, not a reason the pattern shouldn't have been caught in the first place.

Where people run it wrong.
They treat this as a question of whether to trust the stakeholder, instead of a question of whether the photo is representative.
They fix the one photo and stop there, without ever logging it, so the exact same complaint has to happen again from scratch next month.
They let whoever escalates loudest set the sprint order, instead of a stable rate multiplied by how many people it touches.

How to use it live. Before answering, buy yourself a beat by asking out loud: "is this a known pattern, or the first time we're seeing it?" Naming that split is most of the real answer, and it costs you nothing to say while you're still thinking.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
Which framework fits "a stakeholder cherry-picks one bad output as proof the feature is broken"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real test isn't whether the photo is real, it's whether the team can check if it's representative before reacting.
2 · THE PEOPLE
Who are the people this answer names?
Tap to flip
ANSWER
Brennus Kestner, Latent's product manager. Tabitha Castillon, Nitrate Studio's VP of Customer Success. Fintry Larsholm, the customer, and her grandfather Horace, whose photo started it.
3 · THE OLD HABIT
What did the team default to, out of habit, before there was a real process?
Tap to flip
ANSWER
Close the ticket as a one-off, or escalate it based on how upset or senior the person raising it was, with nothing linking the complaint to the real eval-set frequency.
4 · THE TWO-WAY TRAP
What's the two-way trap this question is actually testing?
Tap to flip
ANSWER
Dismiss the photo and you might miss a real 15 percent failure rate. Overreact to it and you might pause a working feature for everyone, while a bigger, quieter problem, the sepia subtype, stays unfixed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
There was no required link between an individual complaint and Latent's own damage taxonomy, so nothing forced anyone to check the real frequency before reacting.
6 · THE NUMBER
Fill in the blank: the crease-over-eye subtype fails ___ percent of the time on ___ eval photos, while the sepia subtype fails ___ percent of the time but causes ___ of the eval set's 96 total failures.
Tap to flip
ANSWER
15 percent on 60 photos. 5.5 percent, but 41 of 96 total failures, more than four times crease-over-eye's 9.
7 · THE REPLAY
Same Tuesday, new process. What changes?
Tap to flip
ANSWER
Brennus answers the leadership channel by 9:40 with both numbers together. Fintry has a corrected photo by 12:15. The Friday rollout ships on schedule, and the sepia subtype, not the eye subtype, becomes the real next sprint's top item.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which one, and who plays the equivalent roles?
Tap to flip
ANSWER
Lucent, Aldergate Veterinary Group's skin-lesion triage tool. Umbriel Achterberg plays Brennus's role. Marlowe Redshaw plays Tabitha's role, escalating one false "urgent" flag on a benign lump shot under yellow exam lighting.

Check yourself Score: 0 / 0

Fill in the blank
1. Latent's golden eval set has ___ photos overall, with a ___ percent pass rate, the number the whole company points to.
Show hint
Check the "Let's learn" section's opening numbers.
Show answer
2,400 photos, 96 percent. That's the aggregate number. It's real, and it's also the number that hid a 15 percent problem inside one small subtype.
Multiple choice
2. Why doesn't it work to just dismiss Tabitha's photo as "one weird case" and move on?
  • A. Because Tabitha is a VP and her judgment always outranks the eval set.
  • B. Because the crease-over-eye subtype genuinely fails 15 percent of the time, a real, elevated rate the eval set already tracks.
  • C. Because any customer complaint has to trigger a feature pause, no exceptions.
  • D. Because Fintry's reunion deadline makes this legally urgent.
Show hint
Check the Groups step in the GUARD recap.
Show answer
B. The photo isn't just upsetting, it's a fair sample of a real, checkable 15 percent failure rate. Dismissing it outright would miss that.
True or false, with why
3. True or false: the sepia skin-tone-shift subtype is a bigger real problem, by raw failure count, than the crease-over-eye subtype Tabitha escalated.
  • True
  • False
Show hint
Check the Unequal step, and the bar chart.
Show answer
True. Sepia shift causes 41 of the eval set's 96 total failures, against 9 for crease-over-eye, even though its own within-subtype rate, 5.5 percent, is lower. It happens on far more volume.
Short answer, name the rejected alternative
4. Nitrate Studio considered one other fix besides logging tickets against the taxonomy. What was it, and why did they reject it?
Show hint
Check the Reduce step in the GUARD recap.
Show answer
Model answer: Auto-pause the feature the instant any complaint reaches an executive channel. Rejected because the team tried an informal version of this last quarter, three times, on anecdotes that turned out to be known rare cases, costing a colorization launch two weeks and teaching support to escalate everything just in case.
Short answer, apply it yourself
5. Think of a product you use that involves some kind of AI judgment call. Name one bad output someone could hold up as "proof it's broken," and what real number you'd want to check before agreeing or disagreeing.
Show hint
Look for a rate broken out by a specific condition, not just the overall average.
Show answer
Model answer: A spam filter that buries one real email. Before agreeing it's "broken," check the false-positive rate for that specific sender type, a first-time sender with an attachment, say, not the filter's overall accuracy across every email it sees.
Fill in the blank, work the number
6. If the crease-over-eye subtype's real fail rate had turned out to be 3 percent instead of 15 percent, would mandatory human review for that subtype still make sense? Fill in the blank: it would ___ make sense, because ___.
Show hint
Compare a 3 percent rate to the 5.5 percent aggregate-adjacent sepia rate, and the six-minute, forty-cent cost of review.
Show answer
Probably not, at 3 percent. A 3 percent rate is close to ordinary variation across subtypes, not a clear outlier, so the extra six minutes and forty cents per photo would be a cost paid for a pattern that was never actually elevated. The 15 percent rate is what earned the review, not the photo alone.
Before you close the answer
Why this works
Tests whether you can hold two things true at once: a single bad output is real evidence of something, and it is not, by itself, proof of how often that something happens. Most candidates pick a side instead of building a way to check.
Follow-up traps
"What if the eval set itself doesn't cover this exact failure yet?" Response: then the photo just became the eval set's first real example of a pattern, and the honest answer is "we don't know the rate yet," which is a different, and equally truthful, thing to tell Tabitha than either "it's broken" or "it's fine."

"Isn't checking the eval set just a way to stall the stakeholder?" Response: it took Brennus thirty five minutes, start to reply, and Fintry had a fixed photo before lunch. A real check that's this fast isn't stalling, it's the difference between a reply built on a number and one built on a guess.
If pressed
The damage classifier that routes crease-over-eye photos to human review is itself imperfect, it catches an estimated 90 percent of true cases, meaning roughly one in ten still slips through to automatic delivery. That's exactly why the fix pairs the classifier with the logging pipeline: a slip-through still shows up as a customer complaint, gets logged against the same taxonomy, and folds back into the real frequency count instead of vanishing.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more