CaseIntermediateQuality, Cost & Token Economics / Eval design for product teams / #20

How do you present eval results to a non-technical decision maker?

A broker owner can't act on a quality score. She can act on five real photos and one plain word: blocked.

The direct answer
Never lead a non-technical decision maker with one blended score. Sort eval results into buckets by what actually happens if a bad one ships, cosmetic and safe, needs a human glance, or blocked for real risk, then open the report with real before and after examples pulled from the worst bucket. On one photo tool's first golden set, a single 96 percent quality score sat on top of 17 photos in 400, about one in 24, where the room in the photo looked bigger than the room really was. Show the decision maker those 17 photos and the plain math behind them, never the average that swallowed them.
Do this, in order
  1. Sort results into buckets by real world consequence, never one blended score.Why: a fine looking average can sit directly on top of the small slice that could get someone a compliance complaint.
  2. Open the report on real before and after examples from the worst bucket, not a summary slide.Why: a non-technical decision maker can judge a photo of furniture that doesn't fit a room. She can't judge what "96 percent" is actually made of.
  3. Translate the bucket counts into her own unit before you translate them into yours.Why: "17 of 400 golden set photos" means nothing to her. "25 of your 600 monthly listings" does.
  4. Check every risky edit against the property's own facts, its floor plan and square footage, before a person ever reviews it.Why: that check is the real guardrail behind the bucket, not the bucket's label.
  5. State the bar as a probability, not a promise.Why: "blocked" clears a catch rate on a golden set, it isn't a promise of zero mistakes forever, and pretending otherwise turns the next expected miss into a broken promise.
  6. Recheck the "safe" bucket on a sample every month, don't trust it once and walk away.Why: new kinds of listings can drift into the safe bucket quietly, the same way the real risk hid inside a fine looking average the first time.

How to answer this, stage by stage

Nobody is grading whether you can describe a report layout. They're grading whether you know a technical score and a real risk are two different things, and whether you can show someone the second one without teaching her the first.

1
Scope it to one real product and one real decision maker
Say it like this
"Let's ground this in Fairlens. It auto enhances real estate listing photos and can stage furniture into an empty room. Genoveva Truscott owns a regional brokerage, about forty agents, and she has to decide whether Fairlens stays on company wide. She isn't technical. She reads listings, not confusion matrices."
Why this works
An abstract "how do you present eval results" answer turns into a slide-design lecture fast. One real decision maker makes it a real design problem instead.
2
Say your structure out loud before diving in
Say it like this
"I'm going to split the results into buckets by what actually happens if a photo goes out wrong, not by how good the pixels look, and I'll open the report with real examples from the worst bucket before she ever sees a number."
Why this works
Tells the interviewer you have a plan and a payoff, not just a dashboard idea.
3
Reframe the question before answering it
Say it like this
"This isn't really asking me how to build a slide deck. It's asking whether I know a technical quality score and a real business risk are two different axes, and whether I can show a non-technical person the second one without making her learn the first."
Why this works
Stops you from describing a dashboard when the real question is about judgment.
4
Give the anchor, the actual design decision
Say it like this
"Bucket every photo by consequence. Cosmetic and safe auto publishes. Anything that could plausibly read as a missing feature gets a human glance. Anything that touches a wall, a window, a room count, or stages furniture at a scale the property's own floor plan can't support gets hard blocked until an agent checks it. Then the report opens on five real blocked photos, not the average."
Why this works
This is the actual answer to the question. Everything else defends it.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without this. Fairlens sent Genoveva a 96 percent quality score before rollout and she approved it. Three weeks in, an agent published a listing where the staged furniture made a nine by eight flex room read like a real home office, a buyer's agent called the photos misleading in public, and the county real estate board opened an informal look. Nothing in that 96 percent number ever told her this photo existed."
Why this works
Shows the real cost of the old artifact, not just the mechanism behind it.
6
Close on the decision, not the story
Say it like this
"So: three buckets, never one score, and the report opens on the photos that would have gone out the door, not the average that hid them."
Why this works
Ending on the rule, not the anecdote, is what makes this sound like a method you'd reuse, not a story you told once.

Let's learn

Fairlens takes a folder of real estate listing photos and cleans them up. It corrects color and light, swaps a grey sky for a blue one, clears clutter out of a room, and can add virtual staging, real looking furniture rendered into a room that's actually empty, all in under a minute.

Before Fairlens, an agent at a mid-size brokerage either paid a freelance photo editor about thirty dollars a photo and waited a day or two, or shot a listing and posted it as is, flat light and empty rooms and all. Genoveva Truscott, who owns the brokerage, about forty agents across six offices, signed off on Fairlens for the whole company after her vendor sent over one number: a 96 percent quality score.

Knowledge spark: what's a golden set? A batch of outputs someone graded by hand, on purpose, and kept aside just for testing. It can hold known problems on purpose, so you can measure whether a check actually catches what it's supposed to catch, instead of guessing.

That 96 percent came from a golden set, four hundred listing photos a design team scored by hand for sharpness, color balance, and how nice a room looked. On that measure, Fairlens really was excellent. The score said nothing about whether an edit had quietly changed something true about the house.

Looking good and being accurate are not the same thing. They usually move together, and pull apart exactly when it matters.

The mistake worth worrying about was never a blurry photo. It was a photo that looked great and wasn't accurate: staged furniture sized for a room bigger than the one it was actually in, or an edit that touched a wall or a window without anyone deciding it should.

The same 400-photo golden set, one blended score vs. split by real world consequence
100% 50% 0% 96% Old blended score 85.5% 10.25% 4.25% Split: safe / glance / blocked
Old score, blendedCosmetic safe, 342 of 400Needs a glance, 41 of 400Blocked, 17 of 400
The 96 percent never disappears, it's still true. It just sits directly on top of 17 photos, about one in 24, that scored fine and misrepresented a room.

Three weeks into the company wide rollout, one of the brokerage's listings staged a nine foot by eight foot flex room with a large sectional sofa and a rug that made it read, in photos, like a spacious home office. A buyer's agent toured it in person, called the photos misleading in a public post, and the county real estate board opened an informal look at whether the marketing had steered a buyer toward space that wasn't actually there. It never became a lawsuit. It cost Genoveva a week of damage control and an uncomfortable call with her insurer.

The choice that mattered The first golden set only asked graders to rate how a photo looked, sharp, well lit, nicely composed, because that's what the photography team cared about when they built it, months before virtual staging existed. Nobody rebuilt the grading criteria the day staging shipped, the one feature that could actually put furniture where a room couldn't hold it.

What I'd leave alone: plain exterior and yard shots. Sky replacement and color correction on a curb photo never touches a wall, a window, or a room's real size, so the blocked bucket check doesn't need to run on them at all. Spend the review time where the risk actually is.

The lesson: a quality score can be completely honest about the thing it measures and still hide the one dimension a decision maker actually needs, because looking good and being accurate are different questions that happen to agree most of the time.

Now here is the same thing as a story

Read the long version below when you want to feel why a fine number and a real risk can point in completely different directions, not just be told that they can.

The report was one slide. One number, ninety six percent, and a green checkmark under it.

Theodoric Nordvik has run quality at Fairlens since before the company had virtual staging at all. He built the very first golden set himself, four hundred photos, scored by hand for sharpness and light and how a room composed, because that was the whole product back then. Better photos. Nothing more.

For the better part of a year, Monday mornings were easy. He'd pull the weekend's score, watch it sit around ninety five, ninety six, and move on to the next thing on his list. Virtual staging shipped quietly into that same golden set, scored the same way it always had been, sharp, well lit, nicely composed. Nobody rewrote the rubric. It had never needed rewriting before.

At first he still opened a sample of staged photos by eye, out of habit, before every release. Then he started trusting the score enough to skim ten instead of fifty. By the third quarter after staging shipped, he was glancing at the number on the dashboard and moving on, the way you nod along in a meeting you stopped following twenty minutes back.

The trigger wasn't a lawsuit. It was Genoveva Truscott's voice on a call, flat and controlled in the way people get right before they stop being polite about it. "I approved a ninety six percent score," she said. "Nobody told me that number had nothing to do with whether a room was the size the photo said it was."

Theodoric pulled the golden set that afternoon and, for the first time, checked every staged photo against the property's own floor plan instead of just looking at it. Seventeen of four hundred had furniture sized for a room bigger than the one it was actually in. Every one of them had scored above ninety on the old rubric. Nobody had ever asked the rubric to check a floor plan, so it never had.

We never lost seventeen photos to a bad model. We lost the one thing the score was supposed to protect: whether a listing told the truth.

The real cost wasn't the seventeen photos. It was that Genoveva had made a company wide decision, forty agents, six offices, off a number that had never once been asked the question that actually mattered to her.

The decision Theodoric would take back happened over a year earlier, in a short meeting nobody wrote up. The photography team building the first golden set asked what to grade for, and the honest answer at the time was: does it look good. Staging didn't exist yet. Nobody in that room could have known the rubric would still be running, unchanged, the day staging shipped.

Run the same rollout again with the bucket system already in place. The seventeen at-risk photos get caught in the golden set before a single one reaches an agent's listing. Genoveva's report opens on five of those actual photos, side by side with the floor plan they don't match, and one line under them: hold these for a manual check. She approves the rollout in the same twenty minute call. Three weeks later, no compliance inquiry, no call with her insurer, because the room that couldn't hold a sectional sofa never got photographed as if it could.

One design asked a rubric whether a photo looked good. The other asked it whether a photo told the truth about a room. Only one of those questions has anything to do with what got Genoveva's brokerage a compliance inquiry.

What I'd tell myself, back in that first rubric meeting: "does it look good" was never a complete question, it was just the only one anyone had thought to ask yet. The moment a product adds a feature that can change what a photo claims about the world, that feature owes its own question, not a spot on the old one.

SPARK, checked against one floor plan

Not a checklist to recite. Each letter has to survive the same near miss the story just walked through.

SSituation. Who is this person, and how does the job get done today, without you?
Genoveva Truscott, forty agents, six offices, deciding whether Fairlens stays on company wide. Before any real eval report existed, the only artifact she got from the vendor was one quality percentage and a green checkmark.
Name the real decision this report has to survive, or the design floats free of the actual job.
Hand sketched labeled parts diagram titled the whole report Genoveva got before rollout. Center icon a document labeled vendor rollout report. Four callouts around it: one quality number, one green checkmark, no real photos, no risk detail.
The whole artifact she had to decide on, before this redesign.
PPayoff. What habit do you want this to build?
Not "trust a friendly looking number." Specifically: look at five real photos, decide for herself whether any of them would embarrass the brokerage, and let that judgment, not a percentage, carry the decision.
A named habit produces a named report layout. A vague goal like "communicate results well" produces a slide nobody can act on.
AAnchor. The one design decision everything else hangs on.
Every eval report buckets each photo by what happens if it ships wrong. Cosmetic and safe auto publishes. Anything that could read as a missing feature gets a human glance. Anything touching a wall, a window, a room count, or a floor plan mismatch gets hard blocked. The report opens on five real blocked photos before a single number appears.
This is the actual design decision. If it doesn't visibly survive the next letter, it's a slogan, not an anchor.
Hand sketched icon list diagram titled how the new report buckets every photo. Row one a gauge icon, cosmetic safe auto publish. Row two a question mark card icon, needs a glance, quick human check. Row three a scale icon, blocked, floor plan mismatch, hard stop.
Three buckets, sorted by what happens next, not by how the pixels score.
RRisk. What breaks the first time you're wrong?
A new kind of listing, a tiny, oddly shaped flex room nothing in the golden set looks like, slips through the cosmetic-safe bucket because nobody built a rule for it yet. The design has to survive that, not pretend it can't happen.
A bucket that only works when every future listing resembles the golden set isn't a design. It's a hope with a label on it.
Hand sketched flow diagram titled the day the safe bucket is wrong. Four boxes in sequence: new tiny flex room, auto published as safe, caught in monthly recheck, emphasized, added to golden set.
The safe bucket is allowed to be wrong once, as long as the monthly recheck catches it before it repeats.
KKeep out. What do you deliberately not build?
No automatic dollar figure for legal exposure per photo, day one, that would look more precise than it actually is. No live self-serve dashboard of raw model confidence scores for Genoveva to check herself. One report, once a month, with real photos in it, not a feed she has to babysit.
A report that hands a non-technical owner a fake-precise number or a raw feed doesn't add clarity. It just moves the confusion somewhere she can't push back on it.

Three things worth stating directly, since this is where the real judgment sits. The alternative the team considered and rejected was a single blended "trust score," zero to a hundred, folding cosmetic quality and compliance risk into one number for simplicity. It lost because it would have recreated the exact same problem the 96 percent score already caused, just rebranded, hiding the risky four percent inside a fine looking number again. The AI specific failure worth naming by name is virtual staging quietly rendering furniture at a scale the real room can't support, a kind of hallucination about physical space, not a text one. The guardrail is checking every staged render's furniture footprint against the property's own floor plan and square footage before it's ever allowed into the cosmetic-safe bucket. The bar was never zero misrepresented rooms forever, a model staging thousands of empty rooms a month will occasionally miss a scale; the blocked bucket has to catch at least 92 percent of known floor plan mismatches on the golden set before any new staging model version ships, a probability bar, not a promise. And the trade-off worth naming too: routing that 4.25 percent to a manual check costs about a day of turnaround instead of Fairlens's usual under-a-minute processing, on just those photos, accepted on purpose because slowing down one photo in twenty four costs far less than another compliance inquiry.

And if you want to be sure it really works, try it somewhere else

Same five letters, a factory floor instead of a listing photo, and this time the ticket is a spec sheet, not a floor plan.

Threadcheck reviews photos of finished garments coming off a production line, flags stitching defects, print misalignment, and label errors before a box gets sealed. Quirin Pellinore owns a small apparel brand and has to decide whether to trust Threadcheck's flags instead of the factory's own manual inspection before this season's biggest shipment. Fenella Ingerslev runs quality on Threadcheck.

S, situation: before Threadcheck, every garment got a 45 second hand check by a floor inspector, and Quirin never saw a number at all, just trusted whichever factory she'd contracted that season.

P, payoff: the habit worth building isn't "trust the pass rate." It's looking at the handful of garments Threadcheck would have shipped wrong and deciding, herself, whether that's a risk worth taking this season.

A, anchor: every garment gets bucketed against the brand's own spec sheet, not a general defect score. Ship safe. Needs a glance. Blocked for anything that contradicts the spec sheet itself, a wrong fabric weight tag, or a size label that doesn't match the pattern actually cut.

R, risk: Threadcheck's first golden set only held last year's fabric prints. This season's new print, a fine houndstooth, confused the stitching check and let three mislabeled-size garments auto ship as safe before anyone noticed the print itself was new.

K, keep out: no live defect feed on the factory floor for Quirin to babysit herself, and no automatic per-garment cost estimate. One report, at the end of each production run, with the blocked garments photographed next to their own spec sheet.

The decision Fenella would take back The golden set never got refreshed with each season's new fabric prints before that season launched. It was built once, early on, and treated as finished, so a genuinely new print had no test coverage the day it hit the line.
Hand sketched decision tree diagram titled bucketing a garment against the spec sheet, not a feeling. Root question does this garment match its own spec sheet. Three branches: matches, cosmetic only leads to ship safe. Small print or stitch doubt leads to needs a glance. Wrong fabric weight or size tag leads to blocked.
Same anchor, a different ticket. The spec sheet plays the part the floor plan played at Fairlens.

Same method, a different weak spot: a floor plan can't move on its own, a room stays the size it is. A fabric print changes every season, so the check has to be refreshed on the same clock as the product line, or the golden set quietly stops representing what's actually shipping.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, bucket by real world consequence and open on real examples from the worst bucket, and give the one number, 17 of 400 scored above ninety and still misrepresented a room.
Cost: there's no budget this quarter to build a fancier live dashboard. Build the one-page monthly report instead, it costs almost nothing and it's the part that actually changes a decision.
The model got better, for real: say Fairlens's staging accuracy improves next quarter, fewer scale mistakes company wide. That's not a reason to go back to one blended number. A rarer mistake still needs a specific, checkable way to catch it when it happens, or it just gets rarer and more surprising when it does.

Where people run it wrong.
They hand a non-technical decision maker a dashboard full of precision, recall, and confidence scores, and call that transparency.
They build one blended "trust score" for simplicity, and quietly recreate the same hidden risk the old percentage had.
They lead with the good news, the 96 percent, and bury the five photos that actually matter at the bottom of the deck, if they include them at all.

How to use it live. Say the real tension out loud before answering: "is this asking me to explain a model, or to help someone make a decision without becoming an engineer first." That buys a beat, and it's almost always the second one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
SPARK: design against the failure before you build. Built for design questions, including designing how you communicate something, not just what you build.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Theodoric Nordvik, who runs quality at Fairlens, an AI tool that enhances and virtually stages real estate listing photos. He built the product's very first golden set.
3 · THE HABIT
What did Theodoric stop doing because the score worked, most of the time?
Tap to flip
ANSWER
He stopped opening a sample of staged photos by eye before every release, and started trusting the dashboard's blended quality score to speak for the whole product.
4 · THE ANCHOR
What's the one design decision the whole eval report hangs on?
Tap to flip
ANSWER
Bucket every photo by real world consequence, cosmetic safe, needs a glance, or blocked, and open the report on real photos from the blocked bucket before any number appears.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building the first golden set to grade only cosmetic quality, sharpness, light, composition, and never rewriting that rubric the day virtual staging shipped and could actually misrepresent a room.
6 · THE NUMBER
Fill in the blank: of the 400 golden set photos, ___ were cosmetic safe, ___ needed a glance, and ___ were blocked for misrepresentation risk.
Tap to flip
ANSWER
342 cosmetic safe, 41 needed a glance, 17 blocked. All 17 blocked photos had scored above ninety on the old, single quality score.
7 · THE REPLAY
Same near miss, new report design, what changes?
Tap to flip
ANSWER
The 17 at-risk photos get caught in the golden set before rollout. Genoveva approves in the same twenty minute call, and three weeks later there's no compliance inquiry and no call with her insurer.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the shared anchor?
Tap to flip
ANSWER
Threadcheck, a garment defect inspection tool for an apparel brand. Same anchor: bucket by consequence against the brand's own spec sheet, not a general defect score.

Check yourself Score: 0 / 0

Multiple choice
1. Which artifact should lead an eval report built for a non-technical decision maker?
  • A. A dashboard of raw model confidence scores for every photo.
  • B. Five real photos from the highest-risk bucket, shown before any number appears.
  • C. A single blended quality percentage, since it's the easiest number to remember.
  • D. An automated monthly email listing every flagged photo's filename, with no images attached.
Show hint
Think about what Genoveva could actually judge for herself, without learning any new vocabulary first.
Show answer
B. A non-technical decision maker can look at a real photo and a floor plan and judge a mismatch herself. She can't judge what a confidence score or a blended percentage is actually made of.
Short answer, name the rejected alternative
2. What alternative did the Fairlens team consider for presenting eval results, and why did it lose?
Show hint
Look at the paragraph right after the K step in the framework recap, where the three closing points are stated directly.
Show answer
Model answer: A single blended "trust score," zero to a hundred, folding cosmetic quality and compliance risk into one number for simplicity. It lost because it would have recreated the exact same problem the 96 percent score already caused, just rebranded, hiding the risky four percent inside a fine looking number again.
True or false
3. True or false: because Fairlens's golden-set quality score was a real, honestly measured 96 percent, the tool was safe to roll out company wide with no further changes.
  • True
  • False
Show hint
Check what the 96 percent rubric was actually built to grade, versus what it never checked at all.
Show answer
False. The rubric only graded cosmetic quality. It never checked a staged photo against the property's own floor plan, so 17 photos in 400 scored above ninety and still misrepresented a room's real size.
Fill in the blank
4. On the 400-photo golden set, ___ photos were cosmetic safe, ___ needed a human glance, and ___ were blocked for misrepresentation risk.
Show hint
Look at the bar chart in Section 1, right after the golden set numbers are introduced.
Show answer
342, 41, and 17. 342 plus 41 plus 17 equals the full 400-photo set, split by what happens if that photo ships wrong, not by how it scored on looks alone.
Short answer, apply it yourself
5. Pick an AI product you use yourself that produces something a non-expert has to approve or trust. Name one place its summary number might be hiding a real risk, and how you'd show that risk instead.
Show hint
Think of an app that shows you one overall rating or score after an AI step, a match percentage, a health score, a safety rating.
Show answer
Model answer: A budgeting app shows one "spending health score" after it auto-categorizes transactions. That score could look fine while it's quietly mis-categorizing rent as discretionary spending for a small slice of users. Instead of the score, I'd show the actual transactions the categorizer was least confident about, sorted by dollar size, so a user can catch the one miscategorized bill that would have thrown off next month's budget.
Multiple choice
6. Threadcheck's cosmetic-safe bucket let three mislabeled-size garments auto ship during a new print's first run. What does that reveal?
  • A. The spec sheet itself was wrong and needs to be rewritten from scratch.
  • B. Bucketing by consequence doesn't work for physical products, only for photos.
  • C. The golden set needs refreshing on the same clock as the product line, since a genuinely new fabric print had no test coverage the day it shipped.
  • D. Quirin should go back to trusting the factory's manual inspection instead.
Show hint
Read the key point block right under Threadcheck's R and K steps, "the decision Fenella would take back."
Show answer
C. The golden set was built once and treated as finished. A fabric print that changes every season needs the same refresh discipline as the product line, or the check quietly stops representing what's actually shipping.
Before you close the answer
Why this works
Tests whether you can separate "explaining a model" from "helping someone make a real decision." Most candidates describe a dashboard with more numbers on it. Few say what makes a single number dangerous to hand a person who can't check what's underneath it.
Follow-up traps
"Isn't showing five photos just cherry picking the scary ones?" Response: the five come straight from a defined bucket, blocked for a stated, checkable reason, floor plan mismatch, not hand-picked for drama. Every photo in that bucket could be swapped in and the report would still make the same point.

"What if she just wants the one number for her board meeting?" Response: give her the one number too, but tie it to her own unit first, "about 25 of your 600 monthly photos would carry real risk," not the vendor's blended percentage, which never told her that number existed.
If pressed
The floor plan check itself runs by comparing the staged furniture's rendered bounding box against the room's square footage from the property's own listing data, not by asking a second model to grade the first one's work, since a grading model can inherit the same blind spot the first model has.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more