How do you present eval results to a non-technical decision maker?
A broker owner can't act on a quality score. She can act on five real photos and one plain word: blocked.
- Sort results into buckets by real world consequence, never one blended score.Why: a fine looking average can sit directly on top of the small slice that could get someone a compliance complaint.
- Open the report on real before and after examples from the worst bucket, not a summary slide.Why: a non-technical decision maker can judge a photo of furniture that doesn't fit a room. She can't judge what "96 percent" is actually made of.
- Translate the bucket counts into her own unit before you translate them into yours.Why: "17 of 400 golden set photos" means nothing to her. "25 of your 600 monthly listings" does.
- Check every risky edit against the property's own facts, its floor plan and square footage, before a person ever reviews it.Why: that check is the real guardrail behind the bucket, not the bucket's label.
- State the bar as a probability, not a promise.Why: "blocked" clears a catch rate on a golden set, it isn't a promise of zero mistakes forever, and pretending otherwise turns the next expected miss into a broken promise.
- Recheck the "safe" bucket on a sample every month, don't trust it once and walk away.Why: new kinds of listings can drift into the safe bucket quietly, the same way the real risk hid inside a fine looking average the first time.
How to answer this, stage by stage
Nobody is grading whether you can describe a report layout. They're grading whether you know a technical score and a real risk are two different things, and whether you can show someone the second one without teaching her the first.
Let's learn
Fairlens takes a folder of real estate listing photos and cleans them up. It corrects color and light, swaps a grey sky for a blue one, clears clutter out of a room, and can add virtual staging, real looking furniture rendered into a room that's actually empty, all in under a minute.
Before Fairlens, an agent at a mid-size brokerage either paid a freelance photo editor about thirty dollars a photo and waited a day or two, or shot a listing and posted it as is, flat light and empty rooms and all. Genoveva Truscott, who owns the brokerage, about forty agents across six offices, signed off on Fairlens for the whole company after her vendor sent over one number: a 96 percent quality score.
That 96 percent came from a golden set, four hundred listing photos a design team scored by hand for sharpness, color balance, and how nice a room looked. On that measure, Fairlens really was excellent. The score said nothing about whether an edit had quietly changed something true about the house.
The mistake worth worrying about was never a blurry photo. It was a photo that looked great and wasn't accurate: staged furniture sized for a room bigger than the one it was actually in, or an edit that touched a wall or a window without anyone deciding it should.
Three weeks into the company wide rollout, one of the brokerage's listings staged a nine foot by eight foot flex room with a large sectional sofa and a rug that made it read, in photos, like a spacious home office. A buyer's agent toured it in person, called the photos misleading in a public post, and the county real estate board opened an informal look at whether the marketing had steered a buyer toward space that wasn't actually there. It never became a lawsuit. It cost Genoveva a week of damage control and an uncomfortable call with her insurer.
What I'd leave alone: plain exterior and yard shots. Sky replacement and color correction on a curb photo never touches a wall, a window, or a room's real size, so the blocked bucket check doesn't need to run on them at all. Spend the review time where the risk actually is.
The lesson: a quality score can be completely honest about the thing it measures and still hide the one dimension a decision maker actually needs, because looking good and being accurate are different questions that happen to agree most of the time.
Now here is the same thing as a story
Read the long version below when you want to feel why a fine number and a real risk can point in completely different directions, not just be told that they can.
The report was one slide. One number, ninety six percent, and a green checkmark under it.
Theodoric Nordvik has run quality at Fairlens since before the company had virtual staging at all. He built the very first golden set himself, four hundred photos, scored by hand for sharpness and light and how a room composed, because that was the whole product back then. Better photos. Nothing more.
For the better part of a year, Monday mornings were easy. He'd pull the weekend's score, watch it sit around ninety five, ninety six, and move on to the next thing on his list. Virtual staging shipped quietly into that same golden set, scored the same way it always had been, sharp, well lit, nicely composed. Nobody rewrote the rubric. It had never needed rewriting before.
At first he still opened a sample of staged photos by eye, out of habit, before every release. Then he started trusting the score enough to skim ten instead of fifty. By the third quarter after staging shipped, he was glancing at the number on the dashboard and moving on, the way you nod along in a meeting you stopped following twenty minutes back.
The trigger wasn't a lawsuit. It was Genoveva Truscott's voice on a call, flat and controlled in the way people get right before they stop being polite about it. "I approved a ninety six percent score," she said. "Nobody told me that number had nothing to do with whether a room was the size the photo said it was."
Theodoric pulled the golden set that afternoon and, for the first time, checked every staged photo against the property's own floor plan instead of just looking at it. Seventeen of four hundred had furniture sized for a room bigger than the one it was actually in. Every one of them had scored above ninety on the old rubric. Nobody had ever asked the rubric to check a floor plan, so it never had.
The real cost wasn't the seventeen photos. It was that Genoveva had made a company wide decision, forty agents, six offices, off a number that had never once been asked the question that actually mattered to her.
The decision Theodoric would take back happened over a year earlier, in a short meeting nobody wrote up. The photography team building the first golden set asked what to grade for, and the honest answer at the time was: does it look good. Staging didn't exist yet. Nobody in that room could have known the rubric would still be running, unchanged, the day staging shipped.
Run the same rollout again with the bucket system already in place. The seventeen at-risk photos get caught in the golden set before a single one reaches an agent's listing. Genoveva's report opens on five of those actual photos, side by side with the floor plan they don't match, and one line under them: hold these for a manual check. She approves the rollout in the same twenty minute call. Three weeks later, no compliance inquiry, no call with her insurer, because the room that couldn't hold a sectional sofa never got photographed as if it could.
One design asked a rubric whether a photo looked good. The other asked it whether a photo told the truth about a room. Only one of those questions has anything to do with what got Genoveva's brokerage a compliance inquiry.
What I'd tell myself, back in that first rubric meeting: "does it look good" was never a complete question, it was just the only one anyone had thought to ask yet. The moment a product adds a feature that can change what a photo claims about the world, that feature owes its own question, not a spot on the old one.
SPARK, checked against one floor plan
Not a checklist to recite. Each letter has to survive the same near miss the story just walked through.
Three things worth stating directly, since this is where the real judgment sits. The alternative the team considered and rejected was a single blended "trust score," zero to a hundred, folding cosmetic quality and compliance risk into one number for simplicity. It lost because it would have recreated the exact same problem the 96 percent score already caused, just rebranded, hiding the risky four percent inside a fine looking number again. The AI specific failure worth naming by name is virtual staging quietly rendering furniture at a scale the real room can't support, a kind of hallucination about physical space, not a text one. The guardrail is checking every staged render's furniture footprint against the property's own floor plan and square footage before it's ever allowed into the cosmetic-safe bucket. The bar was never zero misrepresented rooms forever, a model staging thousands of empty rooms a month will occasionally miss a scale; the blocked bucket has to catch at least 92 percent of known floor plan mismatches on the golden set before any new staging model version ships, a probability bar, not a promise. And the trade-off worth naming too: routing that 4.25 percent to a manual check costs about a day of turnaround instead of Fairlens's usual under-a-minute processing, on just those photos, accepted on purpose because slowing down one photo in twenty four costs far less than another compliance inquiry.
And if you want to be sure it really works, try it somewhere else
Same five letters, a factory floor instead of a listing photo, and this time the ticket is a spec sheet, not a floor plan.
Threadcheck reviews photos of finished garments coming off a production line, flags stitching defects, print misalignment, and label errors before a box gets sealed. Quirin Pellinore owns a small apparel brand and has to decide whether to trust Threadcheck's flags instead of the factory's own manual inspection before this season's biggest shipment. Fenella Ingerslev runs quality on Threadcheck.
S, situation: before Threadcheck, every garment got a 45 second hand check by a floor inspector, and Quirin never saw a number at all, just trusted whichever factory she'd contracted that season.
P, payoff: the habit worth building isn't "trust the pass rate." It's looking at the handful of garments Threadcheck would have shipped wrong and deciding, herself, whether that's a risk worth taking this season.
A, anchor: every garment gets bucketed against the brand's own spec sheet, not a general defect score. Ship safe. Needs a glance. Blocked for anything that contradicts the spec sheet itself, a wrong fabric weight tag, or a size label that doesn't match the pattern actually cut.
R, risk: Threadcheck's first golden set only held last year's fabric prints. This season's new print, a fine houndstooth, confused the stitching check and let three mislabeled-size garments auto ship as safe before anyone noticed the print itself was new.
K, keep out: no live defect feed on the factory floor for Quirin to babysit herself, and no automatic per-garment cost estimate. One report, at the end of each production run, with the blocked garments photographed next to their own spec sheet.
Same method, a different weak spot: a floor plan can't move on its own, a room stays the size it is. A fabric print changes every season, so the check has to be refreshed on the same clock as the product line, or the golden set quietly stops representing what's actually shipping.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, bucket by real world consequence and open on real examples from the worst bucket, and give the one number, 17 of 400 scored above ninety and still misrepresented a room.
Cost: there's no budget this quarter to build a fancier live dashboard. Build the one-page monthly report instead, it costs almost nothing and it's the part that actually changes a decision.
The model got better, for real: say Fairlens's staging accuracy improves next quarter, fewer scale mistakes company wide. That's not a reason to go back to one blended number. A rarer mistake still needs a specific, checkable way to catch it when it happens, or it just gets rarer and more surprising when it does.
Where people run it wrong.
They hand a non-technical decision maker a dashboard full of precision, recall, and confidence scores, and call that transparency.
They build one blended "trust score" for simplicity, and quietly recreate the same hidden risk the old percentage had.
They lead with the good news, the 96 percent, and bury the five photos that actually matter at the bottom of the deck, if they include them at all.
How to use it live. Say the real tension out loud before answering: "is this asking me to explain a model, or to help someone make a decision without becoming an engineer first." That buys a beat, and it's almost always the second one.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if she just wants the one number for her board meeting?" Response: give her the one number too, but tie it to her own unit first, "about 25 of your 600 monthly photos would carry real risk," not the vendor's blended percentage, which never told her that number existed.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Eval design for product teams
- #1 What makes an eval product-relevant rather than research-relevant?
- #2 Design an eval for a feature that drafts email replies.
- #3 How do you decide between automated evals and human review?
- #4 Explain the tradeoffs of LLM-as-judge for a product team.
- #5 How do you validate that your judge model agrees with human raters?
- #6 Describe a rubric that a non-technical reviewer could apply consistently.