Explain how quality perception differs between novice and expert users.
A shopper grades how sure the sentence sounds. A sommelier grades whether the sentence is true. A flat average score can hold steady for months while those two grades pull in opposite directions underneath it.
- Split every satisfaction number by rater expertise before trusting it.Why: a flat blended average can hide two groups moving in opposite directions the whole time.
- Rule out the innocent read: check if the expert score was ever higher, not just recently lower.Why: skip this and you chase a "regression" that was baked in since week one.
- Separate real substance errors from pure style complaints before fixing anything.Why: fixing them the same way can quietly make the real error worse while it "solves" the fake one.
- Ground every sensory claim against a real fact source before it ships.Why: a fluent, wrong claim is the exact failure a novice can't catch and an expert can't unsee.
- Hold only the flagged, out-of-range notes for a person to check, not every note.Why: reviewing everything erases the speed the tool exists to give back.
- Set the bar as a percentage on a graded sample, not zero mismatches.Why: a model that writes fresh sentences will always miss sometimes, so the team needs a real number to ship against.
How to answer this, stage by stage
Nobody is grading whether you know two kinds of users exist. They are grading whether you'd split a number by who's rating it before you trusted it at all.
Let's learn
Corkline is an app. A shop clerk or a home cook types in a bottle's grape, vintage, region and producer, and in about two seconds the app writes a short tasting note and a pairing line, the kind of thing a good clerk might say off the top of their head.
Before Corkline, a busy shop covering 200 bottles a weekend might get real tasting notes written for maybe 20 of them. Writing a good one by hand takes a few minutes and most staff don't have a few minutes times 200. With Corkline, a shop can label the whole case in an afternoon.
Ravindra Tolliver runs quality on what Corkline writes. Every week, a dashboard shows one number: the thumbs-up rate on generated notes, across everyone who rates one. For 22 straight weeks after the trade tier launched, letting professional shop staff use the app too, that number held between 87 and 89 percent. Steady. Reassuring. Exactly what a healthy number looks like.
Then a new hire joined a partner shop, a recently certified sommelier, and pulled up Corkline's note for a house Nebbiolo she knew well: "silky, soft tannins, ready to drink now." She knew that producer's wine needed five more years before it earned the word silky. She left the app a one-star rating and asked Ravindra, in a message, why the tool sounded so sure about a bottle it had never tasted.
Ravindra pulled the ratings apart by who was rating, a self-reported trade flag plus a behavioral check: anyone rating more than 15 bottles a month counted as trade. Home shoppers had been sitting at 96 percent thumbs-up the entire five months. Trade staff, including certified sommeliers, had been sitting at 61 percent, also the entire five months. Both flat. Both steady. Just steady in opposite places, and canceling out in the average because novices rated about 40,000 notes a month and trade rated about 900.
What I would leave alone: the plain, hedged, accurate notes on the consumer app that read a little flat. Not every note needs a sommelier's flourish, and chasing that would mean retraining toward flowerier language everywhere, including on the notes that are already true. The problem was never the vocabulary in general. It was specific, checkable claims being wrong on bottles an expert actually knew.
The lesson: a number can hold steady for five months and still be lying to you, not because it moved, but because it was always the sum of two very different numbers that happened to add up the same way every week. A flat line proves nothing about fairness across groups. It only proves the groups didn't change.
Now here is the same thing as a story
Read this when you want to feel why the split mattered, not just know that it existed.
Ravindra Tolliver built Corkline's quality process from nothing two years ago. Before the trade tier, checking quality meant reading a sample of notes by hand every Friday, comparing them against a wine reference book, and writing down what was wrong. It was slow, and it was honest work.
Then the trade tier launched, and the dashboard got a single new number: a thumbs-up rate on every generated note, from home shoppers and shop staff alike. It climbed fast in the first month, settled near 88 percent, and stayed there. The Friday hand-check quietly stopped. Why spend an afternoon reading forty notes when a live number, updating every hour, said the tool was doing fine?
For five months, that was true, as far as anyone could see. The blended score sat between 87 and 89. Nobody on the team had a reason to open it further. The whole point of a good metric is that you get to stop staring at it.
Then a new hire at a partner shop, a woman who had passed her sommelier exam eight months earlier, pulled up the Corkline note for a bottle she'd sold a hundred times. "Silky, soft tannins, ready to drink now," it said. She knew the producer. She knew the wine needed five more years before anyone honest would call it silky. She rated the note one star, and she typed a question into the feedback box that nobody had ever really answered: why does this sound so sure about something it clearly doesn't know?
Ravindra's first instinct was to check whether the model had gotten worse recently, maybe a bad update, maybe a data issue. The dashboard said no. The blended number hadn't moved in five months. So the next move was the one that actually mattered: pull the ratings apart by who was doing the rating, instead of trusting the blend.
Home shoppers: 96 percent thumbs-up, flat, the whole five months. Trade staff, sommeliers included: 61 percent thumbs-up, flat, the whole five months, going back to week one, before the shop's volume had grown at all. The split was never a new problem. It had been sitting under the average since the day the trade tier opened, quietly outvoted by home shoppers, who rated Corkline notes about 44 times more often than trade staff did.
The decision that let this hide went back to the very first sprint on the ratings system. Someone asked, early on, whether the app should ask raters if they worked in the trade, so scores could be split later. The answer at the time was no, it would slow down onboarding by one extra screen, and the team wanted one clean number for the launch deck anyway. Nobody revisited it once trade partners actually signed up.
Ravindra pulled a sample of the low trade ratings and graded each one against a real wine-profile database: grape, region, producer, expected tannin and body. About two out of every three low trade ratings traced to a genuine factual miss, a tannin or structure claim that didn't match the actual bottle. The other third traced to notes that were factually fine, just plain, and trade reviewers had marked them down anyway for reading more like a spec sheet than a sommelier's pour.
Run the same week again with one change: before a trade-facing note ships, the model's tannin, acid, and body claims get checked against that producer and region's real profile. Inside the expected range, it ships as written. Outside it, the note gets held and a person looks at it before it reaches a shelf. Same Nebbiolo, same producer, same request. The generated note now says "firm, structured tannin, needs several more years," which is true, and it publishes clean. Six weeks after the fix, trade thumbs-up climbs from 61 percent to 84 percent. The plain-but-accurate complaints stay roughly where they were, because that was never actually a quality bug.
One design trusted a single blended score and never asked who was behind each rating. The other checks the claim itself against a real fact source before anyone gets to rate it at all. Those aren't the same design with a new coat of paint. One has a hole an expert falls straight through. The other doesn't.
What I'd tell myself, back in that first sprint: any time a tool starts serving two audiences who know different amounts about the subject, ask whether one clean average is quietly averaging away the group that knows the most. Nobody asked. That's on the room, not on the sommelier who finally noticed.
TRACE, run on a quality score that never moved
Not a diagnosis of a crash. A diagnosis of a number that looked calm while it was hiding two opposite readings the whole time.
Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was retraining the model broadly toward more expert-sounding language, the fix the team reached for first after the sommelier's message. It lost because the evidence test showed two-thirds of the trade complaints were real factual errors, not style complaints, and a model trained to sound more confident and polished everywhere would make a wrong tannin claim sound even more convincing, not less. The AI-specific failure worth naming by name is hallucinated sensory detail: the model generating a plausible, genre-correct descriptor, silky tannins, that isn't actually grounded in that producer's real profile, which is confident wrongness wearing a sommelier's vocabulary. The guardrail is concrete: check every tannin, acid, and body claim against a structured varietal-and-producer profile before the note ships to the trade surface, and hold anything outside the expected range for a person to clear rather than publishing it straight through. That guardrail isn't free. It adds a real lookup, a few hundred milliseconds, to every trade-facing note, and it will occasionally hold back a note that was actually fine, a small speed cost accepted only where a wrong claim reaches a professional buyer's shelf, not rolled out to the casual consumer surface where a soft adjective costs almost nothing. And the bar isn't zero mismatches. A model writing a fresh sentence about a new bottle every time can't promise that. It's a target: trade-facing notes clear the real profile on at least 92 percent of a monthly graded sample, checked by hand against the same reference a sommelier would use, tight enough that the review queue is usually the one catching the rare miss, not a shop's newest hire.
And if you want to be sure it really works, try it somewhere else
Same five letters, a home-inspection report tool instead of a wine app, nothing about tasting or drinking anywhere in sight.
Fieldmark writes plain-language summaries of home inspection reports for buyers, turning a 40-page PDF into a short readout of what actually needs attention. Andres Fenner runs quality for that tool.
T, timeline. The blended "was this summary helpful" score held near 91 percent for six months after Fieldmark opened its report tool to licensed inspectors and appraisers, not just first-time buyers. Nothing moved on the surface.
R, recut. First-time buyers rated summaries helpful 97 percent of the time. Licensed inspectors and appraisers, checking the same summaries against the full report, rated them helpful only 55 percent of the time, the entire six months.
A, assume nothing. Checked the appraiser-only score from the first month the tool launched to professionals: already 51 to 58 percent. Not a recent slide. Baked in from day one.
C, candidates. A buyer's only real check is whether the summary reads clearly and sounds decisive. An appraiser checks whether the summary correctly ranks severity, and one case stood out: a summary called a hairline foundation crack "minor and cosmetic" when the full report flagged it as a structural concern needing an engineer, a genuine miss a buyer had no way to catch.
E, evidence test. Same move as Corkline: grade a sample of low appraiser scores against the source report's own severity flags, not against how the summary reads. Most of the gap traced to real severity misranking, not tone.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the direct answer and the recut, don't describe the two user types in the abstract before you've split a real number.
Cost: there's no budget yet for a structured fact database to ground claims against. Don't skip the check, hand-grade a small stratified sample against a trusted source until one exists.
The model got better, for real: say the blended score climbed to 93 percent after a general upgrade. That's not proof the expert gap closed. A tool can improve on average while a specific, expert-visible error keeps recurring the whole time.
Where people run it wrong.
They read one flat average and conclude quality is fine for everyone, and skip splitting the number at all.
They see the expert complaint and immediately promise a full retrain, before checking whether the gap is style or substance.
They calm down the loudest complaint with an apology and move on, without ever verifying whether it was one rare miss or a pattern the average had been hiding for months.
How to use it live. Say the real question out loud before answering it: "are novice and expert grading the same thing, or running two different tests on the same output." That line buys you a beat to think instead of guessing in front of the interviewer.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't showing trade users a different, more cautious note just hiding the problem from novices?" Response: no, the fix grounds every claim against the real profile before it ships to anyone, novice or trade. Trade users see the flag more often only because they rate enough volume to run into the held cases; the underlying check is the same for everyone.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?