ConceptAdvancedQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #18

Explain how quality perception differs between novice and expert users.

A shopper grades how sure the sentence sounds. A sommelier grades whether the sentence is true. A flat average score can hold steady for months while those two grades pull in opposite directions underneath it.

The direct answer
Novices judge an AI answer by how confident and polished it sounds, because that is the only check they know how to run. Experts judge the same answer by whether its actual claims are true, because they can check the substance underneath the polish. Those are two different tests on one output, so a blended satisfaction score can sit still for months while one group quietly sours. Split the score by who is rating before you trust it at all.
Do this, in order
  1. Split every satisfaction number by rater expertise before trusting it.Why: a flat blended average can hide two groups moving in opposite directions the whole time.
  2. Rule out the innocent read: check if the expert score was ever higher, not just recently lower.Why: skip this and you chase a "regression" that was baked in since week one.
  3. Separate real substance errors from pure style complaints before fixing anything.Why: fixing them the same way can quietly make the real error worse while it "solves" the fake one.
  4. Ground every sensory claim against a real fact source before it ships.Why: a fluent, wrong claim is the exact failure a novice can't catch and an expert can't unsee.
  5. Hold only the flagged, out-of-range notes for a person to check, not every note.Why: reviewing everything erases the speed the tool exists to give back.
  6. Set the bar as a percentage on a graded sample, not zero mismatches.Why: a model that writes fresh sentences will always miss sometimes, so the team needs a real number to ship against.

How to answer this, stage by stage

Nobody is grading whether you know two kinds of users exist. They are grading whether you'd split a number by who's rating it before you trusted it at all.

1
Scope it to one product and one person
Say it like this
"Let's make this real. Corkline is an app. You type in a bottle's grape, vintage, region and producer, and it writes the tasting note and a pairing line in about two seconds. Ravindra Tolliver runs quality on what it writes."
Why this works
A quality-perception question answered in the abstract turns into a lecture on trust. One product, one person, makes it a decision you can defend.
2
Say the method out loud before using it
Say it like this
"I'd run this as a diagnosis. Lay out the timeline, recut the number by who's rating, rule out the boring explanation, name real candidates for the gap, then test the two that survive."
Why this works
Naming the plan up front stops a rambling answer and tells the interviewer you have a method, not just an opinion.
3
Say what the question is actually testing
Say it like this
"This isn't really asking me to describe two kinds of customers. It's asking whether I know a flat score can hide two groups that were never actually agreeing, and whether I'd go looking for that instead of trusting the dashboard."
Why this works
Naming the real question stops you from giving a generic "different users have different needs" answer, which is what most candidates default to.
4
Give the direct answer, cold, before any story
Say it like this
"A novice grades a tasting note on how sure and polished it sounds. An expert grades it on whether the claim is true for that exact bottle. So I'd split the score by who's rating before I believed the average at all."
Why this works
A reader who stops here already knows exactly what you'd do. Everything after this is proof.
5
Run the recut and the rule-out, with real numbers
Say it like this
"I'd split the thumbs-up rate by a trade flag. Say novices sit at 96 percent and trade sits at 61, including week one, before either group's volume changed much. That's not a regression. That's two groups that were never grading the same thing."
Why this works
This is TRACE's strongest move: turning a vague sense of "some people don't love it" into two numbers you can actually check.
6
Split the real error from the style complaint, then name the fix
Say it like this
"I'd grade a sample of the low trade scores against the wine's real profile. Where the note gets the tannin or the structure wrong, that's the model writing a plausible detail it never checked. So I'd ground every claim against the grape and producer's real profile before it ships, and hold anything outside that range for a person to look at."
Why this works
Naming the mechanism, not just the outcome, is what separates a real diagnosis from a shrug.
7
Close with the trade-off, what you'd leave alone, and one line
Say it like this
"That grounding check adds a few hundred milliseconds and it'll hold back the occasional note that was actually fine. I'd only pay that on the trade surface, not on the quick note someone reads before dinner. And I'm not chasing zero mismatches, I want trade notes clearing the real profile on most of a graded sample, not every single one."
Why this works
Closing on a bar you can measure, and a place you'd deliberately skip the fix, is what makes the answer sound like a decision instead of a promise.
If you remember one thing A novice and an expert can look at the exact same sentence and run two different tests on it. One checks how it sounds. The other checks if it's true. A flat average hides that split instead of showing it.

Let's learn

Corkline is an app. A shop clerk or a home cook types in a bottle's grape, vintage, region and producer, and in about two seconds the app writes a short tasting note and a pairing line, the kind of thing a good clerk might say off the top of their head.

Before Corkline, a busy shop covering 200 bottles a weekend might get real tasting notes written for maybe 20 of them. Writing a good one by hand takes a few minutes and most staff don't have a few minutes times 200. With Corkline, a shop can label the whole case in an afternoon.

Knowledge spark: what's a tannin? A compound in wine that gives a drying, gripping feel in your mouth, mostly from grape skins and seeds. Some grapes make wines with soft, gentle tannin. Others, like a serious Nebbiolo, make wines with firm, structured tannin that needs years in a bottle to soften. Getting this wrong in a tasting note is not a small miss to someone trained to notice it.

Ravindra Tolliver runs quality on what Corkline writes. Every week, a dashboard shows one number: the thumbs-up rate on generated notes, across everyone who rates one. For 22 straight weeks after the trade tier launched, letting professional shop staff use the app too, that number held between 87 and 89 percent. Steady. Reassuring. Exactly what a healthy number looks like.

Corkline blended thumbs-up rate, week 1 to week 22
100% 0% holds 87 to 89% the whole stretch Wk 1 Wk 11 Wk 22
Blended thumbs-up rate, all raters
Nothing on this line looks wrong. It never moves more than two points in either direction across five months.

Then a new hire joined a partner shop, a recently certified sommelier, and pulled up Corkline's note for a house Nebbiolo she knew well: "silky, soft tannins, ready to drink now." She knew that producer's wine needed five more years before it earned the word silky. She left the app a one-star rating and asked Ravindra, in a message, why the tool sounded so sure about a bottle it had never tasted.

The blended number never lied. It just never asked the one question that mattered: sure to whom?

Ravindra pulled the ratings apart by who was rating, a self-reported trade flag plus a behavioral check: anyone rating more than 15 bottles a month counted as trade. Home shoppers had been sitting at 96 percent thumbs-up the entire five months. Trade staff, including certified sommeliers, had been sitting at 61 percent, also the entire five months. Both flat. Both steady. Just steady in opposite places, and canceling out in the average because novices rated about 40,000 notes a month and trade rated about 900.

Same 22 weeks, thumbs-up rate split by who's rating
100% 0% 96% 61% Home shoppers Trade / sommeliers
Novice ratersExpert raters
Novices and trade staff were never rating the same experience. Both lines had been flat since week one, the split was never new.
The decision that mattered Corkline shipped one blended satisfaction score because a single number was simpler to build and simpler to put on a dashboard. That was fine before the trade tier existed. It stopped being fine the moment two very different kinds of raters were folded into the same average, and nobody split it back apart until a person asked why.

What I would leave alone: the plain, hedged, accurate notes on the consumer app that read a little flat. Not every note needs a sommelier's flourish, and chasing that would mean retraining toward flowerier language everywhere, including on the notes that are already true. The problem was never the vocabulary in general. It was specific, checkable claims being wrong on bottles an expert actually knew.

The lesson: a number can hold steady for five months and still be lying to you, not because it moved, but because it was always the sum of two very different numbers that happened to add up the same way every week. A flat line proves nothing about fairness across groups. It only proves the groups didn't change.

Now here is the same thing as a story

Read this when you want to feel why the split mattered, not just know that it existed.

Ravindra Tolliver built Corkline's quality process from nothing two years ago. Before the trade tier, checking quality meant reading a sample of notes by hand every Friday, comparing them against a wine reference book, and writing down what was wrong. It was slow, and it was honest work.

Then the trade tier launched, and the dashboard got a single new number: a thumbs-up rate on every generated note, from home shoppers and shop staff alike. It climbed fast in the first month, settled near 88 percent, and stayed there. The Friday hand-check quietly stopped. Why spend an afternoon reading forty notes when a live number, updating every hour, said the tool was doing fine?

For five months, that was true, as far as anyone could see. The blended score sat between 87 and 89. Nobody on the team had a reason to open it further. The whole point of a good metric is that you get to stop staring at it.

Then a new hire at a partner shop, a woman who had passed her sommelier exam eight months earlier, pulled up the Corkline note for a bottle she'd sold a hundred times. "Silky, soft tannins, ready to drink now," it said. She knew the producer. She knew the wine needed five more years before anyone honest would call it silky. She rated the note one star, and she typed a question into the feedback box that nobody had ever really answered: why does this sound so sure about something it clearly doesn't know?

We never had a quality problem the dashboard could see. We had a quality problem hiding inside an average.

Ravindra's first instinct was to check whether the model had gotten worse recently, maybe a bad update, maybe a data issue. The dashboard said no. The blended number hadn't moved in five months. So the next move was the one that actually mattered: pull the ratings apart by who was doing the rating, instead of trusting the blend.

Home shoppers: 96 percent thumbs-up, flat, the whole five months. Trade staff, sommeliers included: 61 percent thumbs-up, flat, the whole five months, going back to week one, before the shop's volume had grown at all. The split was never a new problem. It had been sitting under the average since the day the trade tier opened, quietly outvoted by home shoppers, who rated Corkline notes about 44 times more often than trade staff did.

The decision that let this hide went back to the very first sprint on the ratings system. Someone asked, early on, whether the app should ask raters if they worked in the trade, so scores could be split later. The answer at the time was no, it would slow down onboarding by one extra screen, and the team wanted one clean number for the launch deck anyway. Nobody revisited it once trade partners actually signed up.

Ravindra pulled a sample of the low trade ratings and graded each one against a real wine-profile database: grape, region, producer, expected tannin and body. About two out of every three low trade ratings traced to a genuine factual miss, a tannin or structure claim that didn't match the actual bottle. The other third traced to notes that were factually fine, just plain, and trade reviewers had marked them down anyway for reading more like a spec sheet than a sommelier's pour.

Run the same week again with one change: before a trade-facing note ships, the model's tannin, acid, and body claims get checked against that producer and region's real profile. Inside the expected range, it ships as written. Outside it, the note gets held and a person looks at it before it reaches a shelf. Same Nebbiolo, same producer, same request. The generated note now says "firm, structured tannin, needs several more years," which is true, and it publishes clean. Six weeks after the fix, trade thumbs-up climbs from 61 percent to 84 percent. The plain-but-accurate complaints stay roughly where they were, because that was never actually a quality bug.

One design trusted a single blended score and never asked who was behind each rating. The other checks the claim itself against a real fact source before anyone gets to rate it at all. Those aren't the same design with a new coat of paint. One has a hole an expert falls straight through. The other doesn't.

What I'd tell myself, back in that first sprint: any time a tool starts serving two audiences who know different amounts about the subject, ask whether one clean average is quietly averaging away the group that knows the most. Nobody asked. That's on the room, not on the sommelier who finally noticed.

TRACE, run on a quality score that never moved

Not a diagnosis of a crash. A diagnosis of a number that looked calm while it was hiding two opposite readings the whole time.

TTimeline. When the split actually started, versus when anyone noticed.
The blended score held 87 to 89 percent for 22 weeks. The new hire's question landed in week 23. But the trade-only score, once pulled apart, had been sitting near 61 percent since week one, not since the incident. The timeline the team noticed doesn't match the timeline the data actually shows.
In a perception question, the timeline isn't for finding when quality changed. It's for proving whether a gap is new or was there from the start.
RRecut. Split the score by who's rating, not by when they rated.
Home shoppers: 96 percent thumbs-up. Trade and sommeliers: 61 percent thumbs-up. Same 22 weeks, same product, same notes, two completely different verdicts, because novice volume (about 40,000 ratings a month) buried expert volume (about 900) inside the blend.
This is the whole argument for why a satisfaction average can't stand in for quality by itself: the average only tells you what the loudest group thinks.
AAssume nothing. Rule out the innocent read before diagnosing anything.
The innocent read: a stable average means quality is stable for everyone. Checked against the trade-only score from weeks one through four, before the shop network had grown much: already 58 to 64 percent. The gap was never a regression. It was there from the first bottle the trade tier ever rated.
Skip this and you'll write a "what changed recently" incident report about a gap that was never new.
CCandidates. Two named biases, pulling in opposite directions, plus one rejected fix.
Named: a confident, fluent tone passes a novice's only real check, since they have no way to test the actual claim. And, the opposite bias, an expert sometimes marks down a note that is factually correct but plainly worded, for lacking the flourish they expect, even when nothing in it is wrong. Rejected: retraining the whole model on more sommelier-sounding copy, the instinct the team reached for first, because it would polish the register without fixing the actual tannin errors, and might make wrong claims sound even more convincing.
Naming both biases, and which fix got turned down, is what turns this into a decision instead of a guess dressed up as an explanation.
EEvidence test. The one check that tells the two candidates apart.
Grade a stratified sample of low trade ratings against a real varietal and producer profile database, not against how flowery the note sounds. Result: about two-thirds of the low trade scores traced to a genuine factual miss. The other third traced to accurate notes marked down purely for tone. Different problems, different fixes.
This is the strongest move in the whole method. It turns "trade doesn't love it" into a number you can act on differently depending on which slice it came from.
Hand sketched numbered list titled why the trade score stayed low, three items: a confident tone passes a novice's only real check, an expert checks the real tannin profile and it's wrong, an expert docks a plain but accurate note for no flourish.
All three candidates are real. The evidence test is what tells you how much weight each one actually carries.

Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was retraining the model broadly toward more expert-sounding language, the fix the team reached for first after the sommelier's message. It lost because the evidence test showed two-thirds of the trade complaints were real factual errors, not style complaints, and a model trained to sound more confident and polished everywhere would make a wrong tannin claim sound even more convincing, not less. The AI-specific failure worth naming by name is hallucinated sensory detail: the model generating a plausible, genre-correct descriptor, silky tannins, that isn't actually grounded in that producer's real profile, which is confident wrongness wearing a sommelier's vocabulary. The guardrail is concrete: check every tannin, acid, and body claim against a structured varietal-and-producer profile before the note ships to the trade surface, and hold anything outside the expected range for a person to clear rather than publishing it straight through. That guardrail isn't free. It adds a real lookup, a few hundred milliseconds, to every trade-facing note, and it will occasionally hold back a note that was actually fine, a small speed cost accepted only where a wrong claim reaches a professional buyer's shelf, not rolled out to the casual consumer surface where a soft adjective costs almost nothing. And the bar isn't zero mismatches. A model writing a fresh sentence about a new bottle every time can't promise that. It's a target: trade-facing notes clear the real profile on at least 92 percent of a monthly graded sample, checked by hand against the same reference a sommelier would use, tight enough that the review queue is usually the one catching the rare miss, not a shop's newest hire.

And if you want to be sure it really works, try it somewhere else

Same five letters, a home-inspection report tool instead of a wine app, nothing about tasting or drinking anywhere in sight.

Fieldmark writes plain-language summaries of home inspection reports for buyers, turning a 40-page PDF into a short readout of what actually needs attention. Andres Fenner runs quality for that tool.

T, timeline. The blended "was this summary helpful" score held near 91 percent for six months after Fieldmark opened its report tool to licensed inspectors and appraisers, not just first-time buyers. Nothing moved on the surface.
R, recut. First-time buyers rated summaries helpful 97 percent of the time. Licensed inspectors and appraisers, checking the same summaries against the full report, rated them helpful only 55 percent of the time, the entire six months.
A, assume nothing. Checked the appraiser-only score from the first month the tool launched to professionals: already 51 to 58 percent. Not a recent slide. Baked in from day one.
C, candidates. A buyer's only real check is whether the summary reads clearly and sounds decisive. An appraiser checks whether the summary correctly ranks severity, and one case stood out: a summary called a hairline foundation crack "minor and cosmetic" when the full report flagged it as a structural concern needing an engineer, a genuine miss a buyer had no way to catch.
E, evidence test. Same move as Corkline: grade a sample of low appraiser scores against the source report's own severity flags, not against how the summary reads. Most of the gap traced to real severity misranking, not tone.

Hand sketched two panel comparison titled same gap, two different products. Left panel Corkline: a silky tannins claim passed the shopper, failed the sommelier. Right panel Fieldmark: a confident write up passed the buyer, missed the crack call.
Two different products, two different kinds of expertise, and the same underlying shape: a fluent summary that satisfied the person who couldn't check it, and failed the person who could.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the direct answer and the recut, don't describe the two user types in the abstract before you've split a real number.
Cost: there's no budget yet for a structured fact database to ground claims against. Don't skip the check, hand-grade a small stratified sample against a trusted source until one exists.
The model got better, for real: say the blended score climbed to 93 percent after a general upgrade. That's not proof the expert gap closed. A tool can improve on average while a specific, expert-visible error keeps recurring the whole time.

Where people run it wrong.
They read one flat average and conclude quality is fine for everyone, and skip splitting the number at all.
They see the expert complaint and immediately promise a full retrain, before checking whether the gap is style or substance.
They calm down the loudest complaint with an apology and move on, without ever verifying whether it was one rare miss or a pattern the average had been hiding for months.

How to use it live. Say the real question out loud before answering it: "are novice and expert grading the same thing, or running two different tests on the same output." That line buys you a beat to think instead of guessing in front of the interviewer.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a question about why novices and experts rate the same output so differently?
Tap to flip
ANSWER
TRACE: lay out the timeline, recut the score by who's rating, rule out the innocent read that a flat average means stable quality, name real candidates for the split, then run one evidence test to tell them apart.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ravindra Tolliver, who runs quality for Corkline, an app that writes a tasting note and pairing line from a bottle's basic facts.
3 · THE HABIT
What habit had quietly formed before the new hire's question?
Tap to flip
ANSWER
Trusting the single blended dashboard score and letting the old weekly hand-check lapse, since a live number that looked healthy made the manual check feel unnecessary.
4 · TWO EXPLANATIONS
What are the two competing explanations for the low trade score, and which one held up more?
Tap to flip
ANSWER
Either trade raters were marking down real factual errors the model made, or they were marking down accurate notes just for sounding plain. The evidence test showed real errors explained about two-thirds of the gap.
5 · THE OLD DECISION
What old decision does this answer take back?
Tap to flip
ANSWER
Shipping one blended satisfaction score with no field for rater expertise, decided in the first sprint to keep onboarding short and the launch deck simple, never revisited once trade partners actually joined.
6 · THE NUMBER
Fill in the blank: home shoppers rated Corkline notes helpful ___ percent of the time, while trade and sommelier raters sat at ___ percent, both flat since week one.
Tap to flip
ANSWER
96 percent for home shoppers, 61 percent for trade. Both numbers had been steady the entire five months, which is what proves the gap was never a recent regression.
7 · THE REPLAY
Same Nebbiolo, new design, what changes?
Tap to flip
ANSWER
The tannin and structure claims get checked against the producer's real profile before the note ships. The note now correctly says "firm, structured tannin, needs more years." Trade thumbs-up climbs from 61 to 84 percent within six weeks.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's different about the split?
Tap to flip
ANSWER
Fieldmark, a home-inspection summary tool run by Andres Fenner. Same TRACE steps, but the expert gap traced mainly to severity misranking, a foundation issue called "cosmetic," rather than a tone complaint.

Check yourself Score: 0 / 0

Fill in the blank
1. Home shoppers rated Corkline notes helpful ___ percent of the time, while trade raters sat at ___ percent, the whole five months.
Show hint
Look at the second chart in "Let's learn," the one split by who's rating.
Show answer
96 percent and 61 percent. Both were flat the entire stretch, which is what rules out a recent regression and points at a split that was there from the start.
True or false
2. True or false: because the blended score held steady at 87 to 89 percent, Ravindra could safely conclude the tool's quality was stable for everyone using it.
  • True
  • False
Show hint
A flat average can be the sum of two flat numbers moving in opposite directions.
Show answer
False. The blended score stayed flat because novice satisfaction and trade satisfaction were both flat, just at very different levels, and novice volume drowned out trade volume in the average.
Multiple choice
3. Why does a confident, fluent tasting note pass a novice rater's check, even when a specific claim in it is wrong?
  • A. Novices never read the tasting notes closely enough to rate them at all.
  • B. Tone and readability are the only real quality check a novice knows how to run, so a confident-sounding claim passes even if it's false.
  • C. Novices always give five stars regardless of what the note says.
  • D. The app hides ratings from home shoppers, so their scores don't count.
Show hint
Think about what a novice can actually verify about a bottle they've never tasted.
Show answer
B. A novice has no independent way to check whether "silky tannins" is true for a specific bottle, so a fluent, confident sentence is the only signal they have to grade against.
Short answer, name the rejected alternative
4. What fix did this answer reject, and why did it lose?
Show hint
Look at what the team reached for first, right after the sommelier's message, in the framework recap section.
Show answer
Model answer: Retraining the model broadly toward more sommelier-sounding language. It lost because the evidence test showed most of the trade complaints were real factual errors, not tone complaints, so a more polished-sounding model would have made wrong claims even more convincing instead of fixing them.
Short answer, apply it yourself
5. Pick an AI tool you use yourself. Name one output where a total beginner and a real expert in that subject would likely rate the exact same answer very differently, and say why.
Show hint
Ask what the expert can check that the beginner can't, and what the beginner is grading instead.
Show answer
Model answer: An AI coding assistant explaining a bug fix. A beginner rates it highly because the explanation reads clearly and sounds sure. A senior engineer might rate the same answer low because the fix technically works but papers over the actual root cause, something only someone who understands the system would catch.
Multiple choice
6. Why couldn't Ravindra just retrain the model to sound less confident across the board and call the problem solved?
  • A. Corkline doesn't have enough training data to retrain the model at all.
  • B. Home shoppers would immediately stop using the app if the tone changed even slightly.
  • C. About two-thirds of the trade complaints were real factual errors, so a tone change alone would leave the actual wrong claims in place.
  • D. Trade raters don't actually read the tasting notes, so nothing about the model matters to them.
Show hint
Compare the two-thirds/one-third split from the evidence test in the framework recap.
Show answer
C. The evidence test showed most of the gap was substance, not style. A tone-only fix would leave the real tannin and structure errors in place while making the note sound even more convincing.
Before you close the answer
Why this works
Tests whether you'll treat a flat satisfaction score as proof of stable quality, or go looking for the split underneath it. Most candidates describe novice and expert users in the abstract without ever proposing a way to check the actual gap.
Follow-up traps
"Couldn't the trade score just reflect pickier people, not a real quality gap?" Response: that's exactly why the evidence test graded the low ratings against a real varietal-and-producer profile, not against how picky the raters seemed. Two-thirds of the low scores traced to actual factual misses, not pickiness.

"Isn't showing trade users a different, more cautious note just hiding the problem from novices?" Response: no, the fix grounds every claim against the real profile before it ships to anyone, novice or trade. Trade users see the flag more often only because they rate enough volume to run into the held cases; the underlying check is the same for everyone.
If pressed
The actual bar used at Corkline: trade-facing notes need to clear the real varietal-and-producer profile on at least 92 percent of a monthly graded sample, checked by hand against the same reference a sommelier would use, versus no formal bar at all on the casual consumer surface, because a wrong adjective there costs almost nothing.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more