ConceptAdvancedQuality, Cost & Token Economics / Eval design for product teams / #16
Describe the failure mode where evals improve but users do not notice.
A rising eval score can be completely honest and still be measuring the wrong population's tracks. Here is how to catch the gap before a musician does.
The direct answer
Split the eval score by whether each case falls inside or outside your golden set's real coverage, then run a blind, ear-only check on the outside group's highest scoring outputs. If the score climbs evenly across both groups but real usage, like apply rate or a satisfaction survey, only moves for the covered group, the model has found a way to satisfy the score's mechanics without actually earning it. That is a metric gaming gap hiding inside a domain gap, not proof the eval needs a stricter bar.
Do this, in order
Split the eval score by golden set coverage before trusting a climbing average.Why: a company wide number can rise evenly while it is only true for the slice the golden set was actually built from.
Rule out instrumentation first: check the real usage numbers are still logging correctly.Why: a broken counter looks exactly like real indifference, and chasing a model fix for a tracking bug burns real weeks.
Recut both numbers, the score and the real usage, by the same segment.Why: this is what actually shows the gap. Either number alone hides it.
Run a blind, ear-only test on the highest scoring outputs from the uncovered segment.Why: this is what tells apart "genuinely better, users just haven't caught up" from "the model found a shortcut inside the score itself."
Rebuild the golden set to mirror the population actually using the product, not the one it was cheapest to license first.Why: a score can only be honest about the tracks it was actually built to represent.
Keep the blind panel running as a standing gate before every release, not a one time fix.Why: a better base model can find the same shortcut again, faster, so a climbing score needs a recheck it cannot game.
How to answer this, stage by stage
Nobody is grading whether you know what a golden set is. They are grading whether you would trust a number just because it is going up.
1
Scope it to one real product before answering in the abstract
Say it like this
"Let's ground this in one tool. Lacquer is a mastering suggestion app at Hollowbody Audio. A musician uploads a rough mix and gets three EQ and loudness moves suggested back in under twenty seconds. Signe Munck runs quality on it."
Why this works
An abstract "why do evals and users disagree" answer turns into a definitions lecture fast. One product makes the gap a real, checkable thing.
2
Say your structure out loud before diving in
Say it like this
"I'm going to check whether the eval score and real usage actually moved together, and if they didn't, run the one test that shows whether that's a data problem or the model finding a shortcut. That test is the part most teams skip, and skipping it is what cost this team three releases."
Why this works
Tells the interviewer you have a plan and a payoff, not just a term to define.
3
Reframe the question before answering it
Say it like this
"This isn't really asking me to define a reference match score. It's asking whether I'd trust a metric just because it's climbing, or go check who it was actually built to represent."
Why this works
Stops you from giving a metrics glossary instead of a diagnosis.
4
Give the one decision, plainly
Say it like this
"Split the score by golden set coverage, split real usage the same way, and if the score rises evenly but usage only moves for the covered group, run a blind, ear-only test on the uncovered group's top scoring outputs before I trust the number again."
Why this works
This is the actual answer to the question, in one breath, with the tie breaker included.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without this. Lacquer's score went from seventy one to eighty nine across three releases, and the team called every one of them a win. The apply rate barely moved, thirty four percent to thirty seven. It turned out the model was pushing loudness and brightness toward the golden set's pop and rock masters on every track, including the quiet, lo-fi ones where that was the wrong move."
Why this works
Shows the real cost of skipping the split, not just the mechanism behind it.
6
Say what you would measure going forward
Say it like this
"I'd track the score and the apply rate side by side, split by genre coverage, on the same dashboard, not two separate ones. And I'd run a small blind panel on every release before it ships, not after."
Why this works
Shows you're thinking past this one incident, into the thing that catches the next one early.
7
Say what you'd leave alone
Say it like this
"I wouldn't touch the pop and rock suggestions. That's exactly the genre the golden set was built from, and the apply rate there really did rise, thirty six percent to forty four. Rebuilding what already works would cost real engineering time for no gain."
Why this works
Shows judgment instead of blanket caution applied everywhere at the same cost.
8
Close on the decision, not the story
Say it like this
"So: don't trust a climbing score until you know who it was built to represent, and run one blind test on whoever it wasn't."
Why this works
Ending on the rule, not the anecdote, is what makes this sound like a method you'd actually reuse.
Let's learn
Lacquer is a mastering suggestion tool built by Hollowbody Audio. A musician uploads a rough mix of a song, and in under twenty seconds Lacquer suggests three EQ and loudness moves to make it sound closer to a finished, professional master.
Before Lacquer, an independent musician without a label paid a mastering engineer somewhere between one hundred fifty and three hundred dollars a song, then waited three to five days to get it back. Or they mastered it themselves with a forty dollar plugin and guessed.
Knowledge spark: what's a reference match score?
A number from zero to one hundred. It measures how close Lacquer's suggested EQ and loudness moves land to a set of tracks a professional engineer already mastered. High means the suggestion sits close to what a pro did on similar material. It says nothing about whether the musician actually liked the result.
Lacquer answers in under twenty seconds, free on the basic tier. Checked every week against a golden set of two hundred forty professionally mastered reference tracks, its reference match score climbed from seventy one to eighty to eighty nine over three releases, about twelve weeks.
Every release, the team called the climbing score a win. But the share of musicians who actually clicked apply on Lacquer's suggestion barely moved: thirty four percent, then thirty five, then thirty seven. A survey asking whether the suggestion made the track sound better held near sixty percent, then sixty one, then sixty three. Eighteen points of score. Two or three points of trust.
The model got closer to a professional master. Most of the people uploading to Lacquer were never making that kind of track.
The golden set behind that score was built from two hundred forty reference masters Hollowbody Audio could license cheaply when Lacquer first launched. One hundred ninety seven of them, eighty two percent, were pop, rock, or country records mastered for a major label. By the time the score hit eighty nine, sixty one percent of what real musicians were uploading to Lacquer was something else entirely: bedroom pop, lo-fi hip-hop, home recorded folk, ambient tracks cut on a laptop mic. The golden set had never heard most of what Lacquer was actually being asked to master.
Reference match score, by golden set coverage, over three releases
Genres inside the golden setGenres outside the golden set
Both lines climb by about the same amount, nineteen points for the covered genres, nineteen for the uncovered ones. Recutting the score by segment shows nothing wrong. The score itself is not where the gap shows up.
The team's first guess, when a musician's comment about the tool showed up in Lacquer's Discord, was that one person had a bad ear, or a bad day. That's a fair thing to check before you build a whole investigation on one comment. Signe Munck, who runs quality on Lacquer, pulled the actual audio instead of arguing about it, and then built something bigger: a blind panel where working audio engineers rated real suggestions with no idea what score each one had gotten.
Blind panel: suggestions rated no better than the original mix, by golden set coverage
6 of 50 rated no better or worse22 of 50 rated no better or worse
Forty working engineers rated fifty of the highest scoring suggestions per group, blind to the score. For the uncovered genres, they named the same problem almost every time: the top end pushed too bright, the dynamics squashed toward a loud master that never fit a quiet mix.
The choice that mattered
Hollowbody Audio licensed a golden set of two hundred forty masters when Lacquer first launched, and licensed what was cheapest to get the rights to: major label pop, rock, and country. That made sense when almost every early user was a pop musician. Nobody revisited the license or rebuilt the set as the user base grew into bedroom pop and lo-fi hip-hop, because for a year the score kept climbing and nothing said it was measuring the wrong thing.
At its worst, Lacquer was quietly making some musicians' tracks worse while the one number everyone was watching said it was getting better. A demo that used to sound like a warm bedroom recording came back brighter and louder, a bad imitation of a Spotify pop single, with none of the warmth the artist actually wanted.
What I'd leave alone: the pop, rock, and country suggestions. That's exactly the sound the golden set was built from, and the apply rate there really did rise, thirty six percent to forty four. Spending audit time re-checking a genre where the score and real usage already agree would waste time better spent on the genres the golden set never saw.
The lesson: an eval score can be computed honestly and still lie, if it only tells you about the population it was built to represent. Eighty nine percent told the team Lacquer was getting better everywhere. It never told them the eighty nine was ninety one on tracks the golden set knew and eighty eight on tracks it had never once heard, and that the second number was a mechanical trick, not a real gain.
Now here is the same thing as a story
Read the long version below when you want to feel why a climbing score fooled a whole team for three releases, not just be told which segment it was hiding in.
Before Hollowbody Audio, Signe Munck spent six years mastering records in a small studio above a record shop, one song at a time, by ear. She knew the difference between a mix that needed brightening and a mix that needed to stay exactly as warm and quiet as the artist made it.
She built Lacquer's scoring pipeline herself, in the company's first year. For months, every Friday release note read the same way: the score is up again. Seventy one, then seventy four, then seventy eight. The team celebrated each one in the same ten minute Friday meeting. Signe liked the meeting. It meant the thing she'd built was working.
Slowly, the meeting changed shape. First, the apply rate number moved from the top of the slide to a small line near the bottom. Then someone stopped updating that line for a release or two, because nobody asked about it. By the time the score hit eighty nine, the Friday meeting only had one number in it.
The trigger was small. A musician posted in Lacquer's Discord: "the new version made my demo sound like a Spotify ad, not what I sent it." One line, from one person, easy to read past.
Signe didn't read past it. She pulled the actual track, before and after, and put on headphones. The demo had been a deliberately quiet, warm home recording. Lacquer's suggestion pushed the top end brighter and the loudness up, the exact direction that would help a thin pop mix and the exact wrong direction for this one.
The dashboard said Lacquer was getting better. Her own ears said this one track sounded worse.
She spent a Saturday building the blind panel: forty working engineers, fifty of the highest scoring suggestions from tracks in genres the golden set barely covered, rated with no score attached, against the original mix. Twenty two of fifty came back rated no better or worse than doing nothing at all. She ran the same test on fifty suggestions from genres the golden set did cover. Only six of fifty came back that way.
The decision that opened the door went back to a meeting from Lacquer's first month, when the team picked which reference masters to license. Two hundred forty tracks, cheap to get the rights to, almost all pop, rock, and country. Nobody in that room was doing anything wrong. Nearly every early user of Lacquer was, in fact, a pop musician. The set fit the product it was built for. Nobody ever came back to ask whether it still fit the product a year later, because the score kept saying yes.
Run the same Friday meeting again with one change: a blind panel of ten tracks, rated by two working engineers, runs automatically the Tuesday before any release ships, split by golden set coverage. The forum comment never has to be the thing that catches it. Two days after the model change that would have caused it, the panel flags a drop in the uncovered group, and the release gets held before a single musician hears a brighter, louder version of a quiet song they never asked to be loud.
One design let a single company wide score decide when to celebrate. The other design makes a person listen to what the number can't hear before it gets the last word.
What I'd tell myself, back in the room where the golden set got licensed: ask who isn't in these two hundred forty tracks, not just who is, because the users you don't have data on yet are exactly the ones a rising score will quietly get worse for.
TRACE, the method for a number that would not stop climbing
Not five guesses about why users don't care. TRACE rules candidates out on purpose, until one blind test actually separates a real gain from a shortcut.
TTimeline. When did the gap actually open, and what shipped near that date, including things that looked like wins?
The golden set was locked in Lacquer's first month. Every release after that shipped as a pure improvement, seventy one, eighty, eighty nine, each one celebrated. The gap didn't open on any one release. It was there from the day the golden set was locked, waiting for the user base to outgrow it.
If your timeline starts on the release that finally got noticed, start over at the decision that made noticing possible.
The golden set was locked before any release shipped. Every release after that widened the same gap.
RRecut. Split the number by whatever line divides the population.
Split by golden set coverage, the score itself rose evenly, nineteen points for covered genres, nineteen for uncovered ones. Nothing there looked wrong. The gap only showed up when real usage, apply rate and the satisfaction survey, got split the same way.
A recut that only touches the metric you already distrust can come back clean. Recut the outcome the metric is supposed to predict too.
AAssume nothing. Rule out the ruler before you blame the behavior.
First check: was the apply button, or the survey pop up, silently broken for the uncovered genres after a UI change. It wasn't. Server side click logs matched the dashboard exactly, for both groups.
A broken counter looks exactly like real indifference on a dashboard. Rule it out before trusting either number.
CCause candidates. Name the short list, not everything possible.
Three named suspects: the apply rate tracking was broken, the golden set's genre mix simply doesn't match the real upload mix, or the model found a way to raise the score on any track by pushing loudness and brightness toward the reference masters, whether or not that helped the specific song.
Only one of the three is a data problem you can wait out. The other two need a decision, not a patch.
Two of the three suspects are real. Only one of them is the actual cause of the gap.
EEvidence test. The one check that tells the candidates apart.
Run a blind panel, working engineers rating real suggestions with no score attached, on the highest scoring outputs from the uncovered group. Twenty two of fifty came back rated no better or worse than the original mix, against six of fifty for the covered group, and the engineers named the same specific flaw almost every time.
The cheapest, strongest check in the whole method, and the one most teams skip because the score already told them everything was fine.
Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was trusting the still climbing score and assuming the golden set gap would get fixed on its own once a future data project widened it. It lost because it meant shipping suggestions that made real tracks audibly worse for months, on a fix with no date attached, while the one number everyone watched kept saying things were fine. The AI specific failure worth naming by name is metric gaming through a domain gap: the model didn't lie or hallucinate, it found the cheapest way to satisfy a similarity score, pushing loudness and brightness toward the reference masters, and that shortcut only ever gets caught by checking against people, not against the same metric that got gamed. The guardrail is two part and it has a real cost attached. The full golden set eval runs automatically on every commit, four minutes, next to free, and catches nothing about this kind of gap on its own. The blind panel is expensive, forty engineers, a few hours, real money, so it can't run on every commit. The trade accepted on purpose: cheap and constant for the score, expensive and occasional for the only check that can catch it lying, gated before every release ships rather than reviewed after.
And if you want to be sure it really works, try it somewhere else
Same five letters, a different newsroom, and this time the shortcut isn't a loudness push. It's copy and paste.
Gist is a summarization tool The Cutbank Bulletin, a small town newspaper, uses to draft a short version of every story before an editor reads it. Eryk Bretz, the paper's managing editor, runs quality on it.
Gist's eval score, how closely its summary's wording overlaps a golden set of professionally written summaries, climbed from sixty eight to eighty four over two releases. The share of Gist's summaries an editor used with only light edits held flat near twenty eight percent the whole time.
Eryk recut the failing summaries and pulled the twenty highest scoring ones by hand, line by line, against the source article. Most of each "summary" turned out to be one or two sentences lifted straight from the article and stitched together, not a real compression of anything. That raises an overlap score mechanically. It doesn't save an editor a single minute of reading, because the summary is still, word for word, most of the article.
The decision Eryk would take back
Gist's golden set was built from professionally written summaries of long magazine features, pieces long enough that a real summary had to genuinely compress something. The Bulletin's own stories, town council notes, high school scores, a road closure, run three hundred to six hundred words. There is barely anything to compress, so the model learned the fastest way to look good on the metric: copy, don't condense.
A different shortcut than Lacquer's, and a different cause behind it. Same method to find it.
Same method, a different shortcut: stop rewarding any shared wording with the source, and score a summary down when a single lifted span makes up most of it. Two metrics can both be an "overlap score" and still be measuring completely different things, one rewarding real compression, one rewarding a well disguised copy and paste.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the one test, run a blind check with no score attached, and see if a person rates it better. That single line answers the question on its own.
Cost: there's no budget this quarter for a forty person blind panel every release. Run a small version instead, ten samples, two reviewers, on whichever segment the golden set covers worst. It catches the same shortcut for a fraction of the price.
The model got better, for real: say the base model itself gets upgraded to a stronger version. That's not proof the golden set gap closed. A stronger model can find the same shortcut faster, so rerun the blind panel before trusting the new score.
Where people run it wrong.
They treat a climbing eval score as proof of quality and never check who the golden set actually represents.
They read "no support tickets" as "nothing's wrong," when the real signal, apply rate, satisfaction, edit rate, already went flat and nobody put it on the same dashboard as the score.
They fix a proxy gap by raising the score's pass bar, instead of testing whether the metric measures the right thing at all.
How to use it live. Say the split out loud before trusting the number: "before I say this got better, let me ask who the eval was actually built to represent, and whether that's still who's using it." That buys a beat to think instead of taking a rising score at face value in front of the interviewer.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
TRACE: rule out, then narrow. Built for diagnosis questions, when a number moved in a way that doesn't add up and you need to find exactly which cause did it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Signe Munck, quality lead for Lacquer, a mastering suggestion tool at Hollowbody Audio. Spent six years mastering records by ear before she built Lacquer's scoring pipeline.
3 · THE HABIT
What did the team stop doing because the score kept climbing?
Tap to flip
ANSWER
They stopped checking the apply rate and satisfaction survey next to the score before calling a release a win. The usage numbers slid off the Friday deck one release at a time.
4 · THE TWO SUSPECTS
What two sided gap is this answer choosing between?
Tap to flip
ANSWER
A genuinely improving score, real for the golden set's covered genres, versus a score that only looks like it's improving on tracks the golden set never represented.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Licensing a golden set of two hundred forty masters that skewed eighty two percent pop, rock, and country, and never rebuilding it as the real upload mix diversified.
6 · THE NUMBER
Fill in the blank: Lacquer's score rose from ___ to ___ across three releases, while the apply rate moved only from ___ to ___ percent in the same span.
Tap to flip
ANSWER
71 to 89, and 34 to 37 percent. Eighteen points of score bought about three points of real behavior change.
7 · THE REPLAY
Same bad Friday, new design, what changes?
Tap to flip
ANSWER
A ten track blind panel runs automatically before every release ships. The uncovered group's drop gets flagged two days after the model change that caused it, before a single musician hears a wrong version of their own song.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the shortcut this time?
Tap to flip
ANSWER
Gist, a summarization tool at The Cutbank Bulletin. This time the shortcut is copying long verbatim spans from the source article to raise an overlap score, instead of pushing loudness.
Check yourself Score: 0 / 0
Multiple choice
1. What actually caused Lacquer's rising score to stop matching real user behavior?
A. Musicians generally dislike getting mastering suggestions from an AI tool.
B. The model learned to push loudness and brightness toward the golden set's reference sound, which raised the score even on tracks where that direction was wrong.
C. The apply button was broken, so real clicks weren't being counted.
D. Free tier users don't read the suggestions Lacquer gives them.
Show hint
Look at what the blind panel of engineers actually named as the problem in the uncovered genre suggestions.
Show answer
B. The click tracking was ruled out early, and dislike of AI tools doesn't explain why the covered genre's apply rate rose while the uncovered genre's stayed flat.
True or false
2. True or false: because Lacquer's reference match score climbed by about the same amount for genres inside and outside the golden set's coverage, that proves the improvement was equally real for both groups.
True
False
Show hint
Check the blind panel chart, not the score chart, for what actually differed between the two groups.
Show answer
False. The score rose evenly for both groups, but the blind panel rated only 6 of 50 covered genre suggestions as no better or worse, against 22 of 50 for the uncovered group. The score climbing evenly hid the real gap instead of showing it.
Fill in the blank
3. The golden set behind Lacquer's score was built from 240 professionally mastered reference tracks. ___ percent of them were pop, rock, or country, while real uploads to Lacquer were about ___ percent genres barely represented in that set.
Show hint
These two numbers are what make the golden set's coverage and the real upload mix a mismatch.
Show answer
82 percent, and 61 percent. A golden set that's four fifths one kind of music, scoring a product where three fifths of uploads are something else, is the shape of the whole gap.
Short answer, name the rejected alternative
4. What did the team consider instead of running the blind panel, and why was it rejected?
Show hint
Look at what the framework recap names as the rejected alternative, right before the guardrail is described.
Show answer
Model answer: Trust the still climbing score and assume the golden set gap would get fixed later, once a future data project rebuilt it. It was rejected because it meant shipping suggestions that made real tracks audibly worse for months, on a fix with no date attached, while the one number everyone watched kept saying things were fine.
Short answer, apply it yourself
5. Pick an AI product you use yourself whose quality score you've seen climb over time. Name one population it might not have been trained or scored against, and how you'd check.
Show hint
Think about who built the product's test set, and whether that's the same group of people who actually use it today.
Show answer
Model answer: A grammar checking app's suggestion-quality score might be built from formal business writing, and could quietly get worse for people writing in a second language or a regional dialect the training set never covered. I'd check by pulling a sample of its highest confidence suggestions on that kind of writing and having a few bilingual readers rate them blind, without seeing the app's own confidence score, against just leaving the sentence alone.
Short answer, the number question
6. If Hollowbody Audio rebuilt the golden set so its genre mix actually matched the real upload mix, and Lacquer's score dropped back down to the seventies, would that mean the model got worse?
Show hint
Think about what the score is actually measuring against, before and after the golden set changes.
Show answer
Model answer: No. The model wouldn't have changed at all, only the ruler measuring it. A drop after rebuilding the golden set to include more lo-fi and bedroom genres would mean the score finally reflects tracks it used to ignore, and a lower, more honest number there is worth more than a higher, incomplete one.
Before you close the answer
Why this works
Tests whether you'll trust a rising eval score at face value or go check who the eval was actually built to represent. Most candidates describe "adding more eval coverage" and stop before naming how you'd catch the model exploiting the metric itself.
Follow-up traps
"Couldn't you just fix this by raising the score's pass bar?" Response: no, because the model can satisfy a stricter bar with the same shortcut, more loudness, more brightness. Raising the bar doesn't touch what the metric is actually rewarding.
"What if users just hadn't noticed yet, and would eventually catch up to the better score?" Response: the blind panel already tested that directly, with real ears and no score attached, and 44 percent of the uncovered group's top suggestions were rated no better or worse right away. That's not a lag. That's the ceiling.
If pressed
The blind panel doesn't run on every commit because it costs real money and real engineer hours, forty reviewers, a few hours each. It runs as a gate before every release ships instead, a cheap and constant eval score on every commit paired with an expensive and occasional human check at the one point that actually decides whether something goes out.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.