The direct answer
Four explanations with four different causes: the benchmark and the users are counting different words, the errors got fewer but harder to spot, the real input got harder while the test set stayed frozen, or something else shipped alongside the model. Before any of them, rule out the counting, because a new feedback button raises complaints with nobody's day changing. Then run both models over the same 1,000 recordings and count only the risky errors that got past review, because that one number splits the top two.
Do this, in order
Give four explanations with four different causes, then say which two you would bet on.Why: the question asks for four. Four ways of saying "the model is worse than the score suggests" is one answer told four times, and an interviewer hears that straight away.
Rule out the counting before you doubt the model.Why: if the same release added a thumbs-down button, complaint volume goes up on day one with nothing about the experience changing. A tracking change and a real change look identical on a dashboard.
Split the errors by what they cost, not by how many there are.Why: drug names, doses and numbers are about two words in every hundred, and almost all of the risk. Overall word accuracy can rise while that small part gets more dangerous.
Name the single check that separates your top two, and say what each result would mean.Why: same 1,000 recordings, both models, risky errors only, split into caught in review and reached the doctor. Holding the audio still is what makes the answer readable.
Ask whether the input got harder, not just whether the model got better.Why: once it worked, doctors started dictating in corridors and trusting it with messier cases. The model improved and the average recording got worse in the same month.
Find out what else shipped that week before you blame the model at all.Why: complaints stick to the most visible recent change. A slower screen or an extra confirm step gets reported as "the transcription got worse."
How to answer this, stage by stage
Seven moves. This is a diagnosis question, so most of the work is crossing things off in the right order. Each stage has the words you would actually say.
1
Pin down the scene before you name a suspect
Say it like this
"Let me make this concrete so I'm not guessing in the air. Say it's a tool that drafts a clinical note from a doctor's dictation, and Hana runs the team that reads those drafts before anyone signs them. Word accuracy on our test set went from 96 to 98, and doctor complaints went from about eleven a month to twenty-six."
Why this works
Two numbers and one person, in fifteen seconds. Now every hypothesis you offer has something real to attach to, and you can't drift into a list of generic causes.
2
Rule out the counting before you doubt the model
Say it like this
"The first thing I'd check isn't the model at all. It's whether we changed how complaints get to us. If that release added a thumbs-down button, complaints go up on day one and not one person's day is different. So I want complaints per thousand notes, on the same reporting path, before and after."
Why this works
A tracking change and a behaviour change look exactly the same on a chart. Saying this out loud, unprompted, is the single fastest way to sound like someone who has debugged a real metric.
3
Line up four suspects, and keep it to four
Say it like this
"Okay, four suspects, and they're genuinely different causes. One: the benchmark counts every word the same and the doctor doesn't. Two: fewer errors, but the new ones read cleanly, so they get past review. Three: the audio and the cases got harder while the test set stayed frozen. Four: something else shipped that week and the complaints landed on us."
Why this works
Four causes, not four symptoms. Most candidates list ten things that all reduce to "the model is worse than it looks," which is one idea wearing four hats.
4
Bet on two of them, out loud
Say it like this
"I'd put my money on the first two. Drug names, doses and numbers are roughly two words in every hundred, so a model can gain a whole point of overall accuracy while getting worse right there. And the shape of an error matters as much as the count. A garbled word gets caught by whoever reads the note. A clean, wrong dose gets signed."
Why this works
Listing four and ranking none reads as hedging. Picking two and saying why turns a list into a diagnosis, and it gives the interviewer something to push back on.
5
Name the single check that splits them
Say it like this
"Here's the one check I'd run. Take a thousand recordings we already have, run the old model and the new one over all of them, and count only the errors in drug names, doses and numbers. Then split those two ways: caught in review, or reached the doctor. Same audio both times, so the recordings can't be what changed."
Why this works
It holds three things still and moves one. That is what makes a result readable, and it also quietly rules out the harder-input explanation while you test the other two.
6
Say what you would do either way the check lands
Say it like this
"If the new model makes more risky errors even though word accuracy went up, that's a measuring problem, and I'd change what we score on before I touch the model. If it makes fewer risky errors but more of them get through, that's a spotting problem, and the fix lives in review: flag every drug name and every number for a second read, however tidy the note looks."
Why this works
Committing to what each result would mean, before you have it, is what stops the check turning into "we'd look at the data and see." It also shows both fixes are cheap, which makes the check worth running.
7
Close on the sentence that explains both numbers at once
Say it like this
"So both numbers can be true in the same month. The benchmark counts words and the doctor counts risk, and those two can move in opposite directions. My first check is risky errors that reached a doctor, on the same thousand recordings, because that's the number the complaints are actually about."
Why this works
The question sounds like a contradiction, so the strongest ending is the line that dissolves it. Interviewers remember your last sentence, and this one is the whole answer folded into three parts.
If you remember one thing
Stages 2 and 5 are what the interviewer is really grading. Rule out the counting before the behaviour, then name one check that holds everything still except the thing you're testing. The four explanations are the easy half.
Let's learn
The tool is a draft note on a screen. A doctor talks into a phone for ninety seconds after seeing a patient, and by the time they reach the next room there is a written note waiting to be read, corrected and signed.
Hana runs the review team for a hospital group. Nine people, about a thousand notes a day. Before the tool, the notes were typed by an outside service overnight and came back the next morning. Now they come back in about twenty seconds.
For the first year her team read every draft line by line. About six minutes a note. As the tool got good, that settled into a skim. About two minutes.
Then the new model shipped. On the fixed test set, word accuracy went from 96.2 to 97.9. That is a real gain and nobody made it up. In the same month, complaints from doctors went from eleven to twenty-six.
Hana read all twenty-six. Twenty-three of them were about a drug name, a dose, or a number.
Two in a hundred words, nine in ten complaints
Words in a typical note
2 in 100 words
What the complaints were about
23 of 26
Ordinary words
Drug names, doses, numbers
The part of a note that can hurt somebody is small enough to vanish inside an overall accuracy score. Move it two points in the wrong direction and the score barely twitches.
Here is the turn. The extra complaints are not the problem, and the model getting worse is not the story either. Hana's team ran the old model and the new model over the same thousand recordings and counted only the risky items. The new model made fewer risky errors. It made 41 where the old one made 62.
But review caught 55 of the old model's 62. It caught 24 of the new model's 41.
The new model made twenty-one fewer risky mistakes and sent ten more of them to a doctor.
Review used to catch nearly nine in ten. Now it catches under six in ten. The reason is not that the reviewers got lazy. The old model's mistakes announced themselves. It would drop a syllable, garble a brand name into something that isn't a word, or leave a gap. Any of the nine people would stop on that. The new model writes fluent, confident, plausible medicine, and when it is wrong it is wrong in a sentence that reads perfectly.
Two things about an error matter, and the score only counts one of them
Knowledge spark: what a benchmark score actually counts
A benchmark is a fixed set of recordings where somebody already wrote down the right answer by hand. You run the model over them and count how many words came out different. Every word counts the same. "Patient" counts the same as "40 milligrams." Nothing in the score knows which words a person can be harmed by.
The decision I would take back. We picked one number to ship on, and we picked overall word accuracy, because it is what the vendor reports, what the research papers use, and what every team around us was already tracking. It was the sensible choice at the time and I would still understand anyone who made it. It is also a number that cannot see the difference between a wrong article and a wrong dose.
What I would score on instead
Two numbers, not one. Accuracy on drug names, doses and numbers only. And how many of those errors survive review and reach a doctor. The second one is the one nobody tracks, and it is the one the complaints are made of.
What I would leave alone. Most of a clinical note is prose. History, what the patient said, how they seemed. If the model writes "the patient reports" where the doctor said "the patient says," nobody is harmed and nobody complains. The review team should not spend a single extra second there. Slowing the whole read down to catch an odd phrasing would cost Hana's nine people hours a week and buy nothing.
The lesson. I shipped a number I could not have been hurt by. Overall word accuracy was safe to report and easy to defend, and it was blind to the exact two percent of a note where being wrong matters. The uncomfortable part is that the model really did get better, in the way I asked it to get better. I just never asked it for the thing the doctors were counting.
Five names on the board
One gets crossed off before you touch the model. The other four are the answer to the question.
Two numbers moved. Five things could have done it.
Cross this one off first: the counter
Did the way you count complaints change? A new thumbs-down button, a support form that moved to the top of the menu, a survey that started going out weekly instead of monthly. Any of those raises complaint volume on day one with zero change in what anyone experienced. Check complaints per thousand notes on the same reporting path, before and after. If the release touched the feedback route at all, you have no signal yet and everything below is guesswork.
Explanation 1
Wrong words. The benchmark and the doctor count different things.
The benchmark counts every word the same. "The" and "40 milligrams" are worth one point each. Drug names, doses and numbers are about two words in every hundred, and they are where nearly all the risk sits. So a model can gain a point and a half on the overall score by getting better at the easy ninety-eight percent, and be flat or worse on the two percent that people actually notice.
Nobody lied. The score is real. It is measuring a different thing from the one being complained about.
How you'd test it: re-score both models on drug names, doses and numbers only. If overall accuracy went up and that number went down, this is your answer.
Explanation 2
Hidden errors. The mistakes changed shape, not just count.
Fewer errors, and the new ones are more believable. An old-style error was a garbled word or a missing gap, and a reviewer stopped on it without thinking. A new-style error is a clean sentence with the wrong number in it, and a reviewer reads straight past.
So the count falls and the harm rises. Fewer errors that are better disguised can be worse than more errors that are obvious, and no accuracy score will ever show that, because a score has one axis and this has two.
How you'd test it: take the risky errors each model made and split them into caught in review and reached the doctor. If the total fell but the second column grew, it is this one.
Explanation 3
Harder input. The recordings changed, not the model.
This is the one people forget, and it is the only one where the model is genuinely innocent. When the tool was shaky, doctors used it in a quiet room, for a simple follow-up, speaking carefully. Once it got good, they started dictating in corridors, over a ward round, on a phone in a car park, for complicated patients on eight medications.
The test set never changed. It is the same recordings from two years ago, sitting in the same folder. So the model improved on the frozen set and got a harder job in real life at the same time. Both are true and the score only sees one.
How you'd test it: compare the audio itself, old month against new month. Background noise level, length, how many speakers, and which specialties are using it. If the mix moved, the benchmark stopped representing the job.
Explanation 4
The neighbour. Something else shipped alongside it.
Models rarely ship alone. In the same release there might be a redesigned review screen, a new confirm step before signing, or a change that made the note take nine seconds to appear instead of two. People do not report "the confirm dialog is annoying." They report "the transcription got worse," because the transcription is the thing they have a name for.
Complaints attach to the most visible recent change, and the model is always the most visible recent change.
How you'd test it: read the release notes for that week, then cut the complaints by who actually got each change. If the sites that never received the new screen complain just as much, the screen is innocent and you can drop it.
The one check that splits the top two
Explanations 1 and 2 are the strongest, and they point in opposite directions. One says the new model is worse where it counts. The other says it is better where it counts and sneakier about the rest. You cannot tell them apart from a dashboard, and they need different fixes, so guessing costs you a quarter.
The check is one query, and it holds three things still. Take a thousand recordings the team already has. Run both models over all of them, so the audio is identical and explanation 3 is off the table. Count only errors in drug names, doses and numbers, so explanation 1's blind spot is gone. Then split those errors two ways: caught in review, or reached the doctor.
Same thousand recordings, both models, risky items only
Old model
62 risky errors
55 caught in review
7reached a doctor
New model
41 risky errors
24 caught in review
17reached a doctor
That reads clearly, and it only reads clearly because the audio was held still. Risky errors fell by a third. Errors reaching a doctor more than doubled. So it is explanation 2, not explanation 1, and the fix is not in the model.
The fix is that every drug name and every number gets marked for a second read, however tidy the sentence around it looks. Hana's team stops relying on a mistake looking like a mistake, because with this model most of them no longer do.
TRACE, laid out like a case file
This is a diagnosis question, so the framework is TRACE. A "what if the error rate doubled" question would use FLIPS. Different question shapes get different tools, and forcing a story framework onto a diagnosis gives you a parable where an investigation should be.
TRACE, for diagnosis questions
T, timeline. When exactly did complaints start, and what shipped near that date. Here the model landed the same week complaints doubled, which is suspicious enough to check and weak enough to be a coincidence.
R, recut. Slice it. By specialty, by site, by how long each doctor has used the tool. Twenty-six complaints across the whole group may be four complaints from one ward.
A, assume nothing. Rule out the counting before the behaviour. If that release added a thumbs-down button, complaints rise with nothing else changing.
C, cause candidates. Four, named, and genuinely different: wrong words, hidden errors, harder input, the neighbour. Not a list of everything possible.
E, evidence test. Same thousand recordings, both models, risky errors only, split into caught and reached the doctor. The strongest move in the framework, and the one most people skip.
Why E is the hard step
Anyone can list causes. A check earns its place by holding still everything you are not testing. This one freezes the audio, the reviewers and the note format, and moves one thing. If your check would give the same answer under two different hypotheses, it is not a check yet.
Run TRACE somewhere else
A bank ships a better fraud model. In testing it catches noticeably more real fraud. Two weeks later, branch staff are complaining more than they ever did about the old one. Same question shape, completely different product.
T. Complaints did not start with the model. They started about three weeks later, which is when a separate rules refresh went live. That gap is the first real clue.
R. It is not every branch. It is the branches with the oldest customers, and the complaints are about cards declining at pharmacies and supermarkets on a Saturday morning.
A. The complaint route changed too. Branch staff used to phone the fraud desk. Now there is a two-click form on their screen. Volume up, experience possibly unchanged. Check complaints per thousand blocked payments, not the raw count.
C. Four causes. The model catches more fraud and also blocks more good payments, which is a trade nobody priced. The customer mix shifted after a rival bank closed local branches. The fraud itself changed shape. Or the rules refresh, not the model, is doing the blocking.
E. Replay last month's real transactions through both models. Count good payments blocked, split by customer age. Same transactions, so the customer mix cannot be the reason.
Swap the trigger and it still runs
- Speed: the note now takes nine seconds to appear instead of two. Nothing about accuracy moved, and complaints climb anyway. TRACE still starts at the release date, not the complaint date.
- Cost: the group moves to a cheaper plan and the second clean-up pass gets switched off to save money. On paper the model is unchanged, which is exactly why nobody looks there.
- The model got better: the case on this page. Better model, more trust, harder recordings fed in, more complaints. Improvement is a trigger, and it is the one people never plan for.
Where people run it wrong
- Starting the timeline the day complaints spiked. The change that caused it usually landed weeks earlier, and the habit built on top of it took time to form.
- Listing ten possible causes to look thorough. Ten unranked causes is the same as none, because nothing gets tested. Four you can actually check beats ten you cannot.
- Skipping the counting check, then spending two weeks on the model, and finding out at the end that a support form moved.
If you are asked this cold
Buy yourself ten seconds by saying the two numbers out loud. "So one number went up and one number went up, and they're supposed to move in opposite directions. Let me say what each one is actually measuring." Naming the two numbers is not stalling. It is stage 1, and it is where the answer comes from, because in nearly every version of this question the two numbers are counting different things.
Flashcards (click a card to flip it)
The usual flip-family cards do not apply here, because this is a diagnosis question with no FLIPS run in it. These eight test the TRACE moves and the four explanations instead.
1 · THE FRAMEWORK
Which framework fits "the score went up but users complain more," and why not FLIPS?
Tap to flip
ANSWER
TRACE. FLIPS is for "what if X changed." Here something already changed and you have to find out what, so the job is ruling out and narrowing.
2 · THE FIRST MOVE
What do you rule out before you doubt the model?
Tap to flip
ANSWER
The counting. A new thumbs-down button or a moved support form raises complaint volume on day one with nobody's experience changing.
3 · EXPLANATION 1
How can overall accuracy rise while the doctor's experience gets worse?
Tap to flip
ANSWER
The benchmark counts every word the same. Drug names, doses and numbers are about 2 words in 100 and nearly all the risk. The gain landed on the easy 98 percent.
4 · EXPLANATION 2
What does "fewer errors, better disguised" mean in numbers here?
Tap to flip
ANSWER
Old model: 62 risky errors, 7 reached a doctor. New model: 41 risky errors, 17 reached a doctor. A garbled word gets caught. A clean, wrong dose gets signed.
5 · EXPLANATION 3
How can the model improve and the experience get worse with no change to the model at all?
Tap to flip
ANSWER
The input changed. Once it worked, doctors dictated in corridors and on harder patients. The test set stayed frozen, so the benchmark stopped representing the real job.
6 · EXPLANATION 4
Why ask what else shipped that week?
Tap to flip
ANSWER
Complaints attach to the most visible recent change. A new confirm step or a slower screen gets reported as "the transcription got worse," because that is the thing people have a name for.
7 · THE CHECK
Name the one check that splits the top two, and say what it holds still.
Tap to flip
ANSWER
Same 1,000 recordings, both models, risky errors only, split into caught in review and reached the doctor. Identical audio, so the recordings cannot be the reason.
8 · THE REVERSAL
What decision does this answer take back, and what replaces it?
Tap to flip
ANSWER
Shipping on one overall word-accuracy number. Replace it with two: accuracy on drug names, doses and numbers, and how many of those errors survive review.
Check yourself Score: 0 / 0
Multiple choice
1. You rerun the old and new model over the same 1,000 recordings. Why the same recordings?
- A. It is faster than collecting new ones.
- B. It holds the audio still, so a harder input mix cannot be the reason for whatever you find.
- C. It gives you a bigger sample than the benchmark.
- D. It lets you report the result as an accuracy score.
Show hint
One of the four explanations gets ruled out for free by this choice.
Show answer
B. Explanation 3 says the recordings got harder. Running both models over identical audio takes that off the table without a separate study, which is exactly what makes this check strong enough to act on.
True or false
2. True or false: a jump in complaints always means the experience got worse.
Show hint
Think about how a complaint gets to you in the first place.
Show answer
False. If the same release added a thumbs-down button or moved the support form, complaint volume rises on day one with nobody's day changing. Rule out the counting before the behaviour. Ask for complaints per thousand notes on the same reporting path, before and after.
Multiple choice
3. Three of these are the same explanation in different clothes. Which one is a genuinely separate cause?
- A. The model is worse than the benchmark suggests.
- B. The benchmark was too easy for the model.
- C. A new confirm step shipped in the same release, and complaints landed on the model instead.
- D. The model learned the test set rather than the job.
Show hint
Three of them end with "so the model is not as good as the number says." One does not involve the model at all.
Show answer
C. A, B and D all reduce to the same idea: the score overstates the model. C is a different cause entirely, where the model may be fine and the complaint is about the thing next to it. A list of four needs four causes, not one cause with four labels.
Fill in the blank
4. On the same thousand recordings, the new model made ______ fewer risky mistakes and sent ______ more of them to a doctor.
Show hint
62 and 7 for the old model. 41 and 17 for the new one.
Show answer
21 fewer, and 10 more. 62 risky errors down to 41, and 7 reaching a doctor up to 17. Review caught nearly nine in ten before and under six in ten after, because the new errors read like real medicine.
Short answer
5. Name a part of this same product where a wrong word would cause no complaint at all, and say why that answer helps you.
Show hint
Most of a clinical note is not a number.
Show answer
Model answer: "The prose parts. History, what the patient reported, how they seemed. If the model writes 'the patient reports' where the doctor said 'the patient says,' nobody is harmed and nobody complains, so review should not spend a second there. Naming a place you would deliberately leave alone shows judgment. A candidate who wants every word checked has just asked Hana's nine people to work slower for nothing."
Short answer, apply it yourself
6. Pick a tool you use yourself that got measurably better and felt worse to you. Run stages 2 and 5 on it: what would you rule out first, and what single check would you run?
Show hint
Think of a phone keyboard, a maps app, a spam filter, a music recommender. Something whose owners can honestly say the numbers improved.
Show answer
Model answer: "Take phone autocorrect. It got better at ordinary words and worse for me, because the errors it makes now are real words in the right grammar, so I send them without noticing. First I would rule out the counting: did they add an easy 'report this correction' button that month. Then the one check: take a thousand messages I already sent, run both versions, and count only the corrections I had to undo after sending, not the ones I caught while typing. Same messages, so my own typing cannot be the variable." Any answer works if it rules out the measurement first and then names a check that holds everything still except one thing.