Describe the metric you would use to compare two models on the same task from a user's point of view.
Two models can sit three points apart on a test set and feel nowhere near that close in real use. What predicts which one people actually keep using is a blind vote on real work, not a number from a fixed answer key.
- Run a blind head to head preference test on real documents, not a benchmark score.Why: a fixed benchmark only tells you which model a grader liked once. It never asks whether a real person would trust this deck enough to walk into a room with it.
- Weight the test set to match how hard your users' real documents actually are.Why: a panel built from easy documents lets a model win everywhere while quietly losing on the cases that matter most.
- Score the win rate by document difficulty, not as one blended number.Why: a strong showing on easy documents can hide a real loss on hard ones, the same way it did here.
- Set the go or no go bar on the hard-document win rate, not the overall one.Why: the overall number can clear 60 percent while the slice that actually carries risk sits at 38.
- Flag any AI-written number that is not in the source document for a one-click check.Why: catches a made-up figure before it reaches a boardroom, whichever model shipped.
- Keep the old benchmark only for layout checks, not for picking the winning model.Why: it is still fine for a bullet that overflows its box, it was never built to judge trust.
How to answer this, stage by stage
Nobody is grading whether you know what a benchmark is. They are grading whether you would ship the model that wins on paper, or go find out if a real person would actually pick its work.
Let's learn
Here is what happens when a model wins on paper and loses in the room that actually matters.
Deckwright is a tool that takes a rough document, a memo, a report, a stack of notes, and turns it into a formatted slide deck. About 40,000 teams use it every week, and it builds roughly 210,000 decks in that time. Junhee Alaka is the product manager who decides which model powers the button that does the building.
For two years, Deckwright picked its model the same way every time. A fixed set of 500 documents, gathered when the product first launched, got run through each candidate. A scorer checked how much of each slide's text actually matched the source. The model that scored higher on that old 500-document test won the job.
The current model, call it Model A, costs about 34 cents a deck and takes 14 seconds. A cheaper candidate, Model B, costs 19 cents and takes 8 seconds. On the fixed test, Model B scored 94 percent. Model A scored 91. On paper, Model B was the better model, and a cheaper one too.
Junhee did not trust the fixed test on its own. She built a weekly blind panel: 60 real documents, pulled straight from that week's actual uploads with names stripped out, run through both models, shown to 40 opted-in beta users and three trained graders. Nobody saw which deck came from which model. They just picked the one they would present.
For the first four weeks, the panel was drawn from documents that looked like the old benchmark, short outlines, simple memos. Model B won 69, 67, 70, and 68 percent of the votes. Finance was already pushing for the cheaper model, so the team quietly moved 15 percent of production traffic to Model B in week two.
Then Junhee changed one thing. She rebuilt the panel to match what people actually upload now, not what they uploaded two years ago. About 30 percent of real documents today are dense: quarterly numbers, investor updates, technical reports full of tables and figures. The old benchmark had almost none of those.
Week five, with the real mix in the panel, Model B's win rate fell to 61 percent. Week six, 55. Week seven, 50, and the split underneath explained why: on easy documents, Model B still won 74 percent of the vote. On the dense ones, it won only 38.
By then it was already worse than a slow number. In week ten, a startup's CFO uploaded a report to build a board deck. Model B compressed a table and wrote "$2.4M ARR" on a slide. The real figure in the source document was $4.2M. The number was never in the document Model B actually read. The slide just said it with total confidence. The CFO caught it minutes before the meeting, posted about it, and Deckwright took 43 support tickets in two days.
What it cost, at its worst: a public complaint from a paying customer, a spike in refund requests, and a model already handling 30 percent of production traffic on exactly the documents it was worst at.
What I would leave alone: the same old 500-document set is still useful for one thing, catching layout bugs. Does a bullet overflow its box, does a font shrink to nothing on a crowded slide. That kind of check does not care whether the source document was easy or hard, so there is no reason to retire it, only to stop letting it pick the winning model.
The lesson: a benchmark score can be completely true and still be the wrong number to trust. Ninety four percent told the team Model B was strong. It never had a way to say that its strength lived entirely in the easy half of the work, and that the other half was exactly where a wrong number costs the most.
Now here is the same thing as a story
Read this version when you want to feel why a winning score can still be the wrong call, not just be told that it is.
Junhee Alaka has run model decisions for Deckwright's deck-building feature for two years. She can tell within a day whether a new candidate model is actually better or just better at sounding better, mostly from watching what real users do with it, not from a leaderboard.
When the finance team first pushed for Model B, the case for it looked clean. Cheaper, faster, and three points ahead on the test the company had always used. Junhee did not fight the number. She just refused to move the whole feature on it alone, and started a weekly blind panel to check it against real people instead.
The first month of that panel felt like a formality. Every week, a folder of real documents, stripped of names, ran through both models. Every week, a small group of beta users and graders picked the deck they liked without knowing who made it. Every week, Model B won by a wide, steady margin. Junhee let finance start the rollout. Fifteen percent of new decks, then a plan to grow it.
She kept the panel running anyway, mostly out of habit. Then she noticed something about the folder itself. It was mostly the same kind of document, over and over, short outlines and simple memos. Nobody had picked those on purpose. It was just what the team happened to have lying around from the early days, the same kind of thing the original benchmark was built from.
So she rebuilt the panel on purpose, pulling documents in the same mix real users actually upload, which meant pulling in a lot more of the dense, numeric ones the old test had barely touched. The next week, the winning margin that had looked so steady simply was not there anymore.
It was never really about whether the blended number moved. What moved was something the old benchmark was never built to see: whether a set of 500 documents, gathered two years ago and never revisited, could still answer for a company whose users had quietly started uploading something else entirely.
The decision that opened the door went back to launch. Someone had to pick a starting set to score models against, and the easiest 500 documents to gather were the marketing outlines and memos that made up most of the traffic back then. Nobody planned to keep using that exact set forever. It just never felt urgent enough to replace, especially with a new model candidate always waiting to ship.
Run the same week again with one change: score every candidate on a panel weighted to the real mix of hard and easy work, not the easy leftovers from launch. Same documents, same models, but this time the dense-document gap shows up in week one instead of hiding until a customer finds it. Model B still ships, just narrower, only on the 70 percent of documents where it actually earns the win, with the pricier model kept on the 30 percent where a wrong number is expensive.
One design let a test built from convenience decide which model users would trust with a board deck. The other asks the hard cases the same question it asks the easy ones. Those are not the same test wearing different clothes. One has a hole the highest-stakes work falls straight through. The other does not.
What I would tell myself, back when that first 500-document set got built: the moment a test starts standing in for a whole company's future work, ask whether it still looks like that work, not whether it once did. Nobody asked, for two years. That is on the room, not on either model.
LEAD, the four calls behind the migration decision
Not a diagnosis of a broken model. This is LEAD run on a metric question, with the test set itself as the thing that quietly went stale.
Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was doubling the old benchmark to 1,000 documents, still built the same way, from the same easy-leaning sources, a fix finance actually proposed because it sounded like more rigor for less work. It lost because a bigger sample drawn from the same skewed mix only narrows the error bar around the same misleading average. It would still have shown Model B winning, since the dense-document failure was never in that folder to begin with. The AI-specific failure worth naming by name is confident, ungrounded number generation: when Model B compresses a dense report into a bullet, it can state a specific figure that never appears anywhere in the source, delivered with the same clean tone as a correct one, because it is optimizing for a slide that reads well, not a fact that was checked. The guardrail is concrete: any number on a generated slide gets matched against the source document's text before export, and anything that does not match is held for a one-click human confirm rather than shipped straight through. That check is not free of cost, it adds one interaction for the user on the documents it fires for, a real trade worth taking on the dense 30 percent where a wrong number is expensive, and not worth forcing onto the easy 70 percent where it would just be friction. And the bar for shipping is not zero misses. A model compressing 210,000 documents a week cannot promise that. It is an audited win rate of at least 55 percent on the hard-document slice specifically, checked on a panel that is rebuilt to match real uploads every quarter, not a fixed folder from launch day.
And if you want to be sure it really works, try it somewhere else
Same four letters, a veterinary clinic chain instead of a slide-deck tool, nothing about presentations anywhere in sight.
Alderglen Veterinary Group runs twelve clinics and uses a tool called Vetbrief to turn a vet's spoken or typed visit notes into a summary email a pet owner reads after the appointment. Kiri Baio runs client communications across the group and is choosing between the current model, Model X, and a cheaper candidate, Model Y.
L, the outcome that matters. Not Model Y's 95 percent clinical-detail score against Model X's 92, measured 18 months ago on routine wellness visits. Whether a pet owner reads the summary, understands the plan, and gives the right medicine to the right condition without a confused follow-up call.
E, the early signal. A weekly blind panel: real, anonymized visit notes, both models' summaries side by side, no label, and a mix of vets and front-desk staff (who field the confused calls) pick the one they would actually send.
A, how it gets gamed. The first three weeks, the panel was built from routine visits, the same kind the old benchmark used. Model Y won 71, 73, and 70 percent. Once Kiri rebuilt the panel to match the real caseload, 35 percent of visits involve two or three problems at once, the win rate fell to 58, then 52 percent. Split by case type: 77 percent on routine visits, only 34 on multi-issue ones, where Model Y sometimes attached one condition's dosage to the wrong medicine in the same sentence.
D, the decision. Full switch needed at least 65 percent overall and at least 50 on the multi-issue slice. Model Y missed the second bar badly, so it now only drafts summaries for routine visits, 65 percent of the caseload, while Model X keeps the multi-issue ones, and any summary naming two or more medicines gets a one-tap vet check before it sends.
Swap the trigger and it still runs.
Speed: an interviewer caps you at 90 seconds. Skip straight to the direct answer, run it blind, weight it to the hard cases, do not trust the benchmark alone.
Cost: there is no budget yet for a large beta panel. Use a smaller panel of trained graders instead, still blind, still weighted to real difficulty, and say plainly the read is provisional until it grows.
The model got better, for real: say the overall benchmark score improves again next quarter. That is not proof the hard-document slice improved with it. A model can get better on average while staying exactly as wrong on the rare, expensive cases.
Where people run it wrong.
They read one blended win rate as proof there is no gap anywhere, and never check whether the hard slice got a fair share of the test.
They chase a bigger version of the same test instead of a differently built one.
They fix the one wrong number by hand and never touch the test mix or the routing, so the next hard document sails through the same gap.
How to use it live. Say the real question out loud before answering it: "is this model actually worse here, or did my test never really look at the documents that matter." That buys a beat to think instead of repeating a benchmark score as if it settles the question.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if you don't have enough hard documents yet to build a fair panel?" Response: oversample the ones you do have on purpose rather than waiting for volume, and say plainly that the read stays provisional until the sample grows.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?