Explain why improving accuracy can decrease trust.
Why a fraud model getting genuinely better can be the exact thing that makes a person stop looking, and why that costs more than the mistakes ever did.
- Keep a "first time for this customer" flag visible on every case, no matter how high the score climbs.Why: this is the one flag that would have stopped the real freeze in this story, the exact case a rising score trains people to stop checking.
- Tie any drop in review time to the eval set, never to one good week.Why: precision has to hold for three straight monthly refreshes before a default changes, so one lucky month can't quietly retrain the whole desk.
- Leave the bottom of the score range exactly as it is.Why: cases scored under 60 already force a full look and a timer. That part of the queue was never the problem.
- Don't answer this by putting mandatory review back on every flagged case.Why: that burns back most of the time the better model actually earned, and a busy analyst will find a way around a rule that treats every case the same.
- Watch the rate of opened case panels as its own number, split by score band.Why: it was falling for months before the freeze that reached a real customer, and nobody was watching it on its own.
How to answer this, stage by stage
Nobody is grading whether you can say "trust drops when the model gets better." Anyone can say that line. They're grading whether you can say why, in a way that ends in a design decision, not a warning.
Let's learn
What happens the moment a model finally gets good enough that nobody argues with it anymore?
Tarrow Bank uses a tool called ClearFlag. It scores every card and wire transaction from 0 to 100 for how likely it is to be fraud, and sends the highest scores to a fraud analyst's screen before the money moves.
Two years ago, ClearFlag caught real fraud in 41 out of every 100 cases it flagged. The other 59 were false alarms. Every flagged case took an analyst about six minutes: pull the last five transactions, compare the merchant to the customer's normal spending, and call the customer if the amount was large.
Then the bank retrained the model. ClearFlag now catches real fraud in 91 out of every 100 flagged cases. Review time for most of the queue dropped to under a minute. That is a real win, not a marketing number.
Here is the turn. The extra 50 correct catches are not the problem. The problem is what a person does with a screen that keeps being right. Radomir Ashcroft, who has worked the queue for eight years, used to open every flagged case's history before deciding anything. Once the score kept landing above 95 and kept being correct, week after week, he stopped opening it. He started reading the number instead of the case.
At its worst, this costs more than it looks like it does. A wrong score used to get caught, because someone was always looking behind it. Now the rare wrong one goes through with nobody looking at all. It does not get smaller as the model improves. It gets rarer and bigger, and by the time anyone notices, it has already reached a real customer.
What I would leave alone: cases scored under 60 still force the panel open and hold the analyst there for a two minute minimum. That part of the design never changed, and it doesn't need to. The habit only broke at the very top of the range, where a confident number quietly became certainty.
The lesson: a model that gets better is still a change to the product. Nothing about your design gets to assume the old habit survives it. If it only worked at one exact level of accuracy, it was never really finished, it was just waiting for the day it improved enough to break.
Now here is the same thing as a story
The short version sits above. Read on for the wedding wire that sat frozen for four days.
Radomir Ashcroft can tell a stolen card from a scared customer in under a minute, most days, without opening a single transaction. Eight years on the fraud desk will do that. Analysts younger than him ask how he does it. He usually says he doesn't know, it's just the shape of the thing.
ClearFlag arrived at Tarrow four years ago. For the first two years, the case panel opened by itself under every flag, because two out of every three flags back then were nothing, and Radomir needed the history to sort real fraud from a customer buying a used car. He'd pull the last five transactions, check the merchant against the customer's normal pattern, and if the amount was large, he'd call. Six minutes a case, most afternoons, and the queue never really beat him.
The retrain shipped in the spring. Precision climbed a point or two a month, 41, then 58, then 74, then 91, and stayed there. Somewhere around month six, Radomir noticed he wasn't clicking the panel open on the highest scores anymore. He hadn't decided to stop. It kept being right, so his hand just stopped reaching for it.
By month ten, he was opening maybe 1 in 25 of the score 95-and-up cases. Nobody told him to stop. Nobody told him to keep going either. The number just kept being right, and being right is a hard thing to argue with.
Then came a Thursday in October. A flag landed for Gwendolyn Isakova, a customer of twelve years, an eighteen thousand four hundred dollar wire to a caterer in Lisbon, a country she had never sent money to, a payee ClearFlag had never seen her use. Score: 97. High enough that Radomir, without opening a thing, clicked confirm and froze the account.
It was her daughter's wedding. The caterer needed the payment two days before the event or the booking fell through. Gwendolyn found out her card was declined standing in the caterer's kitchen in Lisbon, three days before the wedding, with her daughter next to her.
It took four days for Tarrow's dispute process to clear the freeze. The caterer nearly walked. Gwendolyn's complaint reached the bank's ombudsman before the wedding even happened, and it did not stay quiet after.
Here's what nobody wants to say out loud about it. ClearFlag was not wrong to score that wire high. A first ever wire, to a new country, to a payee with zero history, genuinely looks like fraud most of the time. The model did its job. Radomir did what eight years of a queue that kept being right had trained him to do. Nobody in that story made a mistake you could point to and call careless.
What was missing was smaller than either of them. A flag, sitting next to the score, saying this exact pattern, this payee, this country, has never happened on this account before. Something that could not disappear just because the number next to it looked confident.
The old decision that set this up goes back to the year the queue tripled, when someone reasonably decided the case panel shouldn't force itself open on every single flag anymore, because a confident analyst working a real backlog needed the option to move fast. Nobody put a date on revisiting that call once the score started climbing. Nobody did, for four more retrains.
Run the same Thursday again with one flag added: a small badge under the score reading "first time: this payee, this country," that stays lit no matter how high the number climbs. Radomir sees it at a glance, the flag alone costs him nothing extra to notice, and it's enough to make him place one call. Ninety seconds. Gwendolyn explains the wedding. The wire clears the same afternoon, the caterer gets paid on time, and nobody outside Tarrow's fraud desk ever hears about it.
What I would tell myself, back in the meeting where the panel default changed: the day you let a screen stop insisting on a look, write down which exact patterns should never be allowed to stop insisting. Somebody has to hold that line on purpose, or the model's own good news quietly erases it.
FLIPS, and the one letter that actually mattered here
Not a diagnosis of one bad Thursday. FLIPS run on the deeper claim: that a model getting better is still a perturbation, and it needs the same five questions any other change does.
Two things worth naming plainly, since this is where the real judgment sits. The alternative I would reject is putting mandatory review back on every flagged case, the way it worked at 41 percent precision. It got floated after the Isakova freeze, and it would burn back most of the time the better model actually earned. An analyst working a full afternoon queue under pressure will find a way around a rule that treats a 97 on a payee ClearFlag has scored five hundred times the same as a 97 on a payee it has never seen. The novelty flag targets the narrow slice that's actually risky, instead of taxing the whole queue equally.
The AI-specific failure worth naming by name: a model can be very sure about a pattern it has almost never seen. A first-time payee in a new country is rare enough that ClearFlag's confidence there is not backed by much real history, even though the number on screen looks exactly like every other confident score. That's the model being confidently wrong on a case sitting outside what it mostly learned from, and a flat score can't tell you which kind of confident you're looking at. The guardrail is the novelty flag itself, one that cannot hide, because a high number on a brand new pattern means something different from a high number on a pattern the model has seen a thousand times.
There's a real cost to the fix, and it's worth saying out loud instead of pretending the flag is free. Forcing a look on every first-time-payee case, even at a score of 99, adds back close to ninety seconds to exactly the cases that would otherwise be an instant confirm. On the busiest afternoon stretch, that can back the queue up by a few minutes. Tarrow took that trade on purpose: a slightly slower afternoon against another multi-day freeze on a stranger's wedding money. And the bar for trusting a precision number enough to loosen any default isn't a single good week. ClearFlag's precision has to clear 90 out of 100 on the fraud team's own monthly labeled eval set, the same roughly two thousand reviewed cases, for three refreshes running, before review time gets shortened for any band of the queue.
And if you want to be sure it really works, try it somewhere else
Same five letters, a veterinary hospital instead of a bank, and a completely different flip, one that fires when the model improves in a different way: it lets a senior hand work down instead of making her stop looking herself.
Falken Veterinary Hospital runs TrustBeacon, a tool that reads chest and limb radiographs and flags each one "likely benign" or "needs review" before a vet signs off. Dr. Costanza Ferren has run the imaging room there for six years, the only vet on staff who reads films herself rather than sending them out.
F, find the person. Dr. Costanza Ferren, six years reading films at Falken, and Alden Marrowbrook, her vet tech of three years, who confirms reads on the floor once she's checked them.
L, locate the habit. She used to review every TrustBeacon-flagged film herself before Alden confirmed a read to an owner. It cost her almost nothing most days, and it caught two real misses in the year before the model improved.
I, identify the flip. This is the delegation flip, not the over-trust flip. Once TrustBeacon agreed with her own reads on 9 out of 10 films for three straight months, she let Alden confirm reads alone, no second look. A hairline fracture came through labeled "likely benign," Alden confirmed it, the dog went home, and it needed a second surgery six weeks later. She now reads every film herself again, and Alden hasn't confirmed a read alone since.
P, pinpoint the old decision. TrustBeacon's screen shows one flat label under the film, "likely benign" or "needs review," nothing about which part of the bone the model weighed most. That felt fine when Alden only ever confirmed what Costanza had already checked. Once he was reading alone, there was nothing on the screen that let either of them see how the call got made.
S, show the replay. Same fracture, same film, run through a screen that shows the region TrustBeacon weighed most under the label. Alden sees the flagged region sits half a centimeter from where the model's confidence was actually built, pulls two more angles, and catches it himself. Costanza reviews 1 film in 10 instead of all of them, and gets her afternoons back.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the flag, and say why it stays lit at every score, not just low ones.
Cost: the fraud team says building a live novelty flag is six weeks out. Don't wait on it, pull "new payee, new country" as a manual daily report until it ships.
The model got better, for real: say precision climbs to 99 next year. That's still not the same claim as "every case is fine now." A model that gets more accurate on average can still be badly wrong on a pattern it almost never saw, and a flat score hides that difference either way.
Where people run it wrong.
They add the flag, then keep the panel collapsed by default anyway, because collapsing it was the habit they'd built for themselves too.
They see the blended precision number climbing and call that proof the whole queue is safer, instead of asking which slice of the queue nobody is looking at anymore.
They respond to one bad freeze by adding review back everywhere, instead of asking whether the real fix is a guardrail on the rare pattern, not a tax on the common one.
How to use it live. Open with the flag, not the story: "I'd keep one signal visible no matter the score, and here's the exact pattern it protects." That buys you the room to walk the interviewer through the failure on your own terms, instead of them walking you into it.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the customer had made international wires before, wouldn't the flag never fire?" Response: then it shouldn't fire, that's the point. The flag tracks whether this exact pattern is new for this exact customer, not whether wires in general are risky, so it stays quiet on a payee ClearFlag has already seen many times.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #6 Describe the calibration problem: what happens when confidence does not match correctness?
- #7 How do you measure whether users over-trust your AI feature?