ConceptAdvancedQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #5

Explain why improving accuracy can decrease trust.

Why a fraud model getting genuinely better can be the exact thing that makes a person stop looking, and why that costs more than the mistakes ever did.

The direct answer
Improving accuracy decreases trust because it removes the friction that kept a person checking the model's work. Once the score is right often enough, the analyst stops opening the case behind it, so the rare wrong call now ships with nobody looking at all. Keep one flag visible no matter how high the score climbs, a "first time for this customer" flag on a new payee, a new country, or a new pattern, so the case most likely to be the rare miss is also the one case a rising score can never quietly train someone to skip.
Do this, in order
  1. Keep a "first time for this customer" flag visible on every case, no matter how high the score climbs.Why: this is the one flag that would have stopped the real freeze in this story, the exact case a rising score trains people to stop checking.
  2. Tie any drop in review time to the eval set, never to one good week.Why: precision has to hold for three straight monthly refreshes before a default changes, so one lucky month can't quietly retrain the whole desk.
  3. Leave the bottom of the score range exactly as it is.Why: cases scored under 60 already force a full look and a timer. That part of the queue was never the problem.
  4. Don't answer this by putting mandatory review back on every flagged case.Why: that burns back most of the time the better model actually earned, and a busy analyst will find a way around a rule that treats every case the same.
  5. Watch the rate of opened case panels as its own number, split by score band.Why: it was falling for months before the freeze that reached a real customer, and nobody was watching it on its own.

How to answer this, stage by stage

Nobody is grading whether you can say "trust drops when the model gets better." Anyone can say that line. They're grading whether you can say why, in a way that ends in a design decision, not a warning.

1
Ground it in one bank, one tool, one person
Say it like this
"Let's ground this. Tarrow Bank runs a tool called ClearFlag, it scores every transaction for fraud risk and flags the ones a person should look at. Radomir Ashcroft has worked that queue for eight years."
Why this works
A general answer about trust in AI is a wish. One person's queue, on one screen, is a decision you can defend.
2
Say your structure out loud
Say it like this
"I'd run this through FLIPS. Find the person whose queue this is, find the habit that quietly built up, name the exact switch that flips, name the old decision behind it, then run the same bad day through a fix."
Why this works
Two seconds of structure stops the answer turning into a general lecture on AI and trust.
3
Reframe what the question is really testing
Say it like this
"This sounds like a contradiction, a better model, worse trust. It isn't really. It's asking whether I know trust lives in a person's habit, not in the model's score, and a habit doesn't move smoothly just because a number does."
Why this works
Naming the real question stops you from giving a vague answer about calibration that never reaches an actual person.
4
Give the decision, plainly
Say it like this
"My answer: keep one flag visible no matter the score, a 'first time for this customer' flag on the payee, the country, or the pattern. A rising score is exactly what makes a person stop looking, so the one thing that should never quietly disappear is the flag on the case most likely to be the rare miss."
Why this works
This is the answer to the question. Everything after it is proof, not more opinion.
5
Prove it with the compressed failure
Say it like this
"Radomir's queue went from catching real fraud in 41 out of 100 flagged cases to 91 out of 100. Genuinely better. But he stopped opening the case panel once a score passed 95, because it kept being right. Then a longtime customer's first ever wire to Portugal scored 97, real, for her daughter's wedding, and he froze it without a look. It took four days to sort out. Nothing was wrong with the score. The one thing missing was any flag saying this exact pattern was new for her."
Why this works
A four day freeze on a wedding payment lands harder than a percentage, and it points straight at the missing flag, not just the failure.
6
Say what you'd leave alone
Say it like this
"I wouldn't touch anything scored under 60. Those cases already force a full look and a two minute timer, and that part of the queue was never where this problem lived."
Why this works
Naming a place the fix doesn't apply shows judgment. Fixing everything makes an answer sound like fear, not a decision.
7
Close on the one line you opened with
Say it like this
"So: a better model doesn't remove the need for a look, it just moves where the look has to happen. Keep the flag on the rare pattern, and let the score handle everything else."
Why this works
Closing on the same decision you opened with is what makes an interviewer believe you actually decided something, instead of thinking out loud.
If you remember one thing A model getting better is still a change to the product. The habit it built at 41 out of 100 does not survive unchanged at 91 out of 100. Something has to be redesigned on purpose, or the improvement quietly redesigns the person instead.

Let's learn

What happens the moment a model finally gets good enough that nobody argues with it anymore?

Tarrow Bank uses a tool called ClearFlag. It scores every card and wire transaction from 0 to 100 for how likely it is to be fraud, and sends the highest scores to a fraud analyst's screen before the money moves.

Two years ago, ClearFlag caught real fraud in 41 out of every 100 cases it flagged. The other 59 were false alarms. Every flagged case took an analyst about six minutes: pull the last five transactions, compare the merchant to the customer's normal spending, and call the customer if the amount was large.

ClearFlag precision, before and after the retrain
100 0 41 91 ClearFlag v2 ClearFlag v4
Old modelRetrained model
Out of every 100 flagged cases, the share that were real fraud more than doubled. That is a genuine improvement, not a smaller version of the same problem.

Then the bank retrained the model. ClearFlag now catches real fraud in 91 out of every 100 flagged cases. Review time for most of the queue dropped to under a minute. That is a real win, not a marketing number.

Here is the turn. The extra 50 correct catches are not the problem. The problem is what a person does with a screen that keeps being right. Radomir Ashcroft, who has worked the queue for eight years, used to open every flagged case's history before deciding anything. Once the score kept landing above 95 and kept being correct, week after week, he stopped opening it. He started reading the number instead of the case.

Knowledge spark: what does "precision" actually mean here? It is how much of what the model flags turns out to be real. If ClearFlag flags 100 transactions and 91 really are fraud, that's 91 percent precision. It says nothing about the 9 that weren't fraud, and nothing about how sure the model was on any single one of them.

At its worst, this costs more than it looks like it does. A wrong score used to get caught, because someone was always looking behind it. Now the rare wrong one goes through with nobody looking at all. It does not get smaller as the model improves. It gets rarer and bigger, and by the time anyone notices, it has already reached a real customer.

We didn't stop the mistakes. We stopped the one thing that used to catch them.
The decision that mattered When ClearFlag first launched, the case panel, the last five transactions, the merchant compare, the call button, opened automatically under every flagged case, because almost every case needed digging into back then. When the queue tripled after Tarrow added overseas transfers, the team set the panel to stay collapsed until clicked, so a confident analyst could move faster. That made sense in year one. Nobody revisited it as the score climbed, because "click to expand if you need it" quietly became "never click," once "need it" started rounding to zero in Radomir's head.

What I would leave alone: cases scored under 60 still force the panel open and hold the analyst there for a two minute minimum. That part of the design never changed, and it doesn't need to. The habit only broke at the very top of the range, where a confident number quietly became certainty.

The lesson: a model that gets better is still a change to the product. Nothing about your design gets to assume the old habit survives it. If it only worked at one exact level of accuracy, it was never really finished, it was just waiting for the day it improved enough to break.

Now here is the same thing as a story

The short version sits above. Read on for the wedding wire that sat frozen for four days.

Radomir Ashcroft can tell a stolen card from a scared customer in under a minute, most days, without opening a single transaction. Eight years on the fraud desk will do that. Analysts younger than him ask how he does it. He usually says he doesn't know, it's just the shape of the thing.

ClearFlag arrived at Tarrow four years ago. For the first two years, the case panel opened by itself under every flag, because two out of every three flags back then were nothing, and Radomir needed the history to sort real fraud from a customer buying a used car. He'd pull the last five transactions, check the merchant against the customer's normal pattern, and if the amount was large, he'd call. Six minutes a case, most afternoons, and the queue never really beat him.

The retrain shipped in the spring. Precision climbed a point or two a month, 41, then 58, then 74, then 91, and stayed there. Somewhere around month six, Radomir noticed he wasn't clicking the panel open on the highest scores anymore. He hadn't decided to stop. It kept being right, so his hand just stopped reaching for it.

Hand sketched two panel sketch titled small move big snap. Left panel a gauge labeled the score, caption climbs slowly forty one to ninety one, a point or two a month. Right panel a broken scale labeled the habit, caption holds holds holds then snaps once, reads the case to trusts the score.
The score moved a point at a time for ten months. The habit behind it did not move at all, until the one week it moved completely.

By month ten, he was opening maybe 1 in 25 of the score 95-and-up cases. Nobody told him to stop. Nobody told him to keep going either. The number just kept being right, and being right is a hard thing to argue with.

Then came a Thursday in October. A flag landed for Gwendolyn Isakova, a customer of twelve years, an eighteen thousand four hundred dollar wire to a caterer in Lisbon, a country she had never sent money to, a payee ClearFlag had never seen her use. Score: 97. High enough that Radomir, without opening a thing, clicked confirm and froze the account.

It was her daughter's wedding. The caterer needed the payment two days before the event or the booking fell through. Gwendolyn found out her card was declined standing in the caterer's kitchen in Lisbon, three days before the wedding, with her daughter next to her.

We didn't take nine emails from her. We took her daughter's wedding week.

It took four days for Tarrow's dispute process to clear the freeze. The caterer nearly walked. Gwendolyn's complaint reached the bank's ombudsman before the wedding even happened, and it did not stay quiet after.

Here's what nobody wants to say out loud about it. ClearFlag was not wrong to score that wire high. A first ever wire, to a new country, to a payee with zero history, genuinely looks like fraud most of the time. The model did its job. Radomir did what eight years of a queue that kept being right had trained him to do. Nobody in that story made a mistake you could point to and call careless.

What was missing was smaller than either of them. A flag, sitting next to the score, saying this exact pattern, this payee, this country, has never happened on this account before. Something that could not disappear just because the number next to it looked confident.

Hand sketched comparison titled switch not dial. Left, a gauge labeled what we assumed, caption a dial, trust ticking up as the score climbs. Right, a plain box labeled what actually happened, caption a switch, only two settings, reads the case or clicks confirm on the score alone.
Everyone designing the screen assumed trust would climb like the score did, a little at a time. It didn't. It held, then snapped, once, with nothing in between.

The old decision that set this up goes back to the year the queue tripled, when someone reasonably decided the case panel shouldn't force itself open on every single flag anymore, because a confident analyst working a real backlog needed the option to move fast. Nobody put a date on revisiting that call once the score started climbing. Nobody did, for four more retrains.

Run the same Thursday again with one flag added: a small badge under the score reading "first time: this payee, this country," that stays lit no matter how high the number climbs. Radomir sees it at a glance, the flag alone costs him nothing extra to notice, and it's enough to make him place one call. Ninety seconds. Gwendolyn explains the wedding. The wire clears the same afternoon, the caterer gets paid on time, and nobody outside Tarrow's fraud desk ever hears about it.

What I would tell myself, back in the meeting where the panel default changed: the day you let a screen stop insisting on a look, write down which exact patterns should never be allowed to stop insisting. Somebody has to hold that line on purpose, or the model's own good news quietly erases it.

FLIPS, and the one letter that actually mattered here

Not a diagnosis of one bad Thursday. FLIPS run on the deeper claim: that a model getting better is still a perturbation, and it needs the same five questions any other change does.

Hand sketched five step list titled FLIPS the five steps. One, F, find the person, whose queue is this at Tarrow Bank. Two, L, locate the habit, what did he stop doing because it kept working. Three, I, identify the flip, reads the case or trusts the score, no middle, shown in red. Four, P, pinpoint the old decision, the panel that used to open on its own. Five, S, show the replay, one badge brings the look back.
The middle step, in red, is the one worth the most time. Everything before it sets up the story. Everything after it proves the fix.
F
Find the person. Whose queue is this?
Radomir Ashcroft, eight years on Tarrow Bank's fraud desk, known for reading a flagged case in under a minute, most days without needing the history at all.
Naming one person with a real skill is what stops this from being a story about "AI trust" in general.
L
Locate the habit. What did he stop doing because it worked?
Opening the case panel before deciding, the last five transactions, the merchant compare, the customer call. It faded in three real steps: he stopped calling first, then stopped comparing the merchant, then stopped opening the panel at all above a score of 95.
The habit is the actual product Tarrow built. The score was only ever the excuse to stop needing it.
I
Identify the flip. What verb snaps?
Opens the case before deciding, or trusts the score alone above a cutoff. No middle setting. Once a score crossed the point that felt obviously right, Radomir stopped weighing how sure he was and started treating the number as a fact.
This is the over-trust flip, the one family that fires on good news instead of bad. Most FLIPS answers reach for the version where the model gets worse. This one is the reverse, and it's the harder one to spot.
P
Pinpoint the old decision. Which choice only made sense before?
Setting the case panel to stay collapsed by default once the queue tripled, instead of forcing it open for the narrow slice of cases where the pattern itself was new. Sensible when the queue was the bottleneck. Nobody revisited it once precision climbed.
A reversible screen default, not a policy or a training. That's what makes it a real decision, not a wish.
S
Show the replay. Same bad day, new design.
Gwendolyn Isakova's wire scores 97 again. This time a "first time: this payee, this country" badge sits under the score and never turns off. Radomir sees it, makes one call, 90 seconds, the wire clears the same afternoon.
The replay ends in a clock, not a feeling. That's what turns a design idea into something you can defend under a follow up question.

Two things worth naming plainly, since this is where the real judgment sits. The alternative I would reject is putting mandatory review back on every flagged case, the way it worked at 41 percent precision. It got floated after the Isakova freeze, and it would burn back most of the time the better model actually earned. An analyst working a full afternoon queue under pressure will find a way around a rule that treats a 97 on a payee ClearFlag has scored five hundred times the same as a 97 on a payee it has never seen. The novelty flag targets the narrow slice that's actually risky, instead of taxing the whole queue equally.

The AI-specific failure worth naming by name: a model can be very sure about a pattern it has almost never seen. A first-time payee in a new country is rare enough that ClearFlag's confidence there is not backed by much real history, even though the number on screen looks exactly like every other confident score. That's the model being confidently wrong on a case sitting outside what it mostly learned from, and a flat score can't tell you which kind of confident you're looking at. The guardrail is the novelty flag itself, one that cannot hide, because a high number on a brand new pattern means something different from a high number on a pattern the model has seen a thousand times.

Knowledge spark: what is a labeled eval set? A pile of past flagged cases where a human already checked and wrote down the true answer, real fraud or not. Once a month, the fraud team scores about two thousand of these fresh with the current model and checks how many it got right. That number, not a single lucky week live, is what precision actually means here.

There's a real cost to the fix, and it's worth saying out loud instead of pretending the flag is free. Forcing a look on every first-time-payee case, even at a score of 99, adds back close to ninety seconds to exactly the cases that would otherwise be an instant confirm. On the busiest afternoon stretch, that can back the queue up by a few minutes. Tarrow took that trade on purpose: a slightly slower afternoon against another multi-day freeze on a stranger's wedding money. And the bar for trusting a precision number enough to loosen any default isn't a single good week. ClearFlag's precision has to clear 90 out of 100 on the fraud team's own monthly labeled eval set, the same roughly two thousand reviewed cases, for three refreshes running, before review time gets shortened for any band of the queue.

And if you want to be sure it really works, try it somewhere else

Same five letters, a veterinary hospital instead of a bank, and a completely different flip, one that fires when the model improves in a different way: it lets a senior hand work down instead of making her stop looking herself.

Falken Veterinary Hospital runs TrustBeacon, a tool that reads chest and limb radiographs and flags each one "likely benign" or "needs review" before a vet signs off. Dr. Costanza Ferren has run the imaging room there for six years, the only vet on staff who reads films herself rather than sending them out.

F, find the person. Dr. Costanza Ferren, six years reading films at Falken, and Alden Marrowbrook, her vet tech of three years, who confirms reads on the floor once she's checked them.
L, locate the habit. She used to review every TrustBeacon-flagged film herself before Alden confirmed a read to an owner. It cost her almost nothing most days, and it caught two real misses in the year before the model improved.
I, identify the flip. This is the delegation flip, not the over-trust flip. Once TrustBeacon agreed with her own reads on 9 out of 10 films for three straight months, she let Alden confirm reads alone, no second look. A hairline fracture came through labeled "likely benign," Alden confirmed it, the dog went home, and it needed a second surgery six weeks later. She now reads every film herself again, and Alden hasn't confirmed a read alone since.
P, pinpoint the old decision. TrustBeacon's screen shows one flat label under the film, "likely benign" or "needs review," nothing about which part of the bone the model weighed most. That felt fine when Alden only ever confirmed what Costanza had already checked. Once he was reading alone, there was nothing on the screen that let either of them see how the call got made.
S, show the replay. Same fracture, same film, run through a screen that shows the region TrustBeacon weighed most under the label. Alden sees the flagged region sits half a centimeter from where the model's confidence was actually built, pulls two more angles, and catches it himself. Costanza reviews 1 film in 10 instead of all of them, and gets her afternoons back.

Share of films Dr. Ferren personally reviewed, before and after delegating to Alden
100% 0% delegated fully the miss Month 1 Month 4 Month 7 Month 10
Costanza's personal review rate
She reviewed nearly every film for the first four months. Once TrustBeacon's agreement rate held for a quarter, she dropped to reviewing about 1 in 10, and stayed there right through the month of the missed fracture.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the flag, and say why it stays lit at every score, not just low ones.
Cost: the fraud team says building a live novelty flag is six weeks out. Don't wait on it, pull "new payee, new country" as a manual daily report until it ships.
The model got better, for real: say precision climbs to 99 next year. That's still not the same claim as "every case is fine now." A model that gets more accurate on average can still be badly wrong on a pattern it almost never saw, and a flat score hides that difference either way.

Where people run it wrong.
They add the flag, then keep the panel collapsed by default anyway, because collapsing it was the habit they'd built for themselves too.
They see the blended precision number climbing and call that proof the whole queue is safer, instead of asking which slice of the queue nobody is looking at anymore.
They respond to one bad freeze by adding review back everywhere, instead of asking whether the real fix is a guardrail on the rare pattern, not a tax on the common one.

How to use it live. Open with the flag, not the story: "I'd keep one signal visible no matter the score, and here's the exact pattern it protects." That buys you the room to walk the interviewer through the failure on your own terms, instead of them walking you into it.

Flashcards (click a card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust. Checks sometimes, then stops checking at all, once the model gets good enough that catching it feels obviously right.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Radomir Ashcroft, a fraud analyst who has worked Tarrow Bank's ClearFlag queue for eight years.
3 · THE HABIT
What did he stop doing because it worked?
Tap to flip
ANSWER
Opening the full case panel, the last five transactions, the merchant compare, and the customer call, before deciding a flagged case.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Reads the case before deciding, or trusts the score alone above a cutoff. Nothing in between.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Setting the case panel to stay collapsed by default once the queue got busy, instead of forcing it open for the rare pattern no matter the score.
6 · THE NUMBER
Fill in the blank: ClearFlag's precision climbed from 41 out of 100 to ___ out of 100 after the retrain.
Tap to flip
ANSWER
91. And the share of score 95-and-up cases where the panel actually got opened fell from about 98 percent to about 4 percent over the same ten months.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
A "first time: this payee, this country" badge stays lit at score 97. Radomir sees it, makes a 90 second call, the wire clears the same afternoon, no freeze, no four day mess.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Falken Veterinary Hospital's TrustBeacon, a radiograph triage tool. The delegation flip: Dr. Costanza Ferren hands reads to her vet tech once the model improves, then takes every read back after one miss.

Check yourself Score: 0 / 0

Multiple choice
1. What is the exact switch that flips in Radomir's story?
  • A. He starts calling customers more often than before.
  • B. He goes from opening every flagged case's history to trusting the score alone above a cutoff.
  • C. He asks Tarrow Bank to lower his workload.
  • D. He starts manually flagging more transactions than ClearFlag does.
Show hint
Look for the exact habit named in the L step, and what it turns into.
Show answer
B. He stops opening the case panel once the score crosses a cutoff that feels obviously right, and starts treating the number alone as the decision.
Fill in the blank
2. ClearFlag's precision climbed from 41 out of 100 to ___ out of 100 after the retrain.
Show hint
Check the bar chart right after the first section's opening numbers.
Show answer
91. A real, large improvement, which is exactly what makes the habit that broke afterward worth taking seriously.
Multiple choice
3. What old decision does this answer take back?
  • A. Setting the case panel to stay collapsed by default once the queue got busy, instead of forcing it open for rare patterns.
  • B. Hiring fewer fraud analysts than the queue actually needs.
  • C. Lowering the score cutoff so more transactions get flagged.
  • D. Turning ClearFlag off during the busiest hours of the day.
Show hint
This is the P step, look at the key point box in the "Let's learn" section.
Show answer
A. A reversible screen default, made for a good reason at the time, that quietly stopped making sense once precision climbed.
True or false
4. True or false: the fix this answer proposes also changes how transactions scored under 60 get handled.
  • True
  • False
Show hint
Look at the "what I would leave alone" paragraph.
Show answer
False. Cases scored under 60 already force a full look and a two minute timer. That part of the design was never the problem, so the fix leaves it untouched.
Short answer, flip versus dial
5. Why couldn't Radomir have just "checked a little more carefully" instead of stopping the check entirely?
Show hint
Ask what a person actually does with their hands when a habit is a switch instead of a dial.
Show answer
Model answer: "Checking more carefully" is a dial, and dials need a reason to move a little at a time. Trust here only had two real settings in his head: this thing needs a look, or this thing doesn't. Once the score kept being right, there was no signal telling him to look "a bit more," only silence, which read as "you don't need to look at all."
Short answer, apply it yourself
6. Pick an AI product you use yourself. What's one habit it built in you that you'd expect to break if the product suddenly got noticeably more accurate?
Show hint
Ask what you currently double check "just in case," and what happens to that habit if the product stopped ever being wrong there.
Show answer
Model answer: A GPS app that reroutes around traffic. Right now I glance at the map before trusting a sudden reroute, because it's occasionally wrong about a closed road. If it got accurate enough that I stopped glancing, the one time it rerouted me into a closed road with a hard deadline, I'd have no warning at all, because I'd have stopped looking at exactly the moment looking mattered most.
Before you close the answer
Why this works
Tests whether you think trust is the model's job or the design's job. Most candidates explain why the model got better and stop there, they never say what the person in front of it does differently because of it.
Follow-up traps
"Isn't the real fix just to lower the score cutoff so more cases get a second look?" Response: that taxes the whole queue again, most of the extra cases it would catch were already fine at 91 percent precision. The novelty flag only slows down the narrow slice that's actually rare, not everything above a fixed number.

"What if the customer had made international wires before, wouldn't the flag never fire?" Response: then it shouldn't fire, that's the point. The flag tracks whether this exact pattern is new for this exact customer, not whether wires in general are risky, so it stays quiet on a payee ClearFlag has already seen many times.
If pressed
The actual rule Tarrow uses: precision has to clear 90 out of 100 on the fraud team's own monthly labeled eval set, about two thousand reviewed cases, for three refreshes running, before review time gets shortened for any score band. One good month never moves a default on its own.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more