CaseAdvancedQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #4

How would you measure trust in an AI feature?

Trust does not show up in a satisfaction score. It shows up in what a person quietly stops double checking, weeks before they would ever say the word "trust" out loud.

The direct answer
Watch how often the front desk still pulls up the raw intake form to check the AI summary against it, on the visits they used to check every time. Pair that number with whether they are still opening the summary at all: a falling check rate next to steady or rising use is real trust, a falling check rate next to falling use is a team that gave up, not one that trusts anything. Set a hard floor on the check rate for the fields that matter most, allergies and red flag symptoms, and force a human read there no matter how high trust runs everywhere else.
Do this, in order
  1. Watch the double check rate, how often staff open the raw form to verify the AI summary, as the main trust signal.Why: it is the earliest thing to move, weeks before a survey or a complaint would show anything at all.
  2. Pair it with the summary's own opt in rate, not the check rate alone.Why: a falling check rate by itself cannot tell you if people trust the tool or if they quietly stopped using it altogether.
  3. Watch the correction rate too, the rate staff flag a summary as wrong, small but never zero.Why: it proves people are still looking closely enough to catch a real miss, not just clicking past every screen.
  4. Set a hard floor on the check rate for allergy and red flag fields, and audit automatically when it is crossed.Why: that is the one field type where a missed error is too expensive to leave sitting on a trend line.
  5. Never let a quarterly self report survey stand in for any of this.Why: it is the number that looks best in the exact week something is about to go wrong, because it measures how people feel, not what they did that morning.
  6. When the correction rate falls together with the check rate, call it fatigue and reset the default, do not just ask people to look closer.Why: a tired team told to try harder stays tired, a reset default actually changes what they see first.

How to answer this, stage by stage

Nobody is grading whether you can name a trust metric that sounds sensible. They are grading whether you know the difference between a number that moves first and a number that only agrees with you after the fact.

1
Scope it to one real product and person before you answer in the abstract
Say it like this
"Let's make this real. Foothill Family Medicine is a primary care clinic. Anteroom reads what a patient typed on their intake form before a visit and turns it into a short summary for the front desk and the doctor, administrative only, it never diagnoses anything. Sylvie Wrenley runs the front desk there."
Why this works
A trust metric named in the abstract is a guess. One tied to one real desk is a decision you can defend.
2
Say what the question is actually testing
Say it like this
"This isn't really asking me to name a metric. It's asking whether I know that trust doesn't show up on a survey until it's already too late to matter. I need the thing that moves first, not the thing that agrees with me a month later."
Why this works
Naming the real question up front stops you from reaching for the obvious, wrong answer: a satisfaction score.
3
Name your structure out loud
Say it like this
"I'd run this through LEAD. Link it to the real outcome trust is standing in for, find the early signal that moves before that outcome does, name how that signal gets gamed, then say what I'd actually do at each level it hits."
Why this works
Two seconds of structure shows the interviewer you have a method for finding a leading number, not just an opinion.
4
Reject the tempting wrong metric
Say it like this
"The easy answer is a quarterly trust survey, ask the front desk to rate Anteroom one to ten. I'd reject that. At Foothill that score hit its highest point, 8.7, the exact week a missed allergy line almost reached a doctor. Self report measures how people feel about a tool, not what they actually did with it that morning."
Why this works
Naming a plausible metric and killing it, with a real reason, is what separates a real answer from a list of ideas.
5
Name the real signal, and how it gets gamed
Say it like this
"The real signal is the double check rate, how often the front desk still pulls up the raw form to check the summary, on visits they used to check every time. But that number alone can lie to you. I pair it with whether they're still opening the summary at all. Falling checks plus steady use, that's trust. Falling checks plus falling use, that's a team that's checked out, not one that trusts anything."
Why this works
This is the answer to the question. Everything else is why this pair, and not some other single number, is the right one.
6
Set the thresholds and what changes at each one
Say it like this
"Above 50 percent, nothing to do, just watch the two lines together. Between 15 and 50 percent, with the correction rate still nonzero, that's healthy trust building, I'd add a light weekly sample audit as routine hygiene, nothing dramatic. Below 15 percent on allergy or new patient visits for two weeks straight, that's a hard trigger, a stratified audit runs on its own, no more waiting for someone to get lucky."
Why this works
A metric nobody acts on is decoration. Naming the exact number that changes your behavior is what makes it a real dashboard.
7
Close on the one line you'd never trade away
Say it like this
"One thing stays fixed no matter how high the trust trend runs. Allergy fields and red flag symptoms never get folded into the summary's prose. They're pulled straight from the patient's own words and shown on their own line, for every patient, every time, forever."
Why this works
Closing on a guardrail that does not move with the trust score is what proves you understand the difference between a metric and a safety rule.
If you remember one thing Trust is not a number a person tells you. It is a habit they stop doing. Watch the habit fall, and watch it against a second number that proves the fall is real trust and not a tired team giving up.

Let's learn

For nine years, before anyone at Foothill Family Medicine had heard of Anteroom, Sylvie Wrenley got to work at 6:45 in the morning, an hour and fifteen minutes before the doors opened, to read intake forms by hand. About thirty patients a day, five minutes a form, checking allergies against the reason for the visit, checking insurance cards against what the system had on file. It took her past 8am most mornings, and she did it anyway, because a form read wrong is a patient handled wrong.

Anteroom is a small AI tool. It reads what a patient typed into their online intake form before a visit and turns it into a short summary, four or five lines, for the front desk and the doctor to glance at. It does not diagnose anything and it does not suggest treatment. It just organizes what the patient already wrote, so a person can prep the visit faster.

Knowledge spark: what does "administrative only" actually mean here? Anteroom never guesses what is wrong with a patient. It only reorganizes what they already told the clinic themselves, the reason for the visit, current medications, listed allergies. The judgment about what to do with that information still belongs to a person. That line matters, because it is also exactly where the risk hides: a tool that only reorganizes text can still drop or misfile one line of it.

Before Anteroom, a full morning of forms cost Sylvie about two and a half hours. After, the same information became something she could read in about fifteen seconds a patient, and a morning that used to eat into her whole day dropped to about ten minutes total.

Front desk double check rate, new patient visits, week 1 to week 20
100% 0% near miss caught Wk 1 Wk 4 Wk 8 Wk 12 Wk 16 Wk 20
Ninety two percent of new patient summaries got checked against the raw form in week one. By week twenty, six percent did. Nothing at Foothill measured this line while it fell.
The extra mistakes were never the real risk. The real risk was whether Sylvie kept checking, and whether that was trust talking or tiredness.
Front desk trust survey, self reported out of 10, Q1 vs Q2
10 0 6.1 8.7 Q1, week 12 Q2, week 18
Earlier quarterRight before the near miss
The survey climbed from 6.1 to 8.7, its best score ever, taken two weeks before the missed allergy line. A metric that only asks how people feel had no way to see the check rate collapsing underneath it.
The decision that mattered Foothill's rollout never built a way to see whether staff were still opening the raw form. It felt like overkill at launch, extra clicks to track a habit everyone assumed would just hold. By month five, nobody, not even Sylvie, could see that habit was nearly gone.

What I would leave alone: the part of Anteroom's summary that suggests which exam room to use. Getting that wrong sends someone to the wrong door for thirty seconds and they turn around. It is not worth the same tracking discipline as a line about what a patient is allergic to.

The lesson: a habit built out of caution does not stay just because it used to matter. If nobody can see it, nobody can tell the difference between a team that trusts a tool and a team that has simply stopped looking.

Now here is the same thing as a story

The short version sits above. Read on for the Tuesday a jammed printer did the job a dashboard should have done six weeks earlier.

Every morning, Sylvie unlocked the front door at Foothill Family Medicine, put the coffee on, and pulled up the day's patient list before anyone else on the team arrived. Nine years at that desk had taught her to read a printed intake form the way a mechanic reads an engine, a symptom that didn't match the stated reason for the visit, a form filled out by a worried parent instead of the patient, a missing insurance card. She caught all of it, by hand, every single morning.

Anteroom arrived in the spring. For the first two months, Sylvie treated it the way she treated a new hire, useful, but not yet proven. She still opened the raw form for every patient, checked it against the summary, and every single time the two matched. By month three she was checking most of them, not all. By month four, Foothill had two more part time front desk hires, both of whom had never known the all manual mornings, and neither of them had ever built the habit of checking at all. Their numbers pulled the whole team's average down fast. And the clinic's own dashboard, built to show how many summaries Anteroom generated each day, never once tracked whether anyone opened the source form to check one. Nobody could see the number falling, because nothing showed it to them.

The good months were real, and they were good. Sylvie picked her daughter up from swim practice at 5:15 instead of 6:00, most days, for the first time in years.

The trigger wasn't a complaint. It was a jammed printer.

On a Tuesday, a print station backup copy jammed, and Sylvie pulled the physical intake sheet for an unrelated patient to clear it. It happened to be a new patient's form. She glanced at it out of habit, the old habit, and the allergy line read "Penicillin, hives." Anteroom's summary for that same patient, sitting open on her screen, said "No known drug allergies." That patient's summary was one of the ninety four percent she had not double checked that morning, on a visit type she used to check without fail.

We didn't almost make a mistake with penicillin. We almost let a habit, quietly worn down, decide who reads an allergy list.

She flagged it before the patient was roomed. Nobody was harmed. But once the team pulled the logs to see how this had happened, they found what nobody had been watching: the check rate on new patient visits had fallen from 92 percent in week one to 6 percent by week twenty, the week of the jam, and not one number on any Foothill dashboard had shown it.

It was never really about Anteroom getting worse. Its accuracy on that field hadn't changed the whole stretch, the model had simply misclassified one line on one form. What changed was how many mornings in a row nothing had gone wrong, and how that quietly rewrote what checking even felt like it was for.

The old decision went back to the rollout meeting, months earlier. Someone said building a log of who opened the raw form, and when, felt like overkill, extra clicks tracked for a habit that had held for nine years and would obviously keep holding. Nobody argued. Nobody built it.

Run the same twenty weeks again with the floor rule the team built afterward: below 15 percent on new patient visits for two straight weeks triggers an automatic audit. That floor would have tripped in week fourteen, six weeks before the jammed printer did the job instead. A stratified pull of twenty summaries against their source forms, an afternoon's work, would have caught the exact same gap on paper, without needing luck to do it.

One design trusted the habit to hold itself up. The other watched the habit, and caught it slipping before someone had to get lucky.

What I would tell myself, back in that rollout meeting: a habit you are counting on to hold does not stay because it used to matter. Write down a way to see it, or you will not know it is gone until something almost proves it for you.

LEAD, run start to finish

Not a diagnosis of one Tuesday. LEAD run on the actual question, which signal to build the whole trust dashboard around, using Foothill's near miss as the proof.

L
Link. The business outcome trust is actually standing in for.
Not "people like Anteroom." Durable, voluntary reliance on it, staff and providers still building the visit around the summary months in, because it holds up, not because a manager told them to use it.
Get this wrong and you end up measuring whether the login screen gets clicked, which happens whether anyone trusts what's behind it or not.
E
Early signal. The thing that moves before the outcome does.
The double check rate: how often the front desk still opens the raw form on visits they used to check every time. At Foothill it fell from 92 to 6 percent across twenty weeks, while the number everyone actually looked at, the quarterly survey, climbed.
This is the hard step, and the whole reason LEAD exists. A survey confirms trust after the fact. This finds it while it's still forming, or still eroding.
A
Abuse. How the signal gets gamed or misread.
A falling check rate looks identical whether it means real trust or a team that's simply given up on checking anything. The fix is pairing it with a second number: whether staff are still voluntarily opening the summary at all, and whether the correction rate, flagging a summary as wrong, stays above zero.
Skip the pairing and you'll read a burned out team as a trusting one, right up until an error nobody was looking for gets through.
D
Decision. What actually changes at each threshold.
Above 50 percent, watch only. 15 to 50 percent with a live correction rate, add a light weekly sample audit. Below 15 percent on allergy or new patient visits for two straight weeks, an automatic stratified audit runs, no manager has to notice first.
A threshold with no action behind it is a chart on a wall. This is the part that actually would have caught Foothill's gap in week fourteen.
Hand sketched comparison. Left panel, a gauge reading trust survey 8.7 out of 10, self reported once a quarter, stayed calm the whole time. Right panel, a document labeled raw intake form, penicillin line, not opened that morning.
The two numbers sitting side by side. One looked healthy the whole time. The other was the one nobody had opened.

Two things worth naming plainly, since this is where the real judgment sits. The rejected alternative that mattered most was not just the trust survey, it was also Anteroom's own extraction confidence score, the number the model reports about how sure it is on each field. That number stayed unremarkable through the whole stretch, because it measures the model's certainty about itself, not whether a human being still checks its work. A model can hedge correctly and still be trusted blindly, or state something plainly wrong with total confidence, and neither shows up in its own score. The AI specific failure worth naming is confident wrongness on a structured field: Anteroom didn't hedge or flag the allergy line as uncertain, it stated "no known drug allergies" as flatly as if the patient had actually written that. The guardrail: allergy fields and red flag symptoms are never compressed into the model's summary prose at all, they're pulled verbatim from the patient's own text and shown on their own line, and any extraction the model itself scores as low confidence on those fields routes to a mandatory human read regardless of what the trust trend is doing elsewhere. There's a real cost to that guardrail, and it's worth saying out loud: showing those fields verbatim makes the summary a few lines longer, and the weekly sample audit costs the front desk real minutes every week they didn't spend before. Foothill accepted both, a slightly longer read and a standing labor cost, against the alternative of a habit nobody could see failing.

And if you want to be sure it really works, try it somewhere else

Same four letters, an insurance claims desk instead of a clinic, and the same shape shows up with nothing medical anywhere near it.

Larch Mutual runs an AI tool called Binder that reads an incoming auto claim, the forms, the photos, the police report if there is one, and writes a short summary an adjuster reads before opening the full file. Odaline Cordray has adjusted claims there for eleven years.

L, link. Not claims processed per day, that number climbs whether adjusters trust Binder or are just clicking through it. The real outcome is adjusters still building their first read of a claim around Binder's summary months in, voluntarily, because it holds up.
E, early signal. How often an adjuster still opens the full claim file to check Binder's summary against it, on claims above a policy's liability limit, the ones they used to check without exception. At Larch that rate fell from 88 to 9 percent over ten weeks while nobody was tracking it.
A, abuse. A backlog spike hit the same quarter, and adjusters started closing routine claims faster without opening the file at all, not because they trusted Binder more, but because they had less time per claim. The check rate alone couldn't tell the difference. Paired against claim volume per adjuster, the drop lined up exactly with the backlog, not with growing trust.
D, decision. Below a 20 percent check rate on above limit claims, for two straight weeks, a supervisor now pulls a random sample by hand. That rule caught two claims that had cleared with a policy limit Binder had summarized wrong, before either one reached a payout decision.

Hand sketched timeline. Backlog looks fine, adjuster survey 7.9 out of 10. Auto close rate climbs, week 6, nobody watching it. Two claims slip past review, week 9, caught by a spot check, marked as the turning point. Policy limits shown verbatim, week 10, never summarized again.
Same shape as Foothill's. A survey that looked fine the whole way through, a number quietly climbing underneath it, and a guardrail added only after a spot check did what a threshold should have done first.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pair: watch the check rate and the opt in rate together, never one alone, and say why in one line.
Cost: engineering says a per adjuster verification log is two quarters out. Don't wait on it, pull the check rate from claim file access logs by hand, once a week, until it ships.
The model got better, for real: say Binder's own accuracy genuinely improved that quarter. That still isn't proof adjusters are checking less because they trust it more. A backlog, a new hire, or plain fatigue can produce the exact same falling number for a worse reason.

Where people run it wrong.
They watch the check rate alone and call a falling number good news, without ever asking whether people are still opening the tool itself.
They let a rising satisfaction survey talk them out of an audit, because the two seem to agree, right up until the week they stop agreeing.
They respond to a bad number by telling the team to "look closer," instead of asking whether the real fix is a floor and a guardrail, not a lecture.

How to use it live. Open with the pair, not the single number: "I'd watch two things together, not one, how often people still check it and how often they still use it, and here's why one alone would have missed what actually happened." That buys you the room to defend the pairing before the interviewer pushes you toward a single tidy metric.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a "how would you measure trust" question like this one?
Tap to flip
ANSWER
LEAD: link the outcome trust really stands in for, find the early signal that moves first, name how it gets gamed, decide what you'd actually do at each threshold.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Sylvie Wrenley, who runs the front desk at Foothill Family Medicine, using Anteroom, an AI tool that summarizes patient intake forms for administrative use only.
3 · THE HABIT THAT FADED
What habit had Sylvie's team quietly dropped before the near miss?
Tap to flip
ANSWER
Opening the raw intake form to double check Anteroom's summary on new patient visits, a habit that fell from 92 to 6 percent over twenty weeks with nothing tracking it.
4 · THE PAIRING
Why isn't the falling double check rate enough on its own to prove trust?
Tap to flip
ANSWER
Because it looks identical whether people trust the tool or have simply stopped engaging with it. Pairing it with the summary's own opt in rate and the correction rate tells the two apart.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Never building a log of whether staff opened the raw form. It felt like overkill at launch, since checking had held for nine years. Nobody revisited it as new hires joined who never built that habit at all.
6 · THE NUMBER
Fill in the blank: the trust survey hit its highest score, ___ out of 10, the same week the check rate had already fallen to ___ percent.
Tap to flip
ANSWER
8.7; 6 percent. The gap between those two numbers is the whole reason a self reported survey can't be trusted as an early warning.
7 · THE REPLAY
Same twenty weeks, with the floor rule built afterward, what changes?
Tap to flip
ANSWER
The 15 percent floor on new patient visits trips in week fourteen, six weeks before the jammed printer forced the discovery. An afternoon audit catches the same gap on paper instead of by luck.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
Binder, an AI claims summary tool at Larch Mutual, used by adjuster Odaline Cordray. Same LEAD steps, a claims backlog instead of new hires, found the same trap: a rising survey and a falling check rate that meant fatigue, not trust.

Check yourself Score: 0 / 0

Multiple choice
1. What is the strongest reason to reject a quarterly self reported trust survey as the main way to measure trust in an AI feature?
  • A. Surveys are too expensive to run every quarter.
  • B. Staff usually give dishonest answers on purpose.
  • C. It measures how people feel about the tool, not what they actually did, so it can look healthy right up until something breaks.
  • D. Surveys can only be given to managers, not front line staff.
Show hint
Look at what the survey score was doing the same week the double check rate had already collapsed at Foothill.
Show answer
C. The survey hit its best ever score, 8.7, the same week the real check rate had fallen to 6 percent. A number that measures feeling, not behavior, has no way to see a habit eroding underneath it.
True or false
2. True or false: a falling double check rate is, by itself, reliable proof that staff trust an AI summary more than they used to.
  • True
  • False
Show hint
Ask what else, besides trust, could make a person stop checking something they used to check every time.
Show answer
False. A falling check rate looks the same whether it comes from real trust or from fatigue and resignation. It only means what you think it means once it's paired with steady or rising voluntary use of the tool itself.
Fill in the blank
3. Foothill's floor rule triggers an automatic audit when the double check rate on new patient visits stays below ___ percent for two straight weeks.
Show hint
Check stage 6 of the walkthrough, "Set the thresholds and what changes at each one."
Show answer
15 percent. That's the point where the answer stops watching and starts acting, rather than waiting for a person to notice by luck.
Multiple choice
4. Why was Anteroom's own extraction confidence score rejected as the main trust metric, even though it's a real number the model produces?
  • A. Confidence scores are too technical for front desk staff to understand.
  • B. Anteroom doesn't actually calculate a confidence score for any field.
  • C. It measures the model's certainty about its own output, not whether a person still checks that output, so it can't tell trust from resignation either.
  • D. The confidence score only updates once a week, too slowly to be useful.
Show hint
This question is about whose behavior the number describes, the model's or the person's.
Show answer
C. A model can state something wrong with total confidence, or hedge correctly and still get trusted blindly. Either way, its own confidence score says nothing about what the human being on the other end actually did.
Short answer, name the guardrail
5. What guardrail did Foothill put in place for allergy fields and red flag symptoms, and why does it hold regardless of the trust trend?
Show hint
Look at the paragraph right after the LEAD rows, where the AI specific failure mode is named.
Show answer
Model answer: Those fields are never compressed into the model's summary prose. They're pulled verbatim from the patient's own words and shown on their own line, and any low confidence extraction on those fields routes to a mandatory human read. It holds regardless of trust because that's the one field type where a missed error is too expensive to leave to a trend line, however healthy that trend looks.
Short answer, apply it yourself
6. Pick an AI feature you use yourself. Name the manual check or fallback you used to do every time, and say what it would mean if you noticed yourself skipping it now, both the good version of that story and the bad one.
Show hint
Think of something you used to double check by hand and now mostly don't, then ask whether that's confidence or just habit fatigue.
Show answer
Model answer: A GPS app that reroutes around traffic. I used to glance at the suggested route on a map before following it blindly, now I mostly don't. Good version: it's been right often enough that checking stopped adding anything. Bad version: I've just gotten used to trusting it and would not necessarily notice if it started sending me somewhere worse, since I stopped comparing it to anything.
Before you close the answer
Why this works
Tests whether you reach for a self reported number because it's easy to collect, or whether you know it lags and can actively mislead. Most candidates answer "trust" with a survey and stop there.
Follow-up traps
"Couldn't a falling check rate just mean the model got better, so there's less to catch?" Response: possible, but that's exactly why it's never read alone. Pair it with the correction rate: if people are still finding and flagging real errors at the same low, steady rate, the model didn't get safer, people just stopped looking as hard.

"Isn't building a verification log just adding more tracking overhead to a tool that was supposed to save time?" Response: it's a few lines in an access log, not a new task for staff. The overhead this answer actually accepts is the audit sampling and the verbatim fields, both named, both worth it against a gap nobody could see for five months.
If pressed
The actual gating rule used at Foothill after the fix: the audit only counts a field as "genuinely missed" if it would have changed what the provider did next, not any wording mismatch between the summary and the source. That kept the audit focused on real risk instead of nitpicking phrasing, which would have buried the team in false alarms within a month.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more