ConceptAdvancedQuality, Cost & Token Economics / Leading vs lagging indicators for AI / #7

Describe the relationship between refusal rate and downstream satisfaction.

A refusal rate that holds steady can still be quietly trading one kind of customer for another. You only see it once you split the coat open.

The direct answer
Refusal rate alone tells you nothing, rising or falling. Split every refusal by its reason code and watch two different numbers: how often the model turns away a customer with thin but real evidence, and how often it approves a request carrying an actual fraud signal. Protect the first number harder by default, because a wrongly refused honest customer says nothing and just stops buying, while a wrongly approved risky one is rare enough to see coming and count.
Do this, in order
  1. Split refusal rate by reason code, then protect thin-evidence, borderline refusals harder than fraud-flagged ones, by default.Why: one blended number cannot tell "we are stopping fraud" apart from "we are driving honest customers away," and the honest customer's cost turns out to be far bigger once you can see it.
  2. Track repeat purchase rate for refused customers, split by reason code, not only for approved ones.Why: satisfaction numbers built only from approved refunds never see the customers who quietly left after a wrong refusal.
  3. Set a kill criterion that flips the default, a fraud-loss threshold, not a feeling.Why: past a fixed dollar line, or the same pattern hitting three times, the hidden cost of over-approving gets bigger than the hidden cost of over-refusing.
  4. Gate any change to the evidence bar behind a golden set of hand-checked borderline cases.Why: a model update that looks like it is catching more fraud can really be quietly refusing more honest customers.
  5. Route thin-evidence refusals to a person before they become a final no.Why: a short human check catches the case a single confidence score would wrongly close.

How to answer this, stage by stage

Nobody is grading whether you can define a false positive. They are grading whether you will commit to a side and then prove, in dollars, that the asymmetry is real. Seven moves get you there.

1
Scope it to one real product, one real number
Say it like this
"Let's ground this. Corymbia sells outdoor gear and home goods online, and Keeper is the model that decides, in under twenty seconds, whether a return request gets a refund or a no. Camryn Falkenrath runs customer ops there and owns Keeper's refusal rate."
Why this works
Grounds the tradeoff in a real product before any number gets debated in the abstract.
2
Say your structure out loud
Say it like this
"Here's how I'd take this apart. I'm going to commit to a position first, then say who feels each kind of mistake, then find which one actually costs more, then say what would make me flip."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion.
3
State the position, no hedging
Say it like this
"A rising refusal rate is not automatically bad, and a falling one is not automatically good. The number that matters is refusal rate split by reason code. If I have to pick one side to protect by default, I protect the honest customer with thin evidence over the customer with a real fraud flag."
Why this works
This is the answer to the question. Everything after it is proof, not more opinion.
4
Name who feels each kind of error
Say it like this
"A legitimate customer wrongly refused feels it in the next thirty seconds, gets one templated denial, and does not file a complaint. She just does not buy from Corymbia again. A risky request wrongly approved costs nothing that day. It shows up weeks later as a chargeback on a report someone actually reads."
Why this works
Naming who feels what, and when, turns a debate about a metric into a debate about a person.
5
Prove the asymmetry with a real number
Say it like this
"At Corymbia's volume, the thin-evidence refusals we get wrong cost about two hundred and five thousand dollars a month in customers who quietly never come back. The fraud we let through by mistake costs about eleven thousand a month. That is not close. Reading refusal rate as one number hides an eighteen-to-one gap."
Why this works
A dollar comparison, not an adjective, is what makes "the hidden cost is bigger" a fact instead of a feeling.
6
Give the kill criteria, a number that flips you
Say it like this
"I would flip this the day fraud loss crosses fifty thousand dollars in a single month, or the day the same pattern hits us three times running, because that is usually one ring working a blind spot, not noise."
Why this works
A number that would change your mind is what makes this a decision instead of a stance you will defend forever.
7
Close on the option you rejected and the price of your own pick
Say it like this
"We looked at just loosening the confidence bar across the board so total refusals drop. We ruled that out, because it loosens the fraud-flagged bucket too, and that is the bucket we actually want tight. What I would trade instead is speed on the thin-evidence bucket: routing it to a person costs hours instead of twenty seconds, and that is a real cost, but it is smaller than two hundred thousand dollars of customers walking away quietly."
Why this works
Naming the alternative you ruled out and the price of the one you picked is what makes this a judgment call, not a guess that happened to land.
If you remember one thing Refusal rate is not one number pretending to be simple. It is two very different numbers sharing one coat. Split the coat before you decide which side to protect.

Let's learn

The decision happens inside a small box on a return form: refund approved, or refund denied, decided before the customer even finishes reloading the page.

Corymbia sells outdoor gear and home goods online, about 40,000 return or exchange requests come through in a normal month. Keeper is the model wired into that box. It reads the reason a customer gives, checks it against photos and order history, and decides, refund or no refund, in under twenty seconds.

Before Keeper, a team of return agents made every call by hand. Average time per case: six minutes. Across forty thousand cases a month that is close to four thousand hours of work, and it showed in the outcomes: two agents looking at the same thin evidence could land on opposite answers, one refusing four percent of the time, another refusing fifteen.

Keeper decides every case the same way, every time. Refusal rate settled at nine percent overall, which read, on a dashboard, as a healthy, boring number.

Knowledge spark: what is a reason code? A short label attached to every refusal, saying why. "No photo, but the story checks out" is a different reason than "same address as three flagged accounts." One is thin evidence. The other is a fraud signal. A raw refusal rate blends both into one line.

Here is the turn. Nine percent hid two very different nine percents. Split by reason code, refusals on requests with an actual fraud signal, repeat abuse, a mismatched address, no order match, held almost exactly where the old team left them. Refusals on requests with thin but real evidence, a blurry photo, a late but honest report, climbed from eleven percent to twenty four percent over eight weeks. Keeper was not catching more fraud. It was getting fussier with the customers who had a real case and just a messy one.

We did not lose to fraud. We lost quietly, one honest customer at a time.
Hand sketched comparison titled Who feels the mistake. Left a small person icon labeled Honest customer wrongly refused, feels it now, says nothing, just stops buying. Right a gauge icon labeled Risky request wrongly approved, rare but shows up as a real loss on the books.
Two mistakes, two very different weights. One is quiet and everyday. The other is rare and shows up on a report.
Cost, by the numbers, at Corymbia's normal volume
$205k $0 $205,128 $11,400 Quiet churn, thin-evidence Fraud loss, wrongly approved
Monthly cost of getting each kind of decision wrong. Roughly an eighteen-to-one gap between the quiet side and the rare, expensive side.

At its worst, this cost more than a bad quarterly number. Over three months, before anyone caught it, that quiet churn added up to about $615,000 in customers who simply stopped shopping at Corymbia, most of whom had done nothing wrong.

The decision that mattered One confidence threshold, shared across every reason code. Fraud-flagged cases and thin-evidence cases both had to clear the same bar to get approved.

The choice I would take back is that shared threshold. When Keeper was built, engineering had one quarter and not nearly enough labeled data to calibrate a separate bar for every reason code, so one bar covered all of them. That was a fair call with the data on hand. It stopped being fair once volume grew and the two reason codes started needing very different amounts of trust.

What I would leave alone: refusals on requests with zero evidence at all, no order number, no photo, no account match. Those should stay high. There is no ambiguity there worth splitting further.

The lesson: a refusal rate that holds steady is not proof that nothing changed underneath it. Two numbers can cancel each other out on a dashboard while doing very different things to real customers.

Now here is the same thing as a story

The short version is above. Read on for the Thursday a slide in someone else's deck showed Camryn what her own number had been hiding.

Camryn Falkenrath has run customer ops at Corymbia for five years. Before Keeper, she could read a return request and tell inside ten seconds whether the story held together, a knack she learned the hard way, once refusing a real customer who had worn a rain jacket on the exact hike she said ruined it.

When Keeper launched, the first months were good. Every Monday she pulled a report split by reason code and hand-checked ten cases against it, five thin-evidence, five fraud-flagged. It always matched. The topline number sat near nine percent, steady, boring, exactly what a healthy number should look like.

So she trimmed the Monday check. Ten cases became three. A few months later, three became a glance at the topline chart and nothing else. The number was still nine percent. Why keep opening the file.

Nothing broke on a Tuesday. There was no single morning where the number jumped. It built up slowly, the way a crack under wallpaper does, until a Thursday quarterly review, called for an unrelated reason, put an eight-week reason-code chart from finance up on the screen. Camryn had not built that chart. She had not asked for it. Borderline refusals: eleven percent, climbing to twenty four. Fraud-flagged refusals: flat the whole time.

She pulled forty of the month's thin-evidence denials that night and read every one herself. Nineteen of the forty, on a second look, should have gone through.

Her topline number had been telling the truth about fraud and lying about everyone else, at the exact same time.

The real cost was not the meeting. Corymbia's finance team, working the same quarter's numbers separately, had already flagged a repeat-purchase dip among refused customers and filed it under normal seasonal noise. Nobody had connected the two charts until Camryn walked hers over.

A year earlier, when engineering built the threshold, the room had good reasons. One bar was simpler to ship, and there was not enough labeled data yet to trust two. Nobody in that meeting expected refusal volume to triple in a year, and nobody put a date on when the shortcut should get revisited.

Run the same quarter through the fixed design. Borderline refusals settle back near twelve percent instead of twenty four. Projected quiet churn drops from roughly two hundred thousand dollars a month to about forty thousand. Fraud loss ticks up slightly, from eleven thousand to about fourteen thousand, because the thin-evidence bar loosened a little on its own side of the split. Net, Corymbia keeps most of the forty thousand honest customers it was about to lose and pays fourteen thousand more a month to a bucket that never touched the fraud-flagged rules at all.

One design watched one number. The other watched two, and only one of them ever needed protecting harder.

What I would tell my past self: I asked whether the refusal rate was healthy. I never asked whose refusal rate it was.

PICK, the four calls this answer is actually making

This is a tradeoff question, so PICK runs the whole thing: a position, stated first, then proof that the two errors do not cost the same, then a line that would change the pick.

P
Position. Say the pick before the reasoning.
Protect thin-evidence, borderline refusals harder than fraud-flagged approvals, by default. Read refusal rate split by reason code, never as one blended number.
This is deliverable 0, restated in the framework's own words.
I
Impact. Who feels each kind of error, and how.
A wrongly refused honest customer feels it in seconds and leaves without a word. A wrongly approved risky request costs nothing that day and shows up weeks later as a chargeback someone actually reads.
One error is invisible and immediate. The other is visible and delayed.
C
Cost asymmetry. Put a real number on each side.
About $205,000 a month in quiet churn from wrongly refused thin-evidence customers, against about $11,000 a month in fraud loss from wrongly approved risky ones. Roughly eighteen to one.
This is the whole case for the position. Without the gap, there is no real tradeoff, just a coin flip.
K
Kill criteria. What would make you flip.
Fraud loss crossing $50,000 in a single month, or the same pattern hitting three times running, which usually means a ring working a blind spot, not ordinary noise.
A pick with no kill criteria is not a decision, it is a habit wearing a decision's clothes.

Two things worth naming directly, since this is where the real judgment lives. First, the rejected alternative: loosening Keeper's confidence bar globally so total refusals drop across the board. Ruled out on purpose, because it loosens the fraud-flagged bucket exactly as much as the thin-evidence one, and the fraud-flagged bucket is the one that needs to stay tight. Second, the failure worth naming by name is silent evidence-bar drift: because Keeper's decision is a probabilistic confidence score, not a fixed rule, a routine update to Corymbia's return policy wording, or a new model version, can quietly shift what counts as "enough evidence" for one reason code without anyone changing a number on purpose. The guardrail is a golden set, roughly two hundred hand-labeled borderline cases, split by reason code, re-run before any prompt or model change ships, with a pass bar of clearing the set at least nine times in ten per reason code, not once, not on average across both. That check costs something too. Routing thin-evidence cases to a person instead of a twenty-second automatic decision adds hours to those cases, and running a separate reason-code classification pass before Keeper's decision adds a small amount of inference cost to every single request. Both are the price of not quietly trading forty thousand honest customers for a cheaper, faster review process.

And if you want to be sure it really works, try it somewhere else

Same four letters, a bank's wire-fraud assistant instead of a refund bot, and this time the numbers land close enough that the pick actually flips.

Marbury Trust runs Finch, a model that holds a business wire transfer for review instead of letting it clear instantly, whenever something about the request looks off. Ulric Bettencourt runs fraud operations for business banking there.

P, position. At first glance the same call applies: protect the quiet side, the legitimate wire that gets held for no good reason, harder than the fraud-flagged one. Marbury's real position ends up the opposite. I, impact. A legitimate business wire wrongly held is loud, not quiet, a payroll run missing its cutoff generates an angry call within the hour. A fraud-flagged wire wrongly released is rare and silent at first, then very expensive. C, cost asymmetry. About $133,200 a month in businesses that quietly move their banking relationship elsewhere after a wrongly held payroll wire, against roughly $82,500 a month in expected fraud loss from wires that should have stayed held. Closer than Corymbia's gap, about one and a half to one, not eighteen to one. K, kill criteria. Marbury does not wait for the monthly average to cross a line. Any single released wire above $100,000 flips the pick immediately, because one loss that size triggers a mandatory suspicious activity report and a review of the bank's own compliance controls, a cost the monthly average never captures.

Same shape, opposite lean Corymbia protects the quiet side because the gap between the two costs is huge. Marbury protects the loud, rare side because one bad incident can be catastrophic and regulated, even though the average monthly math is nearly a tie.
Hand sketched quadrant titled Where each mistake actually lands, plotting cost per incident against how loud the complaint is. Refund borderline refusal and refund fraud approved sit low cost and silent, in the bottom left. Wire legit hold sits low cost but immediate, top left. Wire fraud released sits large cost, moderately loud, on the right.
Four mistakes, two companies, plotted on the same two axes. Cost alone does not decide which side gets protected. How loud the mistake is, and how fast it turns into a regulator's problem, decides too.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the pick: "Split the number by reason code, protect the quiet side by default, unless the rare side has a fat tail, then protect that instead."
Cost: engineering says the reason-code split cannot ship for two months. Do not read the blended refusal rate as healthy in the meantime. Treat it as unmeasured, not fine, and pull a hand sample every week until the real split exists.
The model got better, for real: say Keeper's underlying accuracy genuinely improves next quarter. That is not the same claim as "the blended refusal rate improved." A better model can still drift within a split nobody ever built a way to see.

Where people run it wrong.
They read one refusal rate as one health number, on reflex, in whichever direction is convenient that week.
They fix a bad topline by loosening the whole system at once, instead of the one reason code that is actually wrong.
They wait for a chargeback, or a lost customer, to notice, instead of splitting the metric before either one happens.

How to use it live. Say the position before the reasoning: "A refusal rate is not good or bad on its own, it depends which reason code moved." That buys you room to actually work the tradeoff instead of reciting "false positives are bad" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a tradeoff question like this, and its one-line job?
Tap to flip
ANSWER
PICK. Commit to a position first, then show the real cost asymmetry between the two kinds of error.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Camryn Falkenrath, who runs customer ops at Corymbia and owns Keeper's refusal rate, five years into the job.
3 · THE POSITION
What is the position (P) in this answer?
Tap to flip
ANSWER
Protect thin-evidence, borderline refusals harder than fraud-flagged approvals by default. Read refusal rate split by reason code, never as one number.
4 · THE IMPACT
Who feels each kind of mistake, and how?
Tap to flip
ANSWER
A wrongly refused honest customer feels it in seconds and just leaves, quietly. A wrongly approved risky request costs nothing that day and shows up weeks later as a counted loss.
5 · THE OLD DECISION
What old decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
One confidence threshold shared across every reason code. It made sense at launch, with too little labeled data to calibrate two bars. It stopped making sense once volume tripled and the two reason codes needed different amounts of trust.
6 · THE NUMBER
Fill in the blank: quiet churn from wrongly refused thin-evidence customers cost about ___ a month. Fraud loss from wrongly approved risky requests cost about ___ a month.
Tap to flip
ANSWER
About $205,000, against about $11,000. Roughly an eighteen-to-one gap, which is the entire case for protecting the quiet side by default.
7 · THE KILL CRITERIA
Same numbers, what would make you flip which side you protect harder?
Tap to flip
ANSWER
Fraud loss crossing $50,000 in a single month, or the same pattern hitting three times running, which usually means a ring working a blind spot, not noise.
8 · CROSS PRODUCT TRANSFER
Section 4 runs PICK again on a different product. Which one, and does the position land the same way?
Tap to flip
ANSWER
Marbury Trust's wire-fraud assistant, Finch. No. There the two costs are close to a tie, and Marbury protects the rare, expensive side harder because a single fraud loss can be catastrophic and trigger mandatory reporting.

Check yourself Score: 0 / 0

True or false
1. True or false: a rising refusal rate at Corymbia always means Keeper is getting worse.
  • True
  • False
Show hint
Ask which reason code is actually driving the rise before you judge it.
Show answer
False. A rise concentrated in the fraud-flagged reason code could mean Keeper is working correctly. Only the split tells you which one moved.
Fill in the blank
2. Refusals on thin-evidence, borderline requests climbed from eleven percent to ___ percent over eight weeks, while the aggregate refusal rate barely moved.
Show hint
Check the reason-code chart Camryn saw at the quarterly review.
Show answer
24 percent. Fraud-flagged refusals held steady the whole time, which is exactly why the blended nine percent never moved.
Multiple choice
3. Why does this answer protect wrongly refused honest customers harder than wrongly approved risky ones, by default?
  • A. Refusing a customer is always worse than approving fraud, on principle.
  • B. At Corymbia's actual volume, the wrongly refused side costs far more, and that cost stays invisible until the number gets split.
  • C. Fraud never actually costs the company anything measurable.
  • D. Keeper cannot tell the two kinds of request apart at all.
Show hint
Look at the two dollar figures in "Let's learn," not a general rule.
Show answer
B. The pick is built on a measured cost gap, not a blanket rule that refusing is always worse than approving.
Short answer, apply it yourself
4. Think of a product that decides yes or no about you, a loan app, a rental application, a return policy. Name one "no" you would never bother reporting, and one "yes" a company would notice going wrong right away.
Show hint
One error is quiet and you just walk away. The other shows up on someone's report.
Show answer
Model answer: A food delivery app declining a refund for cold food with no photo attached. You do not call support, you just order less from that app. Compare that to a driver flagged "verified" who steals a package. That shows up fast, as a repeat complaint pattern or a chargeback someone on the fraud team actually reviews.
Multiple choice
5. What would make Camryn flip her position and start protecting fraud-flagged approvals harder instead?
  • A. The aggregate refusal rate ticking up by one point.
  • B. Fraud loss crossing a set dollar threshold in a month, or the same pattern repeating three times running.
  • C. A single customer complaint reaching her inbox.
  • D. Keeper's response time slowing down.
Show hint
This is the K step. Look for a stated number, not a vague feeling.
Show answer
B. Kill criteria have to be a number stated in advance, not a reaction to the first complaint that lands.
True or false
6. True or false: Marbury Trust should apply the exact same default, protect the quiet side, that Corymbia uses.
  • True
  • False
Show hint
Compare the size of the cost gap at each company, not just which side feels quiet.
Show answer
False. Marbury's two costs are close to a tie, and a single released fraud wire can be large enough, and regulated enough, to flip which side gets protected harder. The position is domain-specific math, not a universal rule.
Before you close the answer
Why this works
Tests whether you will commit to a real position with real numbers behind it, instead of hiding behind "it depends" or reciting "false positives are bad" as if that settles a tradeoff question.
Follow-up traps
"Isn't protecting the quiet side just a way of tolerating more fraud?" Response: no, the fraud-flagged bucket's bar never loosens. Only the thin-evidence bucket's bar does, and the kill criteria exists exactly to catch the day that stops being true.

"How do you actually know the $205,000 churn number is real, isn't quiet churn hard to prove?" Response: it is measured off a matched cohort's repeat purchase rate, same customers, same product, before and after the reason-code split existed. Only the split changed what got compared.
If pressed
The reason-code label itself is generated by a separate, smaller classification pass that runs before Keeper's approve or refuse decision. That classifier carries its own golden set, because a fraud-flagged case wrongly bucketed as thin-evidence would quietly defeat the entire design without anyone noticing.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more