Describe the relationship between refusal rate and downstream satisfaction.
A refusal rate that holds steady can still be quietly trading one kind of customer for another. You only see it once you split the coat open.
- Split refusal rate by reason code, then protect thin-evidence, borderline refusals harder than fraud-flagged ones, by default.Why: one blended number cannot tell "we are stopping fraud" apart from "we are driving honest customers away," and the honest customer's cost turns out to be far bigger once you can see it.
- Track repeat purchase rate for refused customers, split by reason code, not only for approved ones.Why: satisfaction numbers built only from approved refunds never see the customers who quietly left after a wrong refusal.
- Set a kill criterion that flips the default, a fraud-loss threshold, not a feeling.Why: past a fixed dollar line, or the same pattern hitting three times, the hidden cost of over-approving gets bigger than the hidden cost of over-refusing.
- Gate any change to the evidence bar behind a golden set of hand-checked borderline cases.Why: a model update that looks like it is catching more fraud can really be quietly refusing more honest customers.
- Route thin-evidence refusals to a person before they become a final no.Why: a short human check catches the case a single confidence score would wrongly close.
How to answer this, stage by stage
Nobody is grading whether you can define a false positive. They are grading whether you will commit to a side and then prove, in dollars, that the asymmetry is real. Seven moves get you there.
Let's learn
The decision happens inside a small box on a return form: refund approved, or refund denied, decided before the customer even finishes reloading the page.
Corymbia sells outdoor gear and home goods online, about 40,000 return or exchange requests come through in a normal month. Keeper is the model wired into that box. It reads the reason a customer gives, checks it against photos and order history, and decides, refund or no refund, in under twenty seconds.
Before Keeper, a team of return agents made every call by hand. Average time per case: six minutes. Across forty thousand cases a month that is close to four thousand hours of work, and it showed in the outcomes: two agents looking at the same thin evidence could land on opposite answers, one refusing four percent of the time, another refusing fifteen.
Keeper decides every case the same way, every time. Refusal rate settled at nine percent overall, which read, on a dashboard, as a healthy, boring number.
Here is the turn. Nine percent hid two very different nine percents. Split by reason code, refusals on requests with an actual fraud signal, repeat abuse, a mismatched address, no order match, held almost exactly where the old team left them. Refusals on requests with thin but real evidence, a blurry photo, a late but honest report, climbed from eleven percent to twenty four percent over eight weeks. Keeper was not catching more fraud. It was getting fussier with the customers who had a real case and just a messy one.
At its worst, this cost more than a bad quarterly number. Over three months, before anyone caught it, that quiet churn added up to about $615,000 in customers who simply stopped shopping at Corymbia, most of whom had done nothing wrong.
The choice I would take back is that shared threshold. When Keeper was built, engineering had one quarter and not nearly enough labeled data to calibrate a separate bar for every reason code, so one bar covered all of them. That was a fair call with the data on hand. It stopped being fair once volume grew and the two reason codes started needing very different amounts of trust.
What I would leave alone: refusals on requests with zero evidence at all, no order number, no photo, no account match. Those should stay high. There is no ambiguity there worth splitting further.
The lesson: a refusal rate that holds steady is not proof that nothing changed underneath it. Two numbers can cancel each other out on a dashboard while doing very different things to real customers.
Now here is the same thing as a story
The short version is above. Read on for the Thursday a slide in someone else's deck showed Camryn what her own number had been hiding.
Camryn Falkenrath has run customer ops at Corymbia for five years. Before Keeper, she could read a return request and tell inside ten seconds whether the story held together, a knack she learned the hard way, once refusing a real customer who had worn a rain jacket on the exact hike she said ruined it.
When Keeper launched, the first months were good. Every Monday she pulled a report split by reason code and hand-checked ten cases against it, five thin-evidence, five fraud-flagged. It always matched. The topline number sat near nine percent, steady, boring, exactly what a healthy number should look like.
So she trimmed the Monday check. Ten cases became three. A few months later, three became a glance at the topline chart and nothing else. The number was still nine percent. Why keep opening the file.
Nothing broke on a Tuesday. There was no single morning where the number jumped. It built up slowly, the way a crack under wallpaper does, until a Thursday quarterly review, called for an unrelated reason, put an eight-week reason-code chart from finance up on the screen. Camryn had not built that chart. She had not asked for it. Borderline refusals: eleven percent, climbing to twenty four. Fraud-flagged refusals: flat the whole time.
She pulled forty of the month's thin-evidence denials that night and read every one herself. Nineteen of the forty, on a second look, should have gone through.
The real cost was not the meeting. Corymbia's finance team, working the same quarter's numbers separately, had already flagged a repeat-purchase dip among refused customers and filed it under normal seasonal noise. Nobody had connected the two charts until Camryn walked hers over.
A year earlier, when engineering built the threshold, the room had good reasons. One bar was simpler to ship, and there was not enough labeled data yet to trust two. Nobody in that meeting expected refusal volume to triple in a year, and nobody put a date on when the shortcut should get revisited.
Run the same quarter through the fixed design. Borderline refusals settle back near twelve percent instead of twenty four. Projected quiet churn drops from roughly two hundred thousand dollars a month to about forty thousand. Fraud loss ticks up slightly, from eleven thousand to about fourteen thousand, because the thin-evidence bar loosened a little on its own side of the split. Net, Corymbia keeps most of the forty thousand honest customers it was about to lose and pays fourteen thousand more a month to a bucket that never touched the fraud-flagged rules at all.
One design watched one number. The other watched two, and only one of them ever needed protecting harder.
What I would tell my past self: I asked whether the refusal rate was healthy. I never asked whose refusal rate it was.
PICK, the four calls this answer is actually making
This is a tradeoff question, so PICK runs the whole thing: a position, stated first, then proof that the two errors do not cost the same, then a line that would change the pick.
Two things worth naming directly, since this is where the real judgment lives. First, the rejected alternative: loosening Keeper's confidence bar globally so total refusals drop across the board. Ruled out on purpose, because it loosens the fraud-flagged bucket exactly as much as the thin-evidence one, and the fraud-flagged bucket is the one that needs to stay tight. Second, the failure worth naming by name is silent evidence-bar drift: because Keeper's decision is a probabilistic confidence score, not a fixed rule, a routine update to Corymbia's return policy wording, or a new model version, can quietly shift what counts as "enough evidence" for one reason code without anyone changing a number on purpose. The guardrail is a golden set, roughly two hundred hand-labeled borderline cases, split by reason code, re-run before any prompt or model change ships, with a pass bar of clearing the set at least nine times in ten per reason code, not once, not on average across both. That check costs something too. Routing thin-evidence cases to a person instead of a twenty-second automatic decision adds hours to those cases, and running a separate reason-code classification pass before Keeper's decision adds a small amount of inference cost to every single request. Both are the price of not quietly trading forty thousand honest customers for a cheaper, faster review process.
And if you want to be sure it really works, try it somewhere else
Same four letters, a bank's wire-fraud assistant instead of a refund bot, and this time the numbers land close enough that the pick actually flips.
Marbury Trust runs Finch, a model that holds a business wire transfer for review instead of letting it clear instantly, whenever something about the request looks off. Ulric Bettencourt runs fraud operations for business banking there.
P, position. At first glance the same call applies: protect the quiet side, the legitimate wire that gets held for no good reason, harder than the fraud-flagged one. Marbury's real position ends up the opposite. I, impact. A legitimate business wire wrongly held is loud, not quiet, a payroll run missing its cutoff generates an angry call within the hour. A fraud-flagged wire wrongly released is rare and silent at first, then very expensive. C, cost asymmetry. About $133,200 a month in businesses that quietly move their banking relationship elsewhere after a wrongly held payroll wire, against roughly $82,500 a month in expected fraud loss from wires that should have stayed held. Closer than Corymbia's gap, about one and a half to one, not eighteen to one. K, kill criteria. Marbury does not wait for the monthly average to cross a line. Any single released wire above $100,000 flips the pick immediately, because one loss that size triggers a mandatory suspicious activity report and a review of the bank's own compliance controls, a cost the monthly average never captures.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the pick: "Split the number by reason code, protect the quiet side by default, unless the rare side has a fat tail, then protect that instead."
Cost: engineering says the reason-code split cannot ship for two months. Do not read the blended refusal rate as healthy in the meantime. Treat it as unmeasured, not fine, and pull a hand sample every week until the real split exists.
The model got better, for real: say Keeper's underlying accuracy genuinely improves next quarter. That is not the same claim as "the blended refusal rate improved." A better model can still drift within a split nobody ever built a way to see.
Where people run it wrong.
They read one refusal rate as one health number, on reflex, in whichever direction is convenient that week.
They fix a bad topline by loosening the whole system at once, instead of the one reason code that is actually wrong.
They wait for a chargeback, or a lost customer, to notice, instead of splitting the metric before either one happens.
How to use it live. Say the position before the reasoning: "A refusal rate is not good or bad on its own, it depends which reason code moved." That buys you room to actually work the tradeoff instead of reciting "false positives are bad" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"How do you actually know the $205,000 churn number is real, isn't quiet churn hard to prove?" Response: it is measured off a matched cohort's repeat purchase rate, same customers, same product, before and after the reason-code split existed. Only the split changed what got compared.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #5 What early signal predicts churn from an AI feature?
- #6 How do you build an early warning system for silent quality degradation?