CaseAdvancedModel Fluency & the AI PM Role / What changes when the product is probabilistic / #20
Describe a product decision you would make differently if the model's error rate doubled overnight.
BOUND · an AI ad-copy and marketing-campaign generator for small businesses
Emberpitch is Quillmarch's campaign generator: give it a business's product brief and it writes back ad headlines, body copy, and a week of social posts and email subject lines. Elowyn Prevette owns Emberpitch's quality numbers. Quillmarch swapped in a cheaper, faster base model to cut inference cost, and the next morning's eval showed the claim-error rate had basically doubled. Iwalani Desroches, who runs brand and legal risk, needed a decision before that day's campaigns went out.
The direct answer
The moment the weekly eval shows Emberpitch's claim-error rate cross the 2.5 percent line, route every claim-bearing line, a number, a guarantee, a certification, a comparison, back through mandatory human review, and pull Quillmarch's top-tier auto-publish accounts back into that review with it. At real volume, the 2.1-to-4.3 jump works out to about 440 extra wrong claim lines a day, and even the kindest read of what the gate already catches still lets 132 to 198 of them reach a live ad. Fix the model's rate on its own timeline, but stop the leak first.
Do this, in order
Route every claim-bearing line back through mandatory human review the moment the weekly eval crosses 2.5 percent, and pull the graduated top-tier accounts back into it.Why: this is the one decision that stops the extra wrong lines from reaching a live ad while the model's rate gets fixed on its own timeline.
Score claim-bearing lines separately from everything else Emberpitch writes.Why: claim-bearing lines are only 35 percent of daily volume, so blending them into one average hides a doubled rate on the risky slice behind an almost-unchanged number on the rest.
Give the shipped-extra-lines estimate a range, not a point.Why: the confidence gate's real catch rate, 55 to 70 percent, comes from quarterly spot audits, not a full count, so the honest exposure is 132 to 198 extra lines a day, not one tidy figure.
Check the doubled rate against both the auto-publish threshold and a real human baseline.Why: 4.3 percent clears the 2.5 percent line that graduated 300 accounts out of review, and it's roughly ten times what Quillmarch's own copywriters measure on themselves.
Watch the confidence gate's catch rate itself, not just the model's error rate.Why: it's the least grounded number in the whole equation, and the one that swings the shipped estimate most if it's wrong.
Reject a full model rollback as the first move.Why: it throws away real gains on the other 65 percent of output and takes longer to redeploy than a gate setting takes to flip.
How to answer this, stage by stage
Nobody's grading whether you can say the error rate doubled and sound concerned. They're grading whether you can turn that into a number of wrong lines a day, and a decision that survives someone asking where 132 to 198 actually came from.
1
Scope it to one real launch decision
Say it like this
"Let me ground this. Emberpitch is Quillmarch's campaign generator, it writes ad headlines, body copy, and full campaign packages for about 2,300 small business accounts. Elowyn Prevette owns its quality numbers. Quillmarch swapped in a cheaper, faster base model to cut inference cost, and the next morning's eval showed the claim-error rate had basically doubled. I'll answer against that."
Why this works
One real product and one real number keeps every claim checkable, instead of a vague point about AI getting worse.
2
Say the shape of the equation before any numbers
Say it like this
"Doubled isn't a decision on its own, it's a multiplier. To turn it into one, I need four things: how many lines a day actually carry a claim, what the error rate was and is on those lines, how many of the extra wrong ones the review gate already catches, and how many of what slips through actually get noticed. Say the shape before naming a single figure."
Why this works
This is the B step, and it stops "the error rate doubled" from being treated as the whole answer.
3
Own each number, with where it came from
Say it like this
"About 20,000 of Emberpitch's 57,000 daily lines carry a claim, a number, a guarantee, a certification, a comparison, that's from the usage logs. The weekly eval sample had the claim-error rate at 2.1 percent before the swap and 4.3 percent the morning after. That's 420 wrong a day before, 860 after, so about 440 extra wrong claim lines a day."
Why this works
Every number traces to a place someone could go check it, not a guess wearing a percent sign.
4
Give the range, not a point
Say it like this
"The gate that routes low-confidence lines to a person isn't perfectly measured, quarterly spot audits put its real catch rate somewhere between 55 and 70 percent. Push the extra 440 through that: somewhere between 132 and 198 of them reach a live ad before anyone but the advertiser's customer sees them. Not one number, a range, because the catch rate itself is an estimate."
Why this works
A single point here would claim more precision than a spot-audited catch rate can back up.
5
Nail the sanity check
Say it like this
"Here's where it actually becomes a decision. Quillmarch graduated its top accounts out of manual review on the promise the claim-error rate would stay under 2.5 percent. 4.3 clears that by a lot. And when we measured our own senior copywriters fact-checking their own drafts, they landed around 0.4 percent. So this isn't the model having an off day, it's running at ten times a careful human's rate, past the line we set for that."
Why this works
The hardest step, and the one that turns a number into a reason to act, not just a number to note.
6
Name the decision, and reject the tempting bigger move
Say it like this
"So here's the call. Tighten the confidence gate specifically on claim-bearing lines, and pull the graduated top-tier accounts back into mandatory review until the rate recovers. I'd reject rolling the whole model back tonight, it's also the version that cut our inference cost and improved everything that isn't a claim, and a full rollback takes longer to redeploy than a gate setting takes to flip."
Why this works
Naming the alternative you didn't take is what makes this a judgment call instead of the only option that occurred to you.
7
Close on the decision, in one breath
Say it like this
"So: about 440 extra wrong claim lines a day, 132 to 198 of them shipping, well past the threshold that let the top tier skip review, so that review comes back for them until two straight weekly evals land back under 2.5 percent."
Why this works
Leaves whoever's grading this with the decision, not just the arithmetic that led to it.
Let's learn
Every Tuesday morning, Elowyn Prevette pulls the same 600-line sample and checks one number. For months, it barely moved.
The number is Emberpitch's claim-error rate, how often a generated ad line gets a fact wrong. Emberpitch is Quillmarch's campaign generator: a small business owner describes what they sell, and it writes back ad headlines, body copy, and a week of social posts and email subject lines, ready to publish.
Before last month, Elowyn's number was reassuring. About 20,000 of Emberpitch's 57,000 daily lines carry a claim, a number, a guarantee, a certification, a comparison, the kind of line that can be factually wrong. At a 2.1 percent error rate on those, that's 420 wrong lines a day, comfortably under the 2.5 percent line Quillmarch had set for its most reliable customers to skip review entirely.
Five numbers hiding behind one word, doubled. Say what each one is before reacting to any of them.
About 300 of Quillmarch's 2,300 accounts, its largest and longest-standing customers, had graduated out of manual review months earlier. The model had been steady long enough that a person checking every line felt like busywork nobody needed.
Then Quillmarch swapped in a newer base model, faster, cheaper to run, better at punchy hooks. Overnight, the next weekly eval read 4.3 percent. Not a drift. A jump.
Emberpitch's weekly claim-error rate, eight weeks
Claim-error rate, weekly eval sampleThe week of the model swap
Seven weeks sat quietly between 1.9 and 2.2 percent. Week eight didn't drift up toward the 2.5 percent line, it jumped straight past it overnight.
Two minutes apart. Same dashboard, same chair, one number that flipped a whole morning.
The number didn't just get worse. It walked straight past the line that had let three hundred accounts stop being checked at all.
Iwalani Desroches, who runs brand and legal risk at Quillmarch, doesn't see a percentage. She sees a landscaping company's ad claiming a certification it doesn't hold, live, published, running, because the account graduated out of review back when 2.1 was the number everyone trusted.
What it costs at its worst is more than a correction email. A fabricated certification or an invented discount sits under a real small business's name, on a platform that can suspend an entire ad account for one unsubstantiated claim, not just pull the one line. The business that hired Quillmarch to save them time ends up spending it explaining to a customer, or a platform reviewer, why their ad said something nobody at the company ever claimed.
The choice I would take back
Graduating an account out of review was a one-time decision, made once, at one error rate, with no expiry and no re-check built in. That was fine while the number held. It stopped being fine the moment the number that earned the exemption was no longer the number in production.
What I would leave alone: the other 65 percent of Emberpitch's output, generic hooks, tone variants, emoji placement, calls to action with no claim in them. A doubled error rate there costs a business owner a slightly flatter headline, not a false claim with their name on it. Gating those the same way just slows down copy that was never the risk.
The lesson: the model didn't fail. Quillmarch built an exemption around a snapshot instead of a standing check. A number is only safe to build an exemption on the day it's treated as a range you keep watching, not a fact you file away.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a doubled number is never really about the number.
For eleven months, the worst part of Elowyn Prevette's Tuesday was deciding what to have for lunch.
She runs quality for Emberpitch, Quillmarch's campaign generator, and every Tuesday she pulls a 600-line sample of claim-bearing copy, a number, a guarantee, a certification, a comparison, and has a person check each one against the business's own brief. For eleven months the number held steady around 2 percent. Good enough that, one by one, Quillmarch's biggest accounts got moved off mandatory review and onto auto-publish. Elowyn signed off on each one. The number said it was fine.
Then the infrastructure team shipped a new base model. Faster. A third cheaper to run. Reviewed for tone and hook quality, and it tested well, funnier headlines, tighter hooks, the kind of copy that makes a small business owner actually want to post it.
Nobody reviewed it for claims. Nobody thought to. The old model hadn't been a claims problem in almost a year.
Elowyn pulled Tuesday's sample the way she always did, coffee first, dashboard second. 2.1 last week. She scrolled. 4.3 this week.
She read it twice.
She didn't panic, and she didn't wait for a bigger sample either. She called Iwalani Desroches, who runs brand and legal risk two floors down, before she'd finished her coffee.
Iwalani pulled up a handful of that week's flagged lines while Elowyn talked. A pressure-washing company's ad claiming a certification the company had never held. A landscaper's headline promising a discount that didn't exist in the brief. Small, specific, wrong, and every one of them auto-published, live, under a real business's name, because the account had graduated out of review back when 2.1 was the number everyone trusted.
This is the fork every extra wrong line hits. The graduated accounts skip the left branch entirely.
Here's the part that isn't really about the model. Elowyn had a number in her head that only had two settings: under the line, or over it. She never carried the actual percentage around, she carried a feeling, calm on a Monday, not calm the moment 4.3 showed up. Two point one to four point three flipped the feeling in one look at a spreadsheet.
So here's the decision Elowyn and Iwalani made that morning, before the second coffee.
Months earlier, when the auto-publish program launched, the rule had been: hit 2.5 percent or under for a full quarter, and the account graduates out of review, for good. Nobody built in a re-check. Nobody asked what happens if the model underneath changes and the account doesn't know to ask again.
I would take that back. Not the graduation itself, the businesses that earned it deserved faster copy. I'd take back making it permanent. Tie it to a live number instead: the moment the weekly eval crosses 2.5 percent on claim-bearing lines, review comes back automatically for every graduated account, no waiting for someone to notice.
Run the same Tuesday again with that rule in place. The eval comes back at 4.3. The system, not Elowyn, pulls the 300 graduated accounts back into the review queue before the day's first campaign goes out. About 440 extra wrong lines get generated that day either way, the model's rate is the model's rate, but somewhere between 132 and 198 of them, the ones that would have shipped straight to a live ad, get a person's eyes on them first instead.
One design waits for a person to notice a spreadsheet before anyone with a lever is told there's a problem. The other design tells the lever the moment the number crosses the line.
What I'd tell my past self, the one who set the graduation rule as a one-time pass: a number that earns someone an exemption today is not a promise about the number tomorrow. Build the check to keep watching, not to file the result away.
BOUND, or doing the arithmetic before the campaign goes live
Not a way to sound more careful about a doubled number. BOUND is what turns "the error rate doubled" into a count anyone can act on, and a decision that survives someone asking where the numbers came from.
BBreak it down. What does doubled actually need?
A doubled error rate isn't one fact, it's a multiplier waiting for four numbers: how many lines a day actually carry a claim, what the error rate was and is on those lines specifically, how many of the extra wrong ones a review gate already catches, and how many of what ships actually gets noticed. Say the shape before naming a single figure, or "the error rate doubled" quietly stands in for arithmetic nobody did.
Skip this step and every number after it is just a guess dressed up as a decision.
OOwn the numbers. Where did each one come from?
About 20,000 of Emberpitch's 57,000 daily lines carry a claim, 35 percent, pulled straight from usage logs. The weekly eval, a 600-line human-checked sample, had the claim-error rate at 2.1 percent before the model swap and 4.3 percent the next morning. That's 420 wrong claim lines a day before, 860 after: about 440 extra a day. Why a 600-line weekly sample and not a full daily audit: full audits don't scale to a person's time, and 600 lines a week is what the review team can actually check by hand without becoming the bottleneck itself.
Owning a number means saying where it came from and why that size, not just stating a figure that sounds specific.
UUse a range, not one point.
The confidence gate that routes low-confidence lines to a person isn't perfectly measured either. Quarterly spot audits put its real catch rate somewhere between 55 and 70 percent, an estimate, not a census. Push the extra 440 through that range and somewhere between 132 and 198 of them ship to a live ad before anyone but the advertiser's customer sees them. A flat "about 165" looks tidier and hides how much this number could move if the next audit lands differently.
This is the direct answer's arithmetic, in one step: the range is the honest version, the point estimate is the comfortable one.
The extra 440 wrong lines a day, split by what the gate catches
Caught by the gate, sent to a reviewerShips to a live ad
Either end of the audited range, the same 440 extra wrong lines a day split into a caught pile and a shipped pile. The shipped pile, 132 to 198, is the honest exposure, not one point figure.
NNail the sanity check. Does it survive a real comparison?
Two comparisons, and they both point the same way. Quillmarch's auto-publish rule promised the graduated accounts a claim-error rate under 2.5 percent, 4.3 clears that by more than seventy percent. And when Quillmarch measured its own senior copywriters fact-checking their own drafts on the same 600-line-style sample, they landed around 0.4 percent. So 4.3 isn't a bad week, it's roughly ten times what a careful human manages on the same task, sitting well past the line the whole auto-publish program was built on.
The hardest step, and the one that turns a number into a reason to act instead of a number to note in a status update.
All four numbers on the same scale. 4.3 doesn't sit near the line, it sits well past everything on it.
DDirection. Which assumption would move it most?
Not the claim-share estimate, 35 percent is steady in the logs and barely moves the answer if it's off by a few points. It's the gate's catch rate. That 55-to-70 range comes from quarterly spot audits, not a full count, and it's the biggest lever in the whole equation: at 50 percent instead of 55, the shipped estimate climbs past the top of the stated range entirely. Watch that number, not the model's raw error rate, for the next real surprise.
Naming the assumption that's both uncertain and consequential, not just the biggest number in the equation, is what a good estimator does that a bad one skips.
The gate's catch rate sits exactly where a good estimator worries most: least grounded, most consequential.
The whole decision in five boxes. The fourth box is the one that turns a number into an action.
Three things worth naming directly, since this is where the real judgment sits. The AI-specific failure mode is a plain hallucination: the new model invents specifics, a certification, a percentage, a limited-time discount, that were never in the business's own brief, because it's optimizing for a punchy line over a grounded one. The guardrail today is the confidence gate; the guardrail worth building next is a check that verifies every claim-bearing line against the fields in the business's actual brief before it's allowed to publish, not just a confidence score. The alternative I rejected is rolling the whole model back overnight, tempting because it looks like the safest move, but it throws away real gains on everything that isn't a claim and takes longer to redeploy than a gate setting takes to flip. And there's a real trade-off in the decision that shipped: tightening the gate means campaigns that used to auto-publish in under a minute now wait for a person, sometimes hours, and Quillmarch's review team costs more hours a week doing it. That's the price, paid in speed and headcount, for not paying it in fabricated claims sitting under a real business's name.
And if you want to be sure it really works, try it somewhere else
Same five letters, a translated appliance manual instead of an ad. This time the rare thing isn't a fabricated discount. It's a mistranslated voltage warning.
Same method, a different label on the box. Both times, the graduated exemption was built on a number that quietly stopped holding.
Marrowdene Appliances ships washing machine and water heater manuals in fourteen languages through Verityscript, a translation tool that also checks safety-critical lines, voltage warnings, chemical handling steps, against a certified terminology glossary before a manual goes to print. Osmund Faircloth owns Verityscript's numbers. Marrowdene swapped its translation model for a cheaper one to cut localization costs, and the next batch's safety-line error rate, measured against the glossary, roughly doubled overnight: 1.8 percent to 3.6.
Osmund runs the same five letters. Break it down: about 50,000 safety-critical lines ship a month, across 40 manuals in 14 languages, and the number that matters isn't the manual-wide average, it's the error rate on lines a certifier actually inspects. Own the numbers: at 1.8 percent, that's 900 wrong safety lines a month before the swap, 1,800 after, about 900 extra. Use a range: bilingual reviewers spot-check and catch an estimated 60 to 75 percent before print, an estimate, not a full count, so somewhere between 225 and 360 of those extra wrong lines reach a printed manual. Nail the sanity check: the certification body's own tolerance for mistranslated safety-critical text is under 1 percent, and even the 1.8 percent baseline, before anything doubled, was already over that line. 3.6 percent is nearly four times it. Direction: same lever as Emberpitch's story, the swing assumption is the reviewer catch rate, the least grounded number in the equation and the one that moves the estimate most if a spot audit turns out to be optimistic.
The decision Osmund would take back
Marrowdene's localization plan assumed a translation model was ready everywhere the moment it cleared testing on ordinary marketing copy. That worked for years. It stopped working the day the safety-critical lines, the ones a certifier actually checks, were riding on a rate nobody had separated out from the average.
Same rank, different lever: the fix for Verityscript isn't a bigger review team either. It's the same habit Elowyn ran on Emberpitch's numbers: pull the risk-bearing category out of the average, range the number that's genuinely uncertain, and route that category, specifically, back to a human, regardless of what the aggregate error rate says.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: claim-bearing volume, times the error-rate delta, times what the gate actually catches, times how often a miss gets noticed, checked against a real threshold.
Cost: no budget to add a grounding step this quarter. Keep the wider human-review net on the graduated segment as the interim cost, it's cheaper than one recalled campaign or one denied ad account.
The model got better, for real: say the next version's claim-error rate came in at 1.0 percent instead of 4.3, half the old rate. The graduated segment doesn't need widening either, it needs the same threshold check, just checked in the other direction, before anyone loosens review somewhere else.
Where people run it wrong.
They report the raw doubling as the finding and stop there, without turning it into a daily count anyone can act on.
They test the whole product's error rate instead of isolating claim-bearing lines, and dilute a real signal into noise.
They treat the confidence gate's catch rate as a fixed fact instead of an estimate from spot audits, and build a false sense of precision on top of it.
How to use it live. Ask "what's the volume this actually happens on, and what already catches it?" before quoting a percentage. That's the two multipliers a doubled rate needs before it means anything.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: turn a doubled error rate into a real daily count, an honest range, and a decision that survives being compared to a real threshold. Built for estimation questions, not a habit-flip story.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Elowyn Prevette, who owns Emberpitch's quality numbers at Quillmarch, and has to turn "the claim-error rate doubled overnight" into a decision Iwalani Desroches, who runs brand and legal risk, can act on.
3 · THE SWING ASSUMPTION
Which single assumption in this answer would move the estimate most if it's wrong?
Tap to flip
ANSWER
The confidence gate's real catch rate, estimated at 55 to 70 percent from quarterly spot audits, not a full count. At 50 percent instead of 55, the shipped estimate climbs past the top of the stated range.
4 · THE EQUATION
What four things does a doubled error rate need before it means anything?
Tap to flip
ANSWER
How many lines a day actually carry a claim, the error rate on those lines before and after, how many of the extra wrong ones a review gate already catches, and how many of what ships actually gets noticed.
5 · THE OLD DECISION
What decision would Elowyn and Iwalani take back?
Tap to flip
ANSWER
Making account graduation out of manual review a one-time, permanent pass instead of tying it to a live number. It made sense while the rate held; it stopped making sense the moment the model underneath changed and nobody had to ask again.
6 · THE NUMBER
Fill in the blank: Emberpitch's claim-error rate went from ___ percent to ___ percent overnight, about ___ extra wrong claim lines a day.
Tap to flip
ANSWER
2.1 to 4.3 percent, about 440 extra wrong claim lines a day. Somewhere between 132 and 198 of those still ship to a live ad even on the kindest read of what the gate catches.
7 · THE REPLAY
Same Tuesday, new rule in place, what changes?
Tap to flip
ANSWER
The system, not a person noticing a spreadsheet, pulls the 300 graduated accounts back into review the moment the weekly eval crosses 2.5 percent. Somewhere between 132 and 198 lines that would have shipped straight to a live ad get a person's eyes on them first instead.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which one, and what's the swing assumption there?
Tap to flip
ANSWER
Verityscript, Marrowdene Appliances' manual-translation tool. Same swing assumption as Emberpitch's story: the reviewer catch rate, 60 to 75 percent, estimated from spot checks, not a full count.
Check yourself Score: 0 / 0
True or false
1. True or false: the real problem with Emberpitch's doubled error rate is that customers will see roughly twice as many awkward or generic lines.
True
False
Show hint
Look at "what I would leave alone" in Let's learn.
Show answer
False. The real problem is narrower and sharper: only claim-bearing lines carry real risk, and the doubled rate on those crossed the 2.5 percent line that let 300 accounts skip review entirely. Generic hooks and calls to action with no claim in them aren't worth gating over.
Multiple choice
2. Why does Elowyn score the error rate on claim-bearing lines specifically, instead of one error rate across everything Emberpitch writes?
A. Because claim-bearing lines are easier to detect automatically.
B. Because blending every line together would dilute the real risk under a much larger pool of low-stakes copy, and hide exactly the lines that can trigger legal or brand risk.
C. Because Quillmarch's contract with advertisers requires it by name.
D. Because the confidence gate can only score claim-bearing lines.
Show hint
Look at the claim-share paragraph in Let's learn.
Show answer
B. With claim-bearing lines at only 35 percent of daily volume, a blended average would hide a doubled rate on the risky slice behind a mostly-unchanged number on the rest.
Fill in the blank
3. Quillmarch's auto-publish threshold was set at ___ percent. Emberpitch's claim-error rate doubled from 2.1 to ___ percent, and somewhere between ___ and ___ of the extra wrong lines a day still ship to a live ad.
Show hint
Check the N step and the U step in the BOUND recap.
Show answer
2.5 percent; 4.3 percent; 132 and 198. The doubled rate cleared the threshold that had graduated the top-tier accounts out of review in the first place.
Short answer, where it wouldn't matter
4. Name a category of Emberpitch's output where this doubled error rate would NOT be worth gating harder over.
Show hint
Look at "what I would leave alone" in Let's learn.
Show answer
Model answer: Generic hooks, tone variants, and calls to action with no factual claim in them, about 65 percent of daily output. A doubled error rate there costs a slightly flatter headline, not a false claim with a business's name on it.
Short answer, apply it yourself
5. Think of an AI feature you've used or built where a number, an error rate, a latency, a cost, changed. What are the equivalent four numbers, volume, rate, catch rate, and notice rate, in your case?
Show hint
Walk the same four the B step names: volume, before and after rate, what catches a miss, how often a miss gets noticed.
Show answer
Model answer: A team running an AI email-summarization tool doubled its "missed key detail" rate after a model swap. The four numbers: how many summaries went out a day, the miss rate before and after, how often a person actually opened the full email anyway, their catch rate, and how often a missed detail led to a real support ticket.
Short answer, work the number
6. If the confidence gate's real catch rate turned out to be 50 percent instead of the estimated 55 to 70 percent range, roughly how many of the 440 extra wrong lines would ship to a live ad a day?
Show hint
Shipped equals extra wrong lines times one minus the catch rate.
Show answer
About 220 a day, 440 times 0.50. That's above the top of the stated 132-to-198 range, which is exactly why the catch rate, not the claim share or the daily volume, is the assumption most worth double-checking.
Before you close the answer
Why this works
Tests whether you'll turn "the error rate doubled" into a daily count anyone can act on, and know that a doubled number only becomes a decision once it's checked against a real threshold, not just noted as worse.
Follow-up traps
"Why not just roll the whole model back tonight, isn't that obviously the safest move?" Response: it also throws away real gains on the other 65 percent of output and takes longer to redeploy than flipping a gate setting. The gate stops the exposure today while the model decision gets made properly.
"A 55 to 70 percent catch-rate range is pretty wide. Isn't that too fuzzy to build a decision on?" Response: no, the width is the finding. It's what shows the catch rate is the assumption most worth checking first, and the decision, reinstating review for the graduated accounts, holds at either end of that range.
If pressed
The confidence gate's score isn't the generation model's own token probability. It's a separate small classifier trained to spot claim-shaped phrasing, numbers, certifications, guarantees, comparisons, scored independently of how sure the generation model "feels" about the line it just wrote.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.