ConceptFoundationalModel Fluency & the AI PM Role / What changes when the product is probabilistic / #6
Why can you not write an acceptance criterion like 'the output must be accurate' for a generative feature?
BOUND · content moderation and flagging for social platforms
Wardline is Hearthloop's generative flagging model. It reads a post and writes back a violation category, a confidence score, and a plain-language reason, for any post a lighter upstream filter has already routed to it as possibly self-harm related. Temperance Bramhall owns Wardline's numbers. Before the newest version ships, Ilse Korhonen, who runs Trust and Safety policy, wants one line in the launch doc: "the output must be accurate." Temperance has to turn that line into something a review board can actually sign.
The direct answer
"Accurate" isn't one number for a generative flagging model, it's hiding four: how big the test set of confirmed real cases is, which kind of mistake you're actually scoring, where the pass line sits, and how wide that line has to be given the sample size. For Wardline that means about 500 confirmed self-harm posts, scored mainly on how much of the real stuff it catches, with a pass line of 88 to 94 percent, not a flat 90. Check that range against something real: two trained reviewers only agree with each other 89 percent of the time on those same posts, so the top of the range already asks Wardline to be surer than a human reviewer ever is.
Do this, in order
Say what "accurate" is hiding, before writing anything else in the launch doc.Why: skip this and one soft word stands in for four decisions nobody actually made.
Build the test set from confirmed real cases, not a slice of everyday traffic.Why: real violations are about 1.2 percent of the queue, so a model that never flags anything would still score 98.8 percent on ordinary traffic and catch nothing real.
Score how much of the real stuff it catches, and how much of what it flags is real, as two separate numbers.Why: missing a genuine at-risk post and wrongly flagging a safe one cost different people different things, and one blended number hides which one you're actually protecting against.
Give the pass line a range, 88 to 94 percent, not a flat point.Why: with 500 real cases behind the measurement, a single point claims a precision the sample size hasn't earned.
Check the range against something real, like two-reviewer agreement.Why: the top of the range already sits above what two trained reviewers agree on with each other, so treating it as a hard floor sets a bar that isn't really where human judgment sits either.
Watch how much of the test set covers coded language, not just its total size.Why: coded language is 34 percent of real violations and climbing, but only 12 percent of the golden set, so the headline number is measuring yesterday's threat more than today's.
How to answer this, stage by stage
Nobody's grading whether you can say "we'd be careful about that." They're grading whether you'll turn a soft word into real arithmetic, and whether you know a confident single number is a warning sign when the sample behind it is only a few hundred cases.
1
Anchor it to one real launch gate
Say it like this
"Let's ground this in one real case. Wardline is Hearthloop's generative flagging model. It reads a post and writes a violation category, a confidence score, and a plain-language reason. Temperance Bramhall owns its numbers. Ilse Korhonen, who runs Trust and Safety policy, wants 'the output must be accurate' written into the launch doc before this version ships, and Temperance has to turn that into something a review board can actually sign off on."
Why this works
One real product and one real gatekeeper keeps every number checkable, instead of a general point about AI quality.
2
Say what "accurate" is actually hiding
Say it like this
"The word accurate sounds like a spec, but it's hiding four unmade decisions: how big the test set is, which kind of mistake we're actually scoring, where the pass line sits, and how sure we can be about that line given how few real cases we've got. Sign off on 'accurate' as written, and you've signed off on none of those."
Why this works
This is the reframe the rest of the answer turns on. Skip it and BOUND looks bolted on instead of the actual fix.
3
Say the equation, then give the decision in one breath
Say it like this
"Here's the shape: build a test set of confirmed real cases, score how much of the real stuff it catches and how much of what it flags is real as two separate numbers, set a pass line, and give that line a range instead of a point. For Wardline: 500 confirmed self-harm posts, catches 88 to 94 percent of the real ones, not a flat 90, and right next to that, two trained reviewers only agree with each other 89 percent of the time on those same posts."
Why this works
Saying the shape before the figures stops "we tested it" from standing in for real arithmetic, and this is the direct answer, said plainly.
4
Own each number, with where it actually came from
Say it like this
"500 confirmed cases, pulled from a year of posts a human reviewer already signed off as real. The 88 to 94 range is the honest math: at 500 cases and a measured catch rate of 91 percent, the real interval runs about three points either way. And 89 percent human agreement is Hearthloop's own quarterly check, where two trained reviewers score the same posts blind and compare notes after."
Why this works
Every number traces to a place someone could go check it, not a guess wearing a percent sign.
5
Name the one assumption that would move it most
Say it like this
"Here's what I'd flag hardest. Only 60 of those 500 cases use coded or evasive language, things like 'unalive' or deliberately misspelled words. Our own data says coded language made up 34 percent of real violations last quarter, and it's climbing. If the test set doesn't grow that slice, 91 percent is a true number about the easy cases and says almost nothing about the ones actually getting through right now."
Why this works
A good estimate names the assumption that would swing the answer most, not just the easiest one to argue about.
6
Close on the decision, in one breath
Say it like this
"So: turn 'accurate' into four real numbers, test it on confirmed cases, give a range instead of one point, check that range against what a human reviewer actually agrees on, and keep watching whether the test set still looks like the posts really getting through."
Why this works
Restates the decision plainly, so whoever's grading this leaves with the answer, not just the story.
Let's learn
What happens the first time a launch doc says "the output must be accurate," and somebody has to make that word actually mean something?
Wardline is Hearthloop's generative flagging model. Give it a post, and it writes back three things: a violation category, a confidence score, and a plain-language reason, for any post a lighter upstream filter has already routed to it as possibly self-harm related. Roughly 50,000 posts a day land in that queue.
Knowledge spark: what's a golden set?
A pile of real, already-confirmed examples you test a model against instead of guessing. Not a random slice of everyday traffic, cases a trained human already checked and signed off as real. It's the ruler you measure the model with, and a ruler that's too short or too old tells you less than you think it does.
Before Wardline wrote reasons, a moderator read every routed post cold, with nothing to start from but their own judgment, and cleared about 40 posts an hour that way. With Wardline's draft judgment as a starting point, a reviewer now checks about 95 posts an hour, because they're correcting a draft instead of writing one from nothing.
Wardline's newest version was asked to do more: not just catch the obvious cases, direct statements, a clear method, a real timeframe, but also catch posts written to dodge a keyword filter, coded words, deliberate misspellings, slang that shifts every few months. Measured against 500 confirmed real cases, its catch rate came out to 91 percent.
One idea got raised early: just report overall accuracy across every post Wardline touches. It sounds reasonable until you run the numbers. Real self-harm violations are about 1.2 percent of the posts even this narrowed queue sends over. A model that flagged nothing, ever, would still be right 98.8 percent of the time. That number would look great on a slide and say nothing true about the one thing anyone's actually trying to catch.
Blended accuracy looks like a real number. On a queue this imbalanced, it's a number that never had to try.
Here's the turn. The 91 percent isn't really the story yet, because a single measured figure like that can't say, on its own, whether it's good enough, or where it's most likely to be wrong. The real question is what "accurate" is supposed to mean well enough that Ilse can sign it, Temperance can defend it, and a review board can act on it.
Wardline's catch rate, honestly ranged
Wardline's honest catch-rate rangeTwo-reviewer agreement, fixed at 89%
Most of the honest range sits at or above what two trained reviewers agree on with each other. The top edge, 94 percent, asks Wardline to be more consistent than two careful humans manage between themselves.
The word accurate wasn't wrong. It just hadn't decided anything yet.
What it costs at its worst: a launch doc that only says "the output must be accurate," with no number behind it, can go wrong in two opposite directions later. Ship on a bar that's actually too low, and a real at-risk post gets missed with nobody able to say afterward why the bar let it through. Or panic after one bad miss and quietly push the bar so high, without saying so, that thousands of ordinary posts get wrongly flagged into a review queue that was never built to handle that much noise.
The decision that mattered
Hearthloop's launch process asks every AI feature for one line: "[capability] must be accurate," signed by policy, no numbers required. That worked fine for older, simpler filters scored against thousands of clean examples. It stopped making sense the day a feature's real evidence was 500 hard-won cases and a rate, not a raw percentage with room to spare.
What I would leave alone: Wardline's spam and duplicate-post detection doesn't need any of this. Those categories have millions of confirmed examples and a threshold that's been stable for two years. A wide range there would just be noise dressed up as rigor.
The lesson: a launch gate that only asks for the word "accurate" isn't being careful. It's skipping the actual decision and hoping nobody asks where the number came from. The real discipline is showing the arithmetic, even when, maybe especially when, the honest answer is a range that makes everyone a little less comfortable than a flat 90 would.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why "the output must be accurate" is the sentence a launch review board should never let stand as written.
Every Thursday, Temperance Bramhall pulls the same report: 500 rows, one per confirmed self-harm post Wardline has ever scored, each one already checked and signed off by a trained human reviewer months before. Ask Temperance which of those 500 nearly got missed, and she can tell you the post number from memory.
For most of the last year, that report was routine. Wardline's older version only flagged the obvious cases, and it caught those well. Ilse's team signed off on "the output must be accurate" every quarter without much debate, because on the obvious cases, it basically was.
Then Wardline's newest version got asked to do more: catch the posts written to dodge a keyword filter too. The sign-off process didn't change with it. First quarter, nobody asked for a number, the launch doc just said "accurate," same as always. Second quarter, someone asked for a rough figure and got a percentage with no source attached. Third quarter, that percentage started showing up in slide decks as though it had always meant something specific.
Then a new hire on Ilse's team, working through a batch of missed posts, asked a plain question in a meeting: "When you say accurate, accurate compared to what?" Nobody in the room could answer her.
This is the exact gap the new hire's question opened up. Nobody could say what the golden set had, or hadn't, seen before.
Temperance went back through the number that had been quoted for two quarters running: 91 percent. Nobody could say where it came from. Not which 91 percent of what. Not how many real cases it was measured against. It had shown up once, sounded fine, and gotten repeated.
The real risk was never that 91 percent was wrong. It might even have been close. The real risk was that nobody could say what would have made it wrong, which meant nobody would notice if it quietly slid to 85, or if the posts it was missing were exactly the coded-language ones the model had barely been tested on at all.
Ninety-one percent wasn't the risk. Not knowing what would prove it wrong was.
The choice Temperance would take back happened two quarters earlier, in a much smaller meeting. Someone had asked for a number to put on a slide, fast, before a leadership review. Temperance pulled a rough figure off a dashboard, said "91 percent, roughly," and moved on, because the meeting had four more items and nobody pushed. That was the moment an estimate became a fact.
Three weeks after the new hire's question, the answer got sharper. A reviewer doing routine spot checks found a post that used a coded phrase Wardline's training data barely covered. It sat unflagged for six days before a user reported it directly.
Run the same new-hire question again, with the report Temperance built afterward already sitting in the room. This time, when someone asks "accurate compared to what," Temperance has it in one breath: 500 confirmed cases, a catch rate of 88 to 94 percent, checked against 89 percent human agreement on the same posts, and a tracked, separate number for coded-language catches alone, currently lower, currently being worked. Fifteen minutes, not a shrug.
One version of that meeting ends in an unsourced number nobody can defend. The other ends in a range everyone in the room can see the edges of.
What Temperance would tell her past self, the day she said "91 percent, roughly" and let the meeting move on: an estimate said out loud, once, under time pressure, doesn't stay an estimate. It becomes a fact the moment nobody writes down where it came from.
BOUND, or turning one soft word into four honest numbers
Not a way to make "accurate" sound more technical. BOUND is what forces the real arithmetic hiding behind that word, and makes you prove the range survives being compared to a real human being.
BBreak it down. What does "accurate" actually need?
A test set of confirmed real cases. A scoring method that keeps "how much of the real stuff it catches" separate from "how much of what it flags is real," because those cost different people different things. A pass line. And a stated range around that line, sized to how many real cases the test actually has.
Say the shape before naming a figure, or "we tested it" quietly stands in for arithmetic nobody actually did.
Four decisions hiding behind one word. Sign off on "accurate" as written, and none of these four actually got made.
OOwn the numbers. Where did each one come from?
Test set: 500 confirmed self-harm posts, pulled from a year of cases a trained reviewer already signed off as real, plus 4,000 ordinary posts pulled at random from the same queue to measure how many wrongful flags show up in real traffic, not just against the hard cases. Why 500, not 50 or 5,000: fifty gives too wide a range to say anything useful, five thousand would mean waiting past the ship date for a category this rare to accumulate that many confirmed cases, five hundred is what a real year of reviewed history actually produced. Threshold: pass on how much of the real stuff it catches, not the blended figure, because missing a real at-risk post and wrongly flagging a safe one aren't the same cost to the same people. And the number gets checked against something that already exists on its own: Hearthloop's quarterly reviewer-agreement audit, run independently of this launch.
Owning a number means saying where it came from, not just stating a figure that sounds right.
UUse a range, not one point.
A measured catch rate of 91 percent, on its own, claims more precision than 500 cases can actually back up. Run the real math: at that sample size, the honest interval sits about three points either side, 88 to 94. Ship the range. A single point on a launch doc looks more decisive and is actually less honest, because it hides how much the number could move if the next 100 confirmed cases came in a little different.
This is the direct answer to the question in one step: a flat "must be accurate" is exactly the false precision this step exists to strip out.
NNail the sanity check. Does it survive a real comparison?
Two trained reviewers, scoring the same 500 posts independently, agree with each other 89 percent of the time. That's not a flaw in the reviewers, genuinely ambiguous posts divide careful judgment. It means the top of Wardline's range, 94 percent, already asks the model to be more certain than two trained humans manage between themselves. Published work on similar detection tasks tends to land in a similar 85 to 90 percent band too, for what it's worth as an outside check. A flat 90 percent floor, taken literally, quietly commits to clearing a bar that isn't really where human judgment sits either.
The hardest step, and the one that gets skipped under deadline pressure. A number that sounds confident can still be sitting on nothing you'd want to defend.
DDirection. Which assumption would move it most?
Not the pass line. Nudging 88 to 94 up or down half a point barely changes the real picture. It's the test set's coverage of coded language. Only 60 of the 500 confirmed cases use it, about 12 percent. The golden set is built from a year of confirmed cases, so it's naturally weighted toward how the world looked for most of that year. Coded language was rarer for most of it, and has surged the last two quarters. That's exactly why the golden set's 12 percent share understates today's 34 percent, not bad luck, history and today have quietly stopped matching.
Naming the assumption that's both uncertain and consequential, not just the biggest number in the equation, is what a good estimator does that a bad one skips.
Coded-language coverage sits exactly where a good estimator worries most: least grounded, most consequential.
Coded language's share of real violations, last five quarters
Coded language, share of confirmed violationsGolden set's own coded-language share
The real share has climbed from 14 percent to 34 percent in five quarters. The golden set, built cumulatively over that same stretch, still sits at 12, closer to where the line started than to where it is now.
Three things worth naming directly, since this is where the real judgment sits. The AI-specific failure mode is a kind of silent drift: people trying to post self-harm content learn what gets caught and adjust their wording, so the exact posts Wardline is weakest on today are the ones getting actively pushed toward. Nothing in the model itself raises a flag saying "this phrasing is new." The guardrail is a refresh cycle, not a one-time test: every quarter, new confirmed cases get pulled specifically from posts that got missed by Wardline and later caught by a user report, and the coded-language catch rate gets tracked as its own number, separate from the headline figure, so a shrinking gap in coverage shows up before it shows up as a real miss. The trade-off is real too: raising the catch-rate floor to be safer pulls down how much of what gets flagged is actually real, which sends more ordinary posts into a review queue built for a smaller, calmer volume, costing reviewer hours and, some of the time, a wrongly flagged post sitting in front of someone who did nothing wrong. Nobody gets a higher catch rate and a quieter queue for free.
The whole answer in one picture. A number that's confident about the easy cases is not the same thing as a number that's confident, full stop.
And if you want to be sure it really works, try it somewhere else
Same five letters, a wheat field instead of a feed. This time the rare thing isn't a coded phrase. It's a disease strain regional labs only started confirming a season ago.
Loamfield Agritech's Blightscan reads a photo of a wheat leaf plus a few field notes and writes back a diagnosis: which disease, how confident, and a plain-language treatment step, before a farmer loses more of the field waiting on a lab result. Blightscan has a strong record on the blight strains it's tracked for years. Wart-leaf rust is different: a newer strain regional agronomists only started confirming eighteen months ago, so barely a season of confirmed cases exist anywhere.
Same method, a different lab. Both times, the golden set's own shape is quietly narrower than the real world it's supposed to stand in for.
Corwin Vessling owns Blightscan's numbers, and runs the same five letters. Break it down: a test set of confirmed wart-leaf rust cases, catch rate scored apart from false-alarm rate, since telling a farmer to spray for the wrong disease costs a season's chemical budget and missing the real one costs the field. Own the numbers: 140 confirmed cases, gathered across two regional extension offices over the one season the strain's been tracked, catch rate measured at 82 percent. Use a range: at only 140 cases, the honest interval runs wider than Hearthloop's, about six points either side, 76 to 88, not a flat 82.
Nail the sanity check: a senior agronomist working from the same photos, with no tool at all, correctly identifies wart-leaf rust about 79 percent of the time, since it looks like two older, more common strains early on. Blightscan's range mostly sits above that human baseline, a genuinely good sign here, the opposite read from Hearthloop's case, and worth saying plainly rather than assuming a wide range always means bad news. Direction: the swing assumption isn't the sample size, next season will roughly double it either way. It's regional spread. All 140 confirmed cases came from two neighboring counties. Wart-leaf rust reacts differently in the drier soil three counties over, where Blightscan has scored exactly zero confirmed cases. A range built entirely on one kind of ground says nothing about ground it's never seen.
The decision Corwin would take back
Loamfield's rollout plan assumed a new disease model was ready everywhere the day it cleared testing in its home region. That worked for strains that behave the same across soil types. It stopped working the day a strain's entire confirmed record came from land with one kind of dirt.
Same rank, different lever: the fix for Blightscan isn't a bigger physical-testing team. It's the same habit Temperance ran on Wardline's numbers: source the test set honestly, range the number that's genuinely uncertain, and watch coverage, not the size of the confirmed count, as the assumption that actually decides whether the range means anything somewhere new.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: test on confirmed real cases, score catch rate and false-alarm rate apart, give the pass line a range sized to the sample, and check it against a real baseline before calling it done.
Cost: no budget this quarter to wait a full extra season for more confirmed cases. Ship the honest interim version, the wider 76-to-88 range, labeled clearly as regionally narrow, updated the moment cases from different soil types come in.
The model got better, for real: say Blightscan's next version doubles its confirmed case count overnight. The range narrows, maybe to 80 to 86. It doesn't earn a flat single number. A bigger sample buys a tighter range, not certainty.
Where people run it wrong.
They report one confident number because a range looks like doubt, the exact habit that turned Hearthloop's 91 percent into an unsourced fact for two quarters.
They test only on the easy, common cases and let the rare, evolving ones ride on a number that was never measured against them.
They treat a wide range as automatically bad news, instead of checking what it's actually being compared to.
How to use it live. Ask the coverage question before naming a number: "what does the test set actually look like, and does it match the cases we're most worried about, or the easy ones we already had plenty of?" That question alone usually tells you whether a number is measuring the real risk or just the risk that was convenient to collect.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: turn a soft word like accurate into real arithmetic, an honest range, and a check against something real. Built for estimation questions, not a habit-flip story.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Temperance Bramhall, who owns Wardline's numbers at Hearthloop, and has to turn "the output must be accurate" into something a launch review board can actually sign.
3 · THE BLIND SPOT
What part of the real threat does the 500-case test set barely cover?
Tap to flip
ANSWER
Coded and evasive language. Only 60 of 500 confirmed cases use it, about 12 percent, while coded language made up 34 percent of real violations last quarter and is climbing.
4 · THE EQUATION
What four things does "accurate" actually need to mean something?
Tap to flip
ANSWER
A test set of confirmed real cases, catch rate and false-alarm rate scored apart, a pass line, and a range around that line sized to the sample.
5 · THE OLD DECISION
What decision would Temperance take back?
Tap to flip
ANSWER
Saying "91 percent, roughly" out loud in a fast meeting with no source attached, letting a rough guess get repeated for two quarters as though it were a measured fact.
6 · THE NUMBER
Fill in the blank: Wardline's honest catch-rate range is ___ to ___ percent. Two trained reviewers agree with each other ___ percent of the time.
Tap to flip
ANSWER
88 to 94 percent. 89 percent human agreement, meaning the top of Wardline's own range already asks it to beat what two careful humans manage between themselves.
7 · THE REPLAY
Same new-hire question, asked again with the real report in the room, what changes?
Tap to flip
ANSWER
Temperance answers in one breath: 500 confirmed cases, an 88 to 94 range, checked against 89 percent human agreement, and a separate tracked number for coded-language catches. Fifteen minutes, not a shrug.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which one, and what's the swing assumption there?
Tap to flip
ANSWER
Blightscan, Loamfield Agritech's crop-disease diagnosis tool. The swing assumption is regional spread: all 140 confirmed wart-leaf rust cases come from two counties with one kind of soil, so the range says nothing about drier ground three counties over.
Check yourself Score: 0 / 0
Fill in the blank
1. Wardline's honest catch-rate range for self-harm posts is ___ to ___ percent, built from a test set of ___ confirmed cases.
Show hint
Check the U step and the O step in the BOUND recap.
Show answer
88 to 94 percent, from 500 confirmed cases. A measured catch rate of 91 percent at that sample size carries an honest interval of about three points either way.
Multiple choice
2. Why does Temperance score "how much of the real stuff it catches" separately from "how much of what it flags is real," instead of one blended accuracy number?
A. Because blended accuracy is harder to calculate.
B. Because missing a real at-risk post and wrongly flagging a safe one cost different people different things, and a blended number can look great by never flagging anything.
C. Because Trust and Safety policy requires two separate metrics by law.
D. Because Wardline's confidence score already includes both numbers.
Show hint
Look at the rejected "report overall accuracy" idea in Let's learn.
Show answer
B. With real violations at about 1.2 percent of the queue, a model that never flags anything would still score 98.8 percent blended accuracy and catch nothing real.
True or false
3. True or false: since 91 percent is Wardline's best single measured catch rate, writing a flat 91 percent as the pass line in the launch doc would have been just as honest as writing the 88-to-94 range.
True
False
Show hint
Look at what the U step says a single number claims that a range doesn't.
Show answer
False. A flat number claims more precision than a 500-case sample supports. The range is what shows the launch board how much the number could move, and that its top edge already sits above human agreement.
Short answer, where it wouldn't matter
4. Name a place in Wardline's own work where this exact range-and-recheck treatment would NOT be worth building.
Show hint
Look at "what I would leave alone" in Let's learn.
Show answer
Model answer: Wardline's spam and duplicate-post detection. Those categories have millions of confirmed examples and a threshold that's been stable for two years, so a wide range there would be noise, not rigor.
Short answer, apply it yourself
5. Think of an AI feature you've used or built where someone wrote a launch requirement like "must be accurate" or "must be reliable." What are the four numbers hiding behind that word in your case?
Show hint
Think about what the test set actually was, which kind of mistake mattered more, where the pass line sat, and how big the test set really was.
Show answer
Model answer: A team shipping an AI resume screener wrote "must be fair and accurate" with no further detail. The four hidden numbers were: a test set of a few hundred past hiring decisions with known outcomes, catch rate for qualified candidates versus false-alarm rate for unqualified ones scored apart, a pass line for each, and a range, since the test set only covered one hiring season.
Short answer, work the number
6. If Wardline's next quarterly test set had 2,000 confirmed cases instead of 500, at the same measured catch rate of 91 percent, would the honest range likely be wider, the same, or narrower than 88 to 94? Roughly how wide?
Show hint
The interval's width shrinks as the sample size grows, roughly by the square root of how many times bigger the sample gets.
Show answer
Narrower, roughly plus or minus one and a half points, about 89.5 to 92.5. Going from 500 to 2,000 cases is four times the sample, and the interval's half-width shrinks by about the square root of four, so about half as wide, from roughly plus or minus three points to roughly plus or minus one and a half.
Before you close the answer
Why this works
Tests whether you'll turn a soft launch requirement into real arithmetic instead of accepting the word at face value, and whether you know a confident single number is a warning sign when the sample behind it is only a few hundred cases.
Follow-up traps
"Isn't 88 to 94 just a wide way of saying you don't know?" Response: no, it's the honest width a 500-case sample actually supports. A flat 90 would look more certain and would be less true, and the range still gives the board a real floor to hold Wardline to.
"If two humans only agree 89 percent of the time, why hold the model to any bar at all?" Response: because 89 percent is itself a number worth watching, not a reason to give up on measurement. It shows that the top of Wardline's range asks for more certainty than the task itself has, which is exactly what a flat "must be accurate" line would have hidden.
If pressed
The 500-case test set isn't drawn as a simple random sample of confirmed violations. It's stratified to include every known evasion pattern at least a handful of times, then weighted back down when the catch rate gets reported, so a single common pattern with a hundred easy examples can't quietly inflate the headline number.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.