ConceptAdvancedEval-Driven Specification / Acceptance criteria for non-deterministic output / #6
Describe acceptance criteria that account for the severity of different error types.
The direct answer
Write the release bar around fire-weather severity, not one pooled accuracy score. Require near-perfect recall on the highest-severity tier, the conditions where a missed flag can start a fire, measured on a set stacked with real high-risk examples, and gate every model update on that number before a pooled score gets reported at all. A single blended number lets thousands of calm, low-stakes days hide the handful of dangerous ones.
How to write the criteria, in order
Set the release bar on the highest fire-weather-severity tier first, not the pooled average.Why: a pooled number blends 999 calm segment-scores against one extreme one, and the extreme one always loses.
Build the extreme-tier test set by stacking it with confirmed high-risk segment-hours, not by sampling it naturally.Why: at a natural rate of about 1 in 1,000, a random sample barely holds enough real extreme cases to measure recall on at all.
Route every segment near the extreme-tier threshold to a person, regardless of score, until the bar is actually met.Why: that is exactly where a missed flag turns into an energized line over dry brush in high wind, not a shrug the next morning.
Say the cost asymmetry out loud in real dollars, not in feelings.Why: "$14,000 against $210 million" is a number someone can argue with. "We should be careful" is not.
Track extreme-tier recall against pooled accuracy every fire season, not once at launch.Why: 68 percent hid inside a 92 percent pooled score for a full season before anyone measured it on its own.
Leave the criteria for calm and moderate days on the pooled bar, untouched.Why: getting a calm-day call slightly wrong costs a few hours of inconvenience, and tightening that bar further only triggers more unnecessary shutoffs for no safety gain.
Six moves for saying the bar out loud
A severity question is a tradeoff wearing a spec's clothes, so PICK does the work here, not a checklist of rubric fields.
1
Ground it in one real system before naming the framework
Say it like this
"Let's say I'm writing the release bar for Emberline, the tool that scores how likely a stretch of power line is to start a fire. That's the feature. That's the bar I'm going to write."
Why this works
Stops the answer floating at "AI evals in general" and gives the interviewer something concrete to push on.
2
Name the four moves before making any of them
Say it like this
"I'm going to commit to a position first, say who feels each kind of mistake and in what units, name why one kind is worse than the other, and then say what would change my mind. Four moves, in that order."
Why this works
Signals a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Reframe what the question is really testing
Say it like this
"This isn't really 'write me a rubric.' It's 'do you know a single accuracy number can hide the one mistake that actually starts a fire.' So I'm not opening with 92 percent overall. I'm opening with the one class of miss that can hurt someone."
Why this works
Shows the interviewer you see past the surface ask to the real judgment being tested.
4
Give the position, with the actual mechanism in it
Say it like this
"My position: before any Emberline update ships, it has to hit at least 99.5 percent recall on the extreme fire-weather tier, Red Flag Warning conditions, tested on a set we deliberately stack with confirmed high-risk segment-hours. Overall pooled accuracy still gets reported, but it never gates a release on its own."
Why this works
PICK rewards committing to a real number and a real mechanism, not a vague promise to "be careful."
5
Name who feels each mistake, then prove the asymmetry with a real case
Say it like this
"Here's the split. A false shutoff on a calm day costs about $14,000, mostly customer credits, and everybody knows about it by dinner. A missed segment during a Red Flag Warning doesn't get caught until a line arcs into dry brush, and that's the shape of the 2019 fire this utility settled for about $210 million. We actually shipped a version that scored 92 percent pooled accuracy and only caught 68 percent of real extreme-tier risk, because those conditions are maybe one segment-score in a thousand, and a dispatcher caught the one that slipped through on a hunch, not because the tool flagged it."
Why this works
The real numbers and the real near miss make the asymmetry checkable, not just asserted.
6
Name the kill criteria and close on the one line
Say it like this
"I'd drop the special bar the day extreme-tier recall, checked every fire season, catches up to the pooled accuracy line. And I'd leave the calm-day criteria exactly where they are, because getting those slightly wrong costs a few hours, not a canyon."
Why this works
Shows confidence without stubbornness, and ends on the line the interviewer should walk away remembering.
One more thing before the walkthrough moves on: this bar covers one severity tier, not the whole spec. Most candidates hear "the worst case matters more" and answer by tightening everything across the board. Say which tier gets the strict bar and which one stays on the pooled average, and you've shown real judgment instead of reciting a checklist.
Let's learn
Every Red Flag Warning morning, grid engineers at Talus Ridge Power & Light used to spend about three hours deciding which of the utility's 1,800 line segments to shut off, checking wind and dryness reports by hand for the riskiest 380 of them.
Then came Emberline, a tool that reads wind speed, humidity, how dry the brush is, and each segment's equipment history, and scores every one of those 1,800 segments for fire risk every fifteen minutes.
One score, one dispatcher, one call before dawn
Knowledge spark: what's a Red Flag Warning?
A weather service call for a day dry and windy enough that almost any spark can turn into a fast-moving fire. Low humidity, strong wind, dry brush, all at once. It's rare. Most days aren't one.
With Emberline sorting segments and flagging only the ones worth a closer look, usually around 40 a morning, that three-hour review fell to about 20 minutes.
Here is the turn. Those 20 minutes were never the risk. What was hiding inside the number that let Emberline ship was. The team had set one release bar: overall fire-risk classification accuracy above 92 percent, pooled across every segment-score, calm days and Red Flag Warning mornings, all averaged together. Emberline cleared it at 92 percent. Nobody asked, separately, how good it was on the mornings that actually matter.
We did not build a bar for how good it looks on an average day. We built one for the one morning it cannot afford to miss.
At its worst, that looks like this: a segment under real Red Flag Warning conditions, dry brush, wind past 35 miles an hour, scores clear, the line stays energized, and a fault arcs into the brush below it. Or, in the version nobody wants to say out loud, nobody catches it until the smoke is already visible from the highway.
The choice I would take back
Early on, the team picked one release number, 92 percent pooled accuracy, because it was simple to explain and simple to gate a launch on. I would take that back. Confirmed Red Flag Warning segment-scores make up about 1 in 1,000 of everything Emberline scores in a season. Any single number that averages across all of them is really just a calm-day score wearing a fire-day score's name.
Cost per event, in dollars
Felt the same afternoon, cheap
Rare, and it starts a fire, expensive
Emberline wrongly flags a calm, low-wind segment
$14,000
Emberline wrongly clears a segment in Red Flag Warning conditions
$210M
The cheap bar is barely there on purpose. $14,000 covers customer credits and crew redeployment for about 1,200 customers, out for a few hours, folded into the shift. $210 million is what the utility's 2019 Hollow Creek Fire cost to settle, the shape of what a missed segment risks again. Roughly 15,000 times the cost, and it doesn't show up until a line has already arced.
What I would leave alone. The criteria for ordinary, low-wind days. Those are common, cheap to get slightly wrong, and a false flag there costs a dispatcher a few minutes and a customer a few hours without power. Tightening that bar to 99.5 percent too would only trigger more shutoffs nobody needed.
The lesson. We wrote the criteria to answer "is this tool accurate," and it answered that well. We never wrote a separate bar for "is it accurate on the one kind of morning that can start a fire." Those are two different questions, and only the second one has a canyon full of houses riding on it.
The wind report Junko almost waved off
You don't need this to answer the question. Read it if you want to feel why the bar has to sit on the rare tier, not the average.
The control room at Talus Ridge Power & Light goes quiet a little before 4 a.m., which is exactly when the wind starts moving over the Widow's Peak corridor. Junko Hirano has run the overnight shutoff desk for eight years. Ask her which segment will be trouble before dawn and she can usually tell you before her coffee's done, corridor 14, always corridor 14, because the canyon funnels the wind straight down it.
Emberline arrived the spring after the 2019 Hollow Creek Fire, and for most of two seasons it was just good. A segment scored in under a second, flagged or clear, and the flagged ones were almost always right. She'd still call the field crew to confirm conditions on anything scored clear near corridor 14, out of habit, the way she'd done it for years before the tool showed up. It agreed with her. Every time.
So she started calling on half of them. Then only the ones Emberline flagged. By the second season, if the screen said clear, she moved on to the next segment.
Same tool, two very different kinds of wrong
Then, one morning in October, she was about to clear corridor 14. Emberline had it scored clear, nothing on the screen looked wrong. But the field radio crackled with a downed-branch report two corridors over, wind gusts already past what the seven a.m. forecast had called for. She paused, and made the call to hold the segment and send a crew out anyway, the way she used to on every uncertain morning.
The crew found a fault on a splice fitting, arcing, sitting eight feet above brush dry enough to have gone up in seconds if the wind had shifted. Junko hadn't caught it because Emberline told her to look. She caught it because eight years of habit hadn't fully left her, and a downed-branch report happened to still be in her ear that morning.
We did not just miss one bad segment. We built a bar that could not tell us we were missing them, because the number it reported was mostly a calm-day score.
Halvard Skeie had owned Emberline's release criteria since before it shipped, six years into a career writing acceptance bars for grid safety equipment before AI ever entered the picture. When Junko's near miss reached him, the header number still looked fine. Pooled accuracy, checked that same week, read 92 percent. He almost logged it as one lucky catch.
Then he pulled the corpus the launch decision had been made against: about 600,000 segment-scores across the prior season, only 600 of them under confirmed Red Flag Warning conditions. One in a thousand. He reran the recall number on those 600 alone. Emberline had correctly flagged 408 of them. Sixty-eight percent.
Sixty-eight percent had been sitting inside a 92 percent average the entire time, because 600 extreme-tier scores can barely move a number built from 600,000.
So Halvard split the bar in two. Before any pooled accuracy number counts for anything, Emberline has to clear 99.5 percent recall on a test set built to actually hold enough confirmed Red Flag Warning segment-scores to measure it, stacked well past the natural one-in-a-thousand rate. Calm and moderate days stay on the old pooled bar, since getting those slightly wrong has never cost anyone more than a few hours without power.
The part he'd take back isn't the launch. It's picking one number, simple to explain, simple to gate a release on, and never asking whether that one number could actually see the mistake that mattered. A number that good on 600,000 segment-scores hid exactly the tier it needed to prove itself on.
Junko's flagged-segment review now runs about 25 minutes a shift, five minutes more than the old average-based bar, because the new one flags a few extra borderline calm-day segments on purpose rather than risk waving off a borderline fire-day one. And corridor 14 can't slip through silently the same way again, because nothing scoring below the new bar on a Red Flag Warning morning reaches her marked as clear.
The four letters, and why this isn't a rubric
This is a tradeoff dressed as a spec, so PICK is the tool. A "how would you measure whether this feature is working" question would reach for LEAD instead.
P, position. Gate release on the extreme fire-weather tier first, recall at or above 99.5 percent on a set stacked with real high-risk examples. Overall pooled accuracy is reported alongside it, but it never gates the release by itself.
I, impact. A dispatcher wrongly told to shut off a calm segment costs about 1,200 customers a few hours without power and the utility about $14,000 in credits, absorbed into the day. A missed extreme-tier segment reaches a live spark under dry brush, and the shape of that cost, based on the utility's own 2019 fire, is about $210 million plus a corridor of evacuated homes.
C, cost asymmetry. Confirmed Red Flag Warning segment-scores are about 1 in 1,000 of everything the model sees in a season. Any single, pooled number will always be dominated by how well the tool handles calm days, never by how well it handles the fire-weather morning that actually matters. Spend the caution on the tier the average can hide, not on the one it already measures well.
K, kill criteria. Drop the special bar the day extreme-tier recall, tracked every fire season, catches up to the overall pooled accuracy line. At that point the model has no real blind spot left for an average to hide, and one bar measures the same thing either way.
Knowledge spark: why can't you just raise the pooled bar instead?
Because the pooled bar is a blend. Raising it from 92 to 99.5 percent barely touches the extreme tier, since those mornings are 1 in 1,000 segment-scores and can't move a pooled average much either way. You'd mostly be squeezing harder on calm days, the part that was never the problem.
Extreme-tier recall, measured each fire season, against the kill line
Extreme-tier recall, on a growing confirmed set
Kill line: matches overall pooled accuracy, 92%
Four fire seasons of retraining closed most of the gap, from 68 percent up to 91 percent, measured against a confirmed extreme-tier set that grew from 600 to 2,300 segment-scores as more real cases came in. It still sits just under the 92 percent line where it would match overall pooled accuracy. Until that line is crossed, the separate bar for the extreme tier stays exactly where it is.
Run the four letters again, on a hospital ward
Lindenfield General Hospital runs a model, VitalArc, that reads a patient's vitals and labs continuously and flags a rising risk of deterioration for the charge nurse. Same shape of question, a different ward, a different kind of harm.
P. Gate any release of VitalArc on its recall for one tier, a patient heading toward septic shock within the next four hours, at 99 percent or higher, on a set stacked with confirmed real cases. General deterioration-score accuracy is reported, but it never gates a release by itself. I. A false alarm on a stable patient costs a nurse about ten minutes to recheck and clear, folded into a normal round. A missed case of oncoming septic shock isn't caught until the patient's own vitals crash on their own, and by then it's an emergency response, days added to the stay, and in the worst case, a death. C. Confirmed septic-shock progressions are a tiny slice of every monitored patient-hour next to ordinary stable readings, the same shape of imbalance as Emberline's extreme fire days. A pooled deterioration-accuracy number will always be dominated by how well the tool reads ordinary vitals, never by how well it catches the rare crash. K. Drop the special bar the day septic-shock-tier recall, tracked quarterly, catches up to the pooled deterioration accuracy. Until then, any patient flagged near the threshold gets a nurse's eyes regardless of score.
What I would leave alone, on the ward
The routine vitals criteria, a heart rate a little high after a walk to the bathroom, a temperature a touch elevated on a warm afternoon. A false alarm there costs a nurse a glance and nothing else. No patient has ever been hurt by that kind of miss.
Swap the trigger and it still runs
Speed: Emberline takes three seconds to score a segment instead of one. Doesn't move the bar, because the bar is about what gets missed, not how fast the tool decides.
Cost: running Emberline gets five times more expensive per segment. Also doesn't move it. A missed extreme-tier segment costs thousands of times a single scoring call, whatever that call happens to cost.
The model gets better: if extreme-tier recall genuinely closes the gap to pooled accuracy, per the kill criteria, move the bar, don't erase it. The two numbers becoming the same thing is exactly the evidence that would let one bar do the job.
Where people run it wrong
Treating a strong pooled accuracy number as proof of safety, when the rare tier was never measured on enough real examples to say anything about it at all.
Fixing it by making the whole model "stricter," which only changes which calm-day segments get flagged, not whether it catches the fire-weather morning.
Writing "add a human review step" without saying which segments get it, so the review time ends up spent re-checking easy calm days instead of the ones that actually needed a person.
If you are asked this cold
Say the reframe out loud before you list a single rubric field. "Give me a second, I want to separate what each kind of mistake actually costs before I pick a number." That's true, it's already stage one of the walkthrough, and it buys the time to find the real asymmetry instead of reciting a generic checklist.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what's its hardest step?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a pooled average will always favor the common, cheap mistake over the rare, dangerous one, and building the bar around the dangerous one instead.
2 · THE PERSON
Who owns the release criteria in this answer, and what did he find that the launch number missed?
Tap to flip
ANSWER
Halvard Skeie, the product manager who owns Emberline's release criteria. He found that a 92 percent pooled accuracy number was hiding a 68 percent recall on real extreme fire-weather conditions, because those mornings were only 1 in 1,000 segment-scores in the corpus behind the number.
3 · THE HABIT
What did Junko Hirano stop doing once Emberline kept agreeing with her?
Tap to flip
ANSWER
Calling the field crew to confirm conditions on any segment Emberline scored clear near corridor 14. She moved from checking all of them, to half, to only the ones Emberline flagged, to trusting a clear screen outright.
4 · THE ASYMMETRY
Name the two kinds of mistake here and what each one costs.
Tap to flip
ANSWER
A false shutoff on a calm segment: felt the same afternoon, about $14,000. A missed segment in Red Flag Warning conditions: rare, but if it starts a fire, about $210 million, based on the utility's own 2019 Hollow Creek Fire settlement.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Gate release on extreme-tier recall at 99.5 percent or higher, tested on a set stacked with real high-risk segment-hours, and only report overall pooled accuracy after that bar is met, never instead of it.
6 · THE NUMBER
When Junko's near miss happened, Emberline's recall on real extreme-tier segments alone was ______ percent, hidden inside a 92 percent pooled score.
Tap to flip
ANSWER
68. Four hundred eight of 600 confirmed Red Flag Warning segment-scores caught, missed by nearly a third, and none of that showed up in the one number the launch decision was made against.
7 · THE KILL CRITERIA
What would make you drop the special bar and go back to one pooled number?
Tap to flip
ANSWER
The day extreme-tier recall, tracked every fire season, catches up to the overall pooled accuracy line. At that point the model has no blind spot left on the rare tier for an average to hide, and one bar would measure the same thing either way.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where's the strict bar there?
Tap to flip
ANSWER
Lindenfield General Hospital's VitalArc, reading patient vitals on a ward. The strict bar sits on recall for patients heading toward septic shock within four hours, not on general deterioration-score accuracy.
Check yourself Score: 0 / 0
Multiple choice
1. Which of these is the actual mechanism behind this answer's criteria bar?
A. Make Emberline score every segment faster so dispatchers see flags sooner.
B. Gate release on recall for the extreme fire-weather tier alone, measured on a set stacked with real examples, before any pooled accuracy number counts.
C. Require dispatchers to call a field crew on every single segment, regardless of what Emberline says.
D. Raise the overall pooled accuracy bar from 92 to 99.5 percent.
Show hint
Three of these either don't touch the rare tier at all, or throw away the time the tool was built to save.
Show answer
B. A wastes effort on the wrong problem, C undoes the whole point of the tool, and D barely moves the extreme-tier number since those mornings are 1 in 1,000 segment-scores in a pooled average. Only B actually targets the mistake that matters.
True or false
2. True or false: once Emberline clears 92 percent pooled accuracy across all conditions, that alone is enough to ship it.
True
False
Show hint
Think about what 92 percent pooled accuracy was actually hiding when Junko's near miss happened.
Show answer
False. That exact number, 92 percent, shipped alongside a 68 percent recall on real extreme fire-weather conditions. Pooled accuracy alone was never enough, and the answer's whole position is that the extreme-tier bar has to be checked and cleared separately.
Fill in the blank
3. Fill in the blank: confirmed Red Flag Warning segment-scores made up about 1 in ______ of everything Emberline scored in a season, which is why a pooled number couldn't see the extreme-tier problem.
Show hint
600 confirmed extreme-tier segment-scores, out of a 600,000-score season.
Show answer
1,000. 600 confirmed Red Flag Warning segment-scores out of 600,000 total is exactly 1 in 1,000, small enough that it barely moves a pooled average no matter how badly the model handles it.
Multiple choice
4. Why couldn't the team just raise the overall accuracy bar to 99.5 percent instead of writing a separate extreme-tier bar?
A. Raising the bar would take too long to retrain for.
B. A pooled bar is dominated by the common, calm-day conditions, so raising it mostly squeezes harder on calm days and barely touches extreme-tier recall.
C. Dispatchers refused to work under a stricter number.
D. Regulators require exactly 92 percent, no more and no less.
Show hint
Ask what a "stricter" version of the same pooled number would actually be measuring more of.
Show answer
B. Turning the pooled dial up just makes the model better at the conditions that already dominate the average, calm and moderate days. It never forces the model to prove anything about a tier that's only 1 in 1,000 segment-scores. The fix has to measure that tier on its own, not tighten the blend it's hiding inside.
Short answer
5. If confirmed extreme fire-weather scores were 1 in 10,000 instead of 1 in 1,000, would splitting the criteria still make sense? Walk through it.
Show hint
Think about how much harder a pooled number gets to move as the rare tier gets even rarer.
Show answer
Yes, even more so. At 1 in 1,000, extreme-tier segments already barely moved a pooled average. At 1 in 10,000, a model could get almost every extreme case wrong and the pooled number would hardly change. The rarer the dangerous condition, the more completely an average metric hides it, which makes the case for a separate, tier-specific bar stronger, never weaker.
Short answer, apply it yourself
6. Pick a place in your own work where one overall score covers several different kinds of mistake. Name the rare, expensive one hiding inside it, and what number you'd pull out to gate on separately.
Show hint
Look for the mistake type that's both the rarest in your data and the one you'd least want to explain to your manager after the fact.
Show answer
Model answer: "Our fraud-screening model scores 96 percent overall accuracy across all transaction types. Buried in that average is wire transfers over $50,000, maybe 1 in 3,000 transactions, where a missed flag means the money is gone before anyone notices. I'd pull recall on that one category out as its own gate, tested on a set stacked with real high-value fraud cases, and never let the 96 percent overall number stand in for it." Any answer works if you can name the rare category the overall score is currently hiding.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.