Write criteria for a feature that must degrade gracefully when confidence is low.
- Add a real confidence gate: below a stated number, a post can't reach full distribution, full stop.Why: right now "safe" is one flat label, so a 54 percent guess gets ranked exactly like a 99 percent certainty.
- Give the held posts a real second look inside a stated window, not an open-ended queue.Why: a hold with no deadline just becomes a slower, quieter version of the same problem.
- Tag anything that clears the hold as "cleared under review," never blended into the same log line as an automatic high-confidence pass.Why: nobody, not a moderator, not a parent request, not a regulator, can ever ask "was this one ever a coin flip" if both cases look identical afterward.
- Report the uncertain rate weekly, split by content category, not as one blended pass rate.Why: a category can sit at barely-better-than-a-guess for months while the overall number looks completely healthy.
- Leave categories that reliably score high confidence exactly as they are.Why: gating content the model is already sure about only slows the feed down for no safety gain at all.
How to answer this, stage by stage
Six moves, from pinning the feature to one product to the line you'd close on.
Let's learn
Here is what happens when a safety check can only say yes or no, and never gets to say "I'm not sure."
Glasswing is a short-video feed built for teenagers. Every post that goes up gets scored by a safety classifier before the ranking system will let it reach anyone's main feed, checking for things like coded self-harm content, extreme dieting content, and dangerous challenges.
Before the classifier existed, Glasswing was small enough that a five-person trust team manually reviewed every post its keyword filters flagged, about 200 a day, each one looked at within twenty minutes before it could go live. The classifier replaced that. Today it auto-clears roughly 50,000 borderline-adjacent posts a day, and anything it scores "safe" gets zero review time, whether the number underneath was 96 or 54.
That speed isn't the problem. Clearing 50,000 posts a day instead of 200 is the whole reason the product works at teenager scale. The problem is what "safe" quietly stopped meaning somewhere in that jump.
Here's the part that actually matters. A 58 percent confidence score is close to a coin flip. Treating that the same as a 94 doesn't just risk one wrong call, it means the system was never honest about which calls it was actually sure of.
At its worst, this doesn't look like a scandal. It looks like a normal Tuesday. A teen who's already been engaging with dieting content gets a coin-flip video pushed to her at full strength, because the category she's drawn to is exactly the category the model is least sure about. She engages, the feed learns from that, and it sends her more. Nobody flagged it, because nothing about "safe" ever changes shape depending on how sure the model actually was.
What I would leave alone. General entertainment and sports content don't need a gate. They sit reliably at 89 to 94 percent confidence, week over week. Holding those posts back for review would slow the whole feed down for content that was never actually the risk.
The lesson. A check that can only say pass or fail is hiding its most useful number. If a system can be unsure, "graceful" has to mean the product actually does something different when it is, not that it fails a little more politely while doing the exact same thing.
Now here is the same thing as a story
The short version is above. Read this one when you want to feel why the ninety-minute window actually matters.
Talia Bregman built the ranking pipeline's safety gate when Glasswing had maybe forty thousand users and one content category worth worrying about: kids reposting stunts they'd seen on bigger platforms. Back then the classifier was almost boringly confident. Ninety-two, ninety-six, ninety-eight percent, week after week. In a meeting with her ranking-eng lead, the two of them agreed to keep the pipeline simple: read the pass/fail label, ignore the confidence float, because wiring the number in added real complexity for a case that basically never came up.
For two years, that held. Talia didn't think about it. Nobody did. The blended daily "safe" rate sat around 96 percent, and 96 percent looked like a number you could trust and move on from.
Glasswing grew. Content diversified, the way it does, into fitness edits, grief pages after a public tragedy went viral, "what I eat in a day" videos with captions that used coded, euphemistic language to slide past keyword filters. None of that showed up as a spike. It showed up as a slow widening of the range of things the classifier had to guess at.
Nobody noticed, because nobody was watching anything except the one blended number.
Then Liesel Corina started. Three weeks into the job as a trust and safety associate, she was doing a routine audit, pulling the highest-reach posts from the under-16 cohort that week to spot-check them by hand. One video stopped her: a "what I eat in a day" edit with a coded, restriction-heavy caption, sitting near the top of several fourteen-year-olds' feeds. She looked up its confidence score out of habit more than suspicion.
Fifty-four percent.
She brought it to Talia in a Thursday standup. "This one's basically a coin flip," she said. "Why did it get full distribution, same as anything at ninety-eight?" Talia didn't have an answer. Not because it was a hard question, but because nobody had ever wired the ranking pipeline to look at that number at all.
The video had reached Marnie Kowalski's feed nine days earlier. Marnie was fifteen, and for a few weeks she'd been searching and lingering on dieting content, the ordinary, quiet way a teenager's attention drifts somewhere and the algorithm notices before anyone else does. Because her own recent activity had trained her personal feed toward exactly the category where the classifier's confidence was worst, she kept getting videos like it, ranked with full weight, no different on her screen from anything the model was actually sure about.
Marnie never reported anything. There was nothing to report. The video looked exactly like every other video in her feed: same layout, same "for you" ribbon, same confident little checkmark of a ranking. She saved two of them that week. The feed, reading that as a signal, sent her three more.
Talia had been in the room when the pipeline was first built to skip that number. It was the right call then. There was one narrow category to worry about, confidence was never anywhere near ambiguous, and wiring in a float the pipeline would never actually need was solving a problem that didn't exist yet.
Run the same nine days again, this time with the gate in place. The coded video scores 54 percent and never touches full-rank distribution. It sits in the holdback pool, gets a second-model pass within ninety minutes, and either clears with a visible "cleared under review" tag or stays capped at a small, non-boosted audience. In the replay, the number of coin-flip posts reaching full reach each week drops from around 1,900 to under 150, the ones that genuinely clear review fast. Marnie's feed that week never gets the push at all, because by the time the video clears, if it clears, her session's already over and the recommendation loop never gets its first tug.
What I'd tell myself, if I could go back to that first meeting: we built something that could say no. We never built it to say "I'm not sure," and that missing sentence was the actual criteria the whole time, not a nice-to-have we could bolt on once the numbers demanded it.
GUARD, for a check that can only ever say yes
This is a risk question, so the framework is GUARD. "Degrade gracefully" sounds like a design nicety, but the real test is who a silent, confident "safe" label catches without meaning to, and whether anyone can even see it happening.
And if you want to be sure it really works, try it somewhere else
An eldercare check-in line calls a homebound senior every morning and a voice model flags "possible confusion" in their answers, sent straight to an adult child's daily digest. Different building, same five letters, same trap.
G, groups. The intended target is real cognitive decline, caught early. The group caught instead is anyone whose voice sample naturally produces lower confidence for other reasons entirely.
U, unequal. Hard-of-hearing callers who mishear a question, and callers with a strong regional accent the model wasn't trained on, both produce the same "low confidence, possible confusion" pattern as someone actually declining.
A, ability to contest. The parent never sees the flag. It goes straight to their child's digest with nothing marking it as a coin-flip call, so there's no way to say "I heard you fine, I was just being careful."
R, reduce. Below 70 percent confidence, don't send an alert on one call alone. Trigger a gentle automatic recheck an hour later. Only escalate to the family digest if the recheck also comes back low, and tag anything that reaches the digest that way as "flagged after recheck," separate from a high-confidence flag.
D, detect. Track the recheck-trigger rate weekly, split by whether the caller uses a hearing aid or has a language or accent note on file, watching for one subgroup's rate climbing while the overall rate looks steady.
Swap the trigger and it still runs
- Speed: a viral trend spikes posting volume overnight, and the same small percentage of low-confidence content turns into a much bigger raw number of coin-flip posts reaching full reach, even though the rate itself never moved.
- Cost: the holdback pool and the second-pass review cost real staffing time to clear within ninety minutes, so it keeps losing to features that show up on a roadmap slide instead.
- The model gets better: a newer classifier version raises average confidence to 91 percent, which reads as pure progress, but the posts still landing under threshold are now almost entirely the hardest, most coded cases, exactly the ones worth catching.
Where people run it wrong
- Building a classifier that outputs a confidence number, then wiring the pipeline to read only the pass/fail label next to it.
- Calling a confidence number "logged" and treating that as the same thing as the system actually acting on it.
- Watching one blended pass rate instead of the same rate split by content category or user cohort.
If you're asked this cold
Ask whether the feature has a way to say "not sure" at all, and if it does, whether anything downstream ever actually reads that number. If the answer to either half is no, that's the whole gap the criteria has to close.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Acceptance criteria for non-deterministic output
- #1 Rewrite this criterion to be testable: the model should not hallucinate.
- #2 How do you express an acceptance criterion as a rate rather than an absolute?
- #3 What is the difference between a threshold criterion and a distributional criterion?
- #4 Write acceptance criteria for an AI feature that extracts fields from an invoice.
- #5 How do you set a pass bar when human performance on the same task is 92 percent?
- #6 Describe acceptance criteria that account for the severity of different error types.