Artifact critiqueAdvancedEval-Driven Specification / Acceptance criteria for non-deterministic output / #16

Write criteria for a feature that must degrade gracefully when confidence is low.

The direct answer
Write a real gate, not a wish. Below a stated confidence number, a post cannot get full reach in the feed. It goes to a small holdback pool and gets a second look inside a stated window. Right now "safe" means the same thing whether the model was almost certain or basically guessing, and it lands hardest on posts about the exact things a struggling teen searches for, because that's the content the model is least sure about in the first place.
Do this, in order
  1. Add a real confidence gate: below a stated number, a post can't reach full distribution, full stop.Why: right now "safe" is one flat label, so a 54 percent guess gets ranked exactly like a 99 percent certainty.
  2. Give the held posts a real second look inside a stated window, not an open-ended queue.Why: a hold with no deadline just becomes a slower, quieter version of the same problem.
  3. Tag anything that clears the hold as "cleared under review," never blended into the same log line as an automatic high-confidence pass.Why: nobody, not a moderator, not a parent request, not a regulator, can ever ask "was this one ever a coin flip" if both cases look identical afterward.
  4. Report the uncertain rate weekly, split by content category, not as one blended pass rate.Why: a category can sit at barely-better-than-a-guess for months while the overall number looks completely healthy.
  5. Leave categories that reliably score high confidence exactly as they are.Why: gating content the model is already sure about only slows the feed down for no safety gain at all.

How to answer this, stage by stage

Six moves, from pinning the feature to one product to the line you'd close on.

1
Pin it to one feature before naming any framework
Say it like this
"Say we build the ranking system behind a teen content feed. Every post gets scored by a safety classifier before it can reach anyone's 'for you' page. That classifier also outputs a confidence number. That's the feature I'd write graceful-degradation criteria for: what happens when that number is low."
Why this works
Grounds "confidence" in one real number on one real pipeline before any framework language shows up.
2
Say your structure out loud
Say it like this
"I'd use GUARD, because 'degrade gracefully' sounds like a UX question but it's really a fairness question. Who the check protects, who it catches by accident, who can't tell it happened, the exact fix, and how I'd know it's failing quietly in production."
Why this works
Two seconds naming the plan, not a recitation of five letters before the real thinking starts.
3
Reframe what "graceful" actually has to mean
Say it like this
"Graceful doesn't mean the feature apologizes nicely. It means the system's behavior actually changes below some number. Right now nothing changes. 'Safe' at 54 percent confidence gets exactly the same reach as 'safe' at 99. That's not degrading, that's just not knowing and not saying so."
Why this works
This is where a checklist answer and a real answer split. Naming the gap plainly earns the rest of the answer.
4
Give the one decision, as a testable criterion
Say it like this
"Below 75 percent confidence, a post is not eligible for full-rank distribution. It goes to a small holdback pool and gets a second pass, model or human, within ninety minutes. If it clears, it's tagged 'cleared under review' in the trust log, never folded into the same line as an automatic high-confidence pass."
Why this works
A number, a fallback behavior, and a visible tag. Not a policy sentence, a ticket an engineer could build this sprint.
5
Prove it with the failure it prevents
Say it like this
"Here's what happens without the gate. A coded dieting video scores 54 percent, gets ranked exactly like a sure thing, and a fifteen-year-old who's already been searching that content gets it pushed to the top of her feed at full weight. She has no way to know the model was ever unsure. Neither does anyone else, because 'safe' only ever gets logged one way."
Why this works
The compressed version of the story below. Four sentences, and the harm is concrete, not hypothetical.
6
Say what you'd watch every week, then close
Say it like this
"Every week I'd pull median confidence by content category, not just the blended pass rate, and watch for one category drifting low while the average looks fine. So: a real threshold, a timed second look, a visible tag, and a weekly number split by category instead of one number that hides all of it."
Why this works
GUARD's D step folded into the close, the line an interviewer actually remembers on the way out.

Let's learn

Here is what happens when a safety check can only say yes or no, and never gets to say "I'm not sure."

Glasswing is a short-video feed built for teenagers. Every post that goes up gets scored by a safety classifier before the ranking system will let it reach anyone's main feed, checking for things like coded self-harm content, extreme dieting content, and dangerous challenges.

Knowledge spark: what a confidence number actually is Most safety classifiers don't just say "safe" or "not safe." Underneath, they also output how sure they are, a number from 0 to 100. This question is about what happens when a product only ever looks at the label and throws that second number away.

Before the classifier existed, Glasswing was small enough that a five-person trust team manually reviewed every post its keyword filters flagged, about 200 a day, each one looked at within twenty minutes before it could go live. The classifier replaced that. Today it auto-clears roughly 50,000 borderline-adjacent posts a day, and anything it scores "safe" gets zero review time, whether the number underneath was 96 or 54.

That speed isn't the problem. Clearing 50,000 posts a day instead of 200 is the whole reason the product works at teenager scale. The problem is what "safe" quietly stopped meaning somewhere in that jump.

Median safety-classifier confidence, by content category
Same week, same classifier, same "safe" label applied to all four.
General entertainment
94%
Sports and challenge videos
89%
Grief and loss-adjacent posts
63%
Body-image and dieting-adjacent
58%
All four categories cleared as "safe." The overall daily pass rate held near 96 percent the whole time, because the two low-confidence categories were a small enough slice to disappear into the blend.

Here's the part that actually matters. A 58 percent confidence score is close to a coin flip. Treating that the same as a 94 doesn't just risk one wrong call, it means the system was never honest about which calls it was actually sure of.

We didn't build a safety check. We built a check that only ever knew how to say yes.

At its worst, this doesn't look like a scandal. It looks like a normal Tuesday. A teen who's already been engaging with dieting content gets a coin-flip video pushed to her at full strength, because the category she's drawn to is exactly the category the model is least sure about. She engages, the feed learns from that, and it sends her more. Nobody flagged it, because nothing about "safe" ever changes shape depending on how sure the model actually was.

The decision I would take back Three years ago, when the classifier launched, confidence was almost always 90 percent or higher, because the app was small and the training content was narrow. So the ranking pipeline was built to read the label only, pass or fail, and ignore the number underneath. Reasonable then. Wrong once the content got wide enough that a real, low-confidence slice showed up every single day.

What I would leave alone. General entertainment and sports content don't need a gate. They sit reliably at 89 to 94 percent confidence, week over week. Holding those posts back for review would slow the whole feed down for content that was never actually the risk.

The lesson. A check that can only say pass or fail is hiding its most useful number. If a system can be unsure, "graceful" has to mean the product actually does something different when it is, not that it fails a little more politely while doing the exact same thing.

Now here is the same thing as a story

The short version is above. Read this one when you want to feel why the ninety-minute window actually matters.

Talia Bregman built the ranking pipeline's safety gate when Glasswing had maybe forty thousand users and one content category worth worrying about: kids reposting stunts they'd seen on bigger platforms. Back then the classifier was almost boringly confident. Ninety-two, ninety-six, ninety-eight percent, week after week. In a meeting with her ranking-eng lead, the two of them agreed to keep the pipeline simple: read the pass/fail label, ignore the confidence float, because wiring the number in added real complexity for a case that basically never came up.

For two years, that held. Talia didn't think about it. Nobody did. The blended daily "safe" rate sat around 96 percent, and 96 percent looked like a number you could trust and move on from.

Glasswing grew. Content diversified, the way it does, into fitness edits, grief pages after a public tragedy went viral, "what I eat in a day" videos with captions that used coded, euphemistic language to slide past keyword filters. None of that showed up as a spike. It showed up as a slow widening of the range of things the classifier had to guess at.

Nobody noticed, because nobody was watching anything except the one blended number.

Then Liesel Corina started. Three weeks into the job as a trust and safety associate, she was doing a routine audit, pulling the highest-reach posts from the under-16 cohort that week to spot-check them by hand. One video stopped her: a "what I eat in a day" edit with a coded, restriction-heavy caption, sitting near the top of several fourteen-year-olds' feeds. She looked up its confidence score out of habit more than suspicion.

Fifty-four percent.

She brought it to Talia in a Thursday standup. "This one's basically a coin flip," she said. "Why did it get full distribution, same as anything at ninety-eight?" Talia didn't have an answer. Not because it was a hard question, but because nobody had ever wired the ranking pipeline to look at that number at all.

Talia hadn't ignored a warning sign. There had never been a number for anyone to warn her with.

The video had reached Marnie Kowalski's feed nine days earlier. Marnie was fifteen, and for a few weeks she'd been searching and lingering on dieting content, the ordinary, quiet way a teenager's attention drifts somewhere and the algorithm notices before anyone else does. Because her own recent activity had trained her personal feed toward exactly the category where the classifier's confidence was worst, she kept getting videos like it, ranked with full weight, no different on her screen from anything the model was actually sure about.

Two figures. On the left, the ranking team's dial sets full reach or hold before anyone sees a post. On the right, a teen holds nothing but the phone, and whatever the gate decided.
Talia's team set the gate. Marnie's feed carries out whatever it decided, with no way to see the gate was ever involved.

Marnie never reported anything. There was nothing to report. The video looked exactly like every other video in her feed: same layout, same "for you" ribbon, same confident little checkmark of a ranking. She saved two of them that week. The feed, reading that as a signal, sent her three more.

A flow of four boxes: post posted, model scores it, a missing confidence check marked in red, full reach anyway.
The step that should sit third, and doesn't

Talia had been in the room when the pipeline was first built to skip that number. It was the right call then. There was one narrow category to worry about, confidence was never anywhere near ambiguous, and wiring in a float the pipeline would never actually need was solving a problem that didn't exist yet.

Run the same nine days again, this time with the gate in place. The coded video scores 54 percent and never touches full-rank distribution. It sits in the holdback pool, gets a second-model pass within ninety minutes, and either clears with a visible "cleared under review" tag or stays capped at a small, non-boosted audience. In the replay, the number of coin-flip posts reaching full reach each week drops from around 1,900 to under 150, the ones that genuinely clear review fast. Marnie's feed that week never gets the push at all, because by the time the video clears, if it clears, her session's already over and the recommendation loop never gets its first tug.

What I'd tell myself, if I could go back to that first meeting: we built something that could say no. We never built it to say "I'm not sure," and that missing sentence was the actual criteria the whole time, not a nice-to-have we could bolt on once the numbers demanded it.

GUARD, for a check that can only ever say yes

This is a risk question, so the framework is GUARD. "Degrade gracefully" sounds like a design nicety, but the real test is who a silent, confident "safe" label catches without meaning to, and whether anyone can even see it happening.

G, groups. Two groups sit inside "safe." The team that ships the feature, whose growth numbers look good every time a post clears fast and nobody asks what the confidence number underneath was. And the teen on the receiving end of the ranking, who never chose any of this and has no idea the classifier was ever unsure about what just landed in her feed.
U, unequal. The harm lands hardest on teens already drawn to the exact categories, grief, body image, dieting, where the classifier's confidence is worst. Their own engagement is what trains their feed toward that category, and that category is precisely where a "safe" label is least likely to mean what it says.
A, ability to contest. Marnie can't tell a confident pass from a coin-flip pass, because both look identical on her screen. Nobody on Talia's own team could either, until Liesel happened to check one video's number by hand. If your own trust team can't see the gap, a fifteen-year-old and her parents never had a chance.
R, reduce. Set a real confidence threshold, 75 percent. Below it, a post can't reach full-rank distribution. It goes to a small holdback pool, gets a second pass within a stated window, and if it clears, that clearance is tagged distinctly in the trust log, not folded into the same line as an automatic high-confidence pass.
D, detect. Track median confidence weekly, split by content category, not as one blended number. Watch for a category's confidence drifting down while its full-rank volume keeps climbing, exactly the shape that let a 58 percent category hide inside a steady 96 percent headline.
Where this answer would fail If the fix here is a softer refusal message, a disclaimer, or a values statement about teen wellbeing, none of it counts. "A stated threshold, a timed second pass, a distinct tag, tracked weekly by category" is a build ticket. Someone can ship it this sprint, and you can check afterward whether the number actually moved.

And if you want to be sure it really works, try it somewhere else

An eldercare check-in line calls a homebound senior every morning and a voice model flags "possible confusion" in their answers, sent straight to an adult child's daily digest. Different building, same five letters, same trap.

G, groups. The intended target is real cognitive decline, caught early. The group caught instead is anyone whose voice sample naturally produces lower confidence for other reasons entirely.
U, unequal. Hard-of-hearing callers who mishear a question, and callers with a strong regional accent the model wasn't trained on, both produce the same "low confidence, possible confusion" pattern as someone actually declining.
A, ability to contest. The parent never sees the flag. It goes straight to their child's digest with nothing marking it as a coin-flip call, so there's no way to say "I heard you fine, I was just being careful."
R, reduce. Below 70 percent confidence, don't send an alert on one call alone. Trigger a gentle automatic recheck an hour later. Only escalate to the family digest if the recheck also comes back low, and tag anything that reaches the digest that way as "flagged after recheck," separate from a high-confidence flag.
D, detect. Track the recheck-trigger rate weekly, split by whether the caller uses a hearing aid or has a language or accent note on file, watching for one subgroup's rate climbing while the overall rate looks steady.

Swap the trigger and it still runs

  • Speed: a viral trend spikes posting volume overnight, and the same small percentage of low-confidence content turns into a much bigger raw number of coin-flip posts reaching full reach, even though the rate itself never moved.
  • Cost: the holdback pool and the second-pass review cost real staffing time to clear within ninety minutes, so it keeps losing to features that show up on a roadmap slide instead.
  • The model gets better: a newer classifier version raises average confidence to 91 percent, which reads as pure progress, but the posts still landing under threshold are now almost entirely the hardest, most coded cases, exactly the ones worth catching.

Where people run it wrong

  • Building a classifier that outputs a confidence number, then wiring the pipeline to read only the pass/fail label next to it.
  • Calling a confidence number "logged" and treating that as the same thing as the system actually acting on it.
  • Watching one blended pass rate instead of the same rate split by content category or user cohort.

If you're asked this cold

Ask whether the feature has a way to say "not sure" at all, and if it does, whether anything downstream ever actually reads that number. If the answer to either half is no, that's the whole gap the criteria has to close.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about writing graceful-degradation criteria, and why?
Tap to flip
ANSWER
GUARD, for risk, safety and fairness. The real question is who a confident-sounding "safe" label catches by accident, exactly what GUARD is built to find.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Talia Bregman, the product manager who built Glasswing's safety gate, back when the classifier's confidence was almost always above 90 percent and the pipeline only ever needed to read pass or fail.
3 · THE HABIT
What did Glasswing's trust team stop doing because the numbers looked fine, and why did that matter later?
Tap to flip
ANSWER
They never watched confidence by content category, only the blended daily pass rate. It held near 96 percent for years, so nobody noticed a real low-confidence slice was quietly growing underneath it.
4 · THE GAP
What's the confidence gap that proves the pass/fail treatment was hiding a real risk?
Tap to flip
ANSWER
General entertainment held 94 percent median confidence; body-image and dieting-adjacent content averaged 58 percent. Both got the identical "safe, full reach" treatment.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
Reading only the pass/fail label and ignoring the confidence number underneath. It made sense when confidence was almost always 90-plus percent and content was narrow. It stopped being safe once a real low-confidence slice showed up every day.
6 · THE NUMBER
Fill in: about ______ posts a week scored under the 75 percent confidence line and still got full-rank distribution, the same as a post scoring 99.
Tap to flip
ANSWER
About 1,900 a week. With the gate in place, the replay drops that to under 150, the ones that genuinely clear a fast second look.
7 · THE REPLAY
Same nine days, new design. What changes, and by how much?
Tap to flip
ANSWER
The 54 percent video never reaches full-rank distribution. It sits in a holdback pool for a 90-minute second pass. Coin-flip posts reaching full reach each week fall from about 1,900 to under 150.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the reduce step become?
Tap to flip
ANSWER
An eldercare check-in line's "possible confusion" voice flag. Reduce: below 70 percent confidence, trigger a gentle recheck call before anything reaches the family digest, and tag anything that escalates that way separately.

Check yourself Score: 0 / 0

Multiple choice
1. Glasswing's daily "safe" pass rate holds near 96 percent overall. What's the real problem with reading that as one healthy number?
  • A. 96 percent is too low a pass rate for a content feed to ship with.
  • B. The classifier needs a bigger training set before anyone should trust it at all.
  • C. A blended pass rate can look healthy while a specific content category sits at barely-better-than-a-guess confidence underneath it, with no threshold ever acting on that number.
  • D. Glasswing should turn the classifier off entirely until it gets more accurate.
Show hint
The problem in this answer was never the size of the overall number.
Show answer
C. A, B, and D worry about strictness or accuracy. The real risk is that the classifier already knew how unsure it was, on a per-post basis, and the pipeline never once looked at that number.
True or false
2. True or false: the classifier itself got worse at telling safe content from unsafe content once dieting and grief-adjacent posts became more common.
  • True
  • False
Show hint
Check what the classifier was actually outputting on those posts, not just the final label.
Show answer
False. The classifier was already telling the truth, a 54 or 58 percent confidence score. The failure was downstream: the ranking pipeline never read that number and treated every "safe" label the same regardless.
Fill in the blank
3. Body-image and dieting-adjacent content held ______ percent median confidence, against ______ percent for general entertainment, in the same week.
Show hint
It's the lowest and highest bars in the "Let's learn" chart.
Show answer
58 percent, and 94 percent. Same classifier, same week, split by content category instead of read as one blended pass rate.
Short answer
4. What old decision does this answer take back, and why did it make sense when Talia's team first made it?
Show hint
Look at what the classifier's confidence looked like at launch, not at who found the gap later.
Show answer
Model answer: "Reading only the pass/fail label and never wiring the confidence number into the ranking pipeline. It made sense at launch because confidence was almost always above 90 percent, so building a threshold for a case that basically never happened would have been solving a problem the product didn't have yet."
Short answer, apply it yourself
5. Pick an AI feature you've used that gives a flat yes/no answer. If the group using it got a lot more varied overnight, what kind of case might start dragging its confidence down without anyone noticing?
Show hint
Look for a case the feature was never trained on much, not a case it's obviously wrong about.
Show answer
Model answer: "A grammar checker that flags 'awkward phrasing' as one flat label. Rolled out to a much larger base of non-native English writers, its confidence on genuinely fine but differently-structured sentences would likely drop, and a flat flag would start reading as 'your English is wrong' instead of 'I'm honestly not sure here.'"
Short answer
6. If the confidence threshold were set at 90 percent instead of 75, would it still catch the coded dieting content, and what would it cost?
Show hint
Compare 90 percent against both the dieting-content number and the sports-content number.
Show answer
Model answer: "Yes, it would still catch it, since 58 percent is nowhere near 90. But it would also pull the 89 percent sports and challenge category into the holdback pool constantly, spending review capacity on content that was never actually the risk. The threshold has to sit between the categories that are genuinely uncertain and the ones that reliably aren't, not just be set as high as possible."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more