ConceptFoundationalEval-Driven Specification / Acceptance criteria for non-deterministic output / #2
How do you express an acceptance criterion as a rate rather than an absolute?
The direct answer
Replace the promise with four things nailed down: what counts as a hit, what the denominator is, how big and how random the sample is, and the number that changes what happens next. For a counterfeit queue: each week, pull 100 listings at random from everything the model flags, and have a second authenticator who hasn't seen the model's reasoning call each one genuine or fake by hand. That percent confirmed fake is the criterion. At 95 percent or higher, the flag goes straight to takedown. Between 85 and 95, a second person has to sign off first. Under 85, auto-takedown pauses for that whole category.
Do this, in order
Rewrite the absolute into a weekly random-sampled rate with a hard number attached.Why: a sentence nobody can fail a build on never actually stops a bad flag from reaching a real seller.
Build the sampling harness, with a blind second read, before the model touches a single takedown.Why: a rate that only exists after the damage is a report, not something that catches drift early.
Make the denominator a random slice of every flag that week, never just the cases a reviewer already agreed with fast.Why: shrinking the denominator to the easy calls is the cheapest way to fake a healthy score.
Set a harder band for takedowns that can escalate to a full store suspension, not just one listing.Why: the overall rate can sit at 95 percent while the handful of repeat-strike sellers, the ones who lose everything, are exactly where it's failing.
Track the lagging outcome too, appeals that get overturned, but act on the rate, not the appeals.Why: appeals lag the real drift by weeks. Waiting for them is waiting for the damage to already be done.
Post the weekly number where the whole trust and safety team can see it, not one engineer's spreadsheet.Why: a number nobody outside the model team looks at is decoration, not a criterion.
How to answer this, stage by stage
Eight moves. Say the numerator and the denominator out loud, by name, or the number you give next just sounds like a guess with more confidence.
1
Say why the absolute can't fail a build
Say it like this
"'The model should never wrongly flag a real seller' sounds like the safe, responsible thing to write down. But there's no number in it, no sample, no way to check it on a Tuesday afternoon. Before I rewrite it, I want to say that plainly, because that's actually the whole job here."
Why this works
Naming the flaw first proves you're not about to bolt a percent sign onto the same vague sentence and call it done.
2
Put one real queue and one real person under it
Say it like this
"Say this is a resale marketplace for designer handbags, sneakers, and watches, and Camila owns the slice of the review queue that catches counterfeits. A model scans every new listing the moment it's posted and flags about 3 in every 100 as likely fake before it ever goes live."
Why this works
A criterion about "the model" is empty until someone would actually act on what it says.
3
Ask what the absolute is really trying to protect
Say it like this
"What we actually want isn't 'the model is never wrong, about anything, ever.' We want that when a flag turns into a takedown, it's actually right, so a real seller doesn't lose her whole store over a bag that just happens to look worn."
Why this works
This is the outcome step. Say it in plain words, or the rate you write next has nothing real underneath it to aim at.
4
Name the numerator and the denominator
Say it like this
"Here are the two words that decide everything. The denominator is the flags the model raises in a week, all of them, roughly 420. The numerator is how many of those, checked independently by a person, turn out to actually be counterfeit."
Why this works
Get either of these wrong, especially the denominator, and the rate can look great while missing exactly the cases it exists to catch.
5
Pick the sample size and keep it honest
Say it like this
"We don't chase all 420 by hand every week. We pull 100 at random from that full 420, not just the easy 100, and a second authenticator who hasn't seen the model's stated reason calls each one genuine or fake. That's the number: percent confirmed fake, out of a random 100."
Why this works
A rate built on a cherry-picked or self-graded sample isn't a rate. It's a number that flatters itself.
6
Set the bands and the action at each one
Say it like this
"At 95 percent or higher, the flag goes straight to takedown. Between 85 and 95, a second person has to sign off before anything comes down. Under 85, auto-takedown pauses for that whole category, and a person reads every single flag until the number recovers."
Why this works
A metric nobody acts on is a chart on a wall. This line is what turns it into a criterion.
7
Say out loud how the number gets played
Say it like this
"And I'd name the cheat before anyone else finds it. The easy way to keep the score high is to only sample the obvious fakes, the ones with a logo spelled wrong, and skip the hard calls, a well-worn but real bag, a brand the model's never seen before. Those hard calls are exactly where a real seller gets hurt."
Why this works
Naming the abuse case yourself is what tells an interviewer you understand metrics, not just picked one off a list.
8
Close on the number that would have caught it early
Say it like this
"That's the whole reason to build this. A score sliding from 96 to 90 over three weeks is something you catch on a Tuesday with a spreadsheet. Finding out from a seller survey two months later, after real stores have already been shut down, is the version where you got lucky it wasn't worse."
Why this works
This ties the rate back to what's actually at stake, and it's the line worth leaving the room with.
If you only get through two stages
Stages 4 and 5 are the answer. Say what the numerator and the denominator actually are, then say the sample size and the threshold. Everything else on this list is how you defend that rate under follow-up.
Let's learn
Say a resale marketplace runs every new handbag, sneaker, and watch listing through a model that checks whether the item looks fake before it ever goes live.
Before the model, a person checked every one of the roughly 14,000 new designer listings that came in each week. That took about two days, so a seller's first sale could sit waiting while a small team worked through the backlog by hand.
Now the model clears 97 of every 100 in under ten minutes. The other 3, about 420 listings a week, get pulled into a closer look before anything goes live.
Here's the part that matters. Those 420 flags a week were never the problem. The problem is what happens to a flag once the team stops checking it as carefully as they used to.
A rate can hold steady on paper for months while the one case it was built to catch never makes it into the sample at all.
A rate can look healthy and still miss the case that mattered
What that costs at its worst: a wrongly confirmed takedown doesn't just delay one listing. Threadloop suspends a seller's whole store after two confirmed counterfeit strikes in 90 days, so a rubber-stamped flag can push a longtime seller over that line and take her entire income with it, not just one bag.
Needs the harder band
Flags that can escalate
A seller's second strike, the one that triggers a full store suspension
A luxury brand partner's own legal escalation on a listing
A seller who has never been flagged before, where the model has no track record to lean on
A worn-but-real item that statistically resembles a fake more than it resembles a fresh one
These are the cases a wrong call actually ends a business. They get the tighter bar.
95 percent is already plenty
The obvious cases
A logo that's visibly misspelled in the listing photo
Hardware or stitching that doesn't match any real production run on file
A listing already confirmed fake by a brand partner's own report
These are the calls the model has been right about for years. A wrong one here costs a re-review, not a business.
Knowledge spark: what "confirmed counterfeit" means here
Not a guess and not the model's own stated reason. It's a person who checks the stitching, the hardware engraving, and the item's serial or box code against the real brand's records, the same way a person would in a shop, and writes down genuine or fake. That's the only thing allowed in the numerator.
The leading edge: flag precision, sampled every week
95 percent or higher, keep auto-takedown
85 to 95, a second person signs off first
under 85, auto-takedown pauses
The rate crossed under 95 in week 3 and under 85 in week 5. Nobody was watching it, because nobody had built anything to watch.
The lagging outcome: sellers wrongly suspended, confirmed later by appeal
2 a week
Weeks 1 to 2, average
19 a week
Week 7, when the survey line surfaced
Appeals lag the real drift by weeks, because a seller has to notice, file, and wait. The flag-precision rate above had already crossed both thresholds before a single one of these appeals was filed.
The choice I'd take back
We signed off on "the model should never wrongly flag a real seller" as the acceptance criterion, and it sounded stricter than a percentage, so nobody argued with it. Nobody could test it either, so nobody built the sampling harness. I would take that back and build the weekly random 100, checked by a second person, before the model ever touched a single takedown.
What I'd leave alone. The hard-block list stays an absolute, and that's correct, not a shortcut. A listing that matches a recalled product's exact model number, or falls in a flat-banned category like live animals, isn't a judgment call. It's an exact match or it isn't. Rates are for the cases with real uncertainty in them. Turning every rule into a rate, including the ones with no uncertainty at all, just adds noise nobody needs to watch.
The lesson. If you can't say the number that would make you pause a feature, you haven't written a real requirement. You've written a hope wearing a rule's clothes. Every acceptance criterion for something the model has to judge needs a sample, a denominator, and a cutoff next to it, or it isn't actually a criterion yet.
Nothing broke. It just drifted.
You don't need this to answer the question. It's here so "95 percent" stops being an abstraction and starts being a Tuesday.
Camila Rocha can tell a real hardware engraving from a good copy of one before she's finished her coffee. Five years in trust and safety, the last two owning the counterfeit slice of Threadloop's review queue, the designer handbags, the sneakers, the watches.
The counterfeit model launched in her second year on the team. For a long stretch after, it was the best part of the job. A listing that used to sit for two days now went live in minutes, or landed in her team's queue with a reason attached, hardware mismatch, price too far under comps, a serial number that didn't match. For months, every authenticator gave a flagged listing the full look: photos, seller history, the item's own paperwork, about six minutes each.
Then, without anyone deciding it on purpose, the look got shorter. The model's stated reason kept checking out, so a full read became a 90-second scan. The scan kept checking out too, so eventually it became a 20-second glance at whatever the model had written and a stamp on the takedown. Nobody wrote that down as a policy. It just happened, the way a habit thins when nothing bad ever seems to come of thinning it.
There was no single bad Tuesday. It didn't arrive as a moment at all. It arrived as one line, three rows down, on page fourteen of Threadloop's quarterly seller survey, the kind of report most people skim once and file. The share of sellers agreeing with "I don't trust that Threadloop's counterfeit checks are fair" had crept from 2 percent to 6 percent over the quarter. Small. Not the kind of number that gets an email of its own. A researcher on the insights team mentioned it to Camila by the coffee machine the following week, more curious than alarmed.
Camila didn't trust the summary. She pulled the raw appeal numbers herself. Appeals that got overturned, meaning a real seller had been wrongly suspended, had gone from about 2 a week at the start of the quarter to 19 a week by the most recent one. Nobody on her team had been tracking that as its own line either. Appeals went to a separate team, case by case, with no shared dashboard connecting them back to the queue.
One of the 19, that final week, belonged to a shop called Junebug Vintage. Six years on Threadloop, reselling pre-loved designer bags with a small, loyal following. Her Chanel flap bag had genuine, well-worn hardware, the kind of soft brassing that comes from a decade of real use, not a fake's crisp new plating. The model flagged it as a likely mismatch. The authenticator on shift that day gave it the 20-second version, agreed with the model's stated reason, and confirmed the takedown. It was her second strike in four months. Her whole store came down that afternoon.
We didn't take one listing from her. We took the storefront she'd spent six years building.
Camila went looking for how the sampling had worked, and found the honest, unglamorous answer: it hadn't. Nobody had ever built a weekly check. The team's occasional spot-checks, back when they still did them, had drifted toward the easy, obvious fakes, the ones a scan could confirm in ten seconds, because those were fast and satisfying to clear. The hard cases, the worn-but-real bags, the brands the model had less history with, almost never got a second look at all, which was exactly the shape of the picture two years of nobody watching had drawn.
She went back to the sign-off meeting from two years earlier, in her head. The requirements doc had one line under quality: the model should never wrongly flag a real seller. Everyone nodded. It sounded like the responsible thing to require, stricter than any percentage anyone could think of on the spot. Nobody could turn it into a number in that room, so it went in unnumbered, and the meeting moved to the next line.
Here's the replay. If a random weekly sample of 100 flags, checked by a second person, had existed from week one, the rate would have shown 96 percent, then 95, then 90 by week 3, already under the line where a second reader should have signed off on every takedown. By week 5 it would have shown 84, under the line where auto-takedown pauses entirely for that whole category. Both of those would have happened two and four weeks before the survey line even got written, let alone noticed on page fourteen.
With that rate running, Junebug Vintage's bag gets the second read the drop already earned it. The takedown doesn't happen. Her store is still open.
What I'd tell myself, back in that sign-off meeting: I let a sentence with no way to fail through, because it sounded stricter than a percentage. It wasn't stricter. It was just untestable, and untestable always loses, quietly, to a percentage someone is actually checking.
LEAD, so you can rebuild the rate from memory
This is a metric question, so the framework is LEAD, not GUARD. GUARD asks who can't push back against a decision. Here the decision hasn't been made wrong yet. The whole question is which number would have said so first.
L, link. The real outcome, not the model's own confidence. Here, it's whether a confirmed takedown is actually right, so a real seller never loses her store over a listing that only looked wrong.
E, early signal. The number that moves before that trust breaks. Here, the percent of a random weekly sample of 100 flagged listings that a second, independent authenticator confirms is really counterfeit.
A, abuse. How the number gets hit without the real problem going away. Sample only the obvious, easy-to-confirm fakes, and the hard judgment calls, worn-but-real items, newer brands, never get checked at all.
D, decision. What changes at each score. At 95 or higher, auto-takedown keeps running. Between 85 and 95, a second person signs off before anything comes down. Under 85, auto-takedown pauses for that category entirely.
The check that keeps the rate honest
Try shrinking the denominator on purpose. Sample only the flags from the categories the model has always been good at. If the score barely moves, the sample was already close to the truth. If it jumps, the denominator was hiding the real problem, and that's exactly the trick a busy team reaches for without meaning to.
And if you want to be sure it really works, try it somewhere else
An auto insurance company flags claims that might be fraudulent before an adjuster approves the payout. Same shape of question. "The model should never wrongly deny a real claim" is just as untestable as the marketplace version.
L. A real claim gets paid without the customer being wrongly accused of fraud and dropped by the insurer.
E. Each week, pull 100 flagged claims at random. An investigator who hasn't seen the model's flag reason checks each one against the real file and confirms fraud or not.
A. Only sampling claims from customers who already have a claims history, the easy calls, and skipping first-time claimants, where the harder judgment calls actually live.
D. Above the bar, auto-deny stands. In the middle band, a second investigator signs off before any denial goes out. Below it, that claim type gets manually reviewed until the model is retrained.
Same rate, a different desk
Swap the trigger and it still runs
It gets slower. Doesn't matter. Sampling 100 claims a week takes the same afternoon whether the model answers in one second or ten.
It gets more expensive. If a new state's data feed triples the flag volume overnight, the weekly sample still stays at 100. You just check it a little more often.
It gets better than planned. If the model starts reading messy scanned police reports reliably, rerun the same weekly check against the wider set of documents it now handles. The process doesn't change. Only what's being sampled does.
Where people run it wrong
Sampling only claims from customers with a prior claims history, never first-time claimants, so the score looks great on the population that was never the risk.
Treating the sample as a one-time certification instead of a weekly habit, so a new fraud pattern slips in unwatched for months.
Averaging every flag into one score instead of scoring the claims that could end in a policy cancellation separately, so a bad run hides inside a good overall number.
How to say it if you're asked this cold
Buy yourself the time to build the number properly. "Before I give you a rate, let me say what the absolute is actually trying to protect." That's not stalling. It's the L step, said out loud, and it gives you somewhere honest to stand while the real numerator and denominator take shape in your head.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and why not GUARD or BOUND?
Tap to flip
ANSWER
LEAD, for a metric question. GUARD asks who can't push back against a decision already being made. BOUND does arithmetic for a fixed estimate. LEAD finds the number that moves before real damage shows up, which is exactly what this question is asking for.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Camila Rocha, a trust and safety product manager at Threadloop who has owned the counterfeit-detection slice of the review queue, handbags, sneakers, and watches, for two years.
3 · THE HABIT
What did the review team stop doing because the model kept being right?
Tap to flip
ANSWER
Giving every flagged listing the full six-minute look. It shrank to a 90-second scan, then to a 20-second stamp on whatever reason the model had written.
4 · THE RATE
What's the actual rate in this story, in plain words?
Tap to flip
ANSWER
Out of a random 100 of the roughly 420 listings the model flags each week, the percent a second, independent authenticator confirms is actually counterfeit.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Signing off on "the model should never wrongly flag a real seller" as the acceptance criterion, with no number and no sampling harness built to check it before it mattered.
6 · THE NUMBER
The rewritten criterion samples ______ listings a week, and needs ______ percent of them confirmed fake to keep auto-takedown running.
Tap to flip
ANSWER
One hundred listings. Ninety-five percent. Between 85 and 95, a second person signs off first. Under 85, auto-takedown pauses for that whole category.
7 · THE REPLAY
Same drift, new design, what changes?
Tap to flip
ANSWER
The rate drops under 95 percent in week 3 and under 85 in week 5, both weeks before the seller survey surfaces anything in week 7. A second reader steps in at week 3, and Junebug Vintage's store never comes down.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
An auto insurance claims desk flagging possible fraud. Its early signal is the percent of a random weekly sample of 100 flagged claims that an independent investigator confirms as real fraud.
Check yourself Score: 0 / 0
Multiple choice
1. Which of these would count as a genuine pass under Camila's rewritten criterion?
A. The model's flag reason names a specific hardware mismatch.
B. The seller has no prior strikes on the account.
C. A second authenticator, who hasn't seen the model's reasoning, opens the listing and confirms it's actually fake.
D. The listing was one of the 3 in 100 the model routed to review.
Show hint
Three of these describe the flag. Only one describes a person independently checking the real item.
Show answer
C. Being flagged, having a stated reason, and a clean seller history are all facts about the process, not proof the item is actually fake. The only real pass is a person checking the item itself.
Fill in the blank
2. The rewritten criterion samples ______ flagged listings a week, out of about ______ the model raises in total, and needs ______ percent of that sample confirmed fake to keep auto-takedown running.
Show hint
The first two numbers are in the walkthrough's stage 5. The third is in the direct answer at the top.
Show answer
100. 420. 95. Between 85 and 95, a second person signs off before any takedown. Under 85, auto-takedown pauses for that whole category.
True or false
3. True or false: this problem gets fixed by telling the model to be more careful with listings that show heavy wear.
True
False
Show hint
Ask who would ever know whether the model actually did that.
Show answer
False. An instruction to the model can't be sampled or audited. You still need a person checking real listings against the real item every week, or there's no way to know the instruction changed anything at all.
Short answer
4. Name a place in this same queue where writing the rule as an absolute is actually fine, not a rate.
Show hint
Look for the rule that's an exact match, not a judgment call.
Show answer
Model answer: "The hard-block list, like an exact match to a recalled product's model number, or a flat-banned category such as live animals. There's no probability there, a listing either matches the exact rule or it doesn't, so 'never allow it' is already a real, testable rule. Rates are for the judgment calls, not the exact matches."
Short answer, apply it yourself
5. Pick a product you use yourself. Name one place it makes an absolute-sounding promise you've never actually seen tested. What would the weekly-sample version of that check look like?
Show hint
Look for a claim with no number attached, like "verified" or "checked" or "never fake."
Show answer
Model answer: "A review site that says it 'removes fake reviews.' The weekly check: pull 100 reviews at random that survived the site's own filter, and have a person try to independently confirm each reviewer actually bought the product. If the confirmed-real rate drifts, I'd know the filter was slipping months before a news story caught it."
Multiple choice
6. Which of these would make the weekly rate look good without actually protecting real sellers?
A. Raising the sample from 100 listings to 300.
B. Only sampling flags from categories the model has always been accurate on, and skipping the newer, harder ones.
C. Having two authenticators check each sampled listing instead of one.
D. Publishing the weekly score to the whole trust and safety team.
Show hint
The dangerous move hides the hardest cases inside an average full of easy ones.
Show answer
B. The overall score can sit at 95 percent while the riskiest cases, the worn-but-real items and newer brands, never get checked at all. That's exactly the abuse the rate has to be designed against.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.