How do you handle acceptance when different customer segments need different bars?
- Publish every segment's acceptance bar side by side, and set a floor tying review rate to error rate.Why: a blended number hides exactly the gap that matters, and without a floor, whichever segment has fewer reviewers quietly gets the loosest bar.
- Split the model's own disagreement rate by segment before trusting any blended number.Why: an average sitting near ninety percent can ride comfortably on top of one segment failing badly for years.
- Fund the floor with real review capacity for the harder segment, not a policy asking reviewers to try harder.Why: the loose bar exists because of headcount, not carelessness, so the fix has to add headcount or route work differently, not add a memo.
- Give the under-reviewed segment a door that reaches the right person inside the appeal window, not a general inbox.Why: right now there's nothing that tells anyone in that inbox a message is a scoring appeal at all.
- Track the override rate by segment, and flag low appeal volume paired with high disagreement as the tell.Why: a segment that never appeals looks calm from the outside and can be the one being shortchanged the most.
- Leave the clearest, highest-confidence files in the well-served segment at the default bar.Why: not every file needs a second look. Spend the added review budget where the model is actually unsure, not everywhere at once.
How to answer this, stage by stage
Seven moves. The middle three are where the actual decision gets made; the rest is scoping it and closing it out.
Let's learn
Every spring, a small financial aid team at one university used to read every scholarship application by hand, transcript, essay, financial forms, the works, about fifty minutes a file.
Say a university builds a tool that reads every scholarship application and hands back a score the financial aid team uses to decide which files get a full human read before a decision goes out.
Before the tool, the team could get through roughly two hundred files carefully in the crunch week before decisions shipped. Files that arrived late, often the international ones waiting on a translated transcript, sat at the bottom of the pile, and sometimes got the fastest read of the whole cycle.
Now the model scores every file in under a second, all of them, no backlog, no pile. But the score alone doesn't decide anything. A confidence number decides whether a person looks at the file again before the letter goes out.
Here's the part that matters. The extra applications aren't the problem, the model handles all of them fine on average. The problem is the two bars sitting behind that one average. Domestic files, standard transcripts, standard grading, get a second read about one file in five. International files, a dozen different grading systems, translated documents, get a second read about one file in twenty, because the reviewers who can actually check a foreign transcript are scarce.
At its worst, this sits behind the product that was never built at all. A team that reads every file the slow, manual way at least applies one bar to everyone. A team that ships a score with two bars, one of them invisible, can deny a qualified student a scholarship for years without a single complaint reaching anyone, because nothing about the denial tells her there's anything to complain about.
What I would leave alone. A domestic file with a strong, clean, standard transcript, sitting well clear of the threshold, model and human already agreeing loudly. Adding a mandatory second read there catches nothing. It just slows down the clearest files in the pile.
The lesson. I used to think one confidence number, checked in aggregate, meant the model was trustworthy. It doesn't. It means the model is trustworthy on whichever pool has the most files in the test set, and the other pool rides along unmeasured until someone goes looking.
Now here is the same thing as a story
Read the short version above if you want it fast. Read this one when you want to feel why the fix matters, not just know what it is.
Suvi Laine can read a disagreement matrix the way most people read a bus schedule. Three years running admissions technology at Corrywood University, and she built the scholarship scorer herself, mostly, with a small team, after years of watching the financial aid office fall further behind every spring.
For the first two cycles, the tool looked exactly like what it was sold as. Every file scored in seconds. The backlog that used to swallow late, translated, international transcripts disappeared completely. Suvi pulled a report every Monday during scoring season and watched the numbers settle: about ninety percent accuracy, blended across the whole pool, cycle after cycle.
By the second cycle she'd stopped pulling the report by segment at all. There wasn't a segment view. There'd never been a reason to build one. The blended number always looked fine.
Then, at a financial aid staff meeting, a counselor mentioned it almost as an aside, going over her own numbers: "You know international appeals basically never happen, right? I don't think we've had one in two cycles." Suvi took that to mean the score was working even better for that pool. She said so out loud. Nobody in the room corrected her.
That sentence sent her digging anyway, mostly out of habit. About a week in, she pulled the raw disagreement numbers apart by pool for the first time since launch. Domestic files: the model and a trained reader disagreed four percent of the time. International files: nineteen percent. The blended number, the one that always looked fine, had been sitting on top of both the whole time, because international files were a small slice of the total.
Then she found the file. Zanele Dlamini, an applicant from South Africa, matric results reported as percentages against a national exam, not a grade point average. The model had misread the conversion and scored her academic record into "needs improvement," a tier that meant automatic denial, no second read, because the file cleared the threshold on its own, easily, on a score built from a misread number.
Zanele did the one thing available to her. She wrote in, the day the letter arrived: "I think there's a mistake in how my marks were read, could someone check?" It landed in the general financial aid inbox, the same queue as bursar questions and login resets. Nothing about her sentence told anyone reading it that it was a scoring appeal. The appeal window at Corrywood is twenty one days. Her message got read, by someone who understood what the scorer actually does, on day twenty nine.
Suvi was in the room, over a year earlier, when the threshold got set. Someone pulled the pilot results, a ninety one percent match on the pilot sample, and everyone signed off, because the sample was mostly domestic files, that's mostly what the university has, and nobody split it out before it shipped. It made sense at the time. Nobody was hiding anything. There just wasn't a reason yet to go looking for the split.
I would take that back. Not the threshold exactly, the decision to set one number for both pools off a sample that was mostly one of them. I'd have pulled the international slice out before launch, seen the model's own error rate on it, and matched the review rate to that number instead of to how many reviewers happened to read foreign transcripts that year.
And the inbox. Zanele wrote the words "marks" and "reading" in a message that sat three weeks in a queue built for password resets. That's not a mystery to fix with more training. That's a routing rule nobody wrote.
Run the cycle again with the fix in place. The acceptance spec already shows both bars side by side, so someone would have asked about the international one before launch, not after a counselor's offhand remark. Zanele's file gets a second read automatically, because international files sit above a review floor now, not below it. And her email, flagged the moment it mentions "marks" or "transcript," reaches the right desk the same day, not on day twenty nine. Second cycle, same kind of file: caught in scoring, corrected before the letter ships, no appeal needed at all.
The thing I'd tell myself, a year back: a blended accuracy number isn't modest, it's quiet. It doesn't announce that it's hiding a segment. You have to go looking, and I only went looking because of a sentence in a meeting that happened to be about something else.
GUARD, aimed at a settings page instead of a launch
This is a risk question, so the framework is GUARD. "Different bars for different segments" reads like an ops tradeoff, which is exactly why nobody flagged Corrywood's threshold as a fairness decision until a counselor said something unrelated out loud.
And if you want to be sure it really works, try it somewhere else
An auto insurer uses AI to score claim severity and decide which claims get auto-approved without an adjuster's second look. Different building entirely, same five letters, same trap.
G, groups. Policyholders with an assigned agent, whose agent flags anything odd straight to a claims handler, and direct-to-consumer policyholders, who bought the cheaper no-agent plan and file through a self-service portal with nobody to escalate for them.
U, unequal. Claims from direct policyholders get auto-approved or auto-denied at a higher rate than agent-assisted claims, not because the claims are simpler, but because there's no agent in the loop pushing the odd ones to a person.
A, ability to contest. A direct policyholder disputing a denial calls a general claims line with long hold times and no assigned rep who already knows the file, while an agent-assisted customer's agent calls a dedicated line the same afternoon.
R, reduce. Give every direct-to-consumer claim the same second-look threshold as agent-assisted claims, and publish both rates side by side instead of letting agent involvement quietly set the bar.
D, detect. Track the denial-overturn rate by channel, agent versus direct, watching for the channel with a high denial rate and a low dispute rate together, the same signature as the scholarship case.
Swap the trigger and it still runs
- Speed: the scorer rolls out from auto claims to home claims in a month, and the same unpublished gap between agent and direct channels ships again, because the launch review only checked the blended number.
- Cost: the insurer cuts direct-channel support staff because those claims are "handling themselves," which is exactly the number that was already quietly wrong.
- The model gets better: overall denial accuracy climbs to ninety six percent, and that's exactly when nobody proposes splitting the number by channel anymore, because the deck already looks finished.
Where people run it wrong
- Writing one acceptance bar into the spec and letting whichever team has more reviewers quietly loosen it in practice.
- Testing a fix against the blended accuracy number instead of the one segment that's actually failing.
- Building a dispute path that exists on paper but isn't routed anywhere a person reads it inside the deadline.
How to use it live
Ask, out loud, who set the number for this segment and why it's different from the other one. If the honest answer is "that's just what our reviewers had time for," say plainly that's not a bar, it's a headcount problem wearing a threshold's clothes, then fix it live by tying the review rate to the error rate instead.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Acceptance criteria for non-deterministic output
- #1 Rewrite this criterion to be testable: the model should not hallucinate.
- #2 How do you express an acceptance criterion as a rate rather than an absolute?
- #3 What is the difference between a threshold criterion and a distributional criterion?
- #4 Write acceptance criteria for an AI feature that extracts fields from an invoice.
- #5 How do you set a pass bar when human performance on the same task is 92 percent?
- #6 Describe acceptance criteria that account for the severity of different error types.