ConceptAdvancedEval-Driven Specification / Acceptance criteria for non-deterministic output / #3

What is the difference between a threshold criterion and a distributional criterion?

The direct answer
A threshold criterion checks one submission against one cutoff number, right now. A distributional criterion checks the whole spread of scores over time, so it catches people learning exactly where that number sits, weeks before any one submission ever crosses it. Keep the threshold as the same day gate on each piece of work, but only trust it long term if the spread of scores behind it stays honest.
Do this, in order
  1. Track the shape of every score over time, not just each score against the cutoff.Why: a flag rate can hold perfectly steady while a whole class learns exactly where the line sits.
  2. Watch the band just under the cutoff, not the average score.Why: gaming clusters right at the edge. The average can stay put while that edge fills up.
  3. Set a real number that trips an alert on the distribution side, and act on it.Why: a chart nobody has to answer to is decoration, not a defense.
  4. Keep the per submission threshold as the same day gate.Why: it is still the fastest way to catch one clearly machine written essay before it gets graded.
  5. When the distribution trips, retrain the detector on fresh cases, don't just tighten the number.Why: tighten the cutoff alone and people find the new line just as fast.
  6. Leave low stakes work on threshold only checking.Why: not everything needs the heavier watch, and treating it all the same buries the courses that do.

How to answer this, stage by stage

Seven moves. Say what each rule protects before you compare them, or the answer sounds like two definitions instead of one real distinction.

1
Scope it to one tool and one person
Say it like this
"Let's make this real. Say a university buys a tool that scores every essay a student turns in for how likely it is that a machine wrote it. Baris Yildirim is the PM who has to tell the school whether the cutoff score can still be trusted."
Why this works
A threshold or a distribution over what stays a vocabulary question until someone's grading decision depends on trusting one of them.
2
Say your structure out loud
Say it like this
"I'll use LEAD here, but sideways. Not to pick one metric, to split two kinds of pass or fail rule: what each one protects, what early signal each one gives you, how each gets played, and what actually changes at each one's warning sign."
Why this works
Naming the plan up front tells the interviewer you're about to compare two real things, not recite two dictionary entries.
3
Reframe what the question is really asking
Say it like this
"This sounds like a definitions question. It's really asking, can one cutoff number, checked one paper at a time, ever tell you the whole system is still working. And on its own, it can't."
Why this works
This is where the answer stops being vocabulary and starts being a decision about what an interviewer can actually trust.
4
Give the one decision
Say it like this
"So here's the split. A threshold criterion is one paper against one line, score over 85, flag it, right now. A distributional criterion is every paper's score, plotted together, week over week, is the whole shape drifting toward that line even while each single paper still clears it. I'd run both, and I'd trust the threshold less over time unless the distribution behind it stays honest."
Why this works
This names the concrete difference plainly, with a real number attached, not a restatement of the question.
5
Prove it with a failure
Say it like this
"Here's why that matters. At Bellhaven, the weekly flag rate at the 85 cutoff held around 3 percent for ten straight weeks, so the dashboard looked fine the whole time. But the share of essays scoring 75 to 84, just under that line, grew from 4 percent to 34 percent over the same stretch. Nobody was tracking that band. A manual review at finals found 19 confirmed cases of contract cheating in essays that had each individually cleared the line."
Why this works
A compressed real failure, with real numbers, does more work than a paragraph of reasoning about what could go wrong.
6
Say how each one gets played, differently
Say it like this
"A threshold gets played one paper at a time, someone edits the wording until the score slides from 92 to 83. A distributional check gets played differently, by hiding in the average. Watch only the mean score, and it can hold steady while a whole cluster piles up right at the edge. You have to watch the band near the cliff, not the middle of the crowd."
Why this works
Naming two different failure modes, not one dressed up twice, shows you actually understand both rules instead of just one of them.
7
Say what changes at each warning sign, and close
Say it like this
"One paper over 85, that goes to a human reviewer, same day, case closed either way. The band under the line growing past somewhere like 15 percent, that's not one more case, that's the cutoff itself going stale. That earns a manual audit of a random sample and a retrain, not another single flag. If someone asks me whether the 85 cutoff still works, I want that band chart in front of me, not just this week's flag count."
Why this works
This ties both rules to a real action at a real number, so the distinction isn't just definitions, it's a plan.
If you only get through two stages Stages 4 and 6 are the answer. Say the one line difference between the two rules, then say how each one gets played, differently. Everything else here is how you defend that under follow up.

Let's learn

Say a university uses a tool called Sourceline that reads every essay a student turns in and gives it a score for how likely it is that a machine wrote it.

Knowledge spark: what's an AI content score? A number the tool gives each piece of writing, its own guess at how likely a machine wrote it. Zero means it reads like a person wrote every word. A hundred means it reads like a machine wrote all of it.

Before Sourceline, an instructor with 90 essays a term read every one by hand and caught the ones that felt off by instinct, maybe two or three cases a term, the kind so obvious they read like nobody had touched them at all.

Now Sourceline scores every essay the moment it's turned in. Anything scoring above 85 gets flagged for a human to look at. For ten weeks in a row, that flag rate held steady around 3 percent, week after week. The dashboard looked fine the entire time.

Here's the turn. That flat 3 percent wasn't proof nothing had changed. It was the exact place the real story was hiding. Underneath the flag count, the share of essays scoring 75 to 84, sitting just under the line, was climbing every single week.

A flag rate that never moves is not proof nothing changed. It can be proof everyone found exactly where the line sits.

At its worst, the tool ends up worse than never building it. Instructors trusted the cleared label completely and stopped reading flagged-clear essays themselves, so work that had quietly learned to sit just under the line traveled all the way to a final grade, in some cases into a transcript, before anyone caught it by hand.

Knowledge spark: what's contract cheating? Turning in work you didn't write, bought, borrowed, or had a machine produce, and handing it in as your own. It's one of the harder things for any single check to catch, because each piece of it can look fine on its own.
A hand-sketch comparison. Left, a lined paper marked eighty three, cleared. Right, a simple grey figure walking away, carrying a laptop under one arm.
Same cutoff, same clean score. The real thing still walks off the page.
Share of essays scoring 75 to 84, just under the cutoff, by week
4% 5% 7% 11% 15% 19% 23% 27% 31% 34% W1 W2 W3 W4 W5 W6 W7 W8 W9 W10, finals audit
under the 15 percent floor, cutoff still trustworthy
over it, the cutoff line no longer means what it did
The band crossed the 15 percent floor in week five, five weeks before finals forced anyone to look. By week ten, over a third of essays were sitting just under the line.
Confirmed misconduct cases found at finals, before and after the drift
2
A normal term, band under 15 percent
19
The term nobody watched the band
Two confirmed cases in a normal term, the same number the manual-only era caught by hand. Nineteen in the term the band ran to 34 percent unwatched. The cutoff had individually cleared every one of those nineteen essays.
The choice I'd take back We only ever charted the flag rate against the 85 cutoff. Nobody plotted the shape of scores underneath it, so a whole class could learn exactly where the line sat and we wouldn't see it until finals. I'd chart the band under the line from day one, so the drift shows up on a graph in week five, not in a manual review after the term is already over.

What I'd leave alone. Low stakes work, a weekly reading response nobody grades hard, doesn't need this. A slightly stale cutoff there barely matters, because no real grade or degree depends on catching every one of them fast.

The lesson. A number that never moves is not proof nothing is wrong. Sometimes it's proof everyone already learned exactly where it sits. The only way to know a cutoff still means something is to keep checking the shape behind it, not just the number it produces.

Now here is the same thing as a story

Use this version when you've got the time. The short version is above. This is for when you want to feel why it mattered.

Baris Yildirim can read a false positive complaint from a professor and tell inside a minute whether the model got it wrong or the professor did. He's run product for Sourceline for two years, and he set the original 85 cutoff himself, after testing it against a stack of sample essays nobody had graded yet.

For the first two terms, that number earned its trust the honest way. Every month, Baris pulled the flag rate next to what the integrity office actually found in random spot checks, and the two agreed. A flagged essay was almost always a real problem. A cleared one almost always was fine.

Then the checking faded, in three small steps, none of them looking wrong at the time. First, because the flag rate held so steady, Baris stopped pulling the raw score list every month and only checked it once a term. Second, seeing the same steady number, the integrity office quietly dropped its extra random audits between terms. Third, instructors, trusting the cleared label completely, stopped even skimming the essays that came back under the line.

Then a teaching assistant, grading a stack of essays late one night, sent Baris an offhand message. A run of essays in one course all sounded the same, oddly polished, in a way she couldn't quite name.

Baris didn't pull the flag count. He pulled the raw scores instead, every single one, plotted by week. The band sitting just under 85 had grown from a sliver to more than a third of the class.

We did not lose nineteen essays that term. We lost the one thing that made eighty five worth trusting.

It would be easy to say the model got worse. It didn't, not really. A class had spent ten weeks learning exactly where 85 sat, one small edit at a time, and the flag count had no way to show that, because none of them individually crossed the line.

So here's the decision Baris would take back.

Two years earlier, in the meeting where the monitoring plan got set, someone had asked whether they should also chart the full spread of scores every week, not just the flag count. It sounded like an extra dashboard for a problem the flag count already seemed to answer. Baris remembers agreeing they could add it later.

I would build that chart from day one instead. Same 85 cutoff, same weekly runs, but with a real floor on the band underneath it, one someone actually has to act on.

Here's the replay. Same ten weeks, new chart running the whole time. The band crosses the 15 percent floor in week five, and that crossing is the trigger, not a teaching assistant's late night message. Baris pulls a fresh batch of that week's actual essays and retrains the cutoff region on them. The band falls back under 15 percent by week seven. Finals close with 2 confirmed cases, the normal number for a term, instead of 19.

One design watches whether today's essay passed. The other watches whether passing still means anything.

And the thing Baris would tell himself, back in that first meeting: the flag count was never the risk. Trusting it with nothing to check it against was.

What LEAD looks like when you're comparing two rules instead of picking one metric

This reads like a definitions question. Underneath it, it's still asking what tells you a pass or fail rule is working, before someone outside the tool finds out it isn't. That's LEAD, run on two rules at once.

L, link. What each rule actually protects. A threshold protects one submission, right now, does this specific essay cross a line we already trust. A distribution protects the population over time, is the line, as a whole, still catching what it's supposed to catch.
E, early signal. The earliest thing each rule can tell you. A threshold answers immediately, but only about the one paper in front of it. A distribution warns weeks before any single paper crosses the line, the shape of every score creeping toward it.
A, abuse. How each gets played, differently. A threshold gets played one paper at a time, edit the wording until the score slides under the line. A distribution gets played by hiding in the average, keep the middle of the crowd normal while a cluster piles up right at the edge, and a check that only watches the mean will miss it completely.
D, decision. What actually changes. One paper over the line, a person reviews it today, case closed either way. The band near the line growing past a real number, that's not one more case, that's the cutoff itself going stale, and it earns a random audit and a retrain, not another single flag.
The check that proves a cutoff is still real Pull one week at random and ask: is the band under the line growing, or holding steady? If it's growing, the cutoff has already started losing its meaning, whatever this week's flag count shows.

And if you want to be sure it really works, try it somewhere else

A circuit board plant runs a camera over every solder joint on the line, scoring how likely each one is to fail, and rejects anything over a set score before the board ships. Same shape of question. A pass rate that holds perfectly steady at that reject line can hide the same thing a flat flag rate hid at Bellhaven: something drifting, just under where anyone's watching.

L. Whether the boards reaching customers actually have reliable joints, not whether each board's number stayed under the reject line.
E. The share of boards scoring just under the reject cutoff, tracked by shift, not the pass rate alone.
A. Nobody has to cheat this one on purpose. Dust settling on the camera lens, or the light rig running a little warmer each month, can push every board's score down a few points without ever failing a single board outright.
D. Reject rate steady, trust it. Near line band growing shift over shift, stop trusting the number and send the camera in for cleaning and recalibration, not just a re-run of the same boards.
A labeled hand-sketch diagram centered on a gauge icon marked Reject Score, with four labels radiating around it: camera lens, dust film, warmer light, reject bin.
Same idea, a solder line

Swap the trigger and it still runs

  • The detector gets faster. Doesn't matter. A faster read on the same 85 line still only tells you about the paper in front of it.
  • The manual review gets cheaper to run. Doesn't matter either. An audit that only ever checks a random handful will still miss a slow, even drift spread across a whole class.
  • The model gets genuinely better at spotting machine writing. Chart it anyway. That's the one case where the near line band should shrink, not grow, and proving that is what makes "it got better" checkable instead of assumed.

Where people run it wrong

  • Treating a flat flag rate as proof nothing changed, with nobody ever pulling the full spread of scores to check.
  • Watching only the average or median score, which can hide a whole cluster building right at the edge.
  • Waiting for a manual audit or an offhand complaint to notice drift, instead of setting a real number that trips a check on its own.

How to use it live

Say the split first. "Before I answer, I want to separate two questions here, is this one paper over the line, and is the line still where I think it is." That's not stalling. That's naming which rule you're actually being asked about, and it buys the room to build the real answer instead of taking the flag count's word for it.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what does it stand for here?
Tap to flip
ANSWER
LEAD. L is the real outcome each rule protects, E is the earliest signal each one gives you, A is how each one gets played differently, D is what actually changes at each rule's warning sign.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Baris Yildirim, the product manager who has run Sourceline, the AI content detector, for two years, and who set its original 85 cutoff himself.
3 · THE HABIT
What did the team stop doing because the flag rate stayed steady?
Tap to flip
ANSWER
Pulling the raw score list. Once the weekly flag rate held around 3 percent, Baris checked it once a term instead of monthly, the integrity office dropped its extra audits, and instructors stopped skimming cleared essays at all.
4 · THE REAL DIFFERENCE
What's the difference this story actually turns on?
Tap to flip
ANSWER
A steady flag count against the 85 line said nothing had changed. The band of essays scoring 75 to 84, just under that line, grew from 4 percent to 34 percent in the same ten weeks. One number was flat. The real thing was moving the whole time.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Agreeing, two years earlier, that tracking the full spread of scores could wait, because the flag count already seemed to answer the question on its own.
6 · THE NUMBER
The band under the cutoff went from ______ in week one to ______ by week ten.
Tap to flip
ANSWER
4 percent to 34 percent. It crossed a 15 percent warning floor in week five, five weeks before finals forced anyone to look.
7 · THE REPLAY
Same 85 cutoff, band tracked from day one, what changes?
Tap to flip
ANSWER
The alert trips in week five instead of going unnoticed. Baris retrains the cutoff region on fresh essays that week. The band falls back under 15 percent by week seven, and finals close with 2 confirmed cases instead of 19.
8 · THE TRANSFER
Section four runs LEAD on a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
A solder joint inspection camera on a circuit board line. Its early signal is the share of boards scoring just under the reject line, tracked by shift, since dust on the lens can drift the whole reading without ever failing one board outright.

Check yourself Score: 0 / 0

True or false
1. True or false: if the weekly flag rate at the 85 cutoff holds steady, the cutoff is still catching machine written essays the same way it always did.
  • True
  • False
Show hint
Ask what a steady flag count can hide underneath it.
Show answer
False. A flag count can hold steady while the whole spread of scores drifts toward the line. Only checking the shape of scores underneath tells you if the cutoff still means what it did.
Fill in the blank
2. In week one, the band of essays scoring 75 to 84 was ______ percent of submissions. By week ten, the week finals found 19 confirmed cases, it had grown to ______ percent.
Show hint
The numbers sit right under the line chart in Section 1.
Show answer
4 percent. 34 percent. The gap is the whole story: a class that once spread out normally had, by week ten, learned almost exactly where the line sat.
Multiple choice
3. Which of these would actually tell you whether Sourceline's 85 cutoff still works?
  • A. The flag rate, tracked week over week.
  • B. How many essays get submitted each week.
  • C. The share of essays scoring just under the cutoff, tracked over time.
  • D. How confident the model is on its most obvious cases.
Show hint
Three of these can look completely normal on a class that has already learned exactly where the line sits.
Show answer
C. Only the band just under the line shows people finding the edge of the rule, before any of them ever crosses it.
Short answer
4. Name a place in this same tool's use where a threshold-only check is fine and doesn't need the full band tracking.
Show hint
Look for work where no real grade depends on catching every case fast.
Show answer
Model answer: "Low stakes work, like an ungraded weekly reading response. A slightly stale cutoff there barely matters, because nothing downstream depends on catching every one of them quickly."
Short answer, apply it yourself
5. Pick a product you use yourself. Name one pass or fail rule it uses that might be getting gamed one case at a time without the average ever moving. What would you check instead?
Show hint
Look for a single cutoff number the product checks each time, with nothing tracking the pattern across many checks.
Show answer
Model answer: "A spam filter that blocks any email scoring over a set number. Spammers can learn to write just under that score every time, and the weekly block rate stays flat. I'd track the share of mail landing just under the block line over time, not just the block rate itself."
Multiple choice
6. If Bellhaven had caught the band crossing 15 percent in week five instead of finding out at finals, what would most likely have happened to the confirmed cases that term?
  • A. No change, it would still be 19.
  • B. Fewer, because the detector would have been retrained on fresh cases before the band grew any further.
  • C. More, because catching it early would have distracted instructors from real grading.
  • D. It depends only on how many essays were submitted that term, not on the band.
Show hint
Think about what the band tracking is actually for: catching the drift before it grows any further.
Show answer
B. Catching it in week five gives the team a chance to retrain the cutoff region before most of the term's essays ever get near the edge, which is exactly what the replay in Section 2 shows happening.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more