ConceptIntermediateEval-Driven Specification / Acceptance criteria for non-deterministic output / #13

What does done mean for a feature whose quality will keep changing?

The direct answer
Done means the feature keeps passing a live, scheduled check, not that it passed one validation set and the ticket closed. Set a threshold, sample real traffic on a calendar, split by the kind of content it's checking, and alert when any slice of it crosses the line for two checks running. The captions did not get worse because the model broke. They got worse because the day everyone called the work finished was also the day everyone stopped watching it.
Do this, in order
  1. Define done as a live, scheduled check with an alert on it, not a one time validation pass that closes the ticket.Why: this is the fix. It puts a live watcher where the closed ticket used to sit, so drift shows up on a calendar instead of in a public clip.
  2. Split the check by content type, not one blended catalog average.Why: the drift lived inside one slang heavy cluster the whole time. A blended average buries it exactly the way the launch validation set did.
  3. Keep the review cadence alive after launch. Do not let closing the epic delete the calendar hold that goes with it.Why: the review did not stop because nobody cared. It stopped because closing a ticket in the tracker took the recurring meeting down with it.
  4. Leave the launch validation gate alone.Why: passing a fixed benchmark before shipping is still the right first step. The problem was never that gate, it was having no gate after it.
  5. Watch the trend line, not just this week's snapshot.Why: the error rate crossed the danger line for weeks before anyone looked. A trend line catches that months earlier than a single spot check would.
  6. Do not fix this with a bigger validation set or a stricter launch number.Why: both are dials on the same one time gate. The flip only stops once done becomes something that keeps happening, not something that gets tested harder once.

How to answer this, stage by stage

Seven moves. This question hides a trap: it sounds like a definitions question, but it's really asking what happens after the definition is met. Each stage has the words you'd actually say.

1
Scope 'done' down to one dashboard
Say it like this
"'Done' is too big a word to answer directly, so let me pick one place it actually lives. I'm going to design what happens the day an AI captioning feature passes its launch bar and someone marks the ticket closed. That's the moment this whole question turns on."
Why this works
A word like "done" has no shape until you pick one real moment where it gets decided. This is that moment.
2
Say your five part structure out loud
Say it like this
"I'll cover five things. Who actually decides this is done. What they stop watching once it's decided. Where that costs us, once quality can still move after the decision. Which old choice I'd take back. And what the same drift looks like once done means something ongoing instead of something finished."
Why this works
Two seconds of structure stop you rambling through an abstract word and show the interviewer you already have a plan for it.
3
Reframe what 'done' is actually testing
Say it like this
"The real question isn't whether the model passed its launch bar. It's whether anyone is still checking it against the world after the day we decided it had. A feature whose output can keep changing doesn't get to have a finish line. It only gets checkpoints."
Why this works
Separates you from a candidate who defines done purely by a launch metric and never asks what happens the week after.
4
Give the one decision, and only one
Say it like this
"Here's what I'd actually do. Keep the launch bar exactly as it is. But make 'done' mean the feature passes that bar once, and then keeps passing a recurring, scheduled sample, split by content type, with an alert if any slice crosses a set line for two checks running."
Why this works
This is the direct answer, said out loud. A concrete recurring check beats "keep monitoring it," which nobody can picture or put on a calendar.
5
Prove it with the failure
Say it like this
"Say Farah signed the caption model off at three point two words in a hundred wrong, and the epic closed, which also cancelled her Friday sampling review. A new true crime series with heavy regional slang comes into the catalog, and by month five its captions are running at eleven percent wrong while the whole catalog average barely moves. Nobody notices until a viewer posts a clip of it."
Why this works
Four sentences, and it's the exact spot where a one time gate quietly turned into a permanent blind spot.
6
Name what you'd watch, and on what schedule
Say it like this
"Before I call anything done, I'd watch word error rate per content cluster, sampled every week, with an alert if any cluster holds above five percent for two weeks running. That number moves for weeks before a single viewer complaint does."
Why this works
Shows you think past launch day, and names a real schedule and threshold instead of a vague promise to keep an eye on things.
7
Land the whole answer in one breath
Say it like this
"So: done means the feature keeps passing a live check, not that it passed one test the day it shipped. Give the checking cadence the same permanence as the feature itself, and quality can keep changing without anyone finding out about it from a stranger's post."
Why this works
Restates the decision and why, in one breath. That's the line an interviewer remembers.
If you remember one thing Stages 3 and 5 carry this answer. Reframe "done" as a checkpoint, not a finish line, then prove it with one specific number that kept climbing while nobody was scheduled to look. Everything else here supports that.

Let's learn

Picture a feature the day after it launches, already marked done in a project tracker somewhere, its ticket closed and its owner already on the next thing.

Say a video streaming service builds a tool that writes the subtitles under every show automatically, instead of a person typing them out by hand.

Before the tool, a captioning vendor typed out subtitles for a new ten episode season by hand. That took about nine days a season, and a backlog of finished shows sat waiting on the list the whole time.

With the tool, captions show up under a new episode within about twenty minutes of it finishing upload. The team ran it against a five hundred clip test set before launch, and it got three point two words in every hundred wrong, comfortably under the four percent bar they had set for themselves. It passed. The ticket closed.

Word error rate for one content cluster, month by month after launch
0% 3% 6% 9% 12% alert line, 5% Launch 3.2% Month 2 4.4% Month 3 5.9%, week 13 Month 5 11.4%, a viewer posts it
A weekly, per cluster sample would have crossed the alert line in week thirteen. Nobody was scheduled to look again until a subscriber did, in week twenty two.
Knowledge spark: what's word error rate? How many words out of every hundred the captions get wrong. Missed, swapped, or made up. Lower is better, and it's usually reported as a percent.

Here is the part that matters. Passing that test on launch day was never the real risk. The risk was that nothing in the design kept checking after the day everyone stopped.

Months later, the catalog picked up a true crime documentary series shot with heavy regional accents and slang the original test set barely touched. It was a small slice of everything Cascade Play captions, about 8% of total volume. Nobody re-ran the validation set against it, because there was no ticket left open to remind anyone to.

Two identical ten by ten grids of one hundred squares. Left grid, labelled whole catalog average, about four squares colored red for wrong captions. Right grid, labelled slang heavy shows alone, about eleven squares colored red.
Same size grid, side by side. The average barely moved.
Knowledge spark: what's a validation set? A fixed batch of clips used to test a model once, before anyone else sees its output. It's a snapshot of one day, not a promise about every day after it.
How long the drift ran before anyone looked
One time gate, closed at launch
Old design
22 weeks
Ongoing, scheduled done
New design
13 weeks
Old design: caught when a subscriber posted a clip, in week twenty two. New design: caught by the scheduled sample the first time it held above the alert line for two weeks running, in week thirteen. Nine weeks earlier, before the error rate ever reached double digits.
The captions did not get slowly worse. The checking got instantly gone, the day someone closed a ticket.

At its worst, that is worse than never shipping the feature at all. The fix, once someone finally saw it, was an emergency recaption of all forty six episodes over one weekend, at about four times the normal per episode rate, plus a public apology and a scramble to answer complaints from deaf and hard of hearing subscribers who had trusted the captions the whole time.

The decision that mattered Bring back an ongoing definition of done: the launch bar plus a recurring, scheduled, per cluster sample with an alert on it. Not a bigger validation set. Not a stricter launch number. The review that used to happen every Friday, built back in on purpose, and no longer tied to a ticket that can be closed.

What I would leave alone. Flat, single host studio recordings in plain American English stayed near two to three percent the entire five months, launch to launch plus five. Running the same weekly, per cluster review on that segment at the same intensity would spend review hours on a slice of the catalog that was never moving.

The lesson. If closing a ticket is the thing that also ends whether anyone is watching, the feature was never actually finished. It was a feature that quietly became nobody's job, on the exact day we called it done.

Now here is the same thing as a story

Use this version when you have room to let it land, not just list it.

Farah Haidari can tell a caption that's garbled from one that's merely stiff before the line even finishes scrolling across the screen. Five years running caption quality for Cascade Play will do that to you.

She was the one who signed off the auto caption model back in the spring. For the first four months after launch, every Friday morning, she pulled forty clips at random from whatever had gone up that week, muted the sound, and read the captions cold, watching for a garbled word, a bad guess at a name, a piece of timing that drifted half a second too late. It was almost always clean. She'd flag one or two small things a week, log them, move on by half past nine.

So she started pulling twenty clips instead of forty. Then it was whichever ten looked interesting on the upload list. Some Fridays, none at all, because the sprint was heavy and the sample had come back clean six weeks running.

Then the ticket that had launched the whole feature, the one that said "auto captions: ship and validate," got marked done and closed in the tracker, the way a finished piece of work is supposed to be. The recurring Friday calendar hold had been created off that same ticket, back when someone set it up. Closing the ticket archived the hold along with it. Nobody meant anything by it. It was just how the tool worked.

Left, a dial with many marks labelled how often to check it, captioned what we assumed. Right, a two position switch labelled watching and done, captioned no middle setting.
People are switches, not dials

Five months after launch, a new true crime documentary series came into the catalog. Three seasons, forty six episodes, filmed almost entirely in a heavy regional accent, thick with local slang the original test clips never had a single line of. The model did what it always does with something it doesn't recognize: it guessed. Confidently, and often wrong.

Nobody was watching that cluster specifically. The whole catalog average, blended across everything Cascade Play captions, moved from 3.2% to about 3.9%, because the drifting series was a small slice of a much bigger pie. It looked, from a distance, exactly like nothing had happened.

We did not lose four sentences of dialogue. We lost the only signal that would have caught it in week thirteen instead of week twenty two.

It surfaced the way these things do now. A deaf reviewer, a regular commentator on caption quality across streaming platforms, stitched together a two minute clip of the worst lines from the series and posted it. A local idiom read out as three random words. A character's name transcribed four different ways in one scene. It moved fast through accessibility circles, then landed on a tech blog with a headline that made "AI captions" sound like a punchline.

I want to say the problem was one bad model update. It wasn't. Farah never had a running number in her head for that cluster. She had a habit, and the habit only had two settings: pull the sample and read it, or don't. There was no setting in between, no "check it a little less carefully this month." Once the ticket closed and took the calendar hold with it, the only thing left standing between a drifting model and the public was whether Farah happened to remember, on her own, to go looking.

So here's the decision I'd take back. In that same spring sprint review, when the launch epic was closing, someone pointed out that the recurring Friday review was tied to a ticket that was about to go away, and asked whether it should live somewhere else. The honest answer at the time was that there was no infrastructure for an ongoing, per cluster sample yet, and building one wasn't in scope for a launch that was already two weeks late. Let the review lapse for now, someone said. We'll figure out monitoring properly once we've shipped a few more things. That was a reasonable call in a sprint review. It stopped being reasonable the day new content arrived that the launch validation set had never once seen.

I'd put the review back, not tied to a ticket that can close, but on its own permanent calendar, split by content cluster, with a number attached: any cluster holding above five percent word error for two weeks running gets flagged automatically, no one has to remember. If that had existed, the documentary series would have tripped the alert in week thirteen, at five point nine percent, long before it reached eleven. Cascade Play would have caught it with a normal, scheduled recaption of one series instead of an emergency weekend job and an apology tour.

That's the whole difference. One design hands the feature a finish line and calls the job over. The other hands it a heartbeat, something that keeps happening on its own schedule, whether or not anyone remembers to ask about it.

And the part I'd tell myself, if I could go back: we asked whether the model was good enough to launch. We never asked who was still going to be checking six months after everyone had stopped talking about the launch.

FLIPS, for a feature that keeps grading itself

This question dresses up as a definitions question. It's actually a Perturbation question: something keeps drifting after ship day, and FLIPS runs straight down the line, F to S.

Five stacked rows, F L I P S, each a letter in a colored box, a step name, and a question. The I row is outlined in red.
FLIPS, in five rows
FFind the person
Whose morning ends the moment the ticket closes?
Not "the QA team." One person, one recurring review, one calendar hold.
In this answer: Farah Haidari, captions quality lead at Cascade Play, the day the launch epic closes.
LLocate the habit
What did they stop watching because it worked, and what did closing the ticket quietly take with it?
The habit is the product working. What a "done" gate covers for once nobody's job is to look anymore.
In this answer: She stopped pulling her weekly forty clip sample. Then the recurring Friday review, tied to the launch ticket, was archived the moment the ticket closed.
IIdentify the flip
What verb snaps, with no middle setting?
Not "more errors." A specific behavior with exactly two settings, and no drift back once the calendar hold is gone.
In this answer: Runs a scheduled, per cluster sample every week, or runs no check at all once the ticket says done.
PPinpoint the old decision
Which choice only made sense before quality could drift?
Small, specific, reasonable at the time. Never "add more review."
In this answer: Defining done as passing the launch validation set once, then letting closing the tracker ticket auto cancel the recurring sampling review that had been tied to it.
SShow the replay
Same drift, an ongoing done. Better ending?
Run the same trigger through the fixed product. End on something you can count.
In this answer: The per cluster alert trips in week thirteen at 5.9%. A normal, scheduled recaption fixes it, weeks before a viewer ever posts a clip.
Two panels. Left, a gently rising line labelled the model's own word error rate, from three point two percent to eleven point four percent. Right, a flat line that holds high then drops straight down and stays low, labelled checked every week to checked never.
A small move in the model. A big snap in who was watching.
Why I is the hard step Anyone can say "the checking got lax." The hard part is naming the exact moment the calendar hold disappeared, and proving there's no setting in between "runs a weekly sample" and "runs nothing." "Checks less often" is a dial. "Stopped checking, permanently, the day the ticket closed" is a switch. If your flip has a middle, keep looking.

And if you want to be sure it really works, try it somewhere else

A commercial real estate firm's AI tool reads lease PDFs and pulls out renewal option deadlines for a lease administrator to track. Same question shape, a completely different product, and a different flip.

A small grid. Two rows, Farah the captions lead and Yusuf the lease administrator. Five columns, F L I P S. Green dots in every cell except the I column, which holds two different short phrases in a red box.
Only one letter changes

F. Yusuf Demir, lease administrator at Corrigan Portfolio Group, six years in, tracking renewal option deadlines across about four hundred commercial leases.
L. He stopped fully re-reading each lease PDF himself once the tool's flagged renewal dates matched his own read, lease after lease, for over a year.
I. A different flip. He doesn't check less carefully. He stops opening the tool at all. After it silently misses a renewal clause on one lease, he goes back to tracking every deadline in his own personal spreadsheet, permanently, even for the leases the tool would have caught correctly. No middle setting between "trusts the dashboard" and "never opens it again."
P. Done for this feature meant passing a one time accuracy test on standard domestic lease formats. Once it passed, the tracker closed the ticket, and nobody scheduled the ongoing sampling that would have caught new formats, like international ground leases, as the portfolio grew. When the tool did miss a clause, it gave no reason why, so Yusuf had no way to trust it selectively.
S. An ongoing done adds a recurring sampled audit across new lease formats, plus a flag on which fields the tool is least sure about. The missed clause on the new ground lease format gets caught by the internal audit within the same month it appears, before Yusuf ever hits it on his own. He keeps using the dashboard for the leases it's actually good at, and no renewal deadline comes within days of lapsing again.

A second decision worth taking back Giving no reason when the tool misses something is a decision, not a limitation. A low confidence flag on the fields it's least sure about would have let Yusuf trust it selectively instead of abandoning it completely the first time it let him down.

Swap the trigger and it still runs

  • Speed: if generating captions took twenty minutes instead of instant, editors would start publishing without them and adding captions later, off the queue, quietly. Same missing ongoing check, a different mechanism.
  • Cost: if Cascade Play started paying per minute of video processed, someone would ration which shows get the accurate model versus a cheaper one, and nobody would widen the sampling to match. Worse captions would move through the same trusted pipeline.
  • The model got better: the case on this page. Passing the launch bar with room to spare is exactly what got the recurring review cancelled in the first place.

Where people run it wrong

  • Blaming the model's accuracy instead of the calendar hold that got deleted along with the ticket.
  • Fixing it with a bigger validation set or a stricter launch number, which still only ever tests one day.
  • Waiting for a viral clip instead of watching the per cluster trend line, which was already climbing for months.

How to use it live

Say the real question out loud before anything else. "So done here can't mean passed a test once, it has to mean still passing a test." Naming that costs five seconds, and it isn't stalling, it's where the real answer starts, because defining "done" only means something once you've said whether it's a gate or a habit.

Flashcards (click a card to flip it)

Eight fixed slots, pulled straight from the answer above.

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip. Checks sometimes, then stops checking at all. It fires when the news is good, a launch passing its bar, not when something breaks.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Farah Haidari, captions quality lead at Cascade Play, five years running caption checks. She can spot a garbled caption before the line finishes scrolling.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
She stopped pulling her weekly forty clip sample, and then the recurring Friday review itself was archived when the launch ticket it was tied to got closed.
4 · THE FLIP, HERE
What's the two setting switch in this story?
Tap to flip
ANSWER
Runs a scheduled, per cluster sample every week, or runs no check at all once the ticket says done. No setting in between once the calendar hold is gone.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Defining done as passing the launch validation set once, then letting the tracker auto cancel the recurring sampling review the moment the epic closed.
6 · THE NUMBER
The whole catalog average moved from 3.2% to only ___%. The slang heavy cluster alone moved from 3.2% to ___%.
Tap to flip
ANSWER
3.9% blended, 11.4% for the cluster alone. The blended number barely moved, which is exactly why nobody watching only the average would ever have caught it.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
A recurring, per cluster sample with a five percent alert catches the drift in week thirteen instead of week twenty two, before it ever reaches double digits, and before a subscriber posts about it.
8 · CROSS-PRODUCT
Section 4 answers this same question for a different product, with a different flip family. Which product, which family?
Tap to flip
ANSWER
A commercial lease abstraction tool, using the abandonment flip: a lease administrator quietly stops opening the tool at all after one missed clause, and goes back to tracking renewals by hand.

Check yourself Score: 0 / 0

Multiple choice
1. What was the flip in Farah's story, and what were its two settings?
  • A. She checks the captions a little less carefully than she used to.
  • B. She runs her weekly, per cluster sample, or she runs no check at all once the ticket says done.
  • C. The model's word error rate went from 3.2% to 11.4%.
  • D. She asks a coworker to double check the documentary series captions before it airs.
Show hint
A flip is a verb the person does, not a change in the model, and it has exactly two settings.
Show answer
B. C describes the model, not a person's behavior. A is a dial, there was no "checks it a bit less" setting she actually landed on, the review either ran or it didn't. D is a fix, not what happened in the story.
True or false
2. True or false: Cascade Play should run the same weekly, per cluster sample at full intensity on every single show, including flat, single host studio recordings that never drifted.
  • True
  • False
Show hint
Look at the "what I would leave alone" paragraph. Which slice of the catalog actually moved?
Show answer
False. Flat, single host studio content held near two to three percent the whole five months. Spending the same weekly review hours there, at the same intensity as a drifting cluster, wastes review time on a slice of the catalog that was never moving.
Fill in the blank
3. The decision this answer takes back is treating ______ as a one time gate instead of an ______ state.
Show hint
It's the exact word this whole question is built around.
Show answer
Done, ongoing. A launch bar passed once is a snapshot of one day. Quality that can keep changing needs a definition of done that keeps checking, not one that closes a ticket and stops.
Multiple choice
4. Why couldn't Farah have just "kept half an eye on it" instead of stopping completely?
  • A. She did keep half an eye on it, that's exactly what happened.
  • B. Once the epic closed, the calendar hold tied to it was cancelled too. There was no scheduled middle ground left, only a recurring review or nothing.
  • C. Cascade Play had a policy banning anyone from checking a closed ticket.
  • D. The model itself flagged that it no longer needed checking.
Show hint
Look at what closing the ticket actually took down with it.
Show answer
B. The recurring review only existed because a calendar hold was tied to the ticket. Once the ticket closed, that hold was gone too, so there was no scheduled "check it a little" option left, only the full weekly review or nothing at all.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one habit it built in you that you'd stop doing completely, rather than just doing more carefully, if it got a little worse?
Show hint
Think of a spam filter, a GPS app, a spell checker. Something whose whole job is letting you stop doing something by hand.
Show answer
Model answer: "Turn by turn directions on my GPS app. I stopped memorizing routes years ago. If it started routing me wrong regularly, I wouldn't 'follow it more carefully.' I'd stop trusting turn by turn directions altogether and go back to checking the whole route on a map before I left, because there's no calibrated middle between trusting it completely and reading the map myself." Any honest answer works if it names a real two setting switch, not just "I'd be more careful."
Fill in the blank, do the math
6. The slang heavy cluster makes up about 8% of everything Cascade Play captions. Its word error rate hit 11.4% while the rest of the catalog held near 3.2%. About what would the blended, whole catalog average come out to?
Show hint
0.08 times 11.4, plus 0.92 times 3.2.
Show answer
About 3.9%. That's close enough to the 3.2% launch number that nobody watching only the blended average would ever have noticed something was wrong. The drift only shows up once you split the number by content cluster.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more