What does done mean for a feature whose quality will keep changing?
- Define done as a live, scheduled check with an alert on it, not a one time validation pass that closes the ticket.Why: this is the fix. It puts a live watcher where the closed ticket used to sit, so drift shows up on a calendar instead of in a public clip.
- Split the check by content type, not one blended catalog average.Why: the drift lived inside one slang heavy cluster the whole time. A blended average buries it exactly the way the launch validation set did.
- Keep the review cadence alive after launch. Do not let closing the epic delete the calendar hold that goes with it.Why: the review did not stop because nobody cared. It stopped because closing a ticket in the tracker took the recurring meeting down with it.
- Leave the launch validation gate alone.Why: passing a fixed benchmark before shipping is still the right first step. The problem was never that gate, it was having no gate after it.
- Watch the trend line, not just this week's snapshot.Why: the error rate crossed the danger line for weeks before anyone looked. A trend line catches that months earlier than a single spot check would.
- Do not fix this with a bigger validation set or a stricter launch number.Why: both are dials on the same one time gate. The flip only stops once done becomes something that keeps happening, not something that gets tested harder once.
How to answer this, stage by stage
Seven moves. This question hides a trap: it sounds like a definitions question, but it's really asking what happens after the definition is met. Each stage has the words you'd actually say.
Let's learn
Picture a feature the day after it launches, already marked done in a project tracker somewhere, its ticket closed and its owner already on the next thing.
Say a video streaming service builds a tool that writes the subtitles under every show automatically, instead of a person typing them out by hand.
Before the tool, a captioning vendor typed out subtitles for a new ten episode season by hand. That took about nine days a season, and a backlog of finished shows sat waiting on the list the whole time.
With the tool, captions show up under a new episode within about twenty minutes of it finishing upload. The team ran it against a five hundred clip test set before launch, and it got three point two words in every hundred wrong, comfortably under the four percent bar they had set for themselves. It passed. The ticket closed.
Here is the part that matters. Passing that test on launch day was never the real risk. The risk was that nothing in the design kept checking after the day everyone stopped.
Months later, the catalog picked up a true crime documentary series shot with heavy regional accents and slang the original test set barely touched. It was a small slice of everything Cascade Play captions, about 8% of total volume. Nobody re-ran the validation set against it, because there was no ticket left open to remind anyone to.
At its worst, that is worse than never shipping the feature at all. The fix, once someone finally saw it, was an emergency recaption of all forty six episodes over one weekend, at about four times the normal per episode rate, plus a public apology and a scramble to answer complaints from deaf and hard of hearing subscribers who had trusted the captions the whole time.
What I would leave alone. Flat, single host studio recordings in plain American English stayed near two to three percent the entire five months, launch to launch plus five. Running the same weekly, per cluster review on that segment at the same intensity would spend review hours on a slice of the catalog that was never moving.
The lesson. If closing a ticket is the thing that also ends whether anyone is watching, the feature was never actually finished. It was a feature that quietly became nobody's job, on the exact day we called it done.
Now here is the same thing as a story
Use this version when you have room to let it land, not just list it.
Farah Haidari can tell a caption that's garbled from one that's merely stiff before the line even finishes scrolling across the screen. Five years running caption quality for Cascade Play will do that to you.
She was the one who signed off the auto caption model back in the spring. For the first four months after launch, every Friday morning, she pulled forty clips at random from whatever had gone up that week, muted the sound, and read the captions cold, watching for a garbled word, a bad guess at a name, a piece of timing that drifted half a second too late. It was almost always clean. She'd flag one or two small things a week, log them, move on by half past nine.
So she started pulling twenty clips instead of forty. Then it was whichever ten looked interesting on the upload list. Some Fridays, none at all, because the sprint was heavy and the sample had come back clean six weeks running.
Then the ticket that had launched the whole feature, the one that said "auto captions: ship and validate," got marked done and closed in the tracker, the way a finished piece of work is supposed to be. The recurring Friday calendar hold had been created off that same ticket, back when someone set it up. Closing the ticket archived the hold along with it. Nobody meant anything by it. It was just how the tool worked.
Five months after launch, a new true crime documentary series came into the catalog. Three seasons, forty six episodes, filmed almost entirely in a heavy regional accent, thick with local slang the original test clips never had a single line of. The model did what it always does with something it doesn't recognize: it guessed. Confidently, and often wrong.
Nobody was watching that cluster specifically. The whole catalog average, blended across everything Cascade Play captions, moved from 3.2% to about 3.9%, because the drifting series was a small slice of a much bigger pie. It looked, from a distance, exactly like nothing had happened.
It surfaced the way these things do now. A deaf reviewer, a regular commentator on caption quality across streaming platforms, stitched together a two minute clip of the worst lines from the series and posted it. A local idiom read out as three random words. A character's name transcribed four different ways in one scene. It moved fast through accessibility circles, then landed on a tech blog with a headline that made "AI captions" sound like a punchline.
I want to say the problem was one bad model update. It wasn't. Farah never had a running number in her head for that cluster. She had a habit, and the habit only had two settings: pull the sample and read it, or don't. There was no setting in between, no "check it a little less carefully this month." Once the ticket closed and took the calendar hold with it, the only thing left standing between a drifting model and the public was whether Farah happened to remember, on her own, to go looking.
So here's the decision I'd take back. In that same spring sprint review, when the launch epic was closing, someone pointed out that the recurring Friday review was tied to a ticket that was about to go away, and asked whether it should live somewhere else. The honest answer at the time was that there was no infrastructure for an ongoing, per cluster sample yet, and building one wasn't in scope for a launch that was already two weeks late. Let the review lapse for now, someone said. We'll figure out monitoring properly once we've shipped a few more things. That was a reasonable call in a sprint review. It stopped being reasonable the day new content arrived that the launch validation set had never once seen.
I'd put the review back, not tied to a ticket that can close, but on its own permanent calendar, split by content cluster, with a number attached: any cluster holding above five percent word error for two weeks running gets flagged automatically, no one has to remember. If that had existed, the documentary series would have tripped the alert in week thirteen, at five point nine percent, long before it reached eleven. Cascade Play would have caught it with a normal, scheduled recaption of one series instead of an emergency weekend job and an apology tour.
That's the whole difference. One design hands the feature a finish line and calls the job over. The other hands it a heartbeat, something that keeps happening on its own schedule, whether or not anyone remembers to ask about it.
And the part I'd tell myself, if I could go back: we asked whether the model was good enough to launch. We never asked who was still going to be checking six months after everyone had stopped talking about the launch.
FLIPS, for a feature that keeps grading itself
This question dresses up as a definitions question. It's actually a Perturbation question: something keeps drifting after ship day, and FLIPS runs straight down the line, F to S.
And if you want to be sure it really works, try it somewhere else
A commercial real estate firm's AI tool reads lease PDFs and pulls out renewal option deadlines for a lease administrator to track. Same question shape, a completely different product, and a different flip.
F. Yusuf Demir, lease administrator at Corrigan Portfolio Group, six years in, tracking renewal option deadlines across about four hundred commercial leases.
L. He stopped fully re-reading each lease PDF himself once the tool's flagged renewal dates matched his own read, lease after lease, for over a year.
I. A different flip. He doesn't check less carefully. He stops opening the tool at all. After it silently misses a renewal clause on one lease, he goes back to tracking every deadline in his own personal spreadsheet, permanently, even for the leases the tool would have caught correctly. No middle setting between "trusts the dashboard" and "never opens it again."
P. Done for this feature meant passing a one time accuracy test on standard domestic lease formats. Once it passed, the tracker closed the ticket, and nobody scheduled the ongoing sampling that would have caught new formats, like international ground leases, as the portfolio grew. When the tool did miss a clause, it gave no reason why, so Yusuf had no way to trust it selectively.
S. An ongoing done adds a recurring sampled audit across new lease formats, plus a flag on which fields the tool is least sure about. The missed clause on the new ground lease format gets caught by the internal audit within the same month it appears, before Yusuf ever hits it on his own. He keeps using the dashboard for the leases it's actually good at, and no renewal deadline comes within days of lapsing again.
Swap the trigger and it still runs
- Speed: if generating captions took twenty minutes instead of instant, editors would start publishing without them and adding captions later, off the queue, quietly. Same missing ongoing check, a different mechanism.
- Cost: if Cascade Play started paying per minute of video processed, someone would ration which shows get the accurate model versus a cheaper one, and nobody would widen the sampling to match. Worse captions would move through the same trusted pipeline.
- The model got better: the case on this page. Passing the launch bar with room to spare is exactly what got the recurring review cancelled in the first place.
Where people run it wrong
- Blaming the model's accuracy instead of the calendar hold that got deleted along with the ticket.
- Fixing it with a bigger validation set or a stricter launch number, which still only ever tests one day.
- Waiting for a viral clip instead of watching the per cluster trend line, which was already climbing for months.
How to use it live
Say the real question out loud before anything else. "So done here can't mean passed a test once, it has to mean still passing a test." Naming that costs five seconds, and it isn't stalling, it's where the real answer starts, because defining "done" only means something once you've said whether it's a gate or a habit.
Flashcards (click a card to flip it)
Eight fixed slots, pulled straight from the answer above.
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Acceptance criteria for non-deterministic output
- #1 Rewrite this criterion to be testable: the model should not hallucinate.
- #2 How do you express an acceptance criterion as a rate rather than an absolute?
- #3 What is the difference between a threshold criterion and a distributional criterion?
- #4 Write acceptance criteria for an AI feature that extracts fields from an invoice.
- #5 How do you set a pass bar when human performance on the same task is 92 percent?
- #6 Describe acceptance criteria that account for the severity of different error types.