CaseAdvancedEval-Driven Specification / Acceptance criteria for non-deterministic output / #24

Explain when you would accept a lower quality bar in exchange for lower latency.

The direct answer
Take the lower quality bar on captions shown live, while the show is still running, and cut the delay to under about two seconds, because a wrong word usually fixes itself on the next line but a late line loses the room for good. Keep the old, careful pass only where a person can't recover the meaning on their own: a safety notice, a gate number, a one-time announcement, or the transcript delivered after the show is already over.
What to actually do, in order
  1. Cut the live caption delay to about two seconds and accept more wrong words.Why: this is the call the rest of the answer hangs on. A late line costs more than a wrong one, every time.
  2. Name which mistake is cheap and which is hidden before setting any number.Why: turns "which one is worse" from a feeling into something a reader can check.
  3. Carve out the content nobody can recover from context, and keep the slow pass there.Why: a safety notice or a one-time number has no next line to fix it on, so speed can't win there.
  4. Leave the after-the-fact transcript on its own slow, careful path.Why: nobody reading it later is racing a room that has already reacted.
  5. Watch meaning-changing errors, not the raw error rate, as the number that would flip the bar back.Why: a wrong word is noise; a wrong word that flips the sentence's meaning is a caption that's actively lying.

How to say this out loud

Eight moves. This is a commit question first, so the position comes before the reasoning, and the one line about where you'd never make this trade comes before the kill number, not after.

1
Scope it to one concrete moment
Say it like this
"Let's say this is Cuepoint, captions on a screen above the stage at a live conference or concert. I'll write the bar for the words shown live, during the show, not for the transcript that goes out afterward."
Why this works
Grounds a broad tradeoff question in something specific before any numbers show up.
2
State the position before any reasoning
Say it like this
"My answer: for the live feed, take the lower quality bar. Get the words on screen in under two seconds even if more of them come out wrong, because a slow, careful line is a line that arrives after the moment it was for."
Why this works
PICK rewards commitment. Hedging with "it depends" fails the question before the reasoning starts.
3
Reframe what the question is really asking
Say it like this
"This isn't really 'fast or accurate.' It's 'which mistake can she fix herself, and which one costs her something she can't get back.' A wrong word, she reads past it. A late line, she's already missed the room laughing."
Why this works
Moves the argument from a vague sense of caution to a specific, checkable question.
4
Name who feels each kind of mistake
Say it like this
"If a word comes out garbled, she reads it, shrugs, and the next line is usually already right. If the line is seven seconds late, she's not reading a mistake at all, she's reading yesterday's news about a joke everyone else just laughed at."
Why this works
This is the impact step. It splits "quality" into two named costs instead of one blurry worry.
5
Put a number on the asymmetry
Say it like this
"At the old setting, captions ran about seven seconds behind and got ninety-seven words in a hundred right. At the fast setting, they run about a second and a half behind and get ninety-one right. That's more wrong words a day, sure, but compare it to what seven seconds actually costs her: every laugh line in the room, every single night."
Why this works
This is C, the heart of the pick, said as something a reader could check, not a feeling about caution.
6
Name where you'd never make this trade
Say it like this
"I wouldn't touch the bar for anything she can't get from context. An evacuation notice. A gate number read out once. Those keep the full, slow pass, because there's no next line to fix it on if it comes out wrong."
Why this works
Shows the position has a real edge, instead of being a rule that applies everywhere without exception.
7
Give the kill criteria as a measurable line
Say it like this
"I'd flip the whole live feed back toward accuracy the day meaning-changing errors, not typos, ones that flip what a sentence actually says, cross about one in twenty lines. Below that, a wrong word is noise. Above it, the caption is actively lying to her."
Why this works
Shows the bet has an expiry date, instead of being a rule copied from a training deck.
8
Close on the line
Say it like this
"So: for the live feed, take the lower bar, get under two seconds, accept more wrong words. Hold the slow, careful pass only where a person can't recover the meaning on their own. And write down the number that sends you back to careful. Don't leave it as a feeling."
Why this works
Restates the position in one breath, the way you want an answer to end, not trail off.

One more thing before the walkthrough moves on: an interviewer asking this wants to see you hold a position and still name exactly where it stops applying. A quality bar that trades away speed everywhere, or trades away accuracy everywhere, isn't a bar. It's a habit.

Let's learn

The captions live on a screen bolted above the stage, a few words behind everything happening under it. Cuepoint puts them there for people who are deaf or hard of hearing, at concerts and conference livestreams, wherever a whole room reacts together in the same second.

For its first two years, Cuepoint held every caption line back for about six seconds before showing it, giving its model time to double-check its own first guess. That patience paid off: 97 words out of every 100 came out exactly right.

Then Cuepoint tried a faster setup. Drop the six-second hold, show each line the moment the model has a guess. Word accuracy fell to 91 out of 100.

Here's the turn. Those extra wrong words are not the real problem. Most of them fix themselves by the next line, the same way a phone keyboard quietly corrects itself after you've already read the wrong word. The real problem is what a person does with six seconds of silence: they stop watching the stage and start watching a screen that's always a half-step behind the room.

It's not that we made the captions harder to read. We made them arrive after the room had already moved on.
What changes when Cuepoint speeds up the live feed
Better number, that row
Worse number, that row
Seconds behind the live room
Old setting
7.2 sec
Fast setting
1.4 sec
Wrong words, per 100
Old setting
3
Fast setting
9
Neither setting wins on both rows. That's the whole shape of a real tradeoff: the old setting is better on words, the fast setting is better on time, and the answer is about which row actually costs more when it's wrong.

Over one three-day conference, at roughly 40,000 words of stage audio a day, the old setting produced about 1,200 wrong words a day. The fast setting produces about 3,600. Three times as many mistakes, and every one of them is still just as fixable as before, on the next line.

At its worst, seven seconds of silence doesn't cost someone a handful of extra wrong words. It costs them the show. A captioned attendee who is always a beat behind stops trusting the feed to be live at all, starts treating it like a delayed recap, and eventually stops asking for captions, or stops buying the ticket, because watching Cuepoint's version of a concert is lonelier than the concert.

The decision that mattered We built a six-second polish buffer into every caption line, live or not, because the only thing the dashboard graded was whether the words were right. Cut that buffer down to under two seconds for the show itself, and keep it on the transcript that comes after. The same model gets a completely different report card depending on who's actually racing the clock.
Knowledge spark: what is a caption buffer? A short hold before a line shows up on screen, so the model has time to fix its own first guess. A bigger buffer makes each line more likely to be right. It also makes every single line later.
What I would leave alone The corrected file that goes out after the event, the one someone reads later or uses for the official record, keeps the full six-second buffer. Nobody reading that is racing a room that has already reacted and moved on.

The lesson. A live caption only has one job while the show is running: land close enough to the room's own timing that a shared laugh is still shared. Get it slightly wrong and on time, and it still does that job. Get it exactly right and seven seconds late, and it has already failed, no matter how correct it turns out to be.

The reveal Behnaz clapped for seven seconds late

Use this one when there's room to feel what seven seconds actually costs. The short version above has the same shape, just none of the room.

Yara Kessab can tell, from the back row of any venue, how many seconds the captions are trailing the stage, without once looking at a clock.

She's the product manager who owns the quality bar for Cuepoint's live captions, and she wrote the original spec herself, back when the whole promise to venues was "every word exactly right." For fourteen months, she kept a habit: fly out for the biggest events, sit in the back row with a stopwatch, and time the lag with her own eyes, line by line.

The numbers held steady the whole time. Seven seconds behind, 97 words in a hundred right, month after month. So the trips got shorter, then rarer, then she let the weekly accuracy report speak for itself instead.

Then came the Northbeam Summit, a three-day conference with twelve thousand people watching the keynotes live. Behnaz Rostami has come to Northbeam every year since it started, and every year she's sat in the accessible-seating section, reading Cuepoint's captions off the screen above the stage.

Day two, the keynote's closing reveal. The presenter says the line that gives it away. The room gasps, then breaks into applause before he's even finished the sentence. Behnaz is still reading the sentence before that one. She looks up at the noise, sees three thousand people on their feet, and has no idea why. She looks back down. The caption for the reveal finally lands, seven seconds after the room already knew.

It happens twice more before the keynote ends. Same seven seconds. Same gap. Same look up at a room that's already moved on without her.

Two boxes of unequal weight: a small calm box for one garbled word, which fixes itself next line, next to a large jagged red-orange box for a caption landing seven seconds late, felt as always being behind the room
Same show, two very different mistakes to make
She wasn't behind on the words. She was behind on the room.

That's what Soren Vikander said, walking out of the hall. Soren audits accessibility for a handful of Cuepoint's venues, was two rows back that day for an unrelated reason, and had spent most of the reveal watching Behnaz's face instead of the stage. He said it to Yara plainly, not as a complaint, just as something he'd noticed.

Yara pulled the year's numbers that afternoon, for the first time since she'd stopped flying out. Word accuracy: still 97 in 100, exactly where it had always been. Lag: still seven seconds, exactly where it had always been. Nothing on the dashboard had moved. Nothing had ever been logged as broken.

But in Cuepoint's own post-event survey, more than half the attendees using captions at Northbeam said the same thing Behnaz did, without being asked: they'd stopped watching the stage during the big moments and started watching other people's faces instead, using the caption screen only to catch words afterward, the way you'd check a friend's recap of a movie you already half saw.

Here's what I'd take back. When we built the live feed, we gave every line the same six-second polish pass the archived transcript gets, because the only number on our dashboard was "percent right." Nothing on it measured how far behind the room a caption user was sitting. Cut the live buffer to under two seconds, even though accuracy drops to 91 in 100, and Behnaz's caption for the reveal lands about a second and a half after the presenter says it, not seven seconds after the room already knew.

The next Northbeam, same keynote slot, fast captions running. Behnaz's screen shows the reveal line 1.6 seconds after the presenter says it. She's on her feet with everyone else within half a second of the person next to her, not seven seconds behind them.

And the thing I'd tell myself, standing in that hallway: we graded a live caption like it was a transcript. A transcript's only job is to be right. A live caption's job is to be right now, and a little bit wrong beats correct and late, right up until the thing being said is something nobody can afford to get wrong at all.

PICK, timed to the room

This is a tradeoff wearing a documentation question's clothes, same as any "when would you accept X" question, so PICK is the tool here, not a list of captioning best practices.

P, position. Take the lower quality bar on the live feed specifically. Get every line under about two seconds, and accept more wrong words in trade.
I, impact. The room itself feels nothing, it isn't reading captions. Behnaz feels every gap between what the room does and what her screen says, over and over, for the length of the whole show. She also feels a wrong word, but it fixes itself a line later.
C, cost asymmetry. A wrong word is cheap and visible: it's obviously off, and the next line usually already caught up. A late line is hidden and expensive: nothing marks it as a mistake anywhere on a dashboard, it just quietly separates her from the room, night after night, until she stops trusting the feed to be live at all.
K, kill criteria. Flip the live feed back toward the slow, careful pass the day meaning-changing errors, not typos, cross about one in twenty lines. Or for content she can't recover on her own: a safety notice, a gate number, a one-time schedule change, which always gets the slow pass regardless of the day's setting.
Knowledge spark: what makes an error "meaning-changing"? A typo like "cat" for "car" is annoying, but the sentence still makes sense. Dropping the word "not" flips the sentence into saying the opposite of what was actually said. One is noise. The other is a caption that's actively lying.
Where meaning-changing errors take over, as word accuracy drops
Meaning-changing errors, per 100 lines
Kill point: 14 percent word error rate
0 2.5 5 kill: 14% 3% 6% 9% 14% 0.3 0.6 1.2, current 4.6
Meaning-changing errors barely move between 3 and 9 percent word error. Past about 14 percent they jump nearly four times over. That bend in the line is the kill point, not a number picked out of the air.

The same pick, in a radiology reading room

Cedarholt Hospital runs an AI flag called First-Read across every CT scan that comes through its emergency department, looking for the kind of blocked vessel that causes a stroke. When it fires, that scan jumps to the top of the radiologist's reading queue, ahead of whatever was already waiting.

Here the position doesn't flip. Take the lower quality bar, more false flags, for the faster one, because what a slow, careful flag costs here isn't measured in an annoyed room. It's measured in brain tissue.

P. Flag fast and flag wide. Accept that a good share of the flags will turn out to be a clean scan.
I. A false flag costs a radiologist about ninety seconds: open the scan, see it's clean, move to the next case. A real stroke that waits its turn costs the patient roughly two million brain cells for every minute it sits unread, the reason emergency teams call stroke care a race against a clock that never stops.
C. The false flag is cheap and visible, over in under two minutes, nobody worse off. The slow, careful read is hidden and catastrophic: nothing about a scan sitting fourth in a queue looks like an emergency, right up until someone finally opens it.
K. This flips only if false flags climb high enough that radiologists stop trusting the jump and start reading in the old order anyway, roughly past one in three flags coming back clean. Below that line, keep flagging wide.

What stays the same, in the reading room. A scan from a patient already mid-treatment for something else, with no new symptoms, doesn't need the same aggressive flag as a first ER visit. Some scans are safe to read in the order they arrived.

What twenty-eight minutes in a slow queue costs, in brain cells
18 min, waiting
6 min, first read
4 min, stroke call
Waiting behind non-urgent scans, about 18 minutes
Time to first read once opened, about 6 minutes
Time to call the stroke team, about 4 minutes
Twenty-eight minutes at roughly two million neurons a minute is about 56 million neurons gone before treatment even starts. First-Read's fast, wide flag gets a read moving in about 3 minutes, roughly 6 million, before the same clock does its damage.

Swap the trigger and it still runs

  • Speed: captions render twice as fast on the same model. Doesn't move the pick, at the venue or the hospital. Speed already won; this just makes winning cheaper.
  • Cost: the model gets three times cheaper to run per hour. Doesn't move it either. The risk was never the price.
  • The model gets better: the fast setting's accuracy climbs from 91 to 95. Moves the bar, doesn't erase the kill criteria. A rare meaning-changing error is still not zero, and a rare missed stroke still isn't either.

Where people run it wrong

  • Writing "keep the delay low" as the whole quality bar, without ever naming what that trades away.
  • Treating the kill criteria as optional, so speed only ever wins and nobody revisits it.
  • Setting one bar for every piece of content, safety notices included, instead of naming where a wrong word can't be fixed on the next line.

If you're asked this cold

Say the reframe out loud before naming a number. "Before I set the bar, I want to know which mistake this person can fix themselves, and which one they can't." That's true, it buys a few seconds, and it's already stage one of the real answer.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's its hardest step?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming which mistake is cheap and self-correcting (a garbled word) and which is hidden and compounding (a caption that lands seconds after the room has already reacted), then setting the bar around the second one.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Yara Kessab, the product manager who owns the quality bar for Cuepoint's live captions. She wrote the original spec and used to fly out and time the lag herself, row by row, at Cuepoint's biggest events.
3 · THE HABIT
What did Yara stop doing because the old bar seemed to be working?
Tap to flip
ANSWER
Flying out to sit in the back row with a stopwatch and time the caption lag by eye. After fourteen months of steady numbers, seven seconds behind, 97 words in a hundred right, she let the weekly dashboard report replace the trip.
4 · THE ASYMMETRY
What's the cost asymmetry in this story?
Tap to flip
ANSWER
A garbled word is cheap and visible: it's obviously wrong, and the next line usually already catches it. Seven seconds of lag is hidden and expensive: nothing on a dashboard marks it as broken, it just quietly costs a captioned attendee the room's shared reaction, over and over, all night.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Take the lower quality bar on the live feed. Get every caption under about two seconds, even if more words come out wrong, because a late line costs more than a wrong one ever does.
6 · THE NUMBER
Fill in the blank: at the fast setting, Cuepoint's live captions run about ___ seconds behind the room instead of seven.
Tap to flip
ANSWER
1.4, roughly a second and a half. Word accuracy drops from 97 to 91 in every hundred, about three times as many wrong words a day, and every one of them is still cheaper than the seven seconds it replaced.
7 · THE KILL CRITERIA
Name the evidence that would flip this bar back toward taking the full six seconds.
Tap to flip
ANSWER
Meaning-changing errors, not typos, crossing about one in twenty lines. Or content a person can't recover on their own: a safety notice, a gate number, a one-time schedule change, which always gets the slow, careful pass regardless of the setting.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and does the pick flip?
Tap to flip
ANSWER
Cedarholt Hospital's First-Read stroke-flag tool, in a radiology reading room. The pick doesn't flip: it still favors the faster, wider flag, because the cost of a slow, careful read isn't an annoyed room, it's brain cells, roughly two million a minute.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these is the cost asymmetry (the C step) in this answer?
  • A. Live captions cost more to run per hour than archived captions do.
  • B. A garbled word is cheap and self-correcting; a caption that lands seconds late is hidden and quietly costs the whole room's shared reaction, over and over.
  • C. Attendees said they would rather have no captions at all than imperfect ones.
  • D. Cuepoint's support team costs more to run than the captioning model itself.
Show hint
One of these has a real number on each side, in units the answer actually measures.
Show answer
B. A, C, and D are never established anywhere in the answer. B is the one with a checkable cost on each side: cheap and self-fixing against hidden and compounding.
True or false
2. True or false: this answer says Cuepoint should run every caption, everywhere, at the fastest possible setting, no matter what's being said. Why or why not?
  • True
  • False
Show hint
Look at the K step, and at what happens to the archived transcript.
Show answer
False. The bar keeps the slow, careful pass for content a person can't recover on their own, and for the transcript delivered after the event, where nobody is racing the room.
Fill in the blank
3. Fill in the blank: under the old setting, live captions ran about ___ seconds behind the room.
Show hint
The same number Yara used to time with a stopwatch from the back row.
Show answer
Seven. The fast setting cuts that to about 1.4 seconds, dropping word accuracy from 97 to 91 out of 100 in trade.
Multiple choice
4. Why couldn't Cuepoint just trim the buffer a little, say from six seconds to four, instead of cutting it all the way to under two?
  • A. Four seconds is still long enough to land after the room's reaction to a joke or a reveal, so the same disconnect from the room keeps happening.
  • B. The model can only run at two fixed speeds, with nothing in between.
  • C. Cuepoint's contract requires captions to render in exactly 1.4 seconds.
  • D. A four-second buffer would cost significantly more money to run than a six-second one.
Show hint
Ask what Behnaz is actually racing: not a stopwatch, the room's own reaction time.
Show answer
A. B, C, and D are never stated in the answer. A room reacts in roughly a second, so trimming the buffer partway still leaves a caption arriving after the laugh or the applause has already started. Only getting close to the room's own reaction time actually fixes it, that's a flip, not a dial.
Short answer, apply it yourself
5. Pick a product you use yourself that gives you information live, live traffic, live sports scores, a live chat. Name one place it's probably tuned to be fast over exactly right, and who eats the cost when it's wrong.
Show hint
Look for the small, constant wrongness you've stopped noticing, and ask what it's trading away.
Show answer
Model answer: "My maps app updates traffic live, and the arrival time it shows is sometimes off by a minute or two. It's tuned for speed, because a slightly wrong time that keeps updating beats a perfectly calculated route that took so long to compute I've already missed the turn it was planned around. I eat that cost, a slightly wrong ETA. If it got the road closure itself wrong instead of just the timing, I'd eat a worse cost: driving into a street that's actually shut." Any answer works if it names a specific mistake the tool is quietly tuned to make, and who pays for it.
Short answer
6. If the fast setting's word accuracy had dropped to 80 out of 100 instead of 91, would the position, take the lower bar for the live feed, still hold? Walk through it.
Show hint
Ask whether a wrong word is still fixable by the next line, or whether whole sentences stop making sense.
Show answer
Probably not at that exact setting, but the direction survives. At 91 in 100, a wrong word is still an occasional, individually fixable slip. At 80 in 100, that's one wrong word in every five, dense enough that whole sentences stop reading cleanly, not just single words. The lag problem doesn't disappear at that point, but the caption itself becomes hard to follow rather than just occasionally rough. The honest fix isn't to abandon speed, it's to find a middle setting, maybe a two-to-three-second buffer, that still cuts most of the seven-second gap while keeping accuracy readable. The pick's direction survives. Cutting the buffer all the way to zero does not.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more