CaseAdvancedModel Fluency & the AI PM Role / Working with ML engineers and researchers / #3

Describe how you would run a planning session when effort estimates are genuinely unknowable.

BOUND · pricing a diarization fix nobody could date, at Tapecourt, an AI podcast transcription and summarization tool

Tapecourt transcribes a podcast episode and pulls out speaker-labeled quotes for the show notes. Nerissa Eichhorn is the product manager for its diarization pipeline, and Vidar Isaksson is the research engineer who owns it. Cobblecast, an independent podcast network, runs Tapecourt across its whole slate, and Gudrun Moreland, Cobblecast's head of production, wants a date for fixing the wrong-guest quotes coming out of The Roundhouse, its five-guest flagship show, after one nearly went out under the wrong name.

The direct answer
Don't ask "how long will this take." Separate what's knowable, the cost of one experiment, the number you can afford to run before checking in, and the bar that counts as good enough, from the one thing that truly isn't knowable: how many experiments it takes. Commit to a checkpoint cadence with a real go or no-go decision at each one, never a single ship date.
Do this, in order
  1. Replace "how long will this take" with "how many iterations can we afford before we check in," and own that number out loud.Why: the exact number of tuning cycles needed to clear a labeling bar on five-way crosstalk is not knowable up front. A budget is.
  2. Price one iteration in real dollars before setting any budget.Why: one diarization tuning cycle costs about $3,560 in Vidar's time and compute. Without that number, "how many can we afford" is just another guess wearing a budget's clothes.
  3. Set the ship bar as a number, not a feeling.Why: a diarization error rate under 9 percent is what keeps a Roundhouse quote trustworthy enough to publish on its own. "Better than it is now" gives nobody a stopping point.
  4. Fix the cadence, not the finish line.Why: a checkpoint every 3 weeks, about 4 iterations, gives Gudrun a real update on a real date, without anyone pretending to know week nine's result in week zero.
  5. Put a confidence floor in front of every quote today, not after the model improves.Why: it's what actually stopped the near miss from becoming a published mistake, and it should keep running while the checkpoints do, not wait for them to finish.
  6. Sanity-check the worst-case budget against what one bad quote actually costs.Why: even 8 iterations, the top of the range, runs about $28,480, a fraction of the $96,000-a-year Cobblecast contract one misattributed quote could put at risk.

How to answer this, stage by stage

Nobody is grading whether you can say "diarization error rate" correctly. They're grading whether you can turn "we genuinely don't know how long this takes" into a real plan, not a guessed date and not a shrug.

01
Scope it to one product and one real session
Say it like this
"Let's make this real. Tapecourt transcribes podcast episodes and labels who said what. Cobblecast's five-guest panel show, The Roundhouse, is where our diarization model struggles most. That's the planning session I'm walking through."
Why this works
Keeps a question about planning under uncertainty tied to one measurable product instead of a generic answer.
02
Name the method before touching a number
Say it like this
"I'd run this as BOUND. Break down what's actually unknowable from what isn't. Own every number I can. Give a range instead of a date. Check it against something real. Then say what would move the estimate most."
Why this works
Two seconds of structure tells the room real numbers are coming, not a guess dressed up as confidence.
03
Split the unknowable from the knowable, the B step
Say it like this
"Nobody can tell you today exactly how many tuning experiments it takes to get diarization error under nine percent on five people talking over each other. That part's genuinely open. But the cost of running one experiment, how many we can afford before checking in, and what 'good enough' means, all three of those we can pin down right now."
Why this works
This is the line that stops the whole session from turning into either a fake date or a shrug.
04
Own the numbers, the O step
Say it like this
"One diarization tuning cycle runs about three engineer days plus around two hundred sixty dollars of compute, so call it three thousand five hundred sixty dollars an iteration. That's not a guess. That's Vidar's last four cycles, averaged, measured against a held-out set of Roundhouse episodes the model has never trained on."
Why this works
A number with a source behind it survives a follow-up question. A round guess doesn't.
05
Give a range instead of a date, the U step
Say it like this
"Based on how this kind of tuning usually goes, I'd expect somewhere between four and eight iterations to clear the bar. At roughly four iterations every three weeks, that's a checkpoint in three weeks and, worst case, another at six. I'm not promising week six. I'm promising you'll know a lot more in week three than you do right now."
Why this works
Turns "we don't know" into something a nervous stakeholder can actually plan a launch around.
06
Run the sanity check, the N step
Say it like this
"Worst case, eight iterations, that's about twenty eight thousand four hundred eighty dollars. Cobblecast's contract with us runs ninety six thousand a year, and one misattributed quote on a debate show is exactly the kind of thing that puts a renewal in question. Spending that much to protect against another one isn't a hard call."
Why this works
Turns a budget number into a decision someone can actually defend in the room, not just a figure on a slide.
07
Name the direction, then close on what the session produces, the D step
Say it like this
"The one thing that would speed this up most is how much overlapping five-guest audio we can get labeled for training, not more headcount and not a longer deadline. So here's what I'd actually walk out of this session with: a checkpoint every three weeks, a real go or no-go number at each one, and a confidence floor that sends shaky quotes to a person starting today, not after we hit the bar."
Why this works
Closes on the actual deliverable of the session, not a feeling that the meeting went well.

Let's learn

Tapecourt listens to a podcast episode, works out who's talking, and turns that into a written summary with quotes pulled out and labeled by speaker.

Hand sketched left to right flow diagram titled How an episode becomes published show notes. Five connected boxes reading Audio lands, Diarization, this box emphasized in red orange with a thicker border, Transcript drafted, Quotes labeled, Notes published.
Five boxes, one pipeline. The second one, working out who's talking, is where a checkpoint actually needs to look.

Before Tapecourt, a Cobblecast producer spent about 90 minutes after each two-guest interview episode, listening back and typing up who said which quote for the show notes. With Tapecourt running, that dropped to about ten minutes of skimming an already-labeled draft. On those shows, the model's diarization error rate, how often it mixes up who's speaking, sits around 5 percent. Close enough to trustworthy that a producer barely has to check it.

Knowledge spark: what is a diarization error rate? It's the percent of an episode's speech that gets tagged as the wrong speaker, or missed entirely. A low number means the transcript's labels can be trusted next to each quote. A high number means a person has to double check who actually said what.

Then Cobblecast pushed Tapecourt onto The Roundhouse, its flagship weekly panel show, five guests and a host, all talking over each other by the second half of most episodes. The model still runs. It still drafts quotes. But its diarization error rate on Roundhouse audio is 24 percent, not 5. Close to one word in four gets attached to the wrong speaker somewhere in a segment.

Hand sketched two panel comparison titled Two shows, two very different error rates, with a VS between them. Left panel, a gauge icon with a green needle, labeled Two-guest interview, caption clean turn-taking, DER about 5 percent. Right panel, a gauge icon with a red orange needle, labeled The Roundhouse, five guests, caption fast crosstalk, DER as high as 24 percent.
Same model, two shows. The gap between them is exactly what a single average diarization score would hide.

Here's the turn. Those extra mistakes were never really the problem. The problem is what a Roundhouse producer does next: she stops trusting the auto-labeled quotes and goes back to checking almost every one by hand, on the one show where the audience, and the sponsors, are paying the closest attention.

Tapecourt didn't save the Roundhouse producer ninety minutes. It cost her the confidence to stop checking.

What it costs at its worst: three weeks before Gudrun called Nerissa asking for a fix date, a wrongly labeled quote nearly went out under a guest's name he never said it under, on a topic that would have embarrassed him publicly. A Cobblecast editor caught it in a final read minutes before the show notes were due to publish. Nobody outside that room ever saw it. Nerissa still calls it the closest her team has come to a real correction.

The choice I would take back For two years, whenever a stakeholder asked "how long will this take" about a research problem with no known scope, Nerissa let Vidar answer with a single date, because that's the answer a status meeting wants and nobody wanted to be the one saying "we don't know" out loud. I'd take that habit back. Ask "how many iterations can we afford before we check in" instead, and own that number.

What I would leave alone: Tapecourt's two-guest interview shows, still running at 5 percent diarization error, still trusted, still saving a producer most of ninety minutes a week. Nothing about the Roundhouse fix should slow those shows down or change how they ship.

The lesson: an estimate you pull out of someone under status-meeting pressure was never a real estimate. It's a placation, and everyone downstream ends up planning against a number that was never actually meant as a promise.

Now here is the same thing as a story

The short version above is what you'd actually say in the room. Read this one for what almost went out under the wrong guest's name.

Nerissa Eichhorn has run product for Tapecourt's diarization pipeline for two years, ever since Cobblecast became its first real network client. For most of that time, the routine planning question was the same one every team asks: how long will this take. And for most of that time, it was a fine question, because the work behind it had a knowable shape. A new show format to onboard. A latency fix. A change to the show-notes editor. Vidar could look at work like that and give a date that held.

Diarization on The Roundhouse was never that kind of work.

Five guests, a host, and a habit of talking over each other by the second half of most episodes. Tapecourt's model, tuned mostly on two-guest interviews, hit a 24 percent diarization error rate there against 5 percent everywhere else. Nerissa asked Vidar the usual question anyway: how long to fix it. He said six weeks, the way you say a number when a room is waiting for one and admitting "I don't actually know" feels like failing at your job.

Six weeks came and went. The model improved, some. It was nowhere near good enough to publish without a person checking almost every quote, which meant the fix had, functionally, not shipped.

Hand sketched two panel comparison titled What Nerissa stopped asking for, with a VS between them. Left panel, a pink square with a large question mark, labeled ONE DATE, caption a single confident guess, six weeks out. Right panel, a gauge icon with a green needle, labeled A CHECKPOINT, caption a priced range, re-aimed every few weeks.
One date is a guess with a deadline attached. One checkpoint is a guess with a price tag and a return date.

Nobody blamed Vidar for missing the date. He'd done the work the date implied. The problem was the date itself. Nobody had ever asked what it was actually built from, and the honest answer was: nothing. It was a guess, offered under the exact kind of pressure that makes guesses sound like plans.

Three weeks after that six-week date quietly slipped, a segment producer flagged something during final review of a Roundhouse episode. A sharp, borderline-defamatory line about a rival podcast host had been auto-labeled as coming from the wrong guest. The actual guest who said it never got flagged at all. The show notes were forty minutes from publishing. A human editor caught it because she happened to remember the moment from listening live. If she hadn't, it would have gone out under Gudrun Moreland's name on the network's biggest show, attributed to a guest who never said it.

We didn't almost publish a transcription error. We almost published a false quote with someone's name on it.
Knowledge spark: what's a confidence floor? The model gives every quote its own certainty score, not just a transcript. A confidence floor is the cutoff below which a quote doesn't auto-publish. It waits for a person to glance at it first, instead of going out looking exactly as sure of itself as a quote it got right.

Gudrun called that afternoon, not angry, but rattled in a way that mattered more. She didn't ask for six more weeks. She asked when this would actually be fixed, and Nerissa realized she was about to give the same kind of answer that had already failed once.

She didn't. Instead she asked Vidar for a different meeting: not "how long," but "how many experiments can we run before we know more than we do today, and what does each one cost." Vidar had never been asked that question by a PM before. It took him one afternoon to price it. One tuning cycle: about three engineer days, around 260 dollars of compute. History with similar tuning problems put the likely range at four to eight cycles before diarization error would clear a 9 percent bar, the point past which a Roundhouse quote was trustworthy enough to auto-publish again.

Nerissa took that to Gudrun instead of a date. A checkpoint in three weeks, four iterations in, showing exactly where the error rate stood. A second checkpoint at six weeks if the first one wasn't enough. A real number, either way, not a promise made just to end a meeting.

Run it forward. Checkpoint one lands at week three: four iterations, diarization error down to 13 percent. Better, not there yet, but the trend line said something worth funding another round on. Checkpoint two lands at week six: eight iterations total, 28,480 dollars spent, diarization error at 8.6 percent. Under the bar. The confidence floor that had been catching shaky quotes since the near miss stays on, for good, for exactly the audio that's still hardest.

What I'd tell myself, back in that first six-week meeting: a date nobody can defend is not a plan. It's a hope wearing a plan's clothes, and it costs exactly as much as a real miss does, just later, and with less warning.

BOUND, for the planning session where nobody can honestly promise a date

Not a way to sound careful about uncertainty. BOUND is what turns "we genuinely don't know how long this takes" into a number Gudrun can actually plan a sponsor launch around, instead of a guess dressed up as a date.

BBreak it down. Separate what's actually unknowable from what already isn't.
The exact number of tuning cycles it takes to get Roundhouse's diarization error under 9 percent is genuinely open. Nobody can price that today with a straight face. But three other things aren't open at all: what one experiment costs, how many the team can afford to run before stopping to look, and what "good enough to ship" actually means as a number, measured against a held-out set of Roundhouse episodes rather than a feeling.
Skip this and every unknown in the room gets treated the same way, which is how a genuinely open research question ends up with a fake ship date stapled to it.
OOwn the numbers. Where did each one actually come from?
One diarization tuning cycle costs about 3 engineer days plus roughly 260 dollars of compute, call it 3,560 dollars, and that number came straight off Vidar's last four cycles, not a guess pulled from the air. The ship bar, diarization error under 9 percent, came from Tapecourt's own two-guest baseline of 5 percent plus a margin the trust and safety review agreed was still safe to auto-publish against.
A number is only owned if you can say where it came from when someone pushes on it. A figure that sounds specific but has no source behind it is a guess wearing a costume.
What the checkpoint budget actually adds up to
$0 $15k $30k One iteration $3,560 Checkpoint 1 · 4 iterations $14,240 Worst case · 8 iterations $28,480
Engineer time, Vidar's daysCompute, GPU and storage
Engineer time is nearly all of the cost at every stage. Even the worst-case budget, eight iterations, still costs less than a third of Cobblecast's annual contract.
UUse a range, not a date.
History with similar tuning problems put this at somewhere between four and eight iterations before diarization error clears the bar. At roughly four iterations per three weeks, that's a checkpoint at week three and, worst case, another at week six, well inside Cobblecast's real sponsor deadline at week ten.
A single number here would have looked exactly like the six-week guess that already failed once. A range with two priced ends is a plan. A single number is a hope.
Hand sketched horizontal timeline titled The range, laid out as a real calendar. Four milestones along the line: Planning session, week 0, budget set. Checkpoint 1, week 3, 4 iterations, DER 13 percent. Checkpoint 2, ship, this point emphasized in red orange, week 6, 8 iterations, DER 8.6 percent. Cobblecast's real deadline, week 10, sponsor launch.
Four weeks of slack sit between the worst-case checkpoint and Cobblecast's actual deadline. Nobody had to invent that room, it was already there.
NNail the sanity check. Does the number survive contact with something real?
Worst case, eight iterations, runs 28,480 dollars. Cobblecast's contract is worth 96,000 dollars a year, and one more near miss like the one that almost went out under Gudrun's name is exactly the kind of thing that ends a renewal conversation. Spending 28,480 dollars, less than a third of the contract's value, to make sure that doesn't happen again isn't a close call.
This is the step a rushed answer skips. It's what turns a budget from a number on a slide into something you'd actually defend in the room.
Where the checkpoint's iterations actually land
0% 25% target: 9% 24% 13% checkpoint 1 8.6%, ships checkpoint 2 iter 0 iter 2 iter 6
Diarization error rate, by iterationShip bar, 9 percent
The steepest gains come early and taper off, which is exactly why a checkpoint sits at four iterations instead of waiting to see the whole line before deciding anything.
DDirection. Which assumption would move this the most, and what does the session actually produce?
Not headcount, and not a longer deadline. The single biggest lever is how much overlapping, five-guest audio the team can get labeled for training, because that's what each iteration actually learns from. So the session doesn't end with a date. It ends with a checkpoint every three weeks, a real go or no-go diarization number at each one, and a confidence floor that routes any shaky quote to a person, starting immediately, not after the model clears the bar.
Naming the one fact that actually swings the outcome, instead of the biggest number in the arithmetic, is what a real estimator does that a rushed one skips.
Hand sketched labeled parts diagram titled What the session actually walks out with. A document icon labeled Checkpoint plan sits in the center, with four labeled callouts around it: Iteration budget, 4 to 8, owned. DER bar, under 9 percent. Cadence, check in every 3 weeks. Confidence floor, shaky quotes go to a person.
Four things, not one date. This is what Gudrun actually left the room holding.
Hand sketched decision tree titled What happens at each checkpoint. Root box reads Checkpoint, DER under 9 percent question mark. Three branches: yes, cleared the bar, leads to Ship the speaker labels. No, but still dropping fast, leads to Fund one more checkpoint. No, and flat two checkpoints running, leads to Cut scope, route to a person instead.
Three honest endings to a checkpoint, not one hopeful one. Only the middle branch buys another three weeks.

One alternative Nerissa considered and rejected: putting a second research engineer on diarization tuning in parallel, to compress the calendar. It lost, because the bottleneck was never idle hands, it was Vidar's evolving read on exactly where the model was failing on Roundhouse audio specifically, and a second engineer would have needed to relearn that same read from scratch before contributing an independent iteration, doubling the cost per cycle without shortening the number of cycles needed. The AI-specific failure sitting underneath all of this: a wrongly labeled quote reads exactly as confident as a correctly labeled one, nothing in the summary itself signals doubt. The guardrail is the per-segment confidence floor, any quote the model itself isn't sure about gets routed to a person before it publishes, live from the day of the near miss, not held back until the checkpoints finish. And the trade-off is real and accepted on purpose: the team is spending real calendar time and up to 28,480 dollars instead of shipping fast and cheap on a guessed date, in exchange for actually knowing, at each checkpoint, whether the fix is real or just later.

And if you want to be sure it really works, try it somewhere else

Same five letters, a shipping terminal instead of a podcast network, and this time the old habit being replaced isn't the question, it's an assumption baked into how every experiment gets priced.

Hullwatch scans photos of shipping containers at the gate and flags damage before a container gets loaded, run by Eberhard Fallowmead, a computer vision researcher. Ironbridge Terminal runs it across every inbound container, and Xantippe Quaid, the terminal's operations lead, wants to know when Hullwatch will reliably catch hairline stress fractures on refrigerated containers, a damage type the model currently catches only about 55 percent of the time, the same way it already catches surface dents and scrapes about 97 percent of the time.

Run BOUND on it. Break it down: the exact number of experiments needed to get the crack classifier above a usable catch rate is unknowable the same way Tapecourt's was, but the cost of one experiment and the bar for "good enough" aren't. Own the numbers: one vision iteration on Hullwatch's setup costs about 2,900 dollars, cheaper than Tapecourt's because the compute-heavy part is shorter and needs less of Eberhard's own time per cycle, sourced the same way, off his last several runs. Use a range: three to six iterations before reassessing, a shorter range than Tapecourt's because the labeled crack data Ironbridge already had on hand was richer than what Cobblecast could offer at the same stage.

Hand sketched numbered icon list titled What Hullwatch's checkpoint owned instead. Three rows: one, a document icon, One vision iteration, about 2,900 dollars, priced from real runs. Two, a gauge icon, Bar, catch hairline stress cracks, not just dents. Three, a funnel icon, Budget, 3 to 6 iterations before reassessing.
Different terminal, same shape of session. The number that changed wasn't the range, it was what one iteration used to assume about itself.
Where Hullwatch's answer genuinely differs Nerissa's old habit was answering "how long" with a single guessed date. Eberhard's team never did that, they'd learned to give ranges years earlier. Their old habit was different: assuming every iteration meant a full retrain from scratch, because that's how the very first version of the crack classifier was built, back when incremental fine-tuning tools weren't mature enough to trust. Nobody had revisited that assumption since, so their iteration budget was roughly double what it needed to be, half of every cycle spent retraining ground the model had already covered.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: separate what's knowable, the cost and the bar, from what isn't, the iteration count, and turn the unknowable part into a range with a checkpoint, never a single date.
Cost: no budget for a second checkpoint this quarter. Shrink the range instead of hiding it, tell the room honestly that four iterations is all the budget allows, and what "good enough" has to mean if the model isn't there by then.
The model got better, for real: say a newer base model needs half as many tuning cycles to hit the same bar. The discipline barely changes, because it was never really about the iteration count, it was about pricing one iteration honestly before promising anything built on top of it.

Where people run it wrong.
They let "how long will this take" stand as the question, instead of splitting it into the parts that are actually knowable.
They give a range with no priced ends, which is just a wider guess, not an estimate.
They treat "we don't know yet" as the end of the planning session instead of the reason to schedule the next checkpoint.

How to use it live. Before answering, ask yourself one plain question out loud: "what's the one number in this problem I could actually defend right now." Whatever answer comes first is usually where your knowable pile starts.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits running a planning session when effort estimates are genuinely unknowable?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, name the direction. Built for splitting a genuinely open question into a knowable part and an unknowable part, instead of guessing at both.
2 · THE CAST
Who is this answer about?
Tap to flip
ANSWER
Nerissa Eichhorn, product manager for Tapecourt's diarization pipeline. Vidar Isaksson, the research engineer who owns it. Gudrun Moreland, head of production at Cobblecast, the podcast network running The Roundhouse.
3 · THE SPLIT
What's the one line that separates knowable from unknowable in this answer?
Tap to flip
ANSWER
The exact number of tuning iterations needed to clear the bar is unknowable up front. The cost of one iteration, the number the team can afford before checking in, and the ship bar itself are all knowable right now.
4 · THE OWNED NUMBER
Fill in the blank: one diarization tuning cycle costs about $___ in engineer time and compute.
Tap to flip
ANSWER
$3,560, made up of about $3,300 in engineer time and $260 in compute, sourced from Vidar's last four cycles.
5 · THE RANGE
What range did Nerissa give Gudrun, and how many iterations did it actually take?
Tap to flip
ANSWER
Four to eight iterations. It took eight, the top of the range, landing at 8.6 percent diarization error in week six, still inside Cobblecast's real week-ten deadline.
6 · THE OLD HABIT
What decision would Nerissa take back?
Tap to flip
ANSWER
Accepting a single confident date from research instead of asking how many iterations the team could afford before checking in. It made sense for ordinary work with a known scope and broke down on genuinely open research.
7 · THE DIRECTION
Which single assumption would swing this estimate the most?
Tap to flip
ANSWER
How much overlapping, five-guest audio the team can get labeled for training. Not headcount, and not a longer deadline, since that's what each iteration actually learns from.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which one, and what's genuinely different about its answer?
Tap to flip
ANSWER
Hullwatch's container damage inspection AI, at Ironbridge Terminal. The old decision there wasn't a guessed date, it was assuming every iteration needed a full retrain instead of a cheaper incremental fine-tune, which quietly doubled the budget.

Check yourself Score: 0 / 0

Multiple choice
1. Why couldn't Nerissa just tell Gudrun exactly how many tuning experiments the Roundhouse fix would take, up front?
  • A. Vidar refused to commit to any number without a formal research proposal.
  • B. The exact number of experiments needed to clear a diarization bar on five-way crosstalk isn't knowable before you see how the first few land, unlike the cost of one experiment or the ship bar, which are knowable right away.
  • C. Cobblecast's contract terms prohibit Tapecourt from sharing internal estimates.
  • D. The compute budget for the quarter hadn't been approved yet.
Show hint
Check the B step, break it down, and what it separates into two piles.
Show answer
B. Iteration cost, iteration budget, and the ship bar were all knowable immediately. Only the exact iteration count needed was genuinely open.
True or false
2. True or false: the checkpoint plan Nerissa proposed gave Gudrun no timeline at all, just an open-ended "we'll get there when we get there."
  • True
  • False
Show hint
Check the U step and the timeline diagram of the actual calendar.
Show answer
False. She got a real checkpoint in three weeks with an actual diarization number attached, and a worst-case second checkpoint at week six, both inside Cobblecast's real week-ten deadline.
Fill in the blank
3. One diarization tuning iteration cost about $___ in engineer time and compute. The team budgeted a range of ___ to ___ iterations before the first checkpoint decision.
Show hint
Check the O step and the stacked bar chart's first bar.
Show answer
$3,560, and a range of 4 to 8. The team ultimately used all 8, spending $28,480 before diarization error cleared the 9 percent bar.
Short answer, name the reversal
4. What old decision would Nerissa take back, and why did it make sense when the team first fell into it?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Accepting a single confident date whenever someone asked "how long will this take," because that's the answer a status meeting wants, and nobody wanted to be the person saying "we don't know" out loud. It worked fine for ordinary engineering tasks with a known scope, and only broke down once the work became genuinely open-ended research.
Short answer, apply it yourself
5. Think of a piece of open-ended work you've been part of, at a job or in school, where nobody could honestly say how long it would take. What would the checkpoint version of that plan have looked like instead of a single date?
Show hint
A checkpoint version prices one attempt, sets a budget of attempts, and names what "good enough" looks like at each check-in.
Show answer
Model answer: Writing a thesis chapter with an unclear argument: instead of "done in three weeks," it could have been "one full draft attempt costs about a week, budget three attempts, check after each one whether the argument actually holds up, not just whether pages got written."
Short answer, work the number
6. If the team had stopped at the low end of the range, four iterations, what diarization error rate would they have shipped at, and would that have cleared the nine percent bar?
Show hint
Check the line chart's marked point at checkpoint 1, iteration four.
Show answer
13 percent, still above the 9 percent bar. Stopping at the low end of the range would have meant either shipping too early, with real risk of another near miss, or extending the checkpoint anyway, which is exactly why the checkpoint needs a real go or no-go number instead of an assumed stopping point.
Before you close the answer
Why this works
Tests whether you can plan genuinely open research without either faking a date or refusing to commit to anything, and whether you know which numbers in an estimate are real and which are guessed. Most candidates can say "we'll do our best." Fewer can say exactly what's knowable today and price it.
Follow-up traps
"What if Gudrun just refuses to accept anything but a hard date?" Response: give her the honest range and the near-miss story alongside it. A partner who almost got a false quote published under her own name on a live debate show has every reason to prefer a priced checkpoint over another lucky guess.

"Isn't four to eight iterations just a wide guess with extra steps?" Response: no, because both ends are priced and dated, not vibes. Four iterations is three weeks and $14,240. Eight is six weeks and $28,480. And the checkpoint has a real number-based go or no-go test, not a mood check.
If pressed
The confidence floor that routes shaky quotes to a person isn't a single fixed number across every show. It's tuned per format, because a two-guest interview's occasional low-confidence segment is cheap to double check, while Roundhouse's five-way crosstalk needs a stricter floor, since a miss there is more likely to be a live, on-air claim getting misattributed rather than a throwaway aside.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more