CaseIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #14

Design a spike to determine whether latency requirements can be met.

SPARKa caption that had to beat the anchor's voice, not a stopwatch

Cape Adair News 9 is piloting CaptionCast, a model that generates live closed captions and a Spanish-language translation overlay for the evening broadcast. Ines Duarte is the AI PM who has to design the spike that decides whether the model can actually keep up with a real broadcast, not a quiet lab.

The direct answer
Build the spike around real broadcast conditions, not a clean bench test: replay an actual past live segment's audio through the model at the same time as the real concurrent system load it would face on air, across multiple simultaneous feeds, and measure the 95th-percentile delay against the hard on-air ceiling, never the average. A latency spike that only tests one clean feed at a time will pass in the lab and fail the first night two anchors talk over each other.
Do this, in order
  1. Test the 95th-percentile delay under real concurrent load, not the average under a single clean feed.Why: a broadcast that looks fine on average can still blow past the ceiling on the specific segment where two feeds overlap.
  2. Use a past real segment's actual audio, not a scripted demo clip.Why: real speech has cross-talk, accents, and pacing a clean script never tests.
  3. Set the hard on-air ceiling before running a single test, not after seeing the numbers.Why: a ceiling picked after the fact bends to whatever number the model happened to hit.
  4. Stage the spike from bench test up to a live shadow run, not straight to air.Why: each stage catches a different failure mode before it reaches a real viewer.
  5. Leave the full multi-language pipeline out of week one.Why: proving the single-language latency case first is what the spike is actually for; the rest can wait.

How to answer this, stage by stage

Nobody is scoring whether you know what "P95" stands for. They are scoring whether the spike you design would have actually caught the failure that a clean bench test would have missed.

Stage 1
Ground it in the actual workflow, before you
Say it like this
"Before any of this, a live captioner listens to the broadcast and types captions in real time, correcting herself as she goes, and pushes them to air a few seconds behind the anchor's voice."
Why this works
SPARK opens on the real, existing workflow so the design decision that follows has something concrete to replace.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as SPARK: situation, payoff, anchor, risk, keep out. The anchor is the actual spike design, and that's the answer to the question."
Why this works
Names a repeatable structure before diving into the design, which tells the interviewer you have a method, not just an opinion.
Stage 3
Reframe: it isn't "can the model caption fast," it's "does it stay fast when the broadcast gets messy"
Say it like this
"This isn't really about whether the model can produce a caption quickly on a quiet test clip. It's about whether it still clears the ceiling once two feeds are talking over each other and the network is doing what it actually does on a Tuesday night."
Why this works
This is where a strong answer separates from someone who just says "test the latency" without saying under what conditions.
Stage 4
Give the one concrete design decision
Say it like this
"Here's the anchor: replay a real past broadcast segment's audio, run it at the same time as the concurrent system load a live night actually produces, and measure the 95th-percentile delay against the hard on-air ceiling, not the average."
Why this works
This is the concrete, arguable decision the whole answer hangs on, exactly what SPARK's anchor step is built to produce.
Stage 5
Prove the anchor survives its own risk
Say it like this
"On a clean single feed, the model averaged 200 milliseconds. Under a real two-feed broadcast load, the 95th-percentile delay hit 1,400 milliseconds, more than a full second past our 900-millisecond ceiling. The average alone would have said we were fine."
Why this works
Naming the exact number the average would have hidden is the strongest single line SPARK's risk step can produce.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't just a normal load test is that a model's response time isn't fixed like a database query, it shifts with input complexity, and real speech under real network jitter is a different input than a clean demo clip. We accepted a slower, staged spike, bench test first, then concurrent load, then a live shadow run, in exchange for never finding out about this gap live on air."
Why this works
This names the load-bearing, AI-specific judgment: a model's latency is a distribution shaped by the input, not a fixed number you can bench-test once and trust.
Stage 7
Say what's out of scope, then close
Say it like this
"I'd leave the full multi-language pipeline out of week one, since proving single-language latency is what actually answers this question. So: real past audio, real concurrent load, 95th percentile against a hard ceiling, staged before it ever touches a live broadcast."
Why this works
Closing with what you deliberately left out shows judgment, and restates the anchor decision in one breath.

Let's learn

Every night at 6:58, two minutes before air, a live captioner at Cape Adair News 9 put on her headset and got ready to listen to the broadcast and type what she heard, a few seconds behind the anchor's own voice, correcting herself in real time when she fell behind.

Before CaptionCast, that captioner produced captions with about a four-second delay behind the anchor, consistently, night after night, for a single English-language feed. With CaptionCast running in the lab on a clean single audio file, that delay dropped to under a quarter of a second on average.

Hand sketched flow diagram titled Today without you: how a caption gets on screen by hand. Four steps left to right, second step emphasized: Listen to feed. Type live. Self-correct. Push to air.
The workflow CaptionCast is meant to speed up. Notice there's no single moment that's obviously slow, just four steps done fast enough to matter.

Here's the turn: a quarter-second average sounds like an easy win. But a broadcast is never one clean feed. On a real night, two reporters talk over each other during a handoff, the network hiccups during a satellite segment, and three other processes are competing for the same server. None of that shows up in a bench test run on a single quiet audio file.

Caption delay: bench test versus real broadcast load
1600ms 800ms 0 900ms ceiling 200ms Bench, single feed, average 1400ms Real load, 95th percentile
The average alone said the model was five times faster than the old captioner. The 95th percentile said it was over the ceiling.

At its worst, CaptionCast ships on the strength of that clean bench number, and the first night two feeds genuinely overlap, captions lag so far behind that the anchor is already three sentences past what's on screen, live, in front of viewers and a sponsor watching that exact segment.

A model that is fast on a quiet Tuesday afternoon test and a model that is fast on a real broadcast are not the same claim, and only one of them was ever tested.
The choice I would take back When the spike plan was first drafted, the team scoped it to test one caption stream at a time, since that was the simplest thing to measure cleanly and the vendor's own demo ran the same way. That made sense when the goal was just proving the model could caption at all. It stopped making sense once the actual question became whether it holds up during the one part of a real broadcast, the handoff, where two feeds genuinely run at once.

What I would leave alone: I wouldn't build the full multi-language translation pipeline into this spike. Proving single-language latency under real load is the whole question here; adding Spanish translation now would just make the failure, if there is one, harder to isolate.

The lesson: we asked "how fast is the model," when the real question was always "how fast is the model when the broadcast is doing what broadcasts actually do." Those are different tests, and we only built the first one.

Now here is the same thing as a story

The short version above is what you'd say defending the spike design in a five-minute standup. Read this one for what it felt like the night a near miss almost became the whole reason the spike existed.

The control room at Cape Adair goes quiet in a specific way right around 6:55, headsets going on, one last check of the rundown, someone's coffee going cold on the console.

Ines had run three model pilots before this one, and each time the pattern was the same: get a clean number from a lab test, show it to the newsroom, get cautious approval, ship it slowly. CaptionCast's lab number was the best she'd seen. Quarter of a second, on a clean single feed.

Hand sketched labeled parts diagram titled The anchor close up: what the spike actually measures. A gauge icon at the center labeled Spike Design, with four labeled callouts around it: Real past broadcast audio, Concurrent real system load, 95th percentile measured, Hard on air ceiling.
The one design decision the whole spike turns on: measure the worst 5 percent, under real load, against a number set in advance.

Nobody had asked for a load test yet. It felt like a later-stage concern, something you'd worry about once the model was already proven fast in principle.

Then came a near miss during an actual live test segment, run quietly during a slow Tuesday broadcast with CaptionCast shadowing the live captioner rather than replacing her. Two reporters, doing a routine field handoff, talked over each other for about four seconds. CaptionCast's captions fell so far behind that by the time they caught up, the anchor had already moved to the next story. The live captioner's captions never lagged past two seconds. Nobody outside the control room ever saw it, because the human captioner was still the one actually on air.

Knowledge spark: why does overlapping speech slow a model down so much? A captioning model has to process audio as it arrives and produce text with almost no lag. Overlapping voices make the input genuinely harder to untangle in real time, the same way a person finds two people talking at once harder to follow. A clean, single-speaker test never puts the model in that position at all.

That near miss is what changed the spike plan. Ines went back to the lab number and realized it had never actually been tested under the one condition that mattered most: two things happening on the broadcast at the same time.

Hand sketched comparison titled The day it's wrong. Left panel, gauge icon labeled Bench test, caption clean single feed, 200 milliseconds average. Right panel, question mark box icon labeled Real broadcast, caption overlapping feeds, 1400 milliseconds at the 95th percentile mark.
Same model. One number came from a quiet room. The other came from the one moment broadcasts actually get hard.

So the spike got redesigned. Instead of one clean feed, they replayed a real past broadcast's actual audio, including an actual field handoff with real overlapping speech, run concurrently with the same server load a live night produces. They measured the 95th-percentile delay, not the average, against a 900-millisecond ceiling the broadcast engineer had set before anyone saw a single result.

Hand sketched decision tree titled What we left for later: how much load the spike includes. Root: what conditions does the spike test. Three branches: one clean feed alone leads to misses the real risk. Concurrent feeds and real network leads to catches the actual failure. Full multi language pipeline leads to overkill for week one.
Three ways to scope the same spike. Only the middle one actually answers the question being asked.

The 95th-percentile number came back at 1,400 milliseconds, more than half a second past the ceiling. Not a small miss. A real one, and one that would have shipped invisibly if the spike had stayed scoped to a single clean feed.

Hand sketched timeline titled The spike plan, staged over one week, third milestone emphasized. Four milestones: Bench baseline day 1. Concurrent replay day 3. Live shadow test day 5. Go or no-go review day 7.
Four stages, each one catching something the stage before it couldn't. Skipping straight to day 7 was the version that almost shipped.

The fix wasn't a faster model. It was moving one part of the pipeline off the shared server so the concurrent-feed case stopped competing for the same resources during a handoff. Re-run, the 95th-percentile delay came back at 780 milliseconds, under the ceiling, and the near miss never had to happen live.

What I'd tell myself, hearing how close that Tuesday came to being a live failure instead of a quiet one: the lab number was never a lie. It was just an answer to a question nobody had actually asked yet.

SPARK, in one screenNot a script for load-testing every feature to death. SPARK is what tells you exactly which real condition your lab number never faced.

S
Situation. Who is this, and how does the job get done today?
A live captioner at Cape Adair News 9, typing captions in real time a few seconds behind the anchor, entirely by hand.
One real person, one real shift, not a segment of "broadcast operations."
P
Payoff. What habit should this build?
The newsroom stops treating a lab latency number as proof the model is ready, and starts asking "under what load."
The habit is the actual product here: a team that questions a clean number instead of shipping on it.
A
Anchor. The one design decision everything hangs on.
Replay real past broadcast audio at real concurrent system load, and measure the 95th-percentile delay against a ceiling set before the test runs.
This is the answer to the question. Everything else in the answer explains and defends this one decision.
R
Risk. What breaks the first time it's wrong?
A single-feed bench test would have shown 200 milliseconds and shipped, then failed live the first night two feeds genuinely overlapped, exactly what the near-miss shadow test caught quietly instead.
Naming the specific broadcast moment that breaks it, not just "it could be slow," is what makes this a real risk instead of a vague worry.
K
Keep out. What stays out of the spike, on purpose?
The full multi-language translation pipeline. Proving single-language latency under real load is what this spike needs to answer first.
Shows judgment about scope instead of trying to test everything at once and learning nothing clearly.

The recap, one line per letter: situation is the live captioner doing it by hand today, payoff is a newsroom that questions lab numbers instead of trusting them, anchor is testing the 95th percentile under real concurrent load against a preset ceiling, risk is the exact broadcast moment, an overlapping handoff, where the clean bench number would have failed live, and keep out is leaving the multi-language pipeline for a later spike.

And if you want to be sure it really works, try it somewhere elseSame five letters, an HVAC field-diagnostic tool instead of a broadcast. Different flip family entirely, the same lab-number trap.

Emory Vantassel is a technician at Grayling Hollow HVAC, testing a mobile app that suggests likely diagnoses to technicians standing in front of a broken unit. Mapped onto SPARK: situation is Emory diagnosing units by ear and gauge readings today, entirely without the app. Payoff is technicians trusting a suggested diagnosis enough to skip a redundant manual check, saving a real visit. Anchor is testing suggestion latency against a hard field ceiling: a technician standing at a unit will not wait past four seconds before trusting their own read instead. Risk here is a different flip: not a scope flip like CaptionCast's, but a pre-editing flip, technicians start describing the problem in short, clean phrases that get a fast suggestion instead of the messy real symptoms, quietly feeding the model an easier version of the actual problem. Keep out is a full parts-ordering integration, which can wait for a second spike entirely.

The same hand sketched decision tree reused: what we left for later, how much load or detail the spike includes, applied here to an HVAC diagnostic tool instead of a broadcast captioning tool.
The same three-way scoping question, asked of a different product entirely: how much of the real, messy input does the spike actually include.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "test the 95th percentile under real concurrent load, against a preset ceiling, before it ever touches air," and stop.
Cost: there's no time to build a full concurrent-load replay rig. Say so honestly, and run the cheapest real version, two archived segments played back at once, rather than skipping straight to a single clean feed.
The model gets faster, for real: even a clear speed improvement still needs the same test, since a faster average can still hide a worse tail under load nobody checked for.

Where people run it wrong.
They test one clean input at a time and call the average good enough.
They pick the latency ceiling after seeing a number they like, instead of before the test runs.
They let one clean pilot convince them the model will hold up under every future condition, instead of testing the next real one too.

How to use it live. The moment an interviewer asks you to design a latency spike, ask yourself: what's the one real condition where this system's inputs actually overlap or compete? Test that condition specifically, and the rest of the design follows.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Scope flip: the spike itself narrowed from testing real broadcast conditions down to one clean feed, the same shrinking-scope pattern a person falls into under time pressure.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ines Duarte, the AI PM at Cape Adair News 9, designing the latency spike for CaptionCast, the live-captioning model.
3 · THE HABIT
What habit did Ines's past pilots build, that stopped working here?
Tap to flip
ANSWER
Trusting a clean lab number as proof a model was ready, without testing it under the one real condition, overlapping feeds, where broadcasts actually get hard.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Testing the whole real broadcast condition, concurrent feeds and all, versus testing only one clean feed at a time. There is no middle scope that catches the actual failure.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Scoping the spike to test one caption stream at a time, because it was the simplest thing to measure cleanly, before anyone asked whether that matched a real broadcast.
6 · THE NUMBER
Fill in the blank: on a clean bench test the average delay was ___ milliseconds, but under real concurrent load the 95th percentile hit ___ milliseconds.
Tap to flip
ANSWER
200 milliseconds on the bench test, 1,400 milliseconds at the 95th percentile under real load, against a 900-millisecond ceiling.
7 · THE REPLAY
Same near miss, redesigned spike already in place. What changes?
Tap to flip
ANSWER
The concurrent-load test catches the 1,400-millisecond delay before air. Moving one pipeline stage off the shared server brings the 95th percentile down to 780 milliseconds, under the ceiling, and the near miss never happens live.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Grayling Hollow HVAC's field-diagnostic app. The flip there is pre-editing: technicians start describing symptoms in short, clean phrases to get a faster suggestion, quietly feeding the model an easier version of the real problem.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: the spike measured the ___ percentile delay under real concurrent load, not the ___.
Show hint
Look at the anchor step in the SPARK recap.
Show answer
95th percentile; average. The average hid the exact problem the 95th-percentile number revealed: a real, rare-but-costly delay under load.
Multiple choice
2. Why did the clean bench test's 200-millisecond average fail to predict the real problem?
  • A. The model was retrained between the bench test and the live segment.
  • B. It never tested concurrent, overlapping feeds, the specific condition where delay actually spiked.
  • C. The captioner typed faster during the bench test.
  • D. The ceiling was set after the bench test ran.
Show hint
Look at the chart comparing bench test to real broadcast load.
Show answer
B. A single clean feed never puts the model under the same load or input complexity as a real broadcast with overlapping speech.
True or false
3. True or false: the near miss during the live shadow test was seen by viewers at home.
  • True
  • False
Show hint
Look at how the shadow test was run.
Show answer
False. CaptionCast was only shadowing the broadcast; the human captioner's captions were the ones actually on air, so the near miss stayed invisible to viewers.
Short answer, where it wouldn't matter
4. Name a part of this same rollout where the concurrent-load latency test wouldn't matter much.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: An overnight, non-live captioning job for archived footage, where nobody is waiting on the output in real time and a few extra seconds costs nothing.
Short answer, apply it yourself
5. Think of an app or tool you use that felt fast in a demo but slowed down in real use. What real-world condition do you think the demo never tested?
Show hint
Think about what's different between a quiet demo and your actual, busy use of the thing.
Show answer
Model answer: A voice assistant that answers instantly in a quiet room but lags in a noisy kitchen with the TV on, a condition a quiet demo room never tests.
Short answer, work the number
6. If the fixed 95th-percentile delay had come back at 950 milliseconds instead of 780, just over the 900-millisecond ceiling, would you ship it?
Show hint
Think about what the ceiling is actually protecting against.
Show answer
Model answer: No. A ceiling set in advance to protect a real viewer experience should hold even by a small margin; shipping 50 milliseconds over it defeats the purpose of setting the ceiling before the test.
Before you close the answer
Why this works
Tests whether you'll design a spike around the real, messy condition that actually breaks a system, or settle for a clean number that looks good and proves nothing.
Follow-up traps
"Isn't this just a normal load-testing question, nothing AI-specific about it?" Response: no, because a model's latency shifts with input complexity in a way a database query doesn't, overlapping speech is a genuinely harder input, not just more traffic hitting the same fixed operation.

"Why not just always test at maximum theoretical load?" Response: that would catch this problem but miss the point, the spike needs to match the load a real broadcast actually produces, not an unrealistic worst case that never happens and never gets fixed for the actual failure mode.
If pressed
The actual fix moved the model's inference process off the same server handling video encoding, since the two were competing for the same GPU memory during a handoff, a resource-contention detail never visible from the latency number alone.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more