CaseAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #20
Your spike succeeded on curated data and failed on real data. What did you learn?
TRACEthe week Lior started relistening to everything
Meridian Wireless runs customer care for about two million wireless subscribers. CallDigest is the AI feature that listens to a support call and drafts a summary for the agent's notes. Btissam Haddad is the AI PM who ran the spike that greenlit it. Lior Ben-David is a senior support agent who has to live with what the summary leaves out.
The direct answer
The lesson is not "the model is worse than we thought." It's that a curated spike and real traffic are two different tests, and nobody checked whether the curated one actually looked like real calls. Before trusting a spike result again, recut it by the segments real traffic will actually contain, rule out anything but the model first, and run one test that tells you whether the gap is the data or the model, before you assume either.
Do this, in order
Recut the failure by real call type, not just the average rate.Why: a 61 percent average can hide one segment at 91 percent and another at 42 percent.
Rule out instrumentation before blaming the model.Why: a broken transcript pipeline looks exactly like a worse model on a dashboard.
Name three specific causes for the gap, not one vague "real calls are harder."Why: a named cause is testable, a vague one just sounds true.
Run the one test that separates your top two causes.Why: rerunning the curated calls through the real pipeline tells you if the data changed or the model did.
Rebuild the eval set to include noise, code-switching, and transfers going forward.Why: the next spike should fail here first, not in a live customer's ear.
Say when a curated-only spike is still fine, like a low-stakes internal draft with a human editing every line.Why: shows judgment about where the extra rigor earns its cost.
How to answer this, stage by stage
Nobody is scoring whether you can say "real data is messier." They're scoring whether you can turn that gap into a specific, ranked list of testable causes.
Stage 1
Scope it to one real gap
Say it like this
"Let's ground this in CallDigest at Meridian Wireless, and the specific gap between its spike score and its real-week score."
Why this works
Keeps the answer from becoming a general essay on curated versus real data.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as TRACE. Timeline, when it actually broke. Recut by segment. Assume nothing about instrumentation. Name cause candidates. Run one evidence test."
Why this works
Signals a repeatable way to investigate a gap, not a single guess dressed up as an answer.
Stage 3
Reframe: it isn't "the spike lied," it's "the spike answered a different question"
Say it like this
"This isn't really a case of a spike being wrong. It's a case of the spike honestly answering 'how good is this model on forty clean calls,' when the question that mattered was 'how good is this model on the calls we'll actually get.'"
Why this works
This is where a strong answer separates from someone who just says "the data was different."
Stage 4
Recut the number by segment
Say it like this
"When I recut the real week by call type, single-issue calls held at 91 percent. Multi-issue calls dropped to 42. The 61 percent average was hiding a segment that had basically failed."
Why this works
This is the direct answer to what the average number was quietly covering up.
Stage 5
Rule out the boring explanation first
Say it like this
"Before I blamed the model, I checked the transcript pipeline on both sets of calls. Same pipeline, same version, no dropped audio. That ruled out instrumentation and left the actual call content as the difference."
Why this works
A tracking or pipeline change looks exactly like a real regression on a dashboard, and ruling it out first is the step most people skip.
Stage 6
Name the cause, and the AI-specific reasoning
Say it like this
"The honest reason this isn't a generic testing question is that a model's score is only ever a claim about the distribution it was scored on, and forty curated calls were never that distribution. I'd rebuild the eval set around noise, code-switching, and transfers, the exact three things the curated set had none of."
Why this works
This is the load-bearing, AI-specific judgment: a benchmark score describes a sample, not the feature.
Stage 7
Say what wouldn't change, then close
Say it like this
"I wouldn't rebuild the eval this heavily for an experimental internal-only summary a team lead reads for their own notes and nobody else relies on. For CallDigest, the lesson holds: a curated spike answers a narrower question than the one you actually asked."
Why this works
Closes with real judgment about where the concern doesn't apply, and restates the direct answer in one breath.
Let's learn
What happens the first time a spike's number and a real customer's week disagree with each other?
Before CallDigest, an agent wrote up call notes by hand after every call, about four minutes each, across roughly 2,000 calls a week. With CallDigest, a draft summary appears the moment the call ends, and an agent reads it instead of writing from scratch.
The five letters, held up as one page. Recut is the step this question is really testing.
Here's the turn: the spike ran on forty curated calls, each a single clean issue, one speaker at a time, and CallDigest scored 93 percent against a human-written summary. It shipped on that number. The real week held calls with background noise, two issues in one call, and customers switching languages mid-sentence, none of which the curated forty ever had.
Mishandled-escalation rate, week by week after launch
Nobody was watching this line until a QA audit in week 6 finally sampled the escalations directly.
At its worst, a customer calls about a billing error and a service outage in the same call, CallDigest's summary only captures the billing part, the agent trusts it and closes the case, and the outage complaint vanishes from every record except the customer's memory of being ignored.
The spike never lied. It answered a smaller question than the one Meridian actually needed answered.
The choice I would take back
The team told agents, at launch, "you can trust the summary, you don't need to relisten to the call." That felt safe when the spike's 93 percent was the only number anyone had. It stopped being safe the moment real call variety, never once represented in that curated forty, started showing up in the mix.
What I would leave alone: I wouldn't rebuild this eval for an experimental, internal-only summary feature a team lead reads purely for their own notes, since nothing downstream acts on it unsupervised.
The lesson: a spike score is a fact about the sample it ran on. Forty curated calls can be a perfectly honest 93 percent and still tell you almost nothing about a real Tuesday.
Now here is the same thing as a story
The short version above is what you'd say defending the fix in a retro. Read this one for what the six weeks before the audit actually felt like from the floor.
The support floor at Meridian gets loud around 2pm, once the lunch-hour call spike rolls in behind the morning's backlog.
The third step is where a curated spike's promise either holds up on a real call or quietly doesn't.
Lior Ben-David had worked the floor for six years and could summarize a call in his head before he finished typing it. When CallDigest launched, he read the first week of summaries carefully, checking them against his own memory of each call. They matched, every time. By week two he stopped checking and just read the summary.
Same word, test call, and only one version ever had a second issue buried inside it.
Knowledge spark: why would the same model score so differently on curated calls versus real ones?
A model's score describes the sample it was tested on, not the feature in general. Forty curated calls with one clean issue each test a narrower skill than "summarize whatever a real customer says," which can include two problems tangled together, background noise, or a call transferred twice before it's resolved.
Three weeks in, a customer called about a billing error, mentioned a dropped-service outage almost as an aside, and CallDigest's summary only captured the billing line. Lior read it, trusted it, and closed the case. The outage complaint never made it into any system Meridian could see.
Four things the curated forty never had a single example of.
By week six, a QA lead pulled thirty recent escalations at random and listened to every one against its CallDigest summary. Nine of the thirty had missed a second issue entirely. The pattern wasn't random: every miss was a call with more than one problem in it.
The gap had been building for five weeks before anyone went looking for it.
Two suspects cleared fast. The real one was hiding in plain sight, in the calls nobody had tested against.
The real question was never whether CallDigest's model had gotten worse. It was whether the forty calls that greenlit it had ever looked anything like the calls it would actually receive.
Real-week accuracy, recut by call type
The 61 percent overall average was three very different numbers wearing one costume.
When CallDigest was first proposed, someone said, "let's spike it on a clean batch of calls so we can see it work," and it sounded reasonable, since a clean batch is the fastest way to see anything work at all.
Rerun the same six weeks with the eval set built from real call variety from day one: the multi-issue gap shows up in the spike itself, at 44 percent, before launch. The team ships CallDigest for single-issue calls only, routes multi-issue and transferred calls to a flagged "review in full" mode, and the missed-outage case never happens, because that call would have been flagged, not summarized and closed.
What I'd tell myself, hearing about the missed outage complaint: forty clean calls were never a test of the feature. They were a test of the easiest version of the job.
TRACE, the gap that turned out to be three numbersNot a script for distrusting every spike that passes. TRACE is what tells you exactly which slice of "real" the spike never touched.
T
Timeline. When did it actually start?
The mishandled-escalation rate began climbing in week 1, right after launch, well before the week 6 audit caught it.
The metric started drifting the same day real call variety started arriving, not the day the audit ran.
R
Recut. Slice it: segment, not average.
Single-issue calls held at 91 percent. Multi-issue calls fell to 42. Transfer and hold-heavy calls landed at 55. The 61 percent average hid all three.
This is the hardest step, and the one that turns a mystery number into a named, testable gap.
A
Assume nothing. Rule out instrumentation first.
Same transcript pipeline, same model version, no dropped audio, checked before blaming the model at all.
A pipeline bug looks exactly like a worse model on a dashboard, and ruling it out is the step most people skip.
C
Cause candidates. Three named guesses.
Instrumentation change, the model itself degrading, or the real call mix simply being harder than the curated forty.
Naming three, not one, is what keeps the investigation honest instead of confirming the first guess.
E
Evidence test. The one query that decides.
Rerunning the original curated forty through the exact production pipeline still scored 92 percent, confirming the model and pipeline were unchanged, and the real call mix was the actual difference.
One clean test, not a debate, is what separated the two live hypotheses.
The recap, one line per letter: timeline is the rate climbing from week 1, recut is 91/42/55 hiding inside one 61 percent average, assume nothing is checking the pipeline before the model, cause candidates is instrumentation, model, or call mix, and evidence test is rerunning the curated forty through production and confirming the call mix was the real difference.
And if you want to be sure it really works, try it somewhere elseSame five letters, a fishing co-op instead of a call center. Different flip family entirely, the same missing segment.
Elke Brandsma runs product at Sturgeon Bay Fisheries Co-op, where a model estimates a catch's weight from a dockside photo. The spike ran on curated studio photos of one fish at a time and scored well. Mapped onto TRACE: timeline is the estimate error creeping up over the first three weeks of real dock use. Recut is a scope flip, not a verification one: workers didn't start double-checking every estimate, they started separating fish onto individual trays and re-photographing them one at a time so each photo looked more like the curated set, quietly adding twenty minutes back to every unloading. Assume nothing ruled out the scale hardware first. Cause candidates named lighting, pile overlap, and wet-net glare; evidence test showed pile overlap alone explained most of the error, since single-fish photos on the same real dock scored nearly as well as the curated set.
A different flip entirely: not relistening to everything, but quietly rebuilding the curated conditions by hand on a real dock.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "recut the failure by segment before blaming the model, real data usually hides one bad slice inside an average," and stop.
Cost: no time for a full re-audit before a fix ships. Say so honestly, and pull even ten real multi-issue calls to confirm the segment before committing engineering time to a fix.
The gap is actually good news, for real: if a recut shows the curated spike underestimated real performance instead, that's still worth reporting honestly, since the goal is an accurate picture, not a flattering one.
Where people run it wrong.
They see one average number drop and look for one cause, instead of recutting by segment first.
They blame the model before ruling out a pipeline or instrumentation change.
They patch the specific bad call they heard about instead of rebuilding the eval set to catch the whole segment.
How to use it live. The moment an interviewer describes a spike that passed and then failed, ask yourself: what slice of real conditions did the passing sample never include? Name that slice out loud, and the rest of TRACE follows on its own.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Verification flip: once a missed outage complaint surfaced, Lior went from trusting every summary at a glance to relistening to every call in full.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Btissam Haddad, the AI PM at Meridian Wireless, who ran the curated spike that shipped CallDigest, and Lior Ben-David, the agent who caught what it missed.
3 · THE HABIT
What did Lior stop doing by week two?
Tap to flip
ANSWER
He stopped checking CallDigest's summaries against his own memory of the call, since every summary he'd checked in week one had matched.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Trusting a summary at a glance versus relistening to the full call before closing it. No middle setting once a missed issue actually reached a customer.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling agents at launch that they could trust the summary and skip relistening, a promise made with no calibration for how quality would vary by call type.
6 · THE NUMBER
Fill in the blank: the curated spike scored ___ percent, the real week averaged ___ percent, and recutting it showed multi-issue calls at only ___ percent.
Tap to flip
ANSWER
93 percent curated, 61 percent real average, 42 percent on multi-issue calls specifically.
7 · THE REPLAY
Same six weeks, eval set built from real call variety from day one. What changes?
Tap to flip
ANSWER
The multi-issue gap shows up in the spike itself, at 44 percent, before launch. Multi-issue and transferred calls route to a flagged review mode, and the missed-outage case never happens.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Sturgeon Bay Fisheries Co-op's catch-weight estimator. The flip is scope: workers started separating and re-photographing fish one at a time to match the curated conditions.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: recutting the real week's 61 percent average by call type showed single-issue calls at 91 percent and multi-issue calls at only ___ percent.
Show hint
Look at "recut the number by segment" in the walkthrough.
Show answer
42 percent. The average of 61 percent was hiding a segment that had nearly failed outright.
Multiple choice
2. What was actually ruled out before the team concluded the real call mix was the cause of the gap?
A. Whether agents liked using CallDigest.
B. Whether an instrumentation or transcript-pipeline change, not the model or the data, explained the drop.
C. Whether the curated forty calls were recorded on the right microphones.
D. Whether Meridian's subscriber count had changed.
Show hint
Look at the "assume nothing" step.
Show answer
B. Checking the pipeline and model version first is what let the team confidently blame the real call mix instead.
True or false
3. True or false: the evidence test found that the model's own weights had degraded since the spike.
True
False
Show hint
Look at the "evidence test" step in the TRACE recap.
Show answer
False. Rerunning the original curated forty through production still scored 92 percent, showing the model and pipeline hadn't changed. The real call mix was the actual difference.
Short answer, where it wouldn't matter
4. Name a situation where a curated-only spike would still be reasonable, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: An experimental, internal-only summary feature a team lead reads purely for their own notes. Nothing downstream acts on it unsupervised, so a curated spike is enough to gauge early promise.
Short answer, apply it yourself
5. Think of a tool that worked great in a demo or trial and then struggled once you used it for real. What real-world condition was probably missing from whatever convinced you to adopt it?
Show hint
Think about what made the demo or trial data cleaner than your actual use.
Show answer
Model answer: A voice assistant that understood every command in a quiet demo room, then struggled at home with a television on in the background, background noise being the missing real-world condition.
Short answer, work the number
6. If multi-issue calls make up 30 percent of Meridian's real call volume instead of a smaller share, does the 42 percent multi-issue accuracy matter more or less to the overall rollout decision?
Show hint
Think about how segment size changes how much a weak segment drags down the whole feature.
Show answer
Model answer: More. At 30 percent of volume, the weak multi-issue segment affects roughly 600 calls a week, which argues strongly for routing that segment to a flagged review mode rather than shipping it unsupervised.
Before you close the answer
Why this works
Tests whether you'll investigate a curated-to-real gap methodically, by segment and cause, or jump straight to "the model must be worse than we thought."
Follow-up traps
"Couldn't you have just used a bigger curated set to begin with?" Response: size wasn't the problem, composition was. A thousand more clean, single-issue calls would still have missed noise, code-switching, and transfers entirely.
"How do you know the recut segments aren't just noise from a small sample?" Response: the multi-issue segment held over 200 real calls in that one week alone, and the evidence test confirming the pipeline and model were unchanged rules out a sampling fluke as the explanation.
If pressed
The rebuilt eval set now pulls a stratified weekly sample by call type, not a random one, specifically so the multi-issue and transfer segments always have enough volume to trust the recut number, instead of being drowned out by the far more common single-issue calls.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.