CaseAdvancedAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #16
Explain how you would use scenario planning for a two-year AI product strategy.
SPARKthe sorting line Cascade almost bought for a future that hadn't happened yet
Picture a two-year plan built on one guess about how fast a computer-vision model will improve, and a warehouse that has already spent the money before anyone checks if the guess was right. Cascade Recycling runs BinVision, a camera system that sorts recyclable material on the line. Bianca Ferraz is the product lead who stopped the company from betting two years of capital on a single number.
The direct answer
Don't write a two-year roadmap that assumes one trajectory for how fast the model improves. Write three: steady, slow, and breakthrough, each with a pre-agreed number that proves you're in it, and a pre-written shift in spending for each one. Check the number every quarter. The plan changes when the evidence says so, not when someone's gut says the future finally arrived.
Do this, in order
Write three named tracks with a pre-agreed trigger number for each, before committing any capital.Why: without a trigger written down in advance, every quarter becomes an argument about whether "now" is the moment to bet.
Test any trigger against your own held-out data, not the vendor's benchmark.Why: a model can clear a general benchmark and still miss badly on the one material category your line actually depends on.
Check the trigger every quarter, on a fixed date, not when something feels urgent.Why: a near miss almost got treated as proof the breakthrough had already arrived.
Build a fourth answer: ambiguous, hold the default track.Why: forcing every quarter into one of three tracks when the evidence is genuinely unclear invites a bet nobody actually has grounds for.
Cap it at three tracks, not ten.Why: more tracks feels thorough but nobody can hold ten pre-written plans in their head when a quarterly check actually happens.
How to answer this, stage by stage
Nobody is scoring whether you can predict the future. They're scoring whether you built three answers for the futures you can't predict.
Stage 1
Scope it to one line, one budget
Say it like this
"I'll ground this in BinVision at Cascade Recycling, and the actual capital decision, a second sorting line, that almost got approved for a future that hadn't happened yet."
Why this works
Keeps the answer from turning into an abstract essay about uncertainty in general.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as SPARK. Situation, how the line runs today without any of this. Payoff, the habit I want the team to build. Anchor, the one concrete decision. Risk, what happens the day the anchor is wrong. Keep out, what I'm deliberately not building yet."
Why this works
Two seconds that show the interviewer you're designing forward, not just reacting to a hypothetical.
Stage 3
Reframe: this isn't predicting the future, it's pre-committing to three of them
Say it like this
"The question isn't 'what will vision models look like in two years.' Nobody knows that. The real question is 'what would we do under each of three plausible answers, decided now, so the decision isn't made under pressure later.'"
Why this works
This is where a strong answer separates from a forecasting exercise dressed up as strategy.
Stage 4
Give the one decision
Say it like this
"Three tracks: steady, slow, breakthrough. Each has a trigger number tested against our own line data, and a pre-written spending shift. We check the trigger every quarter, on the same fixed date."
Why this works
This is the direct answer, stated as a concrete artifact instead of a philosophy about flexibility.
Stage 5
Prove it survives being wrong
Say it like this
"A vendor upgrade looked like the breakthrough track arriving early. We almost approved a second sorting line on that read. Then someone checked film-plastics accuracy specifically, not the vendor's general number, and it had dropped 11 points. The anchor held, and the line order got paused instead of placed."
Why this works
Compresses the whole near miss into the exact moment the pre-committed trigger check paid for itself.
Stage 6
Say what you'd measure past launch
Say it like this
"Every quarter, on the same date, we check the trigger number against our own held-out material stream, never the vendor's published benchmark alone."
Why this works
Shows the plan isn't a one-time document, it's a standing check.
Stage 7
Say what you'd leave alone
Say it like this
"I wouldn't build ten scenario tracks, and I wouldn't re-forecast the whole two-year plan every quarter. Three tracks, checked on a fixed schedule, is the whole method. More than that, nobody actually uses it under pressure."
Why this works
Shows judgment instead of a wish list of extra rigor nobody will maintain.
Stage 8
Close on the one line
Say it like this
"So the two-year plan isn't a guess about the future. It's three pre-written answers and a quarterly check to see which one we're actually living in."
Why this works
Restates the direct answer in one breath and closes on the decision, not the story.
Let's learn
Here is what a two-year AI strategy needs, so a company doesn't spend two years of capital on a guess nobody wrote down.
Before BinVision, workers on Cascade's sorting line pulled recyclable material off a belt by eye, roughly forty items a minute, sorting by material type into separate bins. BinVision adds a camera and a classifier above the belt that flags material type before it reaches the worker, cutting missed items by more than half. Cascade's leadership wanted a two-year plan for how much further to automate the line, and the first draft assumed one number: the model's accuracy would keep improving at roughly the same steady pace it had for the last year.
This is the process the two-year plan is actually about speeding up, one truck at a time.
Here's the turn: the danger wasn't picking the wrong pace of improvement. It was writing a plan that only had one pace in it at all. When a vendor's upgrade looked like it might be arriving faster than expected, the team had no pre-agreed answer for what that would mean, so the temptation was to treat any promising number as proof the fast future had already arrived, capital and all.
BinVision accuracy on film plastics specifically, trailing eight quarters
The general benchmark kept climbing through quarter six. Film plastics, the one category this plan actually depended on, quietly dropped 11 points at the same time.
At its worst, betting on the wrong trajectory doesn't just waste one quarter's planning meeting. Cascade almost placed an order for a second automated sorting line, a six-figure commitment, on the strength of a benchmark number that didn't hold on the material the line actually processes most.
What the false signal would have cost, versus what holding the track actually cost
The line chart above shows the number that almost triggered the wrong track. This one shows why checking it against real data first was worth the four hours it took.
The choice I would take back
The first draft of the two-year plan wrote down one assumed pace of model improvement and built every later decision on top of it. That made sense when the team was just trying to get a plan approved quickly. It stopped making sense the moment a single vendor upgrade could have quietly redirected two years of capital based on the wrong reading.
What I would leave alone: the worker training schedule for the sorting line doesn't need a scenario plan at all. Whether the model improves quickly or slowly, workers still need the same onboarding for how to work alongside the camera system, so that part of the roadmap stays fixed regardless of which track the company ends up in.
The lesson: a two-year AI plan doesn't fail because the forecast was wrong. Forecasts are always a little wrong. It fails because nobody wrote down, in advance, what to actually do about each way it could be wrong.
Now here is the same thing as a story
The short version above is what you'd present to a capital-planning committee. Read this one for how a wall-mounted terminal on the sort-line floor almost triggered a purchase nobody had actually agreed to make.
The terminal that shows BinVision's live accuracy numbers is bolted to a post near the start of the sorting line, where Bianca Ferraz can see it from the catwalk above. Cascade hired her as its only product person eighteen months before this happened, reporting straight to the plant's general manager.
Three named tracks and one fixed check. This is the whole plan, not a summary of it.
For five straight quarters, the terminal told a boring, reassuring story. Accuracy climbed a little each quarter, roughly the pace everyone had assumed. Bianca's scenario plan sat mostly unused, a document nobody had needed to open.
Knowledge spark: why does a general benchmark hide a specific problem?
A vendor's benchmark tests a model against a broad mix of material photos. It says nothing about film plastic specifically, the thin, crinkled, hard-to-classify bags and wraps that make up a large share of what Cascade's line actually handles. A model can genuinely improve overall and get quietly worse on exactly the category a business depends on most.
Then, in the sixth quarter, the vendor pushed a model upgrade. The published benchmark jumped. In the plant's weekly ops meeting, someone floated approving the second sorting line early, since "the model's clearly hitting the breakthrough numbers now." It was the kind of moment that, without a written plan, becomes a decision made on excitement rather than evidence.
The anchor didn't prevent the near miss. It caught it before the order went in.
Bianca's scenario plan had a rule for exactly this moment: no track shift counts until the trigger number is tested against Cascade's own held-out material, not the vendor's general score. The test came back the next morning. Film plastics accuracy, the category driving most of the line's actual sorting decisions, had dropped 11 points even as the overall benchmark rose. The breakthrough track's trigger hadn't been met at all.
Nobody had to argue Cascade out of buying a sorting line for a future that hadn't happened yet. The plan had already written down what "yes, it's real" would need to look like, and this wasn't it.
Four branches, not three. The fourth one, hold and recheck, is what stopped the near miss from becoming a purchase.
So here is the decision I would take back: writing a two-year plan around one assumed pace of improvement in the first place. It made sense when the goal was just getting a plan approved fast. It stopped making sense the moment a single vendor headline could have redirected two years of spending on its own.
The near miss lands right where a fixed quarterly check was already scheduled to look.
With the scenario plan in place, the replay is almost boring. The vendor's benchmark jumps in quarter six. Someone in the ops meeting gets excited. The terminal's numbers get checked against Cascade's own material the next morning, same as every quarter. The plan holds to the Slow track, the sorting-line order stays paused, and two quarters later, once film plastics accuracy genuinely clears the trigger, the Breakthrough track activates on purpose, not on a headline. And the thing I'd tell myself, if I could go back: the plan was never supposed to predict which quarter the future would arrive. It only needed to say, in advance, how we'd know when it had.
SPARK, in one screen
S
Situation. How does the job get done today, without this?
Workers sort recyclable material by eye off a moving belt, and Cascade's leadership wants a two-year plan for how far to automate that line.
Grounds the plan in one real belt, not an abstract "AI roadmap."
P
Payoff. What habit do you want this to build?
A quarterly habit of checking a pre-agreed number against real line data, instead of debating in the moment whether a headline is real.
That habit, not any single forecast, is what the plan is actually for.
A
Anchor. The one design decision everything hangs on.
Three named tracks, steady, slow, breakthrough, each with a trigger number tested on Cascade's own data and a pre-written spending shift.
This is the hardest step, and the direct answer to the whole question.
R
Risk. What breaks the first time you're wrong?
A vendor upgrade looks like the breakthrough track arriving, but the general benchmark hides a real regression on film plastics specifically.
The anchor survives this because the trigger requires Cascade's own data, not the vendor's number, before anything shifts.
K
Keep out. What you deliberately won't build yet.
No ten-track scenario map, no full quarterly re-forecast, no betting on one exact accuracy number two years out.
Three tracks and one fixed check is a plan people will actually use under pressure.
The recap, one line per letter: situation is workers sorting by eye off a belt, payoff is the habit of checking a written trigger instead of reacting to headlines, anchor is the three named tracks with pre-agreed spending shifts, risk is a benchmark hiding a real regression on the material that matters most, and keep out is capping it at three tracks with one fixed quarterly check.
Three tracks and a fixed date beat ten tracks nobody checks under real pressure.
And if you want to be sure it really works, try it somewhere else
Meridian Pharmacy Group runs AuthFlow, a tool that drafts prior-authorization requests for prescriptions, deciding which cases route straight to an insurer and which need a pharmacist's review first. Deshawn Pruitt, Meridian's pharmacy operations manager, faced the same two-year question: bet the roadmap on a steadily improving classifier, or hold room for a frontier model jump that could handle far more cases without a pharmacist touching them at all. Mapped onto SPARK, situation is pharmacists reviewing every borderline case by hand today, payoff is the habit of checking a written trigger before expanding automation, and the anchor is the same three-track structure: steady, slow, breakthrough, each with a number tested against Meridian's own denial-and-appeal data, not a vendor's general accuracy claim. The risk plays out differently here than at Cascade: a new model looked ready to handle appeals unsupervised, clearing a general benchmark easily, but it was quietly worse on the specific insurers whose rules change most often, exactly the cases pharmacists were already spending the most time on. Keep out is the same discipline: three tracks, one quarterly check, no attempt to forecast the exact model six months out.
The breakthrough track always carries the most to lose if the trigger check is skipped, in a pharmacy or on a sorting line alike.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "three tracks, one trigger number each, tested on our own data, checked every quarter," and stop.
Cost: no budget for a formal quarterly review process. Say so honestly, and start with just the three written triggers and a calendar reminder, since writing them down costs nothing and stops the worst version of this mistake.
The model gets better, for real: if the breakthrough trigger genuinely clears on real data, that's the plan working as intended, and the honest move is to activate that track on purpose, not treat a headline as proof it already happened.
Where people run it wrong.
They build the scenario tracks but never write the trigger number down in advance, so the check becomes a debate in the moment anyway.
They trust the vendor's own benchmark instead of testing against their own data, and miss exactly the regression that matters most.
They build ten tracks to feel thorough, and then nobody actually checks any of them once real work gets busy.
How to use it live. The moment someone asks about a multi-year AI strategy, ask yourself: what three futures am I actually planning for, and what number would prove which one we're in? Name both, and the rest of the answer follows.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one-line job?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Its job is designing against the failure before you build, not describing a feature after the fact.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bianca Ferraz, Cascade Recycling's only product person, who built a three-track scenario plan for BinVision's two-year roadmap.
3 · THE HABIT
What habit did the scenario plan build?
Tap to flip
ANSWER
Checking a written trigger number against Cascade's own line data every quarter, instead of deciding in the moment whether a headline is real.
4 · THE ANCHOR
What's the one concrete decision this whole answer hangs on?
Tap to flip
ANSWER
Three named tracks, steady, slow, breakthrough, each with a pre-agreed trigger number and a pre-written spending shift.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing the first draft of the two-year plan around one assumed pace of improvement, which made sense only because it got the plan approved quickly.
6 · THE NUMBER
Fill in the blank: film plastics accuracy dropped ___ points in quarter six even as the vendor's general benchmark rose.
Tap to flip
ANSWER
11 points, which is why the breakthrough track's trigger wasn't actually met despite the exciting headline number.
7 · THE REPLAY
Same vendor headline, plan already in place, what changes?
Tap to flip
ANSWER
The trigger gets tested against Cascade's own data the next morning, comes back short on film plastics, and the sorting-line order stays paused instead of getting placed.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what varied in the risk?
Tap to flip
ANSWER
Meridian Pharmacy Group's AuthFlow. The risk shifts to a model that clears a general benchmark but is quietly worse on the specific insurers whose rules change most, the exact cases pharmacists spend the most time on.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: Cascade's two-year scenario plan uses ___ named tracks, not ten.
Show hint
Look at the direct answer and "keep out."
Show answer
Three. Steady, slow, and breakthrough. More than three, and nobody checks all of them under real pressure.
Multiple choice
2. Why didn't the vendor's rising benchmark in quarter six count as clearing the breakthrough track's trigger?
A. Because the vendor's benchmark had been discontinued that quarter.
B. Because the trigger required a test against Cascade's own film-plastics data, which had actually dropped 11 points despite the general benchmark rising.
C. Because Bianca hadn't approved the vendor's upgrade yet.
D. Because the breakthrough track had already been used once that year.
Show hint
Look at the knowledge spark and the line chart.
Show answer
B. The rule was that no track shift counts until it's tested against Cascade's own data, and that test came back short.
True or false
3. True or false: this answer recommends re-forecasting the entire two-year plan from scratch every quarter.
True
False
Show hint
Look at "keep out" and the icon list of what was left for later.
Show answer
False. The plan checks a fixed trigger number each quarter against three pre-written tracks, it doesn't rebuild the whole plan from zero.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Writing the first plan around one assumed pace of improvement. It made sense because it got a plan approved quickly, and stopped making sense once one headline could redirect two years of spending.
Short answer, apply it yourself
5. Pick a plan you're making yourself, over months or years. What are two different ways it could actually go, and what number would tell you, in advance, which one you're living in?
Show hint
Think of a savings goal, a fitness plan, or a career move with more than one likely path.
Show answer
Model answer: A savings plan with a "steady" and a "windfall" track, where a specific bonus amount received is the pre-agreed number that shifts spending from the steady plan to the windfall one.
Short answer, where it wouldn't matter
6. Name a part of Cascade's roadmap where the scenario plan genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The worker onboarding schedule for the sorting line. It stays the same regardless of which track the model improvement lands in.
Before you close the answer
Why this works
Tests whether you'll treat a two-year AI plan as a forecast to defend, or as three pre-committed answers and a check to see which one is real.
Follow-up traps
"What if the evidence lands right between two tracks?" Response: that's exactly what the fourth branch, hold and recheck in 90 days, is for; forcing an ambiguous quarter into a clean track invites a bet with no real grounds.
"Isn't three tracks still just a guess about which futures are plausible?" Response: yes, and that's fine, the goal isn't a perfect map of the future, it's having a written answer ready instead of deciding under pressure when a headline lands.
If pressed
The actual trigger Cascade settled on requires a 5-point or larger accuracy gain, sustained across two consecutive quarterly checks, on Cascade's own held-out film-plastics stream specifically, not a single quarter's number and not the vendor's aggregate score.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.