ConceptIntermediateAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #9

What is the right planning horizon for an AI product team, and why?

PICK the horizon question is really a question about which error you can't unwind

Verdalane Stream is a streaming service. Wren Ashbrook leads product for its personalization team, and the question on the table is a simple-sounding one that turns out not to be: how far out should an AI product team actually plan?

The direct answer
Give two different horizons to two different kinds of work: plan in years for infrastructure that doesn't care which model wins, and plan in weeks for anything whose value depends on one specific, unshipped model capability. Never use one horizon for both, and agree in advance what evidence would make you cut the long bet short.
Do this, in order
  1. Split the roadmap into capability-dependent work and everything else.Why: these two kinds of work fail in completely different ways, so one horizon can't serve both.
  2. Give infra and reusable pipelines a long horizon.Why: this work pays off whichever model wins, so committing to it early is close to free.
  3. Give any bet on one unshipped model capability a short horizon, in weeks not quarters.Why: the hidden, expensive error is a long promise that outlives the model assumption it was built on.
  4. Write down, in advance, what evidence would shorten a long bet.Why: without a stated kill line, a team keeps believing a bet is fine right up until it isn't.
  5. Keep a standing check-in on any assumption a long bet depends on.Why: the habit of checking is exactly what quietly stops once early results look good.

How to answer this, stage by stage

Nobody is scoring whether you can name a number of weeks. They're scoring whether you can tell which kind of roadmap item deserves which horizon, and why.

Stage 1
Scope it to one team, one bet
Say it like this
"I'll ground this in Verdalane Stream's personalization team, and the specific bet: a year-long roadmap for mood-based playlists, tied to a long-context model that hadn't shipped to general availability yet."
Why this works
Keeps the answer from turning into a generic essay about agile versus waterfall.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as PICK. Position, my actual pick. Impact, who feels each kind of error. Cost asymmetry, which error is hidden and which is cheap. Kill criteria, what evidence would change my mind."
Why this works
Signals a repeatable way to reason about time horizons, not a personal preference.
Stage 3
Reframe: not "how long," but "which error can't you take back"
Say it like this
"The real question isn't 'should we plan for a year or a quarter.' It's 'which kind of roadmap item, if we guess wrong about the model, leaves us with no way to walk it back.' That's the thing the horizon should actually track."
Why this works
This is where a strong answer stops sounding like a scheduling opinion and starts sounding like a risk argument.
Stage 4
Give the position
Say it like this
"Long horizon for the embeddings pipeline and the UI shell, because those pay off no matter which model wins. Short horizon, six weeks at a time, for the mood-playlist feature, because it depends on one vendor's context window actually shipping."
Why this works
This is the direct answer, stated as a specific split instead of a general philosophy.
Stage 5
Prove it with the compressed failure
Say it like this
"Verdalane's twelve-month roadmap put the mood-playlist bet and the infra work on the same clock. The infra shipped fine. The playlist feature depended on a context-window SKU that was still stuck in procurement eight months in, and by then the team had no shorter fallback to fall back on."
Why this works
Compresses the whole failure into the one structural mistake, mixing the two kinds of work onto one horizon.
Stage 6
Close on the one line
Say it like this
"So: the right horizon isn't one number. It's long for what survives either model future, short for what depends on one, and a written kill line so nobody has to notice the bet went bad by accident."
Why this works
Restates the direct answer in one breath, tying the whole method back to the actual question asked.

Let's learn

How far out should an AI product team actually promise something?

Verdalane Stream's personalization team wanted to ship mood-based playlists, generated from a listener's entire history instead of just their last few sessions. Before this project, the team ran on six-week cycles, and every cycle included a short check-in: does the thing we're building still assume what we think it assumes about the model underneath it?

Hand sketched comparison titled The asymmetry, drawn. Left panel, a document icon labeled Short horizon churn, caption a few replanning meetings, cheap and visible. Right panel, a question box icon labeled Long horizon overcommit, caption a promise nobody can unwind when the bet doesn't land, shown in a different color.
Both errors cost something. Only one of them is expensive in a way nobody notices until it's too late to undo.

Leadership wanted something bigger: a twelve-month roadmap, timed to the fiscal year, promising whole-catalog mood playlists by Q4. The plan leaned on a vendor's preview API for a much larger context window, one big enough to hold a listener's full history in one pass. Early tests on a small internal sample looked excellent. Here's the turn: the twelve-month horizon wasn't wrong because it was long. It was wrong because it applied the same clock to a piece of infrastructure that would pay off regardless of the model, and a feature that only existed if one specific, unshipped capability arrived on schedule.

Roadmap items delivered on time, by horizon type
100% 50% 0 92% Short-horizon, infra-independent 34% Long-horizon, capability-dependent
The infra-independent work delivered almost every time. The work betting on one unshipped capability missed two years out of three.

At its worst, a team can spend a year telling leadership a single confident number while quietly having no idea whether the one thing the whole plan depends on will exist on schedule.

The choice I would take back Verdalane set the roadmap horizon at twelve months because that's what matched the fiscal planning calendar, not because it matched how fast the underlying model capability was actually moving. That default made sense for the normal feature work the team had always planned that way. It stopped making sense the moment one line item depended on a vendor SKU still being negotiated.

What I would leave alone: I wouldn't touch the twelve-month horizon for the embeddings pipeline or the recommendation UI shell. Both of those pay off no matter which model the team ends up using, so there's no asymmetry to fix there.

The lesson: the right planning horizon isn't a company-wide setting. It's a property of each individual bet, and the bets that depend on one unproven capability need a much shorter leash than everything else on the same roadmap.

Now here is the same thing as a story

The short version above is what you'd say defending a roadmap review to leadership. Read this one for how a habit of checking quietly stopped, and what it took to notice.

Every Thursday, before the mood-playlist project started, Wren Ashbrook ran a fifteen-minute check-in with the data science lead: does anything we've built this cycle still assume what we think it assumes about the model. It was a small ritual. Almost nobody outside the team knew it happened.

Hand sketched metaphor scene titled Bends or snaps. Left, a scale icon labeled SHORT HORIZON, caption bends without breaking. Right, a box icon labeled LONG HORIZON, caption rigid, snaps if the guess is wrong, shown in a different color.
A short horizon can absorb a wrong guess. A long one, tied to a single unproven assumption, cannot.

When the twelve-month mood-playlist roadmap got announced, the first two months were genuinely good. The vendor's preview API handled small test cases beautifully. Wren kept running the Thursday check-in, and it kept coming back clean: yes, the assumption still holds, yes, the timeline still looks right.

Knowledge spark: what's a context window, in plain terms? It's how much text a model can look at in one pass. A small window means it can only "remember" a little at once, like reading a book one page at a time with no memory of the last page. A larger window lets it hold a listener's entire history in view all at once.

Somewhere around month five, the check-in stopped happening. Nobody decided to cancel it. It just kept confirming the same good news, and eventually the calendar invite got skipped once, then twice, then it quietly disappeared from the recurring schedule altogether.

Hand sketched timeline titled How the assumption check habit thinned out, fourth milestone emphasized. Four milestones: 12 month bet announced, tied to fiscal calendar not model speed. Good months, early prototype impresses everyone. Check ins quietly stop, kept confirming yes so nobody asked. A colleague's remark, wait is that SKU even shipping, shown in a different color.
Nothing dramatic happened at any single step. By the fourth, there was no habit left to catch the problem.
The Thursday check-in, held or skipped, month by month
100% 50% 0 Mo. 1-4: 100% Mo. 5: 60% Mo. 7: 0% Mo. 8: back to 100%
The check-in didn't stop all at once. It thinned out over three months, quietly, before a single offhand remark brought it back.

In month eight, a data scientist mentioned something in passing, in the middle of an unrelated planning meeting: "wait, are we still assuming that context-window SKU ships in general availability? I heard procurement's still negotiating the pricing tier." No alarm went off. Nobody had done anything wrong. It was one sentence, said almost as an aside.

The twelve months didn't fail because the team stopped checking. They failed because the roadmap had been built as one long bet with no shorter piece underneath it, so once the checking stopped, there was nothing left standing on its own.

Wren went looking for a fallback plan, something the team could ship on the current, smaller context window while the vendor sorted out the larger one. There wasn't one. The whole feature had been designed around the large window from day one, no chunked version, no interim design. Four months of engineering time had gone into a single, all-or-nothing bet.

Hand sketched labeled parts diagram titled What earns a long horizon. A document icon at the center labeled Long-Horizon Bet, with four labeled callouts: Benchmark flat two releases running, not tied to one vendor, infra reusable either way, kill criteria agreed up front.
None of the mood-playlist bet's four boxes were checked. All four of the embeddings pipeline's were.
Hand sketched quadrant titled Which roadmap items get which horizon. X axis capability maturity, early unproven to stable mature. Y axis strategic value, nice to have to core bet. Embeddings pipeline placed high maturity, moderate value. Mood playlists two million token feature placed low maturity, high value. UI redesign placed high maturity, low value. Real time voice search placed low maturity, low value.
The mood-playlist feature sits exactly where a bet needs a short leash: unproven, and the whole reason for the roadmap.

What I'd tell myself, watching Wren search for a fallback that was never built: the twelve-month number itself was never the mistake. The mistake was letting one unproven assumption ride on the same clock as everything else, with no smaller, provable piece underneath it.

PICK, choosing which horizon each promise deservesNot a scheduling preference. PICK is what tells you which roadmap item can survive a long promise and which can't.

P
Position. The pick, before any reasoning.
Long horizon for infra that pays off either way. Short horizon, six weeks at a time, for anything riding on one unshipped model capability.
This is the hardest step, and the one the whole answer is actually about.
I
Impact. Who feels each kind of error.
A too-short horizon costs the team a few extra planning meetings. A too-long horizon costs leadership a public commitment they can't walk back and an engineering team four months of work with nothing shippable.
Naming both costs, in real units, is what makes this a decision instead of a preference.
C
Cost asymmetry. Which error is hidden.
Short-horizon churn is cheap and visible, everyone feels the extra meeting immediately. Long-horizon overcommit is hidden for months, then arrives all at once as a missed date with no fallback.
Optimize against the hidden one, since nobody catches it early on their own.
K
Kill criteria. What would change the pick.
If a capability's benchmark performance has held flat across two major model releases in a row, and it isn't tied to one vendor's contract terms, it earns a longer horizon. Until then, it stays on the short leash.
A stated kill line is what separates a real decision from a hope that the vendor timeline holds.

The recap, one line per letter: position is splitting the roadmap by what depends on a model capability and what doesn't, impact is a few meetings against months of unshippable work, cost asymmetry is the hidden overcommit against the cheap and visible churn, and kill criteria is two stable release cycles and no single-vendor dependency before a bet earns a longer leash.

And if you want to be sure it really works, try it somewhere elseSame four letters, a container terminal instead of a streaming service. Different flip family entirely, the same one-bet roadmap.

Northquay Terminal committed to a nine-month rollout of an AI berth-scheduling optimizer, betting on a vendor's next-generation solver that promised to handle the terminal's full seasonal mix of vessel sizes. Mapped onto PICK: position is the same split, long horizon for the yard-mapping infrastructure underneath it, short horizon for the solver-dependent scheduling logic itself. Impact is felt by yard planners, who lose hours to bad berth assignments if the solver isn't ready. Cost asymmetry is the same shape: short replanning cycles are a minor annoyance, but a nine-month promise the whole terminal is counting on is a hidden, expensive bet if the solver slips. Kill criteria would have been the solver's benchmark holding flat across its last two releases, which it hadn't. The flip here is different: workaround, not over-trust. Rasheed Colby, a yard planner uncertain the nine-month bet would land on schedule, quietly started keeping his own spreadsheet predicting berth conflicts by hand, a shadow system nobody asked him to build and nobody at Northquay could see, support, or learn from.

Hand sketched decision tree titled How long should Northquay commit? Root: Is the solver's benchmark stable across releases? Three branches: still shifting release to release leads to plan in 6 week slices. Flat for two releases running leads to commit the full season. Vendor contract still being negotiated leads to build the fallback manual tool now.
The same question, asked before the nine-month commitment went out: what's actually been proven, and for how long.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "long horizon for what survives either model, short horizon for what depends on one, with a written kill line," and stop.
Cost: no time to split the roadmap into two tracks this quarter. Say so honestly, and at minimum tag every item with which kind of bet it is, so the risk is visible even before the process catches up.
The model gets better, for real: if the vendor's capability ships early and holds steady, the honest move is to extend the horizon on that item and say so, not treat every capability bet as permanently short-leashed out of habit.

Where people run it wrong.
They pick one horizon for the whole roadmap because it's simpler to plan around, instead of asking which items can survive being wrong.
They let a check-in habit fade quietly because it kept confirming good news, with nobody noticing the habit was the whole safety net.
They build the model-dependent bet with no smaller, provable interim version, so there's nothing to fall back on when the timeline slips.

How to use it live. The moment someone asks about planning horizons, ask yourself: which of these roadmap items would still be worth having built even if the model never improves again? Give those the long horizon. Give everything else the short one.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip: the team's Thursday assumption check-in stopped happening because it kept confirming good news, until nobody was checking the one thing the whole roadmap depended on.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Wren Ashbrook, who leads product for Verdalane Stream's personalization team.
3 · THE HABIT
What did Wren's team stop doing as the pilot looked stronger?
Tap to flip
ANSWER
The weekly fifteen-minute check-in asking whether the roadmap's model assumption still held, which quietly disappeared from the calendar around month five.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Checking the model assumption sometimes, versus not checking it at all, with nothing in between once the good months made checking feel unnecessary.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Setting the roadmap horizon at twelve months to match the fiscal calendar, a default that made sense for ordinary feature work and stopped making sense once one item depended on an unshipped vendor SKU.
6 · THE NUMBER
Fill in the blank: short-horizon, capability-independent items were delivered on time 92 percent of the time. Long-horizon, capability-dependent items were delivered on time only ___ percent of the time.
Tap to flip
ANSWER
34 percent, the exact gap that shows why one horizon can't serve both kinds of work.
7 · THE REPLAY
Same year, the two-track horizon split in place from day one. What changes?
Tap to flip
ANSWER
The infra ships on the twelve-month clock as planned. The mood-playlist feature runs in six-week slices with a chunked, smaller-context interim version shipping to real users well before month eight, so the vendor's slip costs a delay, not a missing product.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Northquay Terminal's berth-scheduling optimizer. The flip is workaround: yard planner Rasheed Colby built his own private spreadsheet predicting berth conflicts, a shadow system Northquay couldn't see or support.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: long-horizon, capability-dependent roadmap items at Verdalane were delivered on time only ___ percent of the time.
Show hint
Look at the grouped bar chart comparing the two kinds of roadmap items.
Show answer
34 percent. Compared to 92 percent for the short-horizon, infra-independent items.
Multiple choice
2. According to this answer, what actually determines whether a roadmap item deserves a long or short horizon?
  • A. Whether the item is a customer-facing feature or an internal tool.
  • B. How large the engineering team assigned to it is.
  • C. Whether the item's value depends on one specific, unproven model capability, or would pay off regardless of which model wins.
  • D. How excited leadership is about the feature at the time it's proposed.
Show hint
Look at the direct answer and the quadrant diagram.
Show answer
C. Infra and reusable pipelines earn a long horizon. Anything betting on one unshipped capability gets a short one, regardless of how exciting it is.
True or false
3. True or false: this answer argues every AI product team should plan in short, six-week cycles across the entire roadmap.
  • True
  • False
Show hint
Look at "what I would leave alone" and the labeled-parts diagram.
Show answer
False. Infra like the embeddings pipeline keeps its long, twelve-month horizon. Only the capability-dependent bet gets shortened.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Setting the whole roadmap's horizon at twelve months to match the fiscal calendar. It made sense for ordinary feature work and stopped making sense once one item depended on an unshipped vendor SKU with no fallback plan.
Short answer, where it wouldn't matter
5. Name a part of Verdalane's roadmap where the long-versus-short horizon tension genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The embeddings pipeline and the recommendation UI shell. Both pay off no matter which model the team ends up using, so there's no asymmetry to manage there.
Short answer, apply it yourself
6. Think of a product roadmap you've seen or worked on. Was there an item that quietly depended on one unproven capability, riding on the same timeline as everything else? What would splitting it onto its own shorter horizon have looked like?
Show hint
Think about a launch date that assumed a specific model or vendor feature would be ready in time.
Show answer
Model answer: A support-automation rollout tied to a vendor's promised multilingual model, planned on the same six-month timeline as an unrelated internal dashboard rebuild that had nothing to do with the vendor at all.
Before you close the answer
Why this works
Tests whether you'll answer "how far out" with a single number, or split the roadmap by which items can actually survive being wrong about the model.
Follow-up traps
"Isn't splitting the roadmap into two tracks just extra process?" Response: it's one tag per item, not a new team. The cost is a meeting; the thing it prevents is four months of unshippable engineering work.

"What if leadership insists on one big public number?" Response: give them one, built from the short-horizon items only, and name the capability-dependent work as a separate, explicitly hedged track rather than baking it into the same promise.
If pressed
The version that shipped afterward required any roadmap item flagged as capability-dependent to carry a named, dated re-check owner, not just a note in a doc, so the habit of checking couldn't quietly disappear a second time.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more