ConceptAdvancedAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #21
What is the difference between a roadmap and a bet portfolio for AI products?
PICKthe document that promised a date for something nobody could actually promise
A roadmap and a bet portfolio can look identical on the page, right up until one of them is wrong. Deep Current Fisheries Cooperative runs TideCast, a tool that predicts where fish will be, zone by zone, so boats get sent somewhere worth the fuel. Ilse Marrant is the AI PM who has to explain why her planning document was never really a roadmap at all.
The direct answer
A roadmap promises a date for something you already know how to build. A bet portfolio prices a set of uncertain outcomes and expects some of them to lose. Anything on your plan whose success is a threshold on an eval set, not a fixed rule, is a bet, not a roadmap item, and it should never be handed a firm date until it clears that bar.
Do this, in order
Sort every plan item into "known" or "uncertain" before assigning it a date.Why: a date on an uncertain item is a promise nobody can actually keep.
Give every bet a confidence tier and a cost if it's wrong, not just a name.Why: without those two numbers, a hopeful guess and a near-certain plan look identical on the page.
Only publish a firm date once an item clears its eval bar, twice running.Why: clearing it once could be luck. Clearing it twice is a pattern worth a promise.
Set a kill criterion for every bet before it starts, not after it disappoints.Why: deciding what would make you cut a bet, in advance, is what separates a portfolio from a wish list.
Keep the roadmap-shaped summary for whoever needs a plan to act on.Why: captains and regulators still need something to plan around, just not a promise the model can't actually back.
How to answer this, stage by stage
Nobody is scoring whether you know the word "portfolio." They're scoring whether you can tell a real commitment from a hopeful one.
Stage 1
Scope it to one real planning document
Say it like this
"I'll use TideCast at Deep Current Fisheries, and the actual planning document Ilse Marrant has to defend to captains and a fishing regulator."
Why this works
Keeps the answer from turning into a dictionary definition of two words.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as PICK. Position, my actual pick. Impact, who feels each kind of error. Cost asymmetry, which error is the expensive one. Kill criteria, what would flip my mind."
Why this works
Signals you're about to commit to something, not hedge between two definitions.
Stage 3
Take a position, before any reasoning
Say it like this
"My pick: run the plan internally as a bet portfolio with confidence tiers, and only let the parts that have already cleared their eval bar show up as roadmap-style dates to anyone outside the team."
Why this works
Commits to an answer before the reasoning, which is exactly what the framework's first letter is for.
Stage 4
Name who feels each kind of error
Say it like this
"If I roadmap an uncertain bet and it misses, a captain planned a crew schedule around a date that quietly slipped, and a regulator was promised an accuracy number that never arrived. If I bet-portfolio something that turns out solid, nobody's hurt, they just wanted more certainty than the label gave them."
Why this works
Names both errors in real terms instead of staying abstract about "risk."
Stage 5
Say which error is the expensive one
Say it like this
"Treating a bet as a roadmap promise is the expensive error. It's invisible until launch, and by then you've already made an external promise you can't take back. Treating a roadmap item as a bet just costs you a slightly less confident sentence in a planning doc."
Why this works
This is the heart of PICK: naming which cost is hidden and which is cheap.
Stage 6
Close on the kill criteria
Say it like this
"What would flip me back toward a straight roadmap: the day a zone's prediction clears its accuracy bar two seasons running. At that point it's not a bet anymore, it's a known capability, and it earns a real date."
Why this works
Shows this isn't a permanent stance, it's a call that updates as evidence comes in.
Let's learn
Here is what happens when a plan built on an uncertain model gets written up in the same language as a plan for something you already know how to do.
Before TideCast, a dock master decided where boats went each week from tide charts, radio chatter with other crews, and a lifetime of noticing patterns nobody wrote down. With TideCast, a per-zone catch prediction comes back overnight, built from satellite temperature data and years of logged hauls. In the zones it knew well, it beat the dock master's own guesses by a wide margin.
Both are plans. Only one of them should ever get a promised date attached.
Here's the turn: Deep Current's planning document listed "sub-zone prediction live in all 12 zones by Q3" the same way it listed "boat maintenance completed by June," one line, one date, no distinction between a thing that was scheduled and a thing that was hoped for. The zones TideCast knew well got the same confident language as the zones it barely understood yet.
TideCast's own confidence, by zone, before the roadmap was published
All four zones got the same "Q3" on the roadmap. Only two of them had actually earned it.
At its worst, a plan that promises the same date for a proven capability and a hopeful one doesn't fail everywhere at once. It fails exactly where confidence was lowest, and it fails loudest, because that's where a captain trusted the date most.
The choice I would take back
The planning template Ilse inherited had one column for the feature and one for the date, nothing for confidence, nothing for what would happen if the number came in low. That made sense when the document only tracked things like boat maintenance, where a date really is just a date. It stopped making sense the moment the same document started carrying model predictions with a real, varying chance of missing.
What I would leave alone: south banks prediction, at 84 percent accuracy on two straight seasons of its own eval set, earns a real roadmap date. Not everything on the plan needs a confidence tier bolted on; some of it has genuinely graduated.
The lesson: the words "roadmap" and "bet portfolio" aren't a branding choice. They're a promise about how sure you actually are, and using the wrong word is how a hopeful guess ends up believed like a certainty.
Now here is the same thing as a story
The short version above is what you'd say defending the plan to the cooperative's board. Read this one for how one sentence in a meeting gave the whole problem away.
Casimir Norlund has run dock scheduling at Deep Current for nine years. He can look at a week's forecast and know, before anyone tells him, which two boats will fight over the same good zone.
When TideCast's roadmap came out, Casimir built his whole quarter's crew rotation around the promised dates, same as he always had for maintenance and dry-dock schedules. South banks, open channel, reef shallows, north banks, all live by Q3, all with the same weight on the page.
One box is small on purpose. The other one is the box that actually costs something.
Reef shallows and north banks both slipped past Q3, not because anyone missed a deadline, but because the model's own accuracy on those zones had never actually cleared a bar worth promising. Nobody had written that bar down. The roadmap just said "live," the same word it used for the zones that were genuinely ready.
Knowledge spark: what makes something a bet instead of a plan?
A plan item is a bet the moment its success is a threshold on an eval set instead of a fixed rule you can just build to. "Ships by June" is a plan. "Clears 80 percent accuracy on this season's catch data" is a bet, because the model might not get there no matter how hard anyone works.
In a fleet-ops meeting, one of the boat captains looked at the roadmap and asked, plainly: "Wait, so north banks prediction is as sure as the maintenance schedule being done by June?" Nobody had a good answer. It wasn't. It had never been.
The roadmap didn't lie about a date. It lied about how sure anyone actually was, by using one column for two very different kinds of promise.
Deep Current split the document that quarter: a real roadmap for maintenance, staffing, and anything already proven, and a bet portfolio, with a confidence tier and a cost-if-wrong number, for every zone prediction still short of its bar. Casimir stopped scheduling crews around north banks entirely, and started planning that zone week to week instead, the way he always had before TideCast existed.
North banks is the one zone that should never have shared a column with boat maintenance.
PICK, the commitment that admits what it doesn't knowNot a definition to recite. PICK is what stops a hopeful guess from being believed like a certainty.
P
Position. The pick, before any reasoning.
Run the plan as a bet portfolio internally, and only let proven items graduate to roadmap-style dates externally.
This is the hardest step to commit to plainly, since "it depends" always feels safer.
I
Impact. Who feels each kind of error.
Captains and a regulator absorb a broken roadmap promise. Nobody is really hurt by an over-cautious bet label on something solid.
Naming both sides in real terms is what keeps this from staying abstract.
C
Cost asymmetry. Which error is the expensive one.
A false roadmap promise is invisible until it breaks, and by then it's already been repeated to a regulator. An over-cautious bet label just costs a slightly less confident sentence.
This is the whole answer to the question, once it's said plainly.
K
Kill criteria. What would flip the pick.
A zone earns a real date once its prediction clears its accuracy bar two seasons running, not once.
Without this, "bet portfolio" becomes a permanent excuse instead of a temporary label.
The recap, one line per letter: position is running the plan as a bet portfolio with a roadmap-shaped summary layered on top, impact is naming captains and a regulator as the ones who absorb a broken promise, cost asymmetry is that the false-certainty error is the expensive, hidden one, and kill criteria is a zone graduating to a real date only after clearing its bar twice.
And if you want to be sure it really works, try it somewhere elseSame four letters, a regional airline's turbulence-prediction tool instead of a fishing fleet. A different sort of promise breaks this time.
Vasterlund Regional Air runs a model that predicts turbulence risk along each route, feeding into how crews plan service timing mid-flight. Its planning document listed "smooth-service window prediction, all routes, by spring" with the same confidence as "new galley carts installed by spring." Mapped onto PICK: position is treating each route's prediction as its own bet, not a fleet-wide roadmap line. Impact lands on flight attendants who plan service timing around a window that sometimes closes without warning, and on passengers who get a drink knocked into their lap when it does. Cost asymmetry is the same shape: a route wrongly roadmapped as reliable costs a real injury report; a route wrongly bet-portfolio'd that turns out solid just costs an attendant double-checking a window that didn't need it. Kill criteria is a route graduating once its prediction holds across two full seasons of weather variation. The reversal here isn't a missing confidence column, it's a packaging choice: Vasterlund had bundled all routes into one fleet-wide accuracy number for its investor updates, which hid the fact that mountain-corridor routes were carrying nearly all the error while flatland routes were already reliable enough for a real date.
The same four branches decide a fishing zone and a flight route alike.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a roadmap promises a date for something known, a bet portfolio prices something uncertain, and mixing the two into one column is how a hopeful guess gets believed like a promise," and stop.
Cost: no time to build a formal confidence-tier system this quarter. Say so honestly, and start by just splitting the document into two columns, since that alone stops the worst failure.
The model gets better, for real: if a zone's prediction genuinely clears its bar two seasons running, that's the honest trigger to move it into the roadmap, not a reason to keep hedging out of habit.
None of these four things belong on a maintenance schedule. All four belong on a model prediction.
Where people run it wrong.
They let one column hold both proven capabilities and hopeful bets, with the same confident word attached to each.
They treat "bet portfolio" as a permanent hedge instead of a temporary label with a real graduation test.
They publish an external, regulator-facing promise before the internal eval bar has actually been cleared.
How to use it live. The moment someone asks you this question, ask yourself: which items on this plan could reasonably lose, and does the document actually admit that. If it doesn't, it's not a roadmap. It's a bet wearing one.
Any one of these four signs is reason enough to move the item off the roadmap and into the portfolio.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits an "A or B, which one and why" tradeoff question?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It commits to a side first, then shows the asymmetry that justifies it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Casimir Norlund, a nine-year dock scheduler at Deep Current Fisheries who builds crew rotations around whatever the roadmap promises.
3 · THE HABIT
What did Casimir stop doing once north banks got split into the bet portfolio?
Tap to flip
ANSWER
He stopped scheduling crews around north banks' predicted date and went back to planning that zone week to week, the way he did before TideCast.
4 · THE COST ASYMMETRY
Which error is the expensive one, and which is cheap?
Tap to flip
ANSWER
Roadmapping an uncertain bet is expensive, since it's invisible until it breaks a promise already made externally. Bet-labeling a proven item is cheap, just an overly cautious sentence.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
The planning template's single feature-and-date column, with no field for confidence, which made sense for maintenance schedules and stopped making sense once model predictions joined the same document.
6 · THE NUMBER
Fill in the blank: north banks prediction accuracy sat at ___ percent, the lowest of the four zones, yet got the same Q3 date as south banks at 84 percent.
Tap to flip
ANSWER
52 percent, well short of a bar worth promising to a regulator.
7 · THE REPLAY
Same four zones, same Q3 target, but the split document already exists. What changes?
Tap to flip
ANSWER
South banks and open channel get real dates. Reef shallows and north banks show up as priced bets with a confidence tier and a cost if wrong, so Casimir never builds a crew rotation around a number that wasn't earned yet.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Vasterlund Regional Air's turbulence-prediction tool. The reversal is a packaging choice: bundling all routes into one fleet-wide accuracy number hid that mountain-corridor routes carried nearly all the error.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: south banks prediction, at 84 percent accuracy across two straight seasons, earns a real ___, while north banks at 52 percent should stay a priced bet.
Show hint
Look at "what I would leave alone."
Show answer
Roadmap date. Not everything on the plan needs a confidence tier; south banks has genuinely graduated.
Multiple choice
2. Per this answer, what actually makes a plan item a "bet" instead of a roadmap item?
A. It involves a model at all.
B. Its success is a threshold on an eval set, not a fixed rule you can just build to.
C. It costs more money than other roadmap items.
D. It was requested by leadership instead of engineering.
Show hint
Look at the knowledge spark about what makes something a bet.
Show answer
B. "Clears 80 percent accuracy" is a threshold the model might not reach. "Ships by June" is a rule you can just build to.
True or false
3. True or false: this answer argues every model-dependent feature should permanently stay a bet and never earn a roadmap date.
True
False
Show hint
Look at the kill criteria step.
Show answer
False. A zone graduates to a real date once its prediction clears its accuracy bar two seasons running.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: The planning template's single feature-and-date column with no confidence field. It made sense when the document only tracked things like maintenance, where a date really is just a date.
Short answer, apply it yourself
5. Think of a plan or announcement you've seen that used roadmap language for something that was really still uncertain. What would a bet-portfolio version of that same plan have said instead?
Show hint
Think about a product launch date tied to a feature that depended on a model working well enough, not just being built.
Show answer
Model answer: Instead of "voice search launches in March," a bet-portfolio version says "voice search launches once accuracy clears 90 percent on our test set, currently at 78 and rising."
Short answer, where it wouldn't matter
6. Name something on Deep Current's plan where the roadmap-versus-bet distinction genuinely doesn't apply.
Show hint
Look at "what I would leave alone," or think about non-model line items on the same document.
Show answer
Model answer: Boat maintenance completed by June. Nothing about it depends on a model clearing a threshold, so it's a real roadmap item with no bet involved.
Before you close the answer
Why this works
Tests whether you can tell a genuine commitment from a hopeful one, in language that a captain, a regulator, or a board member could actually act on without getting burned.
Follow-up traps
"Doesn't calling something a 'bet' just give the team an excuse to miss deadlines?" Response: no, because every bet still carries a kill criterion and a cost-if-wrong; it's accountable to a threshold, just not to a calendar date it can't control.
"What if leadership just wants one number to report to the board, not two documents?" Response: give them one number, a weighted portfolio value across all the bets, but keep the underlying confidence tiers visible underneath it, not flattened away.
If pressed
The eval bar that followed set a specific rule: a zone needed 75 percent accuracy or higher across two full fishing seasons, evaluated on catch data withheld from training, before its prediction could carry a firm date on the public roadmap.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.