ConceptAdvancedAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #3
What does a roadmap look like when it is organized by uncertainty rather than by quarter?
SPARKthe Q3 promise a bill-negotiation model never actually agreed to keep
Every day between fares, Marisol Duarte checks whether the bill she asked an app to fight for her actually got fought. Rovengale is a consumer app that negotiates bills and cancels subscriptions on a customer's behalf. Callum Iverson is Head of Product, and the question on the table is what the roadmap should look like once "when will this ship" stops being the same question as "will this ever actually work."
The direct answer
Replace the quarter columns with three confidence lanes: Committed, work that runs on today's model and gets a real date; Directional, a capability bet with a named threshold and a recheck date, no ship date; and Exploratory, a research question with no date and no promise, revisited each time the threshold check runs. Every item carries what would move it up a lane, not just when it's supposedly due.
Do this, in order
Replace the quarter column with a confidence lane.Why: a date tells you when, not whether. Uncertainty needs a structure that answers whether first.
Give every Directional and Exploratory item a named threshold, not a placeholder quarter.Why: without a bar to clear, "still in progress" can mean anything for years.
Design against the failure the format itself creates: a promise leaking downstream as fact.Why: sales, marketing, and eventually the customer only ever see the promise, not the lane it lived in.
Keep the lane review manual, not automated, on day one.Why: an automatic recalculation off a dashboard hides exactly the judgment call this whole structure exists to force into the open.
Set a revisit cadence for Exploratory items, or they become a graveyard nobody reads.Why: a lane with no return trip is just a nicer-looking place to bury a promise.
How to answer this, stage by stage
Nobody is scoring whether you can draw three boxes instead of four. They're scoring whether you can say what happens to a real customer once the roadmap format itself starts lying by omission.
Stage 1
Anchor the person before naming the fix
Say it like this
"I'll ground this in Rovengale, a bill-negotiation app, and one user, Marisol, who has no way of knowing whether a promise on the roadmap was ever meant to be a sure thing."
Why this works
Stops the answer from staying a slide-design exercise with nobody actually affected by it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as SPARK. Situation, how the roadmap works today. Payoff, the habit I want the new format to build. Anchor, the actual structure. Risk, what breaks the first time a bet is wrong. Keep out, what I won't automate yet."
Why this works
Signals a repeatable design method instead of a one-off opinion about slide layout.
Stage 3
Reframe: not "when," but "which promises can be checked"
Say it like this
"A quarter answers 'when.' It says nothing about whether the thing in that box depends on the model doing something it has never done yet. The real question a roadmap should answer is which of these are promises and which are bets."
Why this works
This is where the answer stops being about slide design and starts being about what the format hides.
Stage 4
Give the anchor
Say it like this
"Three lanes. Committed: works on today's model, gets a date. Directional: a bet with a named threshold and a recheck date, no ship date. Exploratory: a research question, no date at all. Every item says, in one line, what would move it up."
Why this works
This is the direct answer, made concrete enough that someone could actually build the slide from it.
Stage 5
Prove the anchor survives the model being wrong
Say it like this
"If a Directional item leaks downstream as 'coming Q3' before it clears its threshold, the lane structure didn't help; it just moved the lie one layer up. So the anchor also has to include a rule: nothing in Directional or Exploratory gets a customer-facing date, full stop, no exceptions for a good demo."
Why this works
Shows the anchor was designed against its own most likely failure, not just described in isolation.
Stage 6
Say what's deliberately left out
Say it like this
"I wouldn't automate lane placement off a live eval dashboard on day one. That hides the judgment call inside a script instead of forcing a person to look at the number and decide. Start with a monthly manual review, tied to the same threshold everyone already agreed on."
Why this works
Shows judgment about scope instead of reaching for the most impressive-sounding version first.
Stage 7
Close on what changes for the person holding the phone
Say it like this
"Once the lanes exist, Marisol never sees 'coming Q3' for a bet that hasn't cleared its bar. She sees 'still learning your provider,' which is a smaller promise, but it's one Rovengale can actually keep."
Why this works
Restates the direct answer in one breath, tied back to the person the whole question is really about.
Let's learn
Every Tuesday, between fares, Marisol Duarte used to spend about twenty minutes on hold with her internet provider, reading a script off her phone to talk her bill down. It worked, most months, for about forty dollars off.
She signed up for Rovengale eight months ago. The app calls providers on her behalf, and for well-known ones, a negotiation now takes about ninety seconds of her time: approve the ask, wait for a text. For the two smaller, regional providers she and her neighbors use, Rovengale's roadmap listed "negotiate any provider, no matter how new" as a Q3 feature, the same kind of line as every other item on the slide.
This is the workaround the whole anchor is designed to make unnecessary: a person keeping her own ledger because the app won't.
Here's the turn: Q3 arrived, and the "any provider" capability hadn't cleared its own bar. It still only worked reliably on the handful of national providers whose phone trees Rovengale's model had seen thousands of times. For Marisol's regional provider, a negotiation would start, sit in a "processing" state for days, then quietly fail with no explanation. The problem was never the failure rate. It was that the roadmap gave that bet the exact same shape, a box with a quarter in it, as a feature that already worked.
Negotiation completion rate, national providers versus Marisol's regional provider
Marking the bet "shipped" changed nothing about what the model could actually do. It only changed what Marisol was told to expect.
Negotiations stuck in "processing" for more than 48 hours, regional providers, month over month
The stuck rate climbed for three straight months before anyone connected it to a bet that had been marked done too early. The lane fix, not a model change, is what brought it back down.
At its worst, a roadmap format can make a hope look exactly as solid as a fact, and the person who pays for that resemblance is the one holding the phone, not the one holding the slide deck.
The choice I would take back
Rovengale's roadmap template put every item in a quarter column, with no separate field for confidence. That made sense back when almost everything on the roadmap really was a feature commitment, built for a model that already did the job. It stopped making sense the moment "any provider" went on the slide as a bet on something the model had never actually done.
What I would leave alone: the quarter labels themselves for Committed items. A confirmed, working feature genuinely benefits from a real date, and adding a confidence lane on top of something that's already solid would just be noise.
The lesson: a roadmap doesn't just tell your team what's coming. It tells your sales page, your support scripts, and eventually your users, what's already true. Get the format wrong and all three learn the wrong thing at once.
Now here is the same thing as a story
The short version above is what you'd say scoping this live in an interview. Read this one for how a near miss on a Thursday afternoon turned a slide-format complaint into an actual product decision.
Marisol drives mornings and late nights, six days a week, and checks Rovengale in the gaps: a red light, a slow pickup, the ten minutes before a shift starts. She's careful with money in a way that comes from years of it mattering; she reads every text the app sends twice.
For her first seven months on Rovengale, every negotiation the app started for her national providers, internet, phone, streaming, ended the same clean way: a text saying how much she'd saved, and a line in her account history she could scroll back through any time. She trusted the app the way you trust a scale that's never once given you a strange number.
This is the anchor made concrete: the same routing logic that should decide a lane, applied to one bill at a time.
Then her regional provider raised rates, and she asked Rovengale to fight it, the same way she always had. The app showed "negotiating" for three days. On day four, distracted between two airport rides, she almost approved a full payment on the same bill through her bank's own app, assuming Rovengale's attempt had quietly failed and nobody had told her. She caught it with about ninety seconds to spare, comparing two screens at a red light.
Knowledge spark: why does a stuck "processing" status happen at all?
A model that's confident says so fast. A model that's genuinely unsure, because it's never handled this exact provider's phone tree before, can sit in a strange, in-between state: not failed, not done, just uncertain. Without a way to show that uncertainty honestly, "processing" ends up meaning two very different things wearing the same word.
Nothing about that near miss was Marisol's fault. She did exactly what a careful person does: she double-checked. The gap she fell into wasn't a mistake in her judgment. It was a roadmap slide, made eight months earlier, that had put "any provider" in a quarter box with no way to say the model hadn't actually agreed to it yet.
Each step made sense alone. Together, they built the exact private workaround a confidence lane was supposed to make unnecessary.
Rovengale didn't cost Marisol one stuck negotiation. For three days, it cost her the one thing the app was supposed to give back: not having to track her own bills by hand anymore.
After the near miss, Marisol started keeping her own running note on her phone, one line per bill, per attempt, per provider, exactly the ledger Rovengale should have been keeping for her all along.
The product-facing version of the same anchor: what Marisol sees instead of a silent "processing" screen.
Rerun the same rate hike with the lane structure already in place: "any provider" never leaves Directional, because its threshold, ninety percent completion across a rotating sample of regional providers, was never actually cleared. Marisol's app shows "still learning your provider, human review within 24 hours" instead of a stuck spinner. No near miss, no private spreadsheet, because the promise she was given matched the promise the model could actually keep.
Same unresolved bet, two very different amounts of honesty about what it actually is.
What I'd tell myself, watching Marisol catch her own near miss at a red light: the quarter-column roadmap wasn't wrong to want a plan. It was wrong to let a hope sit in the same box as a fact, with nothing on the slide to tell the two apart.
SPARK, the lanes that survive their own betNot a slide redesign. SPARK is what turns "organize by uncertainty" into a structure a team can actually run.
S
Situation. How the roadmap works today, without this.
Every item, whether it's a working feature or an unproven bet, gets the same box: a quarter, a name, a line on a slide.
Naming the current, flattening format is what makes the new one feel earned instead of decorative.
P
Payoff. The habit you want this to build.
Stakeholders, and eventually customers, stop asking "is this still on track" for things that were never date-bound, and start asking "what would move this up."
The habit, not the slide's new look, is the actual thing being designed.
A
Anchor. The one concrete design decision.
Three lanes, Committed, Directional, Exploratory, each item carrying a named threshold or a real date, never both, never neither.
This is the hardest step, and the one the whole answer is actually about.
R
Risk. What breaks the first time you're wrong.
A Directional item leaks downstream as a customer-facing date before it clears its bar, and the lane structure becomes decoration instead of a real guardrail.
Naming the failure up front is what makes the anchor a design instead of a wish.
K
Keep out. What you won't build on day one.
Automated lane placement off a live dashboard waits, since it would hide the judgment call the whole format exists to force into view.
Saying what's left out is what shows judgment instead of a wish list.
The recap, one line per letter: situation is every roadmap item flattened into the same quarter box today, payoff is stakeholders learning to ask "what would move this up" instead of "is it still on track," anchor is the three confidence lanes with a threshold or a date but never both, risk is a bet leaking downstream as a customer promise before it clears its bar, and keep out is a manual monthly review instead of an automated one, at least for now.
And if you want to be sure it really works, try it somewhere elseSame five letters, a public library's cataloguing system instead of a bill-negotiation app. Different flip family entirely, the same undated bet.
Quarrytown Public Library uses an AI assistant that suggests subject headings for new acquisitions, meant eventually to catalogue books without a librarian checking every suggestion. Mapped onto SPARK: situation is Agnes Petrosyan, a cataloguing librarian, checking every suggested heading by hand today, same as before the tool existed. Payoff is Agnes trusting the model enough on well-documented titles to only spot-check, freeing her for the genuinely obscure ones. Anchor is the same three lanes: heading suggestions for mainstream nonfiction sit Committed, suggestions for small-press or self-published work sit Directional with a named accuracy threshold, and suggestions for material in translation sit Exploratory with no promise at all yet. Risk is a Directional suggestion getting auto-applied to satisfy a backlog metric before its threshold clears. Keep out is any plan to fully automate translated-work cataloguing this year; it stays hand-reviewed. The flip here is concealment, not workaround: once a patron complained publicly about a mis-catalogued translated novel, Agnes stopped mentioning to her cataloguing team that she was using the AI's suggestions at all for anything outside the safest, most mainstream titles, quietly cutting off the exact feedback that would have shown which suggestions were actually safe to trust.
The same lane logic, one library shelf at a time: what's solid, what's a bet, and what still needs a flag.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "never let a bet share a box with a fact, and never give a bet a customer-facing date," and stop.
Cost: no time to build named thresholds for every roadmap line this quarter. Say so honestly, and start with the three highest-risk Directional items only, expanding coverage as the habit proves out.
The model gets better, for real: if a Directional item's threshold actually clears on schedule, that's the lane working exactly as designed, and the honest move is promoting it to Committed with a real date, not quietly leaving it in Directional out of habit.
Where people run it wrong.
They add a confidence label to the slide but keep giving every item a date anyway, which defeats the entire point.
They let Exploratory items sit with no revisit cadence, so the lane becomes a polite graveyard instead of a live status.
They let a good demo pull a Directional item's promise downstream to customers before its threshold has actually been cleared.
How to use it live. The moment someone asks what a roadmap should look like under uncertainty, ask yourself: if I deleted every date on this slide, would anyone actually know less about what's real? If the answer is no, the dates were never doing the job you thought they were.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Workaround flip: Marisol started keeping her own running note of every negotiation attempt, since Rovengale kept no memory of past tries she could check herself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Duarte, a rideshare driver who checks her bills in the gaps between fares and reads every text twice.
3 · THE HABIT
What did Marisol stop doing once Rovengale worked reliably?
Tap to flip
ANSWER
She stopped double-checking whether a negotiation had actually gone through, trusting the app's text confirmations the way she'd trust a scale that never gave a strange number.
4 · THE ANCHOR, IN THIS STORY
What's the one concrete structure this answer proposes?
Tap to flip
ANSWER
Three roadmap lanes, Committed, Directional, and Exploratory, where every item carries either a real date or a named threshold, never both and never neither.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Putting every roadmap item in a quarter column with no separate confidence field, a call that made sense when almost everything on the roadmap was already a working feature.
6 · THE NUMBER
Fill in the blank: after Q3, negotiation completion for Marisol's regional provider only reached ___ percent, barely above the 68 percent it started at.
Tap to flip
ANSWER
71 percent, while national providers held steady around 91 to 92 percent the entire time.
7 · THE REPLAY
Same rate hike, same regional provider, but the lane structure already built. What changes?
Tap to flip
ANSWER
"Any provider" stays in Directional since its 90 percent completion threshold was never cleared. Marisol sees "still learning your provider, human review within 24 hours" instead of a stuck spinner, and never builds her own private tracking spreadsheet.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Quarrytown Public Library's cataloguing assistant. The flip is concealment: Agnes stopped mentioning she used the AI's suggestions on anything but mainstream titles once a patron complained publicly.
Check yourself Score: 0 / 0
True or false
1. True or false: this answer recommends keeping quarter dates on every roadmap item, just adding a confidence label next to each one.
True
False
Show hint
Look at the anchor step and "where people run it wrong."
Show answer
False. Directional and Exploratory items get a threshold instead of a date, not a date plus a label. Giving both defeats the point.
Multiple choice
2. Why did marking "negotiate any provider" as shipped in Q3 fail to help Marisol?
A. Because the app crashed whenever she opened it that quarter.
B. Because labeling a bet as shipped didn't change what the model could actually do for her regional provider.
C. Because Rovengale removed the feature entirely after Q3.
D. Because her provider stopped accepting negotiated rates altogether.
Show hint
Look at the grouped bar chart and "here's the turn."
Show answer
B. Completion for her regional provider barely moved, from 68 to 71 percent, even after the feature was marked done.
Fill in the blank
3. Fill in the blank: national provider completion rate held around ___ to 92 percent both before and after Q3.
Show hint
Look at the grouped bar chart.
Show answer
91 percent. Steady the whole time, unlike the regional provider's near-flat, much lower rate.
Short answer, where it wouldn't matter
4. Name a part of Rovengale's roadmap where this confidence-lane structure genuinely wouldn't matter.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Committed items for national providers. They already work on today's model, so a plain date serves them fine with no confidence lane needed.
Short answer, apply it yourself
5. Think of a product update you were told was "coming soon." Could you tell, from how it was announced, whether it was a sure thing or a bet on the tool getting better?
Show hint
Think of a voice assistant, a translation app, or a recommendation feature that was announced with a date that later slipped with no explanation.
Show answer
Model answer: A smart-home app's "understands any phrasing" update announced with a specific month, that quietly disappeared from release notes without ever being addressed again.
Short answer, work the number
6. If Rovengale runs about 40,000 regional-provider negotiations a year and the completion rate is 71 percent instead of the 91 percent national rate, roughly how many extra failed or stalled negotiations is that gap worth in a year?
Show hint
The gap is 20 percentage points; apply it to 40,000.
Show answer
Model answer: Roughly 8,000 extra stalled or failed regional negotiations a year, each one a moment like Marisol's near miss waiting to happen.
Before you close the answer
Why this works
Tests whether you'll treat a roadmap as a scheduling document, or notice that its format is itself a promise to real people, one that can quietly turn a bet into a fact nobody agreed to.
Follow-up traps
"Won't executives just demand a date for Directional items anyway?" Response: give them the threshold and the recheck date instead, which is a real answer to "when will we know," just not a fake answer to "when will it ship."
"Isn't three lanes just relabeling the same optimism?" Response: not if Directional items are barred from getting a customer-facing date at all, since that's the actual rule that stops a bet from leaking downstream as a fact.
If pressed
The version that shipped required any item moving from Directional to Committed to clear its threshold across three consecutive weekly evals, not just one good week, before a real date could be attached.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.