CaseAdvancedAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #1

How do you build a twelve-month roadmap when model capability in month six is unknown?

The direct answer
You don't plan the capability. You plan the interface to it. Put everything that stays true whatever the model can do (data access, the eval set, permissions, writing results into the scheduler the customer already uses) in the first six months, then make month six a stated fork with two named options instead of one guessed feature.
Do this, in order
  1. Split the work into what the model can't change and what it can, and do the first pile first.Why: none of that work can be wasted by a surprise in month six, because you would have built it for any model at any level.
  2. Build the eval set in months two to four, before any feature ships.Why: 400 scored past breakdowns is what turns a model swap from eleven weeks into four days. That speed is worth more in month six than any feature you could have shipped in month two.
  3. Get results into the scheduler the dealer already books jobs in.Why: a suggestion on a screen nobody opens is worth nothing at any capability level. Everything about autonomy sits on top of this one piece.
  4. Write both month-six branches down now, with names and with what picks between them.Why: "if it can read a tech's notes we do A, if not we do B" is a plan. "We'll see where models are by then" is a shrug wearing a roadmap's clothes.
  5. Commit to dates for months one to six. Commit to options for months seven to twelve.Why: a date you're guessing at costs you more trust than no date. Sales can sell a fork if you tell them what each side actually looks like.
  6. Keep the model choice off the roadmap entirely.Why: once the eval set exists it's a four-day decision. Putting a four-day decision on a twelve-month sheet makes everyone treat it as the big one, and it isn't.

How to answer this, stage by stage

Six moves. Say each one out loud once before you use it. This question pulls people toward big words, so keep every sentence flat and specific.

1
Cut the year into two piles before you plan a single month
Say it like this
"Before I lay out any months, I'd split the work into two piles. Pile one is everything that's true whatever the model can do: getting the data, building the eval set, permissions, writing results back into the customer's own system. Pile two is the one thing that actually depends on capability, which is how much the feature is allowed to do on its own. Then I do pile one first."
Why this works
Most people answer this by guessing at the model, and then defending the guess. This move says out loud that the guess is the small part of the plan. It also gives the interviewer a frame to hang every later thing you say on.
2
Put one real machine and one real person under it
Say it like this
"Let me make it concrete. We sell telematics to farm equipment dealers, about ninety of them. A combine out in a field sends fault codes back to us, and the dealer's service manager decides who gets a truck sent out. The feature we want is: tell that service manager which machines are going to break in the next two weeks, before harvest."
Why this works
A roadmap answer with no product in it turns into a talk about planning. Name one machine and one person and every later stage has something solid to sit on. It takes fifteen seconds.
3
Fund the eval set first, and say why it isn't a testing job
Say it like this
"The first real thing I'd fund isn't a feature. It's 400 real breakdowns from the last two seasons, with what actually went wrong written next to each one, scored by two service managers. About three weeks of work. That set is what lets me try a new model on a Monday and know by Thursday whether it's better than what we've got."
Why this works
Everyone else in the room says "stay flexible." This is the concrete thing flexibility is actually made of, and it's specific enough to argue with, which is what you want an interviewer doing next.
4
Price the month you skipped it
Say it like this
"Say I skip that and ship a feature in month two instead. Month six comes, a better model lands, and I've got no way to tell if it's actually better. So I find a friendly dealer, wait for real machines to break, then argue about it for a month. That's eleven weeks. With the eval set it's four days. That gap is the whole difference between the two roadmaps."
Why this works
Four days against eleven weeks is a number an interviewer repeats to the next person in the room. A principle isn't. This is also the moment your first stage stops sounding clever and starts sounding paid for.
5
Say both month-six branches out loud, including the losing one
Say it like this
"Month six is a fork and I'd write both sides down now. If the model can read a tech's handwritten notes well enough, branch A: the system books the visit and the service manager approves it in one tap. If it can't, branch B: it just flags the machine and we spend the back half adding two more equipment brands. Either way, months seven to twelve are full."
Why this works
It kills the obvious follow-up before it's asked. Naming the branch you don't want is what proves you actually thought about it, and it means nobody on your team is sitting idle in month seven whichever way it lands.
6
Draw the line between what you promise and what you offer
Say it like this
"So what I'd sign up for is: months one to six, real dates and real outcomes, hold me to them. Months seven to twelve, two options and the one test that picks between them. I'd rather tell a dealer 'one of these two, and here's what decides it' than hand them a date I already know I'm guessing at."
Why this works
This is the line that turns the answer into something an exec could sign. It's also the last thing they hear, and last lines stick.
If you remember one thing Stages 1 and 3 are the answer. Split the year into two piles, then put the eval set in the first pile ahead of every feature. If the interviewer cuts you off after ninety seconds, those two are the ones you needed to get out.

Let's learn

The roadmap is a single sheet Britt prints out and pins next to her desk. Twelve columns, one per month. She runs product at a forty-person company that sells telematics to farm equipment dealers: tractors, combines, sprayers, all sending fault codes and engine hours back over a cell connection.

The dealers buy it for one reason. When a customer's combine dies in the middle of harvest, that farmer loses a day, and the dealer hears about it for the next three years. So the dealer's service manager wants to know, in the two weeks before harvest, which machines are about to fail.

Today that manager works it out himself. He scrolls a list of fault codes, thinks about which farms are furthest out, and books his four trucks. It takes most of a Friday. He's good at it. He also misses things, because a fault code from a 2014 sprayer means something different than the same code from a 2022 one, and nobody has all of that in their head.

So the AI feature writes itself: read the codes, read the service history, read what the tech typed after the last visit, and tell him which twelve machines to worry about. Easy to describe. The problem is that whether the model can read what the tech typed is a month-six question, and Britt has to hand a twelve-month sheet to her CEO on Monday.

Two panels: work that is true anyway, versus work that depends on the model
Sort the work before you sort the months

Here is the move. Don't plan the capability. Plan the interface to it. Almost everything on a twelve-month AI roadmap is work you'd do at any capability level, and it's sitting in the same list as the one item that isn't.

Pile one

True whatever the model can do

  • Fault codes and service history from the three big equipment brands, in one place
  • The eval set: 400 past breakdowns with the real answer next to each
  • Who at the dealer is allowed to approve a truck being sent
  • Suggestions written back into the scheduler the dealer already books jobs in
  • The accept and reject step, and logging what the manager did
None of this changes if the model turns out twice as good, or half as good, in month six.
Pile two

Depends on the model

  • How much the feature does on its own
That's the whole pile. One line. It's the only place where "we don't know yet" is a real answer, and it's the last thing you build.
The plan doesn't go stale in month six. Your ability to change it does.

So the roadmap gets built backwards from that. Months one and two, get the data flowing. Months two to four, build the eval set. Months three to five, get suggestions landing inside the dealer's own scheduler. Month five, the approve and reject step. Month six is the fork. Months seven to twelve go to whichever branch month six picked.

Timeline: data access, eval set, write-back, approve step, branch point at month six, then A or B
Everything before the fork is work you'd do for any model
Knowledge spark: what an eval set actually is A pile of real past cases with the right answer already written next to each one. Here it's 400 machines that broke, what the codes said beforehand, and what the tech found when he opened it up. You feed the same 400 to any model and count how many it gets right. Nobody has to wait for a real combine to break to find out.

That eval set is the load-bearing item on the whole sheet, and it looks like the least exciting one. Three weeks of a data person plus about two days of two service managers sitting in a room saying "no, that one was a bearing, not the sensor." It ships no feature. A CEO looking at the sheet will ask why it's there.

Here's why. In month six a better model shows up, and the only question that matters is how fast Britt can find out whether it's better and get it in front of dealers.

Swapping in a better model, month six
With the eval set
4 days
Without it
55 days
Set it up (1 day / 10 days finding a pilot dealer)
Get results (2 days reading misses / 25 days waiting for machines to break)
Decide (1 day / 20 days arguing about it)
Drawn at true scale on purpose. Same model, same dealers, same month. The only difference is whether somebody spent three weeks in month three writing down what already happened.

Eleven weeks in this business is not eleven weeks. Harvest starts, and no dealer changes how they book service trucks in the middle of harvest. So the swap doesn't happen in month nine. It happens next spring, which means the better model sits unused for a season while a competitor with a worse model and a working eval set ships twice.

The choice I'd take back Putting "pick the model" on the roadmap as a month-four milestone with an owner and a review meeting. It made the sheet look decisive. What it actually did was make a four-day decision feel like a big one, so everyone treated it as settled once it was made, and nobody asked again for eight months. Take that row off the sheet. Put the eval set there instead, and let the model choice be something you redo whenever you feel like it.

What I'd leave alone. Months one to six. Plan them like nothing is going to change, because in them nothing does. There's a habit of hedging every early row ("build it so we can swap this out later") and on a twelve-month plan that habit turns six months of finished work into nine months of half-finished choices. The write-back into the dealer's scheduler is the same job whether the model is brilliant or useless. Build it once, in month three, and stop discussing it.

The lesson. The thing that goes out of date is never the roadmap. It's the cost of changing the roadmap. If that cost is four days, you can be wrong about month six and it doesn't matter. If that cost is a quarter, you have to be right about month six, and nobody is.

Month six, played twice

You don't need this to answer the question. It's here so the four-days-against-eleven-weeks line stops being a statistic.

First version. Britt puts features on the sheet, the way everyone does. Month two, the suggestion screen ships. Month four, coverage for a second equipment brand. Month six, a model comes out that reads the techs' free-text notes properly for the first time, and one of her engineers sends her the link on a Tuesday afternoon.

She wants to say yes. She can't, because she has no way to know. The old model and the new one both produce lists of machines, and both lists look reasonable. So she does the only thing available: calls a dealer in Iowa she trusts, gets him to run both for a while, and waits for machines to actually break so there's something to compare. That takes until late August. By then the argument inside her own company is about whose fault the delay was, and it eats most of September. She makes the call on the first of October. Harvest is running. Nobody is touching the service scheduler until December.

The new model goes live in April. Ten months after her engineer sent the link.

She didn't lose ten months to a hard decision. She lost ten months to not being able to make an easy one.

Second version, same Tuesday. Same link, same engineer. Britt has 400 scored breakdowns sitting in a repo since month four. She runs both models against all 400 that afternoon. Wednesday and Thursday she reads the thirty cases the new model got wrong and the twenty the old one got wrong, sitting with one of the service managers who scored them. Friday morning she decides.

It's better on notes, worse on old sprayers, and the fix for the sprayers is two lines in the prompt. She takes branch A. Auto-booking with one-tap approval ships in month eight, six weeks before harvest, which is exactly when a service manager will try something new because he's about to be buried.

Same model. Same company. Same forty people. One of those two Britts spent three weeks in month three writing down what already happened, and the other one shipped a feature in month two instead.

ORDER, laid out on the actual sheet

This is a prioritisation question, so the framework is ORDER. Not FLIPS. FLIPS finds the moment a person's behaviour snaps, which is the right tool for a "what if the model got worse" question. Here nothing has changed yet. The whole question is what goes first.

ORDER: outcome, reversibility, dependency, evidence, rank
ORDER, for prioritisation questions
O, outcome. Everything on this sheet competes to move one number: how many machines break in season with no warning, across ninety dealer groups. Not model accuracy. A dealer never sees model accuracy.
R, reversibility. Data deals and a service manager's trust are the hardest things to undo. Once he's stopped reading the suggestions, he doesn't start again. A model choice is the easiest thing to undo, as long as you can measure it. So the hard-to-undo work goes first and the easy-to-undo decision goes last.
D, dependency. Nothing about autonomy can ship until suggestions land in the dealer's own scheduler, and nothing can be judged until the eval set exists. That's forced by reality, not by taste, which makes it the cheapest part of the argument to win.
E, evidence. Four days and a few hundred dollars of scoring buys you the real ceiling on what any model can do with this data. Buy that answer before you promise anyone a date, not after.
R, rank. Data, eval set, write-back, approve step, fork. Autonomy last, because it's the only item the month-six unknown actually touches.
Month six, branch A If the model reads a tech's free-text notes well enough The system books the service visit itself and the manager approves it in one tap. Months seven to twelve go to approval flows, undo, and the argument with the three dealers who won't allow it.
Month six, branch B If it can't It flags the machine and the manager books the visit himself, same as now, just faster. Months seven to twelve go to two more equipment brands, which is worth real money and needs no capability we don't already have.
The check that makes ORDER honest Swap the outcome in step O and see if the ranking moves. Change it from "fewer machines break in season" to "sell more parts," and the ranking really does change: parts recommendations jump ahead of the approve step, and the eval set is now built out of parts orders, not breakdowns. If your ranking survives any outcome you plug in, you ranked by gut and wrote the outcome afterwards.

Run ORDER on something completely different

A freight brokerage, sixty people, wants a twelve-month plan for a tool that reads a shipper's email and turns it into a price for the load. Same question: nobody knows what the models will do six months out. Same split.

O. How many loads get quoted within ten minutes of the email arriving. That's what wins the load, and it's the only number the owner cares about.
R. Access to the load board and the carrier rate history. Lose that and you start over. A model swap is a weekend.
D. Nothing quotes on its own until quotes can be written into the brokerage's own system and a broker can change one in a single click.
E. 600 old emails with the quote that actually went out next to each one. Two weeks to put together. After that, any model is a two-day test.
R. Email access, then the 600-email set, then write-back, then the override, then the fork: quote on its own, or draft it and let a broker send.

Swap the trigger and it still runs

  • It gets slower. Doesn't matter. Nothing in pile one was ever about speed, and a service manager booking Friday's trucks doesn't care whether the list took two seconds or twenty.
  • It gets more expensive. The price per call triples. With an eval set you move to a cheaper model in four days, because you can prove it's close enough. Without one you just pay.
  • It gets better than you planned. Branch A becomes possible in month four, not six. You can take the fork early only because the write-back and the approve step are already built. This is the case people forget, and it's the one where the two-pile roadmap pays out biggest.

Where people run it wrong

  • Building the eval set out of cases the current model already handles. Easy cases score high and teach you nothing. Pull the 400 from breakdowns that surprised somebody.
  • Writing "TBD" over months seven to twelve. TBD reads as no plan. Two named branches and the test that picks between them reads as a plan with a switch in it, and it's the same amount of certainty.
  • Treating the eval set as a one-time job. It has to grow every season, or by year two you're scoring this year's models against machines nobody sells any more.
  • Hedging months one to six as well. Everything gets built to be swappable, nothing gets finished, and you arrive at the fork with no working product to fork.

If you're asked this cold

Buy yourself ten seconds by saying the split out loud before anything else. "Let me split this into what's true whatever the model does, and what actually depends on it." That's not stalling. It's stage one, and it hands you a shelf to put every later thought on, which is the real reason to say it first.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and why not FLIPS?
Tap to flip
ANSWER
ORDER, for prioritisation. FLIPS finds the moment a person's behaviour snaps after something changes. Here nothing has changed yet. The whole question is what goes first.
2 · THE CORE MOVE
What's the first thing you do with a twelve-month roadmap you can't predict?
Tap to flip
ANSWER
Split it into two piles: work that's true whatever the model can do, and work that depends on capability. Do pile one first. You plan the interface to the capability, not the capability.
3 · PILE ONE
Name four things that belong in the capability-independent pile.
Tap to flip
ANSWER
Data access, the eval set, permissions and the approve step, and writing results back into the scheduler the dealer already uses.
4 · THE ONE DECISION
What goes in months two to four, ahead of every feature?
Tap to flip
ANSWER
The eval set. 400 real breakdowns from two past seasons with the true answer next to each, scored by two service managers. Three weeks of work, no feature shipped.
5 · THE NUMBER
Swapping in a better model takes ______ with an eval set and ______ without one.
Tap to flip
ANSWER
Four days, against fifty-five days (eleven weeks). And eleven weeks lands you inside harvest, so the real delay is a full season.
6 · THE FORCED ORDER
Which dependency isn't a judgement call?
Tap to flip
ANSWER
Nothing about autonomy can ship until suggestions land in the dealer's own scheduler, and nothing can be judged until the eval set exists. Reality forces that order, not taste.
7 · THE TWO BRANCHES
What are the month-six branches, and what picks between them?
Tap to flip
ANSWER
A: the system books the visit, manager approves in one tap. B: it flags the machine and the manager books it, and the back half goes to two more equipment brands. What picks: whether the model can read a tech's free-text notes.
8 · THE TRANSFER
The last section runs ORDER on a different business. Which one, and what's its eval set?
Tap to flip
ANSWER
A freight brokerage quoting loads from shipper emails. Its eval set is 600 old emails with the quote that actually went out next to each one. Two weeks to build, then any model is a two-day test.

Check yourself Score: 0 / 0

Fill in the blank
1. You don't plan the ______. You plan the ______ to it.
Show hint
Six words total. It's the first line of the direct answer.
Show answer
Capability. Interface. Data, evals, permissions and write-back are all interface to whatever the model turns out to be. Only the autonomy level is the capability itself.
Multiple choice
2. Which of these belongs in the capability-dependent pile?
  • A. Getting fault codes from the three big equipment brands into one place.
  • B. Writing suggestions into the scheduler the dealer already books jobs in.
  • C. Whether the system books the service visit itself or just flags the machine.
  • D. Deciding who at the dealer can approve a truck being sent out.
Show hint
Three of these are the same job whether the model is brilliant or useless.
Show answer
C. How much the feature does on its own is the whole of pile two. It's one line. A, B and D get built exactly the same way at any capability level, which is why they go first.
True or false
3. True or false: the main risk in a twelve-month AI roadmap is guessing the wrong capability for month six.
  • True
  • False
Show hint
Britt guessed wrong in both versions of month six. Only one of them cost her a season.
Show answer
False. Everyone guesses wrong about month six. The risk is being slow to act once you know. Four days versus eleven weeks is the actual difference between the two roadmaps, and neither one predicted the model correctly.
Multiple choice
4. Your CEO wants a date for auto-booking so sales can put it in a contract. What do you say?
  • A. Give month nine. It's the middle of the range and you can always move it.
  • B. Say you can't give a date for anything past month six.
  • C. Give month eight if the notes test passes in month six, and say what ships instead if it doesn't.
  • D. Give month nine and add a caveat in the appendix of the plan.
Show hint
Sales can sell a fork. They can't sell a shrug, and they really can't sell a date that slips.
Show answer
C. Commit to outcomes and dates for the near term, commit to named options for the far term. A and D are dates you know you're guessing at, which costs more trust than no date. B is honest and useless, because it leaves the back half of the year with nothing in it, and there's plenty you can promise: branch B is fully planned work.
Short answer
5. Name something on this roadmap you would deliberately NOT hedge, and say why hedging it would hurt.
Show hint
Look for the part of the plan where the month-six unknown genuinely doesn't reach.
Show answer
Model answer: "Months one to six. I'd plan them like nothing is going to change, because nothing in them does. The write-back into the dealer's scheduler is the same job whether the model is brilliant or useless. If I hedge it, six months of finished work turns into nine months of half-finished choices, and I arrive at the fork with no working product to fork." Naming a place you won't hedge is what stops the answer reading as blanket caution.
Short answer, apply it yourself
6. Pick a product or team you know. Split its next six months into the two piles, then name the one thing in pile two and the test that would settle it.
Show hint
Pile two is almost always smaller than it feels. If yours has five items, most of them are really pile one wearing a disguise.
Show answer
Model answer: "Internal contract review tool. Pile one: get the signed contracts out of the shared drive, build a set of 150 past contracts with the clauses our lawyer actually flagged, write findings back into the ticket the deal team already opens, and decide who can mark a clause cleared. Pile two: one item, whether the tool flags clauses for a human or drafts the redline itself. The test that settles it: run the 150 set and check whether the redlines it writes need editing more than half the time. If they do, we stay on flagging and spend the back half on a second contract type." Any answer works if pile two is one line and there's a real, cheap test attached to it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more