ConceptIntermediateAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #19
How much of an AI roadmap should be eval and infrastructure work?
ORDERthe roadmap line item nobody wanted to fund until it was the only thing that mattered
Every roadmap has a line for new features and, if there's room, a line for eval and infrastructure work. Haldenport Telecom ships TriageAssist, a tool that reads incoming support tickets and routes them to the right queue with a drafted reply. Wren Kastellan is the AI PM who has to tell leadership how much of the next four quarters should go to something that produces no new feature at all.
The direct answer
Give eval and infrastructure work a standing floor, not a leftover budget: about a third of every quarter's capacity, more like two-thirds in the first quarter when nothing exists yet. Fund that percentage before a single feature gets a name, because every feature after it depends on being able to tell, automatically, when it made things worse.
Do this, in order
Set eval and infrastructure work as a fixed share of every quarter, not a task you get to later.Why: a task always loses to a feature with a launch date. A standing share can't.
Rank it first because it's the hardest thing to undo once skipped.Why: a missed feature just ships next quarter. Weeks of silent drift with no record of what changed can't be recovered after the fact.
Treat it as a dependency, not a parallel track.Why: no other roadmap item is actually safe to ship without a way to catch what it broke.
Buy a cheap two-week spike before committing the whole quarter.Why: a small sample of real tickets tells you how bad the gap already is, for far less than building the full harness blind.
Rank everything else behind whatever the golden set says is currently weakest.Why: the roadmap should follow evidence about where it's actually breaking, not a wishlist written in January.
How to answer this, stage by stage
Nobody is scoring whether you know what an eval set is. They're scoring whether you'll actually fund it before a feature exists to defend it.
Stage 1
Scope it to one real roadmap
Say it like this
"I'll ground this in TriageAssist at Haldenport Telecom, an AI PM's actual next-four-quarters roadmap, not a general split every AI team should copy."
Why this works
Stops the answer from turning into a percentage pulled out of the air.
Stage 2
Say your structure out loud
Say it like this
"I'll rank this with ORDER: outcome, what everything is actually competing to move. Reversibility, which decision is hardest to undo. Dependency, what unblocks what. Evidence, what I could learn cheap first. Rank, the actual order and why."
Why this works
Shows a repeatable ranking method instead of a gut-feel number.
Stage 3
Reframe: this isn't a percentage question, it's a dependency question
Say it like this
"The real question isn't 'how much of the roadmap.' It's 'which of these items can I safely ship without a way to catch what it breaks.' Framed that way, the answer stops being a number I made up and starts being the thing every other feature needs first."
Why this works
Separates a strong answer from someone who just names a round number like twenty percent.
Stage 4
Give the one decision
Say it like this
"A standing third of every quarter's capacity goes to eval and infrastructure, permanently, not just at launch. Higher in quarter one, since there's nothing to build on yet."
Why this works
This is the direct answer, stated as a number with a rule attached, not a value statement about quality mattering.
Stage 5
Prove it with the compressed failure
Say it like this
"When Haldenport shipped a prompt change meant to help French-language tickets, it quietly broke billing-dispute routing. Misrouting on that category climbed from 4 to 22 percent over five weeks, and nobody knew, because there was no golden set to compare against, only a shift lead who started re-checking every route by hand."
Why this works
Compresses the whole cost of skipping evals into one number a two-week spike would have caught immediately.
Stage 6
Name the trade-off you're accepting
Say it like this
"This means a real feature, like sentiment-based escalation, slips by a month. I'd rather ship it a month late than ship it on top of a routing system I can't verify."
Why this works
Says plainly what gets given up, instead of pretending the thirty percent is free.
Stage 7
Close on the one line
Say it like this
"Fund eval and infrastructure work as a fixed floor, not a leftover, because it's the one line item every other feature quietly depends on to be safe to ship."
Why this works
Restates the direct answer in one breath, which is what a live interview actually rewards.
Let's learn
Here is what happens when a roadmap treats eval and infrastructure work as a task you get to once the real features are shipped.
Before TriageAssist, a Haldenport support agent read each incoming ticket, decided billing or technical or retention, and typed a reply from scratch, about six minutes a ticket. With TriageAssist, a route and a drafted reply come back in under two seconds. The tool launched handling one simple flow, and a shift lead could eyeball a handful of routes each morning and trust the rest.
Every box after the second one depends on the second one existing at all.
Here's the turn: two years in, Haldenport's roadmap added a new language model, then a sentiment-escalation feature, then a routing tweak for high-value accounts, each one a fresh prompt change shipped straight to production. None of it is the problem by itself. The problem is that not one of those changes had anything to compare itself against, so nobody could tell a real improvement from a quiet regression until a person noticed the queue looked wrong.
Misrouting rate, billing-dispute tickets, five weeks after the language-model prompt change
The line moved for five straight weeks before anyone noticed. A golden set built in week one would have flagged it at week two.
At its worst, a roadmap with no standing eval floor doesn't fail loudly. It fails as a slow climb nobody is watching, until the cost of catching up is far bigger than the eval work would ever have cost.
The choice I would take back
Early on, the team decided evals were a launch-week checklist item, something you build once and revisit "once things stabilize." That made sense when TriageAssist handled one flow and one queue. It stopped making sense the moment a second and third prompt change started shipping every quarter, each one able to break something the last one didn't touch.
What I would leave alone: I wouldn't put an eval floor under a purely cosmetic change, like reordering which fields show on the agent's screen. Nothing about routing accuracy depends on that, so the thirty percent doesn't need to cover it.
The lesson: eval and infrastructure work isn't the thing you fund after the roadmap is real. It's the thing that decides whether anything else on the roadmap can be trusted at all.
Now here is the same thing as a story
The short version above is what you'd say defending next quarter's roadmap to leadership. Read this one for how a small, sensible decision turned into three hours nobody had.
The night desk at Haldenport's contact center runs from six to two. Marisol Trench has worked it for four years, and part of the job, before TriageAssist, was reading the first line of a ticket and knowing within a breath which queue it belonged in.
For the first eighteen months after TriageAssist launched, Marisol pulled ten routes at random each shift and checked them against what she'd have called herself. They matched, almost every time. She stopped pulling ten. She pulled two. Eventually she stopped pulling any at all, because the tool had never once been wrong in a way that mattered.
Everything in the top right has to exist before anything else is safe to ship.
Then the roadmap moved fast for two quarters straight: a language-model update, a sentiment-escalation feature, a routing tweak for high-value accounts. Each one shipped behind a launch date. None shipped behind a regression check, because no one had ever built one, and nobody had put "build one" on a roadmap that already felt full.
Knowledge spark: what's a golden set?
A pile of real past cases with the right answer already known, kept on purpose so a new version of the tool can be checked against it before it goes live. Without one, you're comparing a new model to a memory instead of a record.
Five weeks after the language update shipped, Marisol noticed the billing queue looked heavier than usual. She pulled a sample. Almost a quarter of billing-dispute tickets were routed wrong, most of them landing in technical support instead, where nobody could actually resolve a billing hold.
Nobody decided to make routing worse. The roadmap just never funded a way to notice when it did.
Marisol went back to checking every route by hand, the way she had at launch, except now the volume was three times higher. What used to be a five-minute spot-check at the start of shift became three hours a night, every night, until someone finally rebuilt the harness that should have existed from the start.
Four small parts. None of them is a new feature. All four are the reason a feature is safe to ship.
ORDER, the number that decides what ships firstNot a wishlist ranked by excitement. ORDER is what tells you which quiet decision every other decision depends on.
O
Outcome. What all the candidates are competing to move.
Correct routing and fast resolution, without a new failure mode showing up unannounced.
Without naming this, ranking is just opinion dressed up as a plan.
R
Reversibility. Which decision is hardest to undo.
Skipping the eval floor is the hardest to undo, because the damage compounds silently for weeks before anyone can even see it happened.
This is the hardest step, and the one this whole answer actually turns on.
D
Dependency. What unblocks what.
No prompt or model change is safe to ship without a golden set to check it against first.
Eval work isn't a parallel track to features. It's underneath every one of them.
E
Evidence. What you could learn cheaply first.
A two-week spike, building a 200-ticket golden set from real past routes, tells you how bad the current gap already is before committing a whole quarter.
Cheap evidence beats a guess about how urgent this really is.
R
Rank. The actual order, defended.
Eval and infrastructure work first, as a standing third of capacity, every quarter. New features are sequenced behind whatever the golden set says is weakest right now.
The rank follows evidence, not whichever feature leadership is most excited about this month.
The recap, one line per letter: outcome is correct routing without a silent new failure, reversibility is that skipped eval work compounds and can't be undone after the fact, dependency is that eval work sits underneath every feature rather than beside it, evidence is a two-week golden-set spike before committing a full quarter, and rank is a standing third of every quarter going to eval and infrastructure, permanently.
And if you want to be sure it really works, try it somewhere elseSame five letters, a warehouse's pallet-scanning model instead of a support queue. A different kind of decision gets taken back this time.
Corvale Freight runs a model that reads photos of pallets at intake and flags damage before a shipment gets accepted. Mapped onto ORDER: outcome is catching real damage without slowing intake down. Reversibility is highest for the model's confidence calibration, since a badly calibrated model that's been trusted for months is far harder to walk back than a feature that just ships late. Dependency runs the same way, nothing about adding a new damage category is safe until the existing thresholds are known to hold. Evidence is a one-week sample of manually re-inspected pallets against the model's calls. Rank puts calibration work first, ahead of the two new damage categories requested for next quarter. The reversal here isn't an absent golden set, it's a merged step: Corvale had combined "model flags damage" and "shipment accepted" into one automatic action with no pause for a human glance, because pausing felt like it defeated the point of automating intake at all. Once the model's confidence drifted slightly high, that merged step meant nobody caught anything until a damaged pallet was already three states away.
The same four things fund both a support queue and a pallet scanner.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "fund eval work as a standing third of every quarter, because it's the one thing every feature after it depends on," and stop.
Cost: leadership says there's no headcount for a dedicated eval owner this quarter. Say so honestly, and start with the golden set alone, since a comparison point costs almost nothing and everything else builds on it.
The model gets better, for real: if a new model version genuinely routes better across the board, that's still a change that needs the same golden set to confirm it, not a reason to skip the check because the news is good.
Where people run it wrong.
They treat eval work as a one-time launch task instead of a standing percentage of every quarter after.
They let the loudest feature request set the ranking instead of asking what it actually depends on.
They wait for a person to notice the queue looks wrong instead of building the alarm that would have caught it in week two.
How to use it live. The moment someone asks how much of a roadmap should be eval work, ask yourself: which of these items becomes impossible to trust if I skip it. Rank that one first, and the percentage follows.
Eval work isn't squeezed in before features. Features get sequenced behind it.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "what should we build first" or "how much should go where" roadmap question?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. It ranks candidates by which one is hardest to undo if skipped, not by excitement.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Trench, a four-year night-desk shift lead at Haldenport Telecom who used to spot-check TriageAssist's routes and trust the rest.
3 · THE HABIT
What did Marisol stop doing because the tool kept being right?
Tap to flip
ANSWER
She stopped pulling any routes at random to check, since the tool had never been wrong in a way that mattered, until it quietly was.
4 · THE DEPENDENCY
Why is eval work not a parallel track to features?
Tap to flip
ANSWER
Because no prompt or model change is actually safe to ship without a golden set to check it against first, so every feature depends on eval work existing.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating evals as a launch-week checklist item instead of a standing share of every quarter, a call that made sense with one flow and stopped making sense once multiple prompt changes started shipping.
6 · THE NUMBER
Fill in the blank: billing-dispute misrouting climbed from 4 percent to ___ percent over five weeks with no golden set in place.
Tap to flip
ANSWER
22 percent, and a golden set built in week one would have flagged the drift by week two.
7 · THE REPLAY
Same language-model update ships, but the standing eval floor is already funded. What changes?
Tap to flip
ANSWER
The golden set flags the billing-routing drop within days, not five weeks. Marisol never goes back to checking every route by hand, and the fix ships before the queue backs up.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Corvale Freight's pallet-damage scanner. The reversal is a merged step: flagging damage and accepting the shipment got collapsed into one automatic action with no human glance in between.
Check yourself Score: 0 / 0
Multiple choice
1. Per this answer, why does eval and infrastructure work rank ahead of new features, not just alongside them?
A. Because eval work is cheaper to build than most features.
B. Because no feature is actually safe to ship without a way to catch what it silently breaks, so eval work is a dependency, not a parallel option.
C. Because leadership always prefers infrastructure work over visible features.
D. Because eval work has no real cost to the roadmap's timeline.
Show hint
Look at the dependency step and the flow diagram.
Show answer
B. Eval work sits underneath every other roadmap item, not beside it, which is why it gets ranked and funded first.
True or false
2. True or false: this answer argues eval and infrastructure work should get every quarter's full capacity for as long as the product exists.
True
False
Show hint
Look at the direct answer's actual number.
Show answer
False. The recommendation is roughly a third of capacity as a standing floor, higher in the first quarter, with the rest going to features.
Fill in the blank
3. Fill in the blank: Marisol went from checking a handful of routes at random each shift to checking ___ of them, once the drift became visible.
Show hint
Look at "at its worst" and the paragraph right after the highlight block.
Show answer
Every single one. What used to be a five-minute spot-check became three hours a night at the higher ticket volume.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Treating evals as a launch-week checklist item. That made sense when the tool handled one simple flow and there was nothing yet to compare a change against.
Short answer, apply it yourself
5. Think of a tool you use that gets updated regularly. Would you know if an update quietly made it worse at one specific thing, or would you only notice once it piled up?
Show hint
Think about an app whose recommendations, search results, or autocomplete changed after an update, with no note about what changed.
Show answer
Model answer: A navigation app that started suggesting a slower route after an update, with no way to compare it to how it used to route the same trip.
Short answer, where it wouldn't matter
6. Name a change to TriageAssist where this eval floor genuinely wouldn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A cosmetic change to which fields show on the agent's screen. Nothing about routing accuracy depends on it, so it doesn't need the same regression check.
Before you close the answer
Why this works
Tests whether you'll fund eval work as a real, standing line item, or as a nice-to-have that always loses to a feature with a launch date attached to it.
Follow-up traps
"A third sounds arbitrary. Why not ten percent?" Response: it's set by the two-week evidence spike, not picked in advance; ten percent might be right for a stable single-flow product, a third fits a roadmap shipping multiple prompt changes a quarter.
"What if leadership just won't approve that much time with nothing visible to show for it?" Response: show them the five-week misrouting climb as the cost of not doing it; the ask stops being abstract once it's a number leadership already recognizes as expensive.
If pressed
The actual golden set that followed held 240 real tickets stratified by category, refreshed quarterly with newly contested routes, and the drift alarm fired on any category whose match rate fell more than four points week over week.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.