CalculationAdvancedAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #20
Your roadmap assumes costs fall 50 percent this year. How do you plan if they do not?
BOUNDthe roadmap that priced in a discount the vendor never gave
Do the math out loud, then plan for the version where the number you were counting on never shows up. Stridewell runs Coach, a feature inside its running app that turns a runner's watch data into a short, personal voice review after every run. Baz Ondiek is the AI PM who built the free-tier rollout on a 50 percent cost drop that hadn't happened yet.
The direct answer
Don't put a calendar date on a price you don't control. Roll the free tier out in gated bands tied to the actual price you're paying that quarter, and cap the token budget per session yourself, since a longer coaching narrative costs you more than the vendor's price ever will. That way the plan survives even if the discount never arrives.
Do this, in order
Replace the date-based promise with a price-gated trigger.Why: a fixed date commits you before the evidence exists. A trigger only fires once the number actually supports it.
Cap the token budget per session, not just the model choice.Why: the arithmetic shows session length swings the monthly cost more than the vendor's price does.
Run all three scenarios, not just the hopeful one.Why: a single number implies confidence you don't have; a range is what a real plan is built on.
Check the result against what a free user is actually worth.Why: cost per user only means something next to revenue per user, and that comparison is what caught this before full rollout.
Tell support the trigger, not the date.Why: a vague "it's coming eventually" answer trusts less than an honest, specific condition.
How to answer this, stage by stage
Nobody is scoring whether you memorized a pricing trend. They're scoring whether you show real arithmetic instead of a confident guess.
Stage 1
Scope it to one real number
Say it like this
"I'll use Stridewell's Coach feature, and the actual assumption behind the roadmap: that per-session inference cost falls 50 percent this year, letting Coach go free-tier-wide by Q3."
Why this works
Keeps the whole answer anchored in one real cost model instead of talking about pricing in the abstract.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as BOUND. Break it down: state the cost equation. Own the numbers: where each assumption comes from. Use a range: three scenarios, not one. Nail the sanity check: does it survive a smell test. Direction: which assumption swings it most."
Why this works
Tells the interviewer you're about to show arithmetic, not a story dressed up as an estimate.
Stage 3
Break it down, out loud
Say it like this
"Cost per free user per month equals tokens per coaching session, times sessions per month, times price per thousand tokens. Right now that's roughly 2,000 tokens, 20 sessions, at a blended price of six-tenths of a cent per thousand."
Why this works
Shows the equation before touching a single scenario, which is exactly what BOUND's test demands.
Stage 4
Give the one decision
Say it like this
"I wouldn't promise a date at all. I'd roll the free tier out in bands tied to the actual price each quarter, and cap tokens per session myself, so the plan doesn't depend on a discount I don't control."
Why this works
This is the direct answer, stated as a mechanism instead of a hope that prices behave.
Stage 5
Show the range, not one number
Say it like this
"At plan, that's twelve cents a user a month. Actual so far, with price down only fifteen percent and sessions running longer than we scoped, it's forty cents. Worst case, flat pricing and continued growth, fifty-five cents."
Why this works
A range this wide is the whole reason a fixed date was the wrong commitment to make.
Stage 6
Nail the sanity check
Say it like this
"A free user is worth about eighteen cents a month to us from ads and upsell. At forty cents in cost, that feature loses money on every single free user, at scale. That number alone should have stopped the calendar promise before it was ever made."
Why this works
This is the move that separates a real estimate from an estimate nobody checked against reality.
Stage 7
Name the direction that matters most
Say it like this
"It's tempting to blame the vendor's price. But session length swinging longer moves the number more than the price miss does, and session length is the one thing we can actually control with a cap."
Why this works
Shows you know which lever is worth pulling, not just which one is easiest to complain about.
Stage 8
Close on the one line
Say it like this
"Don't bake a hoped-for discount into a public promise. Bake a trigger into the plan instead, and control the one number that's actually yours to control."
Why this works
Restates the direct answer in one breath, which is exactly what a live follow-up rewards.
Let's learn
Here is what happens when a roadmap bets on a price that hasn't fallen yet, and announces the bet as a promise.
Before Coach, a runner finished a session and got nothing but raw numbers: pace, heart rate, distance. With Coach, a short voice review comes back within a minute, built from the same watch data plus a model that turns it into plain coaching language. It launched as a premium feature, priced to cover its own cost with room to spare.
One of these two ideas can survive a price that doesn't move. The other one can't.
Here's the turn: the roadmap that planned to bring Coach to the free tier by Q3 assumed the same 50 percent cost fall the vendor's own pricing trend seemed to point toward. Marketing, working from that roadmap, told users publicly: free Coach, this year, guaranteed. The price fell. Just not by fifty percent, and not for the reason anyone expected.
Cost per free user per month, built up from price and session length
Plan had no session-length overrun built in at all. Both later scenarios do, and it's the bigger of the two bars in each.
At its worst, a public promise built on a hoped-for discount doesn't just cost more than planned. It costs more per user than that user is worth, at a scale where nobody notices until finance runs the number.
The choice I would take back
Marketing published a firm Q3 date for free Coach based on the roadmap's cost model, months before real usage data existed to confirm it. That made sense when the vendor's public pricing trend looked like a straight line down. It stopped making sense the moment real sessions started running longer than the model that produced the fifty percent number ever assumed.
What I would leave alone: I wouldn't gate the premium tier the same way. Premium users already pay enough to cover the actual cost at any of the three scenarios, so there's no trigger needed there.
The lesson: a roadmap number isn't a promise until it's been checked against a range, not a single hopeful line, and against what the thing is actually worth to the person receiving it for free.
Now here is the same thing as a story
The short version above is what you'd say defending the rollout plan to finance. Read this one for how a specific promise turned into a vaguer, less trusted one.
Soren Duclos runs support for Stridewell's app. His headset is on from eight to four, and for eleven months after Coach launched, one question came in almost every shift: when is it coming to the free tier?
Soren's team had a clean, confident answer, straight from marketing's own page: Q3, guaranteed. Users liked that answer. It was specific, and it sounded like a company that knew exactly what it was doing.
Marketing only ever planned for the leftmost branch.
Two weeks before the scheduled Q3 flip, finance flagged the real number: token spend per session had crept up for months as users asked Coach follow-up questions after their run, turning a short review into a longer back-and-forth. The vendor's price had only moved fifteen percent. Someone caught it just in time, before the free tier went fully live and the burn became irreversible.
Knowledge spark: why does output length matter more than price?
A model charges by the token, in and out. A longer answer costs more no matter what the price per token is. If people start asking for longer answers, the bill grows even if the vendor never changes a thing.
Leadership scaled the Q3 rollout back to a smaller, invite-only slice of the free tier. Soren had to walk the promise back. And here's the part that actually cost something: rather than give users the honest, specific reason, Soren told his team to stop giving any date at all, worried a second broken promise would be worse than a vague one.
The company didn't lose trust because the price didn't fall. It lost trust twice, once when the date broke, and again when the honest reason got replaced with nothing.
Users who once heard "Q3, guaranteed" started hearing "it's rolling out gradually," a sentence that answers nothing and reassures nobody. The specific promise had been wrong. The vague one, meant to protect against being wrong again, ended up trusted even less.
The assumption in the top right is the one the promise should have been built around.
BOUND, the arithmetic behind a promise nobody checkedNot a guess dressed up in confident language. BOUND is what makes a plan survive the scenario where the hoped-for number never shows up.
B
Break it down. State the equation.
Cost per free user per month equals tokens per session, times sessions per month, times price per thousand tokens.
Without this, a percentage claim like "costs will fall fifty percent" has nothing to actually mean.
O
Own the numbers. Say where each one came from.
2,000 tokens a session and 20 sessions a month, from current usage logs. A price fall to half of six-tenths of a cent per thousand tokens, from the vendor's own twelve-month pricing trend.
A number with no source is a guess wearing a decimal point.
U
Use a range. Three scenarios, not one.
Plan: twelve cents a user. Actual so far: forty cents. Worst case: fifty-five cents.
This is the hardest step, and the one that turns "the price might not fall" into a plan you can actually act on.
N
Nail the sanity check. Compare it to something real.
A free user is worth about eighteen cents a month. At forty cents in cost, the feature loses money on every free user at scale.
The smell test is what should have stopped the calendar promise before marketing ever published it.
D
Direction. Which assumption swings it most.
Session length, not vendor price, moves the monthly number more, and it's the one Stridewell can actually control with a token cap.
A good estimator names the lever worth pulling. A bad one just blames the vendor.
The recap, one line per letter: break it down is the cost equation stated before touching a single number, own the numbers is stating exactly where 2,000 tokens and a fifty-percent price fall came from, use a range is running plan, actual, and worst case side by side, nail the sanity check is comparing cost per user to what a free user is actually worth, and direction is naming session length, not price, as the lever that matters most.
And if you want to be sure it really works, try it somewhere elseSame five letters, a warehouse robotics vendor's service-call model instead of a fitness app. A different old decision gets taken back this time.
Kelmarsh Robotics runs a model that reads a technician's photo of a jammed sorter arm and drafts a repair note before the visit. Its roadmap assumed GPU rental costs would fall 40 percent this year, letting the tool expand from one regional depot to all forty. Mapped onto BOUND: break it down is cost per repair call equals image tokens plus draft length, times calls per month, times GPU price per hour. Own the numbers states the current call volume and the vendor's published GPU roadmap. Use a range runs plan, a 10 percent fall, and a flat-price case. Nail the sanity check compares the cost per call to the labor hour it replaces. Direction finds that call volume, not GPU price, swings the number most, since technicians started sending three photos per jam instead of one once the tool got a little slower to reject blurry ones. The old decision here isn't a promise, it's an absent state: nobody logged how many images an average call actually used, so the volume assumption was never checked against real data until the estimate came in three times too low.
The same three-scenario shape works whether the unit is a coaching session or a repair call.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "gate the rollout to the actual price, cap the token budget, and don't put a date on a discount you don't control," and stop.
Cost: there's no time to build a full pricing dashboard this quarter. Say so honestly, and start with the token cap alone, since it's the lever inside your own control and costs nothing to add.
The model gets better, for real: if the price genuinely falls fifty percent as planned, that's still worth confirming against the range before flipping the switch, not a reason to skip the check because the news is good.
Where people run it wrong.
They treat a vendor's published pricing trend as a commitment instead of a forecast with a range around it.
They blame the vendor's price when the arithmetic actually points at something the team controls, like output length.
They replace a broken specific promise with a vaguer one, instead of an honest, specific trigger.
How to use it live. The moment someone gives you a roadmap built on a future price, ask yourself: what does the plan do if that number simply never arrives. If you don't have an answer, neither does the roadmap.
None of these four parts existed before the near miss. All four exist now.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a cost-assumption, "show your math" roadmap question?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, direction. It's built for estimation questions where FLIPS doesn't fit.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Soren Duclos, Stridewell's support lead, who gave users a firm Q3 date for free Coach access straight from marketing's own page.
3 · THE HABIT
What did Soren's team stop doing after the rollout got scaled back?
Tap to flip
ANSWER
They stopped giving any specific date at all, replacing a confident promise with a vague "rolling out gradually," to avoid breaking a second promise.
4 · THE EQUATION
What's the cost equation this answer builds everything on?
Tap to flip
ANSWER
Cost per free user per month equals tokens per session, times sessions per month, times price per thousand tokens.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Marketing publishing a firm Q3 date based on a hoped-for price fall, months before real usage data existed to confirm the assumption held.
6 · THE NUMBER
Fill in the blank: at plan, cost was 12 cents a user a month. Actual so far, it's ___ cents.
Tap to flip
ANSWER
40 cents, more than double what a free user is worth in ad and upsell revenue.
7 · THE REPLAY
Same price miss, but the trigger-based plan was already in place. What changes?
Tap to flip
ANSWER
No public date ever existed to break. The rollout gates itself to the actual price each quarter, tokens are capped per session, and support gives users an honest condition instead of a vague non-answer.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Kelmarsh Robotics's repair-note tool. The reversal is an absent state: nobody logged real image volume per call, so the cost estimate was never checked against actual usage.
Check yourself Score: 0 / 0
True or false
1. True or false: this answer's main recommendation is to wait and see if the vendor's price eventually falls 50 percent before changing the rollout plan.
True
False
Show hint
Look at the direct answer's actual mechanism.
Show answer
False. The plan gates rollout to the actual price each quarter and caps token spend directly, instead of waiting on the vendor.
Fill in the blank
2. Fill in the blank: a free user is worth about ___ cents a month to Stridewell from ads and upsell.
Show hint
Look at the sanity check stage of the walkthrough.
Show answer
18 cents. Against a 40-cent actual cost, that's a loss on every free user at scale.
Multiple choice
3. According to the direction step, which assumption actually swings the monthly cost the most?
A. The vendor's price per token.
B. How long each coaching session's output runs.
C. How many total free users sign up.
D. Which model version is running that month.
Show hint
Look at the quadrant diagram and the direction step.
Show answer
B. Session length sits highest on both axes of the quadrant, and it's the lever Stridewell can actually control.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Publishing a firm Q3 date based on the cost model. It made sense when the vendor's pricing trend looked like a straight line down, before real session-length data existed to check it.
Short answer, apply it yourself
5. Think of a product or plan you've seen that promised a price or feature "once costs come down." What would you ask to check if that promise was actually safe to make?
Show hint
Think about a subscription price, a hardware discount, or a free-tier expansion tied to some future cost drop.
Show answer
Model answer: Ask what happens to the plan if the cost never falls, and whether anyone checked the promise against a range instead of one hopeful number.
Short answer, the number question
6. If session length had stayed exactly at the planned 2,000 tokens, but the price still only fell 15 percent instead of 50, would the actual cost of 40 cents a user still hold? Why or why not?
Show hint
Look at the stacked bar chart's two components.
Show answer
Model answer: No, it would land closer to 24 cents, since the session-length component is what pushes actual cost from 24 up to 40. The price miss alone doesn't explain the full overrun.
Before you close the answer
Why this works
Tests whether you'll show real arithmetic under a cost assumption, or retreat into a confident-sounding paragraph with no numbers a listener could actually check.
Follow-up traps
"Isn't a token budget cap just going to make the coaching worse?" Response: not if it's set above what a good review actually needs; the cap targets the small share of sessions running long on follow-up chatter, not the review itself.
"What if finance hadn't caught this before the Q3 flip?" Response: the free tier would have gone fully live at a loss on every user, which is exactly why the plan needed a price trigger built in from day one instead of relying on someone catching it in time.
If pressed
The token cap that shipped afterward limited each session to 1,600 total tokens, with a second short follow-up allowed only if the runner asked a direct question, cutting the session-length overrun by about two-thirds without users noticing a shorter review.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.