How do you forecast cost for a feature with no usage history?
One guess dressed up as certain is how a $1,800 forecast turns into a $6,400 bill. Two guesses that have to agree is how it doesn't happen again.
- Build two independent estimates, one borrowed from a real feature and one built from raw unit cost, and only trust the number where they agree.Why: a single guess has nothing to check itself against, and that's exactly how a bad assumption survives all the way to launch.
- Publish a range, a low number and a high number, never one point figure.Why: a single number tells finance you're sure. Before any real user has touched the thing, you aren't, and pretending otherwise is what makes the miss feel like a surprise later.
- Set a fixed date to recheck the range with real data, at first usage, not at the next budget cycle.Why: catching a wrong assumption in two weeks costs an email. Catching it in a quarter costs a scramble.
- Build the unit cost from what actually drives it, minutes of audio talked, not a per-seat headcount.Why: real-time audio bills by the minute someone is actually talking, so a forecast built on login counts will be wrong in a completely different direction than the real cost.
- Chase down any gap between the two estimates before the number ever reaches finance.Why: when a borrowed curve and a from-scratch build land far apart, that gap is usually a wrong assumption hiding, not noise worth averaging away.
- Leave the fine-grained breakdown, by language pair, by call length bucket, for later, once real volume exists.Why: building that model now spends a week modeling something real usage data will hand you for free in a month.
How to answer this, stage by stage
Nobody is grading whether you can build a spreadsheet. They're grading whether you know a single guess and a checked estimate are two different things, and whether your method has a built-in way to catch itself being wrong.
Let's learn
Voxferry sits inside a live video or phone call between two people who don't speak the same language, and turns what one person says into the other person's language, out loud, in close to real time, while they're both still talking.
Voxferry already had one working translation feature before this: a text-chat translator, live for fourteen months, on all twenty three hundred of the company's enterprise seats. It had real history. By day ninety after launch, twenty eight percent of those seats were using it at least once a week, and it settled into a monthly infrastructure cost of about thirteen hundred dollars.
The new feature, live speech interpretation for real phone and video calls, has none of that. It hasn't launched. There is no curve to look at, no cost to point to, nothing but a product spec and a launch date on a roadmap.
Before this method existed, Voxferry forecast new features with a single number from a vendor slide: real-time voice pipelines cost about five times what a text message costs, per interaction. That number went straight into a spreadsheet, no second check, and became the budget line finance approved.
The mistake worth worrying about was never getting the average cost per call wrong by a little. It was never modeling what happens when usage doesn't ramp evenly, one huge customer flips a feature on for their whole team at once, and the flat guess has no way to see that coming because it was never built from real behavior in the first place.
Three weeks before this feature shipped, the old captions incident was still fresh enough that Wulfric wasn't willing to repeat it. So this time, the top-down estimate borrowed the text feature's own curve, twenty eight percent of seats active weekly by day ninety, discounted forty percent for the extra friction of a live, mic-on feature. The bottom-up estimate started from nothing: nine cents a minute of two-way audio, times a twenty two minute average call, times how many calls an active seat makes in a week. The two landed within twenty dollars of each other, around forty seven hundred a month. That agreement, not either number alone, was what let Wulfric take a range to finance instead of a guess.
What I'd leave alone: small, low-volume settings changes, like adding a new toggle to an admin panel, don't need two independent estimates and a recheck date. The cost of being wrong about a toggle is a rounding error. Save the two-method process for anything that bills by usage, where a wrong guess compounds every single call.
The lesson: a forecast for something with no history isn't wrong because the number is off. It's wrong because nothing in it could have told you it was off before the bill did. Build the check in before you build the number.
Now here is the same thing as a story
Read the story below when you want to feel why two guesses that land close together earn more trust than one guess that sounds sure, not just be told that they do.
Wulfric Feuerstein had forecast the cost of every feature Voxferry shipped for three years. He was good at it. Rarely off by more than fifteen percent, and finance had stopped asking him to show his work, because his work had always held up.
In the early years, that trust was earned honestly. Every new feature was close cousins with something already live, a new chat mode next to an old chat mode, a bigger context window on a text pipeline he already understood cold. He'd pull the closest comparison, do the math by hand, and land within a few percent almost every time.
Then the roadmap sped up. Three features shipping a quarter instead of one. Wulfric started leaning on a shortcut a vendor had handed the team eighteen months earlier, in a pitch deck slide near the end of a sales call: real-time voice pipelines cost about five times what a text interaction costs. He wrote it into the internal forecasting template as a constant, the way you'd write in a tax rate. For a while it worked well enough that nobody looked twice.
Finance started approving his numbers on sight. Not because the method had earned it fresh each time, but because Wulfric had. His name on the forecast became the check, instead of the forecast itself being checked.
The trigger wasn't a lawsuit or a board meeting. It was a new engineer, three weeks into building the cost dashboard for the live captions launch, asking a question in a stand-up almost as an aside: "Wait, why five times? Where does that number actually come from?" Wulfric opened his mouth to answer and realized he didn't have one. It was a slide. Not a measurement.
Live captions launched anyway, on the old number. The forecast said eighteen hundred dollars a month. In week six, one customer, four hundred seats, turned captions on company-wide two days before their CEO's quarterly earnings call, wanting every internal meeting captioned for a hearing-impaired executive. Real cost that month: sixty four hundred dollars. Finance called an emergency review. The feature got rate-limited for new signups for eleven days while everyone figured out what had actually happened, which made the product worse for every other customer just to buy time to understand a number that should have been checked before launch.
The real cost wasn't the overage itself. It was that finance stopped trusting Wulfric's number on sight, the exact thing that had let every earlier launch move fast. Every forecast after that needed a second reviewer, which added two weeks to every launch for the rest of the year.
The decision Wulfric would take back happened in a five-minute moment near the end of a vendor call, over a year before captions shipped. Someone on the call mentioned the five-times figure almost in passing, a rough industry number meant to set expectations, not a promise about Voxferry's own pipeline. Nobody in that call wrote down where it came from. Somebody just typed it into the template afterward, and it sat there long enough to look like a fact.
Run the interpretation feature's launch the old way, and it repeats the same mistake with a bigger number attached, since live calls process far more audio than captions ever did. Run it the new way: Wulfric builds the top-down number from the text feature's real curve, and the bottom-up number from the actual pipeline cost, minute by minute. They land within twenty dollars of each other, both near forty seven hundred a month. He publishes a range, forty seven hundred to seventy five hundred, wide enough to survive one large customer flipping the switch all at once, and puts day fourteen on his calendar as the day the guess gets replaced. Day fourteen arrives. Real cost: about forty one hundred a month, comfortably inside the range. No emergency review. No rate limiting. Finance reads one line and moves on with their day.
One design trusted a single number because the person attached to it had been right before. The other design trusted a number because two different paths to it agreed, and set a date to stop trusting it the moment real data existed.
What I'd tell myself, back on that vendor call: a number nobody can trace to a measurement is not a fact just because it's the only number in the room. The moment a forecast has to stand on its own with nothing to check it, that forecast owes you a second, independent way to build it, not a bigger font on the slide.
SPARK, built for a number nobody has yet
Not a checklist to recite. Each letter has to survive the same customer flipping four hundred seats on at once that the story just walked through.
Three things worth stating directly, since this is where the real judgment sits. The alternative Wulfric considered and rejected was trusting the vendor's flat multiplier alone, with no independent bottom-up build to check it against, exactly what caused the captions miss. It lost because a number nobody can trace to a real measurement isn't a forecast, it's a rumor with a decimal point. The AI specific failure worth naming by name is the borrowed curve itself: assuming a brand-new feature's adoption will shadow an existing feature's shape, when a live, mic-on interpretation feature might get adopted in sudden bursts a text feature never showed, a form of distribution shift between what history you have and what you're actually trying to predict. The guardrail is the day-14 recheck, collapsing the range with real numbers before a full quarter can pass. And the quality, latency, and cost trade-off worth naming too: Voxferry's pipeline keeps a rolling ninety-second window of each call's audio as context for the translation model, long enough to catch a pronoun that refers back to something said a minute earlier, short enough to keep re-processing cost and delay small; a longer window would translate more accurately but cost more per minute and add a beat of lag two people on a live call would actually notice.
And if you want to be sure it really works, try it somewhere else
Same five letters, a drive-thru speaker instead of a phone call, and this time the closest comparison is a mobile app, not a chat feature.
OrderVale is a voice ordering assistant that takes an order at a drive-thru speaker and sends it straight to the kitchen, no cashier needed to type it in. Crispline, a four hundred location quick-service chain, is piloting it at twelve stores and has to forecast what it would cost to run at every location before signing a chain-wide contract. Rosamund Meurer runs finance for operations at Crispline.
S, situation: before OrderVale, every drive-thru order got typed in by a cashier wearing a headset, and Rosamund had never had to forecast an AI feature's usage cost at all, only headcount and hours.
P, payoff: the habit worth building isn't "trust OrderVale's own sales estimate for cost per store." It's building a number Crispline can check itself, from its own order volume, before four hundred locations are all committed to it.
A, anchor: top-down, borrow the adoption curve of Crispline's own mobile ordering app, which took nine months to reach sixty percent of transactions company-wide, discounted for the fact that a drive-thru customer doesn't choose to use OrderVale, it's just how the speaker works now. Bottom-up, add up the real cost of one order: audio capture, intent parsing against the menu, a spoken confirmation back to the customer, about six cents an order. Publish the range, and recheck at the end of the twelve-store pilot's first full week.
R, risk: a regional promotion pushed one pilot store's order volume up sixty percent for eleven days, and the bottom-up number, built on an average order, briefly underpriced the real cost of a promo-week rush before the recheck caught it.
K, keep out: no per-menu-item cost breakdown before the chain-wide rollout, and no live per-store dashboard for every store manager to watch. One report, at the end of the pilot's first week, with the real number next to the range it was supposed to fall inside.
Same method, a different weak spot: a phone call's cost scales with minutes talked. A drive-thru order's cost scales with how busy the lot is that hour, so the bottom-up build needed a promotion-week case, not just an average week, or the range quietly stopped covering the real range of what could happen.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, two estimates, a published range, a fixed recheck date, and give the one number, forty seven hundred to seventy five hundred a month, both paths landing within twenty dollars of each other.
Cost: there's no budget this quarter for a fancier live dashboard. Build the one-page range with a recheck date instead, it costs almost nothing and it's the part that actually changes a decision.
The model got better, for real: say the translation pipeline gets thirty percent cheaper per minute next quarter. That's not a reason to go back to a single guess. A cheaper unit cost still needs its own bottom-up rebuild, or the range just becomes wrong in the other direction, an overestimate nobody catches because it never gets checked either.
Where people run it wrong.
They trust a vendor's rule of thumb because it's the only number in the room, and never build their own second check.
They publish a single point number to look confident, and quietly turn a wide uncertainty into a promise nobody can keep.
They set the recheck date for the next quarterly budget review instead of the first real week of usage, so a wrong guess gets three months to compound before anyone looks at it again.
How to use it live. Say the real tension out loud before answering: "is this asking me for a number, or for a method that catches itself being wrong." That buys a beat, and it's almost always the second one.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't publishing a range just a way to avoid committing to a number?" Response: no, because the commitment is the recheck date itself, a fixed point where the range gets replaced by a real number. A hedge with no trigger would be avoiding commitment. This one has a deadline built into it.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Cost modeling and unit economics
- #1 Build the cost-per-interaction model for a feature with a 2,000-token prompt and a 500-token response.
- #2 What cost drivers exist for an AI feature beyond model tokens?
- #3 Explain how a RAG pipeline's cost structure differs from a single model call.
- #4 How does prompt caching change your unit economics, and when does it not help?
- #5 Model the monthly cost of a feature used by 50,000 users averaging 12 interactions each.
- #6 What is the cost impact of moving from a single call to a five-step agent?