CalculationAdvancedQuality, Cost & Token Economics / Cost modeling and unit economics / #20

How do you forecast cost for a feature with no usage history?

Cost modeling and unit economics

One guess dressed up as certain is how a $1,800 forecast turns into a $6,400 bill. Two guesses that have to agree is how it doesn't happen again.

The direct answer
Build the cost forecast two different ways at the same time. Borrow the adoption curve from a feature that already has real users, and separately add up the raw unit cost of one call from scratch. Publish the range those two numbers create, a low and a high, never a single point guess, and set a fixed date, the first real two weeks of usage, to collapse that range with real numbers instead of waiting for a budget review to notice it was wrong.
Do this, in order
  1. Build two independent estimates, one borrowed from a real feature and one built from raw unit cost, and only trust the number where they agree.Why: a single guess has nothing to check itself against, and that's exactly how a bad assumption survives all the way to launch.
  2. Publish a range, a low number and a high number, never one point figure.Why: a single number tells finance you're sure. Before any real user has touched the thing, you aren't, and pretending otherwise is what makes the miss feel like a surprise later.
  3. Set a fixed date to recheck the range with real data, at first usage, not at the next budget cycle.Why: catching a wrong assumption in two weeks costs an email. Catching it in a quarter costs a scramble.
  4. Build the unit cost from what actually drives it, minutes of audio talked, not a per-seat headcount.Why: real-time audio bills by the minute someone is actually talking, so a forecast built on login counts will be wrong in a completely different direction than the real cost.
  5. Chase down any gap between the two estimates before the number ever reaches finance.Why: when a borrowed curve and a from-scratch build land far apart, that gap is usually a wrong assumption hiding, not noise worth averaging away.
  6. Leave the fine-grained breakdown, by language pair, by call length bucket, for later, once real volume exists.Why: building that model now spends a week modeling something real usage data will hand you for free in a month.

How to answer this, stage by stage

Nobody is grading whether you can build a spreadsheet. They're grading whether you know a single guess and a checked estimate are two different things, and whether your method has a built-in way to catch itself being wrong.

1
Scope it to one real product and one real forecast
Say it like this
"Let's ground this in Voxferry. It sits inside a live call and turns what one person says into the other person's language, out loud, while they're still talking. Wulfric Feuerstein is the PM who owns the cost forecast for it, and it hasn't launched yet, so there's no usage history at all."
Why this works
A vague "how would you forecast cost" answer turns into a spreadsheet lecture fast. One real feature with zero history makes it a real design problem instead.
2
Say your structure out loud before diving in
Say it like this
"I'm going to build two estimates that don't lean on each other, one borrowed from a feature that already has real users, one built straight from unit cost, and I won't trust the number until I see how close they land."
Why this works
Tells the interviewer you have a method, not just a guess with confidence attached to it.
3
Reframe the question before answering it
Say it like this
"This isn't really asking me for a number. It's asking whether I know a forecast with no history behind it needs a way to check itself, because the alternative is one confident-sounding guess that nobody questions until the bill shows up."
Why this works
Stops you from reciting a formula when the real question is about judgment under uncertainty.
4
Give the anchor, the actual design decision
Say it like this
"Top-down, I take the text-chat translation feature's own adoption curve, twenty eight percent of seats active weekly by day ninety, and I discount it forty percent because voice needs a mic and a live moment, more friction than typing a message whenever you feel like it. Bottom-up, I add up one call's real cost, minute by minute: speech to text, live translation, and voice synthesis for both sides of the conversation, about nine cents a minute, times a twenty two minute average call. If those two paths land within about double of each other, I publish the range they make, and I put a hard date on my calendar, day fourteen of real usage, to throw the guess out and use the real number instead."
Why this works
This is the actual answer to the question. Everything else defends it.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without this. Voxferry's live captions feature launched on one guess, a vendor's rule of thumb that captions cost about five times what a text message costs, no second check. The forecast said eighteen hundred dollars a month. By week six, one customer had turned captions on for all four hundred of their seats ahead of a public earnings call, and the real bill was sixty four hundred. Nothing in that single guess ever asked what happens if one customer moves all at once."
Why this works
Shows the real cost of a single unchecked guess, not just the mechanism behind it.
6
Say what you'd watch after launch
Say it like this
"I'd watch real cost per day against the range starting day one, not day ninety. If it's tracking near the low end, fine, keep going. If it's climbing toward the high end before day fourteen even arrives, that's the signal to pull the recheck forward instead of waiting for the date on the calendar."
Why this works
Shows you think about the forecast as a living check, not a number you set once and forget.
7
Close on the decision, not the story
Say it like this
"So: two estimates that don't lean on each other, a published range instead of a point number, and a hard date to swap the guess for reality. That's what forecasting cost with no history actually means."
Why this works
Ending on the method, not the anecdote, is what makes this sound like something you'd reuse, not a story you told once.

Let's learn

Voxferry sits inside a live video or phone call between two people who don't speak the same language, and turns what one person says into the other person's language, out loud, in close to real time, while they're both still talking.

Voxferry already had one working translation feature before this: a text-chat translator, live for fourteen months, on all twenty three hundred of the company's enterprise seats. It had real history. By day ninety after launch, twenty eight percent of those seats were using it at least once a week, and it settled into a monthly infrastructure cost of about thirteen hundred dollars.

The new feature, live speech interpretation for real phone and video calls, has none of that. It hasn't launched. There is no curve to look at, no cost to point to, nothing but a product spec and a launch date on a roadmap.

Knowledge spark: what's top-down versus bottom-up estimating? Top-down starts from something you already know, a bigger number or a similar product, and scales it down to fit the new thing. Bottom-up starts from nothing and builds the number back up from its smallest real pieces. Used alone, each one can be quietly wrong. Used together, they catch each other.

Before this method existed, Voxferry forecast new features with a single number from a vendor slide: real-time voice pipelines cost about five times what a text message costs, per interaction. That number went straight into a spreadsheet, no second check, and became the budget line finance approved.

A guess that nobody can check against anything is not a forecast. It is a hope with a dollar sign in front of it.

The mistake worth worrying about was never getting the average cost per call wrong by a little. It was never modeling what happens when usage doesn't ramp evenly, one huge customer flips a feature on for their whole team at once, and the flat guess has no way to see that coming because it was never built from real behavior in the first place.

Voxferry's interpretation feature: two independent estimates, checked against the real day-14 number
$8,000 $4,000 $0 $4,680 Top-down $4,700 Bottom-up $4,700-$7,500 Published range $4,100 Real, day 14
Top-down, text curve scaled for voiceBottom-up, built from a minuteRange published to financeReal day-14 number
The two guessed numbers landed twenty dollars apart, which is what earned the range its trust. The real number, three weeks later, sat comfortably inside it.

Three weeks before this feature shipped, the old captions incident was still fresh enough that Wulfric wasn't willing to repeat it. So this time, the top-down estimate borrowed the text feature's own curve, twenty eight percent of seats active weekly by day ninety, discounted forty percent for the extra friction of a live, mic-on feature. The bottom-up estimate started from nothing: nine cents a minute of two-way audio, times a twenty two minute average call, times how many calls an active seat makes in a week. The two landed within twenty dollars of each other, around forty seven hundred a month. That agreement, not either number alone, was what let Wulfric take a range to finance instead of a guess.

The choice that mattered The captions forecast used the vendor's flat multiplier as if it were a fact about Voxferry's own pipeline, instead of a rough number from a sales deck. Nobody ever built Voxferry's own bottom-up number to check it against, so there was nothing to catch it being wrong before the bill arrived.

What I'd leave alone: small, low-volume settings changes, like adding a new toggle to an admin panel, don't need two independent estimates and a recheck date. The cost of being wrong about a toggle is a rounding error. Save the two-method process for anything that bills by usage, where a wrong guess compounds every single call.

The lesson: a forecast for something with no history isn't wrong because the number is off. It's wrong because nothing in it could have told you it was off before the bill did. Build the check in before you build the number.

Now here is the same thing as a story

Read the story below when you want to feel why two guesses that land close together earn more trust than one guess that sounds sure, not just be told that they do.

Wulfric Feuerstein had forecast the cost of every feature Voxferry shipped for three years. He was good at it. Rarely off by more than fifteen percent, and finance had stopped asking him to show his work, because his work had always held up.

In the early years, that trust was earned honestly. Every new feature was close cousins with something already live, a new chat mode next to an old chat mode, a bigger context window on a text pipeline he already understood cold. He'd pull the closest comparison, do the math by hand, and land within a few percent almost every time.

Then the roadmap sped up. Three features shipping a quarter instead of one. Wulfric started leaning on a shortcut a vendor had handed the team eighteen months earlier, in a pitch deck slide near the end of a sales call: real-time voice pipelines cost about five times what a text interaction costs. He wrote it into the internal forecasting template as a constant, the way you'd write in a tax rate. For a while it worked well enough that nobody looked twice.

Finance started approving his numbers on sight. Not because the method had earned it fresh each time, but because Wulfric had. His name on the forecast became the check, instead of the forecast itself being checked.

The trigger wasn't a lawsuit or a board meeting. It was a new engineer, three weeks into building the cost dashboard for the live captions launch, asking a question in a stand-up almost as an aside: "Wait, why five times? Where does that number actually come from?" Wulfric opened his mouth to answer and realized he didn't have one. It was a slide. Not a measurement.

Live captions launched anyway, on the old number. The forecast said eighteen hundred dollars a month. In week six, one customer, four hundred seats, turned captions on company-wide two days before their CEO's quarterly earnings call, wanting every internal meeting captioned for a hearing-impaired executive. Real cost that month: sixty four hundred dollars. Finance called an emergency review. The feature got rate-limited for new signups for eleven days while everyone figured out what had actually happened, which made the product worse for every other customer just to buy time to understand a number that should have been checked before launch.

We didn't lose forty six hundred dollars to a bad guess. We lost three weeks where nobody, including Wulfric, could tell finance what any future forecast was actually worth.

The real cost wasn't the overage itself. It was that finance stopped trusting Wulfric's number on sight, the exact thing that had let every earlier launch move fast. Every forecast after that needed a second reviewer, which added two weeks to every launch for the rest of the year.

The decision Wulfric would take back happened in a five-minute moment near the end of a vendor call, over a year before captions shipped. Someone on the call mentioned the five-times figure almost in passing, a rough industry number meant to set expectations, not a promise about Voxferry's own pipeline. Nobody in that call wrote down where it came from. Somebody just typed it into the template afterward, and it sat there long enough to look like a fact.

Run the interpretation feature's launch the old way, and it repeats the same mistake with a bigger number attached, since live calls process far more audio than captions ever did. Run it the new way: Wulfric builds the top-down number from the text feature's real curve, and the bottom-up number from the actual pipeline cost, minute by minute. They land within twenty dollars of each other, both near forty seven hundred a month. He publishes a range, forty seven hundred to seventy five hundred, wide enough to survive one large customer flipping the switch all at once, and puts day fourteen on his calendar as the day the guess gets replaced. Day fourteen arrives. Real cost: about forty one hundred a month, comfortably inside the range. No emergency review. No rate limiting. Finance reads one line and moves on with their day.

One design trusted a single number because the person attached to it had been right before. The other design trusted a number because two different paths to it agreed, and set a date to stop trusting it the moment real data existed.

What I'd tell myself, back on that vendor call: a number nobody can trace to a measurement is not a fact just because it's the only number in the room. The moment a forecast has to stand on its own with nothing to check it, that forecast owes you a second, independent way to build it, not a bigger font on the slide.

SPARK, built for a number nobody has yet

Not a checklist to recite. Each letter has to survive the same customer flipping four hundred seats on at once that the story just walked through.

SSituation. Who is this person, and how does the job get done today, without you?
Wulfric Feuerstein, PM at Voxferry, owns the cost forecast for a live speech interpretation feature that hasn't launched. Before this method, a brand-new feature with no history got one number, borrowed from a vendor's rough rule of thumb, written into a spreadsheet as if it were measured.
Name the real decision this forecast has to survive, or the method floats free of the actual budget line it feeds.
Hand sketched labeled parts diagram titled the whole forecast before this method. Center icon a document labeled one number no check. Four callouts around it: vendor's flat multiplier, no adoption curve, no range just a point, no recheck date.
The whole forecast Wulfric used to hand finance, before this redesign.
PPayoff. What habit do you want this to build?
Not "trust the number because the person attached to it has been right before." Specifically: finance and product together look at a range built two independent ways, and treat the forecast as something to keep checking against real usage, not a figure set once at launch and never touched again.
A named habit produces a named forecasting method. A vague goal like "get better at estimating" produces nothing anyone can actually do on a Tuesday.
AAnchor. The one design decision everything else hangs on.
Build every cost forecast with no usage history two ways at once. Top-down: borrow the adoption curve of the closest live feature, discounted for whatever makes the new one harder to start using. Bottom-up: add up the real unit cost of one instance of use, from its smallest real pieces. Publish the range those two make, P50 to P90, and set a fixed date, first real usage, to collapse the range with actual numbers.
This is the actual design decision. If it doesn't visibly survive the next letter, it's a slogan, not an anchor.
Hand sketched icon list diagram titled the anchor, two guesses that have to earn each other. Row one a gauge icon, top-down, borrow the text feature's curve scaled for voice friction. Row two a scale icon, bottom-up, add up one call's real cost minute by minute. Row three a funnel icon, publish the range the two numbers make, P50 to P90. Row four a document icon, recheck at day 14 with real calls, not next quarter.
Two independent paths, one published range, one hard date to replace it.
RRisk. What breaks the first time you're wrong?
Real usage never ramps as smoothly as either curve assumes. One large customer turns the feature on for their whole team at once, the way it happened with captions, and actual cost jumps past what either estimate alone predicted. The design has to survive that, not pretend usage always arrives evenly.
A range that only works when every customer adopts gradually, like the average, isn't a design. It's a hope with a band drawn around it.
Hand sketched flow diagram titled the day one customer flips on 400 seats at once. Four boxes in sequence: one customer turns on 400 seats, real cost jumps past P50, day 14 recheck catches it, emphasized, range updates budget holds.
The range is allowed to be tested hard once, as long as the day-14 recheck catches it before a quarter passes.
KKeep out. What do you deliberately not build?
No per-language-pair, per-call-length breakdown on day one, a model nobody has time to maintain before a single real call exists. No live, cent-by-cent cost dashboard promising precision that early volume can't actually support. Just the range, the two methods behind it, and the recheck date.
A forecast that hands finance a fake-precise breakdown before real volume exists doesn't add clarity. It just moves the guessing somewhere harder to spot.

Three things worth stating directly, since this is where the real judgment sits. The alternative Wulfric considered and rejected was trusting the vendor's flat multiplier alone, with no independent bottom-up build to check it against, exactly what caused the captions miss. It lost because a number nobody can trace to a real measurement isn't a forecast, it's a rumor with a decimal point. The AI specific failure worth naming by name is the borrowed curve itself: assuming a brand-new feature's adoption will shadow an existing feature's shape, when a live, mic-on interpretation feature might get adopted in sudden bursts a text feature never showed, a form of distribution shift between what history you have and what you're actually trying to predict. The guardrail is the day-14 recheck, collapsing the range with real numbers before a full quarter can pass. And the quality, latency, and cost trade-off worth naming too: Voxferry's pipeline keeps a rolling ninety-second window of each call's audio as context for the translation model, long enough to catch a pronoun that refers back to something said a minute earlier, short enough to keep re-processing cost and delay small; a longer window would translate more accurately but cost more per minute and add a beat of lag two people on a live call would actually notice.

And if you want to be sure it really works, try it somewhere else

Same five letters, a drive-thru speaker instead of a phone call, and this time the closest comparison is a mobile app, not a chat feature.

OrderVale is a voice ordering assistant that takes an order at a drive-thru speaker and sends it straight to the kitchen, no cashier needed to type it in. Crispline, a four hundred location quick-service chain, is piloting it at twelve stores and has to forecast what it would cost to run at every location before signing a chain-wide contract. Rosamund Meurer runs finance for operations at Crispline.

S, situation: before OrderVale, every drive-thru order got typed in by a cashier wearing a headset, and Rosamund had never had to forecast an AI feature's usage cost at all, only headcount and hours.

P, payoff: the habit worth building isn't "trust OrderVale's own sales estimate for cost per store." It's building a number Crispline can check itself, from its own order volume, before four hundred locations are all committed to it.

A, anchor: top-down, borrow the adoption curve of Crispline's own mobile ordering app, which took nine months to reach sixty percent of transactions company-wide, discounted for the fact that a drive-thru customer doesn't choose to use OrderVale, it's just how the speaker works now. Bottom-up, add up the real cost of one order: audio capture, intent parsing against the menu, a spoken confirmation back to the customer, about six cents an order. Publish the range, and recheck at the end of the twelve-store pilot's first full week.

R, risk: a regional promotion pushed one pilot store's order volume up sixty percent for eleven days, and the bottom-up number, built on an average order, briefly underpriced the real cost of a promo-week rush before the recheck caught it.

K, keep out: no per-menu-item cost breakdown before the chain-wide rollout, and no live per-store dashboard for every store manager to watch. One report, at the end of the pilot's first week, with the real number next to the range it was supposed to fall inside.

The decision Rosamund would take back The pilot's bottom-up estimate was built from an average order across a normal week, with no version of the number for a promotion week. Nothing in the plan asked what a busy week would do to the average, so the first one nearly slipped past unnoticed.
Hand sketched decision tree diagram titled which forecast does Crispline actually trust. Root question do the two estimates agree. Three branches: within about 2x of each other leads to ship the P50 number. More than 2x apart leads to investigate before shipping a number. No analog feature exists at all leads to widen the band, recheck at day 1.
Same anchor, a different kitchen. The mobile app's curve plays the part the text feature played at Voxferry.

Same method, a different weak spot: a phone call's cost scales with minutes talked. A drive-thru order's cost scales with how busy the lot is that hour, so the bottom-up build needed a promotion-week case, not just an average week, or the range quietly stopped covering the real range of what could happen.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, two estimates, a published range, a fixed recheck date, and give the one number, forty seven hundred to seventy five hundred a month, both paths landing within twenty dollars of each other.
Cost: there's no budget this quarter for a fancier live dashboard. Build the one-page range with a recheck date instead, it costs almost nothing and it's the part that actually changes a decision.
The model got better, for real: say the translation pipeline gets thirty percent cheaper per minute next quarter. That's not a reason to go back to a single guess. A cheaper unit cost still needs its own bottom-up rebuild, or the range just becomes wrong in the other direction, an overestimate nobody catches because it never gets checked either.

Where people run it wrong.
They trust a vendor's rule of thumb because it's the only number in the room, and never build their own second check.
They publish a single point number to look confident, and quietly turn a wide uncertainty into a promise nobody can keep.
They set the recheck date for the next quarterly budget review instead of the first real week of usage, so a wrong guess gets three months to compound before anyone looks at it again.

How to use it live. Say the real tension out loud before answering: "is this asking me for a number, or for a method that catches itself being wrong." That buys a beat, and it's almost always the second one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
SPARK: design against the failure before you build. Here, that means designing a forecasting method against the failure of trusting one unchecked guess, before a single real user exists.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Wulfric Feuerstein, PM at Voxferry, an AI real-time interpretation tool for live calls. He's forecast cost for every feature the company has shipped for three years.
3 · THE HABIT
What did Wulfric stop doing because his forecasts kept working, most of the time?
Tap to flip
ANSWER
He stopped building a second, independent check on his numbers, and started trusting a vendor's flat multiplier, written into the template once, as if it were a measured fact.
4 · THE ANCHOR
What's the one design decision the whole forecasting method hangs on?
Tap to flip
ANSWER
Build two independent estimates, top-down from a borrowed adoption curve and bottom-up from raw unit cost, publish the range they make, and set a fixed date to recheck it with real data.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing a vendor's rough rule of thumb, mentioned once on a sales call, into the forecasting template as if it were a measured constant, and never building Voxferry's own number to check it against.
6 · THE NUMBER
Fill in the blank: the live captions forecast said ___ a month. By week six it actually cost ___ a month.
Tap to flip
ANSWER
Eighteen hundred dollars forecast, sixty four hundred actual, after one customer turned captions on for all four hundred of their seats at once.
7 · THE REPLAY
Same kind of near miss, new forecasting method, what changes?
Tap to flip
ANSWER
The interpretation feature's range, $4,700 to $7,500 a month, holds through a customer flipping on hundreds of seats at once. Day 14's real number lands around $4,100, comfortably inside it, no emergency review.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the shared anchor?
Tap to flip
ANSWER
OrderVale, a drive-thru voice ordering assistant at Crispline. Same anchor: two independent estimates, a published range, and a fixed recheck date, this time against a mobile app's curve instead of a chat feature's.

Check yourself Score: 0 / 0

Multiple choice
1. A feature has never shipped, so there's no usage history at all. What should the cost forecast look like?
  • A. A single point number, since giving a range would look like you don't know your own product.
  • B. A range built from two independent estimates, top-down and bottom-up, with a fixed date to recheck it against real data.
  • C. Whatever number the vendor's pricing page quotes, since they built the pipeline and know it best.
  • D. No forecast at all until real usage exists, since any number before that is just a guess.
Show hint
Think about what let Wulfric catch a wrong assumption in two weeks instead of a quarter.
Show answer
B. Two independent paths that have to agree catch a bad assumption before launch. A single number, even a vendor's, has nothing to check itself against, and no forecast at all leaves finance with nothing to plan around.
Short answer, name the rejected alternative
2. What alternative did Wulfric consider for forecasting the interpretation feature's cost, and why did it lose?
Show hint
Look at the paragraph right after the K step in the framework recap, where the three closing points are stated directly.
Show answer
Model answer: Trusting the vendor's flat multiplier alone, with no independent bottom-up build to check it against, exactly what caused the earlier captions miss. It lost because a number nobody can trace to a real measurement isn't a forecast, it's a rumor with a decimal point.
True or false
3. True or false: because the top-down and bottom-up estimates for the interpretation feature landed close together, Wulfric skipped the day-14 recheck.
  • True
  • False
Show hint
Check what the anchor step says about the recheck date, and whether it's conditional on the two estimates agreeing.
Show answer
False. The recheck at first real usage always happens, even when the two estimates agree. Agreement lowers the risk, it doesn't prove the guess was right, only real data does that.
Fill in the blank
4. The interpretation feature's published range was $___ to $___ a month, and the real day-14 number came in around $___.
Show hint
Look at the chart in Section 1, right after the two estimates are introduced.
Show answer
$4,700 to $7,500, with the real day-14 number landing around $4,100. The real number sat inside the range, near the low end, which is what let finance read one line and move on instead of calling an emergency review.
Short answer, apply it yourself
5. Pick an AI product you use yourself. Name one feature in it that's new enough to have no real usage history, and how you'd forecast its cost.
Show hint
Think of a feature that just shipped or is still in beta, one that bills by usage rather than a flat subscription.
Show answer
Model answer: A photo app just added an AI video-from-photos feature with no usage history yet. Top-down, borrow the adoption curve of its existing AI photo-editing feature, discounted for the extra step of picking multiple photos first. Bottom-up, add up the real render cost per video, seconds of output times the per-second model cost. Publish the range, and recheck it against real renders after the first week live.
Multiple choice
6. Crispline's bottom-up estimate for OrderVale briefly underpriced one pilot store during an eleven-day promotion. What does that reveal?
  • A. Bottom-up estimating doesn't work for physical locations, only for software features.
  • B. A unit cost built from one average week has no way to represent a busy week, so the estimate needs its own promotion-week case, not just an average.
  • C. Crispline should cancel the pilot and go back to human cashiers.
  • D. The top-down estimate was the only one that mattered, so the bottom-up number should be dropped entirely.
Show hint
Read the key point block right under OrderVale's R and K steps, "the decision Rosamund would take back."
Show answer
B. The bottom-up build was based on a normal week's average order, with no version of the number for a busy week, so a real promotion nearly slipped past the range unnoticed.
Before you close the answer
Why this works
Tests whether you know a number with nothing to check it against isn't a forecast, just a guess with confidence attached. Most candidates describe one estimation method. Few build in a way for the method to catch itself being wrong before real money is on the line.
Follow-up traps
"What if you don't have any comparable feature to borrow a curve from at all?" Response: fall back to the bottom-up build alone, but widen the published range further, say P50 to P95 instead of P50 to P90, and move the recheck trigger up to day one of real usage instead of day fourteen, since you've lost your second independent check.

"Isn't publishing a range just a way to avoid committing to a number?" Response: no, because the commitment is the recheck date itself, a fixed point where the range gets replaced by a real number. A hedge with no trigger would be avoiding commitment. This one has a deadline built into it.
If pressed
The interpretation pipeline keeps a rolling ninety-second window of each call as context for the translation model, long enough to resolve a pronoun referring back to something said a minute earlier, short enough to keep re-processing cost and lag small. A longer window would translate more accurately but cost more per minute and add a delay two people on a live call would actually feel.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more