Describe the hidden cost of context window growth over a long conversation.
A long negotiation does not cost twice as much because it runs twice as long. It costs close to four times as much, because every new turn also pays to resend everything said before it, and the average across every conversation hides that completely, right up until one deal's bill is already enormous.
- Watch cost per open session live, not the blended monthly average across every session.Why: the average is dragged around by thousands of short, cheap sessions, so a small, growing slice of long ones can double in cost without moving it much.
- Put a real alert on any single session that crosses a clear cost multiple, one that pages someone while the conversation is still open.Why: a number nobody gets paged on is the same as a number nobody is watching, and by the time a long session closes, the bill is already spent.
- Before trusting a monthly sample, check what share of sessions actually run long.Why: if long sessions are a small slice of the total, a random sample can miss them for months on end.
- Scope the cost check to the single session, never to the account or the customer as a whole.Why: folding one marathon conversation into a customer's whole month of normal usage dilutes it right back into looking fine.
- Do not fix this by shrinking every session's memory by the same amount.Why: that saves nothing on the short sessions that were never the problem, and quietly makes the long, high-stakes ones worse.
- Recheck the alert threshold whenever the underlying model's price changes.Why: a cheaper model changes the dollar number, not the shape of the curve, so the old threshold stops meaning what it used to mean.
How to answer this, stage by stage
Nobody is grading whether you know what a context window is. They are grading whether you know its cost is a curve that bends upward, not a straight line, and whether you would watch the one number that shows the bend early.
Let's learn
Here's what happens when a cost that looks fine on paper only looks fine because of how it's being averaged.
Parlay is the negotiation and deal coaching chat inside Holmcrest, a sales software company. A rep keeps Parlay open during a live call, and when a customer pushes back on price, the rep types a quick question and Parlay answers, using everything said in the negotiation so far, so its advice never contradicts something the customer already heard.
Before Parlay, reps re-read old email threads and call notes before every follow-up call, and still forgot which number they'd already offered about one time in six. Parlay ended that. It answers in seconds and never forgets a single thing the customer has said.
For most calls, this costs almost nothing. A typical negotiation runs about eighteen back-and-forth exchanges and costs Holmcrest about eighteen cents to run start to finish. But Parlay's real selling point is that it remembers a deal across calls, not just within one. A big renewal, with legal review, procurement, and a change of mind or two, can stretch across eleven separate calls over three weeks, all stitched into one continuous conversation so Parlay never has to be re-briefed.
That's the part nobody outside engineering was watching. A single renewal, stitched across eleven calls into three hundred and forty turns, ends up costing about nine dollars and sixty cents by the time it closes, more than fifty typical calls put together, and the price does not climb evenly. It climbs slowly for a while, then fast, then very fast.
At its worst, this quietly eats the margin on exactly the accounts worth protecting most. Holmcrest prices Parlay as a flat monthly fee per rep. Every enterprise account whose renewals run long, slow, and lawyer-heavy costs far more to run through Parlay than the flat fee ever priced in, and it's precisely the biggest, stickiest customers who negotiate that way.
In the review Baashir Vantreight, the revenue operations lead who was pulling Holmcrest's ten most expensive Parlay sessions for a gross-margin slide, found the one renewal deal sitting at nine dollars sixty cents, more than the entire month's Parlay bill for fifty three ordinary customers combined.
What I'd leave alone: ordinary single-call sessions, ninety eight percent of everything Parlay runs. Almost none of them cross even forty cents, so a live cost meter there catches nothing and just adds noise.
The lesson: a growing context window doesn't raise the price a little at a time. It resends more of the past on every single turn, so the price bends upward, and a monthly average blends that bend away until one conversation's bill is already large enough to show up on a board slide.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why Vivika's twenty five session sample kept coming back clean for seven months while the real number was already climbing.
Vivika Oxendine could smell a pricing model about to go wrong before the numbers proved it. She'd built cost forecasts for two earlier Holmcrest products before Parlay ever shipped, and she was the one who set Parlay's launch-day budget: thirty cents a session, comfortable, with room to spare.
For the first several months, every check she ran agreed with her. On the first Monday of the month, she pulled a random sample of twenty five Parlay sessions out of the thousands that had run, checked their cost against the model, and moved on. Eighteen cents, twenty two cents, nineteen cents. Nothing near the budget line. It took her about two hours a month, and every single time, the sample said Parlay was fine.
By month four, she'd started trusting the sample enough to skip the write-up and just glance at the total. By month six, she was pulling the sample mostly out of habit, because it had never once turned up anything worth a second look.
What Vivika's sample could never see: Parlay's marathon sessions, the deals that ran across several calls instead of one, were still under two percent of everything Parlay ran. Pull twenty five sessions at random out of a few thousand, and there's a real chance you land on zero of them, month after month, purely by luck of the draw. And even when her sample did catch one, it usually caught it early, on call two or three, back when the running cost still looked like any other session's.
The renewal that finally surfaced had nothing dramatic about it. A returning logistics customer, slow procurement, a legal team that wanted three separate rounds of redlines. Eleven calls, three weeks, three hundred and forty turns in one Parlay deal file, because that continuity was the whole point, the customer never had to repeat a number they'd already given. By the last call, the running conversation was so long that a single reply from Parlay cost more than an entire ordinary negotiation, start to finish.
Eight months earlier, in the meeting where Parlay's cost reporting first got designed, someone asked whether to build a live counter on every open session or just total everything up once a month. A live counter meant a new system to build and maintain. The monthly total was nearly free, since the billing data already existed for invoicing. The monthly total won, reasonably, because back then almost every session finished the same day it started, so a monthly average was a faithful enough picture of what any one session actually cost.
Vivika didn't catch the renewal deal. Baashir Vantreight did, three months after it closed, while pulling Holmcrest's ten most expensive Parlay sessions for a gross-margin slide ahead of the board meeting. He found one session sitting at nine dollars sixty cents and asked Vivika why a single conversation had cost as much as fifty three ordinary ones.
Run the same Tuesday again, with one change: a live running-cost meter on every session that carries past its first call, set to alert the moment one session's running cost crosses two dollars, about ten times a typical session. The same renewal ships, the same eleven calls happen, and the alert fires on call four, nine days in, at two dollars and five cents, while the deal is still open and there's still time to look at what's driving it. Twenty one days of quiet compounding becomes nine days of a page someone actually sees.
One design waited for the invoice to add everything up once a month. The other watches the one number that was climbing the whole time, on the one conversation where it mattered.
What I would tell myself, back in that first reporting meeting: the moment a system remembers a whole conversation and resends it every turn, ask what a sample can and can't see, because a rare, slow-building cost hides perfectly inside a random sample and a monthly total, right up until someone pulls the extremes on purpose. Nobody asked that question in the room. That's on the room, not on Vivika.
FLIPS, and the one letter that has no middle setting
Not five guesses about why a bill got big. FLIPS names the one habit that snapped, and asks which old choice made snapping the only option.
Three things worth stating directly, since this is where the real judgment sits. The alternative Holmcrest could have tried first was capping every Parlay session at a fixed number of turns, say sixty, and forcing a fresh session past that point. It loses because Parlay's whole value is remembering the deal; a rep re-explaining a customer's own objections back to Parlay defeats the reason reps trust it on their hardest negotiations, so this trades away quality to chase a cost problem that visibility alone can fix. The AI-specific failure worth naming by name is silent cost compounding from an unbounded context window: nothing about Parlay's advice to the rep ever looked wrong, so the conversation-level view stayed clean while the resend cost underneath it kept climbing, invisible to anyone reading only a blended average. The guardrail is the live per-session alert itself, plus a rule that any session crossing its second call gets the meter attached automatically, not opted in by someone remembering to flip a switch. That guardrail is not free: building and maintaining a live per-session counter costs real engineering time that the old monthly batch job never needed, a trade accepted on purpose, because nine dollars sixty cents discovered three months late costs Holmcrest far more in surprised board slides than the counter ever will. And the bar it enforces was never zero growth. A deal that runs long is allowed to cost more; that's expected. It's a probability bar, checked against a real jump: the alert pages when a single session's running cost clears ten times a typical session, not a promise that no conversation is ever allowed to grow.
And if you want to be sure it really works, try it somewhere else
Same five letters, an industry that has never coached a single sales call, and this time the cost isn't hidden in a finance report. It's hidden in which trucks a dispatcher decides to ask about.
Ravensdale is a regional freight company. Wayfarer is the AI dispatch copilot its overnight dispatchers keep open in one running chat for the whole shift, coordinating dozens of trucks as delays, breakdowns, and reroutes come in. Ferdinanda Coldbrook is the dispatch systems engineer who watches how dispatchers actually use it.
The case for trusting it as built: for routine adjustments, a truck running twenty minutes late, a delivery window that needs pushing, Wayfarer cut a dispatcher's decision time from about six minutes to under one. Cheap, fast, and by shift hour ten, the running chat already carries eight hours of the night's decisions behind it.
The case against it: about one shift in five includes a genuinely hard reroute, a highway closure forcing four trucks to be rebalanced across three routes at once. Those queries are the longest, because they need the whole shift's context reasoned over together, and by hour ten that context is enormous, so they're also the most expensive single queries of the night by a wide margin.
That's exactly backwards from where the real cost sat. Dispatchers started skipping Wayfarer on routine swaps, since those were fast enough to just do from memory, and saving it for the hard multi-truck reroutes instead, rationing their query count toward the cases that were already the longest, priciest, and most likely to get a rushed read once the answer finally came back. Wayfarer never got worse. The mix of what it was being asked simply shifted toward its most expensive, highest-pressure moments, and errors started concentrating there too.
Ferdinanda's fix wasn't a smarter model. It was the same kind of decision as Parlay's: replace the flat per-query cap with a cost-weighted budget, so a routine check barely dents a depot's monthly allowance and a deep reroute counts for what it actually costs. Dispatchers went back to asking about routine swaps too, because the count no longer punished them for it.
Same rank as before, different family: know what's actually driving the expensive cases before you cap usage by a number that doesn't reflect the real cost, or you'll ration people straight into the exact cases where a wrong answer costs the most.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix: watch cost per open session live, never the blended monthly average alone.
Cost: there's no budget this quarter to build a live meter. Hand-pull the ten longest-running open sessions each week and eyeball their running turn count against a typical session, a poor man's version of the same idea.
The model got better, for real: say the underlying model gets cheaper per token. The per-turn price drops, but the shape doesn't change. A session with ten times the turns still resends ten times as much history on every one of them, just at a lower headline number, so the blended average still hides the same bend, and the alert threshold needs recalibrating down, not removing.
Where people run it wrong.
They watch the blended monthly average as the whole picture, and a small rise in it gets waved off as noise instead of investigated as a mix shift.
They scope a cost alert to the account or the customer instead of the single open session, so one marathon deal still gets diluted across a whole month of that customer's ordinary usage before anything trips.
They "fix" it by shortening every session's memory by the same amount, saving nothing on the ninety eight percent that never needed it, and quietly hurting quality on the long deals where the full history mattered most.
How to use it live. Say the split out loud before answering: "is this cost climbing because volume is going up, or because individual conversations are getting longer." Naming that split buys a beat to work out which one is actually happening, instead of guessing out loud in front of the interviewer.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a 34 cent average still cheap? Why build a whole alert system for that?" Response: the average was never the risk. The risk was the tail underneath it, one session at $9.60, and the fact that nothing would have caught the next one until it was just as far along.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Cost modeling and unit economics
- #1 Build the cost-per-interaction model for a feature with a 2,000-token prompt and a 500-token response.
- #2 What cost drivers exist for an AI feature beyond model tokens?
- #3 Explain how a RAG pipeline's cost structure differs from a single model call.
- #4 How does prompt caching change your unit economics, and when does it not help?
- #5 Model the monthly cost of a feature used by 50,000 users averaging 12 interactions each.
- #6 What is the cost impact of moving from a single call to a five-step agent?