CaseAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #17
How do you use a spike to de-risk a cost assumption rather than a quality assumption?
BOUNDa chatbot that answered every question correctly, and still went over budget
Cue Row is a ticketing platform for regional concerts and venues. Devraj Anand is the AI PM piloting a customer-support chatbot, and it passed every quality check the team threw at it. The number nobody had spiked yet was what it would actually cost to run.
The direct answer
Run the cost spike against real customer conversations, not clean synthetic ones, and measure cost per resolved conversation end to end, including the long, messy ones with pasted screenshots and repeated questions. Compare that number to the support cost it's meant to replace, before ever asking whether the answers are good enough. A model that answers perfectly at ten times the support cost it was supposed to save is not a win, no matter how good the answers are.
Do this, in order
Break the cost equation into its real parts before running any numbers.Why: input tokens, output tokens, retries, and human escalation each behave differently, and blending them hides which one actually drives cost.
Sample real conversation length and messiness, not a curated demo set.Why: a spike built on short, clean test chats will always under-price what real customers actually send.
Give a range for the cost estimate, not one confident number.Why: conversation length varies a lot, and a single number hides how wide that range really is.
Sanity-check the estimate against the support cost it's replacing.Why: a cost per conversation higher than a human agent's cost per ticket means something in the estimate is wrong or the plan needs to change.
Name which single assumption would swing the estimate the most if it's wrong.Why: that's the one worth re-checking before committing budget, not the one that happens to be easiest to measure.
How to answer this, stage by stage
Nobody is scoring whether you can define "cost per token." They're scoring whether you'd have caught the gap between a clean cost estimate and what real, messy customers actually do to a bill.
Stage 1
Scope it to one real number
Say it like this
"Let's ground this in Cue Row's support chatbot, where the quality checks all passed and the cost checks hadn't really been run yet."
Why this works
Keeps the answer from turning into a lecture about cost management in general, with nothing specific behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as BOUND: break it down, own the numbers, use a range, nail the sanity check, name the direction it could swing. This is an estimation question, not a story question, so I want the arithmetic visible."
Why this works
Naming BOUND up front signals you know this question doesn't want a narrative frame bolted onto it.
Stage 3
Break down the equation before touching a single number
Say it like this
"Cost per conversation equals input tokens times input price, plus output tokens times output price, plus the cost of retries, plus whatever fraction still escalates to a human agent."
Why this works
Saying the equation out loud before any numbers is what separates BOUND from a guess with a confident tone.
Stage 4
Own the assumptions behind each part
Say it like this
"For the spike, I'd assume a real conversation runs 8 to 20 turns, not the 3-turn demo script, because that's what actual ticket refund and reschedule questions take once someone's frustrated and pasting a confirmation email."
Why this works
Naming where each assumption comes from is what makes the estimate defensible instead of a number pulled from nowhere.
Stage 5
Give the range, then the sanity check
Say it like this
"Our clean spike sample estimated 6 cents per conversation. Against real week-three traffic, the honest range was 40 to 90 cents. That's higher than the 35 cents per ticket a human agent costs us today, which means something in the plan has to change before this scales."
Why this works
This is the strongest line BOUND can produce: comparing the estimate against something known, and catching that it doesn't survive the comparison.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't just a normal budgeting exercise is that a model's cost scales with how much a customer chooses to type and paste, unlike a fixed-price software seat. We accepted a wider, messier spike sample, real traffic instead of a clean script, in exchange for a cost number we could actually trust before scaling to every venue."
Why this works
This names the load-bearing, AI-specific judgment: token-based cost is a variable, usage-driven number, not a flat per-seat license fee.
Stage 7
Say what moves the estimate most, then close
Say it like this
"Conversation length is the single biggest lever, more than the model's own per-token price. So: break down the equation, own real conversation length as the assumption, range it honestly, sanity-check against the human cost it replaces, and watch length specifically before scaling."
Why this works
Naming the one assumption that swings the estimate most is what a good estimator says and a rushed one skips.
Let's learn
Picture a chat window on Cue Row's help page. A customer types "my show got rescheduled, can I get a refund," and a model reads the ticket, checks the policy, and answers in a few seconds.
Before the chatbot, a Cue Row support agent handled that same question by hand, roughly 35 cents in labor cost per resolved ticket, averaged across a typical shift of quick and slow tickets alike. The chatbot's early spike, tested on fifty short, clean sample conversations the team had written themselves, came back at about 6 cents per conversation.
Four parts that make up the real number. The spike sample only really tested the first two.
Here's the turn: the quality of the chatbot's answers was never really in question, it passed every accuracy check the team ran. The real doubt was always the cost assumption, and the fifty curated sample conversations the spike used were nothing like what actual customers, frustrated and pasting confirmation emails, were about to send.
Cost per conversation: build-up, spike estimate versus real week-3 traffic
Same model, same questions, thirteen times the real cost, once escalation and retries from actually messy conversations entered the picture.
At its worst, the chatbot rolls out to every venue on the platform, the invoice arrives at the end of the month at a number nobody had modeled, and the "cost-saving" AI feature turns out to cost more per resolved ticket than the human agents it was meant to replace.
A model can answer every question correctly and still lose the business case, if nobody spiked what a real customer's actual typing habits do to the bill.
The choice I would take back
When the pilot launched, the team kept no ongoing record of cost per conversation, since the early spike number had looked settled and nobody expected it to move. That made sense when the sample was small and clean. It stopped making sense the moment real traffic, with its longer, messier conversations, started quietly costing far more than that first number, and nobody had a trend to notice it drifting.
What I would leave alone: I wouldn't build a full cost-monitoring dashboard for a low-volume internal tool used by five employees. The stakes of an undetected cost creep there are tiny compared to a feature facing every customer on the platform.
The lesson: the chatbot was never expensive because it was bad. It was expensive because nobody had asked what a real, messy conversation costs, only what a clean one does.
Now here is the same thing as a story
The short version above is what you'd say defending the budget in a five-minute finance review. Read this one for what it felt like the week a new hire asked the one question nobody else had thought to ask.
Before Cue Row's chatbot, a support agent handled tickets from a shared tablet at the help desk, three people rotating through it across a shift, each one clearing refund and reschedule questions by hand.
Devraj had watched the chatbot pass every quality gate the team threw at it: accurate refund eligibility checks, correct show-reschedule policy answers, a tone reviewers liked. The rollout went smoothly across the first two weeks. Nobody was worried.
The spike and real traffic were never testing the same thing, even though everyone assumed they were.
Then, in the third week, a new hire on the finance side, reviewing the month's cloud spend for the first time, asked a plain question in a shared channel: "why did the chatbot line item triple since week one?" Nobody could answer her right away. That was the actual trigger, not a dashboard alert, not an outage, just a question nobody senior enough had thought to ask sooner.
Knowledge spark: why does a longer conversation cost so much more, not just a little more?
A model is charged by how much text it reads and writes, both what the customer sends and what it generates back. A short, clean test question might be 40 words. A real frustrated customer pasting a confirmation email and repeating themselves twice can easily be ten times that, and every one of those extra words gets billed, on both sides of the conversation.
Devraj pulled the real week-three logs and found the gap immediately: real customers were pasting entire confirmation emails, asking the same question two different ways when the first answer felt uncertain, and about 1 in 5 conversations were escalating to a human anyway after several model turns, meaning the platform was now paying for both the model and the agent on those tickets.
Three honest reasons a clean spike sample will always look cheaper than the real thing.
The real question was never whether the model's answers were good. It was whether anyone had ever tested what a genuinely messy, ten-turn, screenshot-heavy conversation actually cost, and nobody had, because the spike sample had been written by the team itself, calm and tidy, nothing like a real Tuesday afternoon at the help desk.
One question that would have caught the gap in week one instead of week three.
The fix wasn't a cheaper model. It was a length cap with a graceful handoff: after six turns without resolution, the chatbot summarized the conversation and handed it to a human agent immediately, instead of continuing to burn tokens on a conversation already heading toward escalation anyway.
Four weeks, and the fix only happened because someone finally asked the plain question out loud.
Re-measured after the fix, real cost per resolved conversation settled at about 31 cents, just under the 35-cent human baseline, and the finance channel got a standing weekly number instead of a one-time spike result nobody ever looked at again.
What I'd tell myself, hearing that plain question land in the channel: the number was never hidden on purpose. It just never got asked a second time after the first, clean answer looked fine.
BOUND, the estimate that survives the invoiceNot a script for distrusting every cost estimate forever. BOUND is what tells you exactly which part of a clean number a real customer will blow past.
B
Break it down. State the equation out loud.
Cost per conversation equals input tokens times input price, plus output tokens times output price, plus retries, plus the fraction that still escalates to a human.
Saying the equation first is what makes every number after it checkable.
O
Own numbers. State each assumption and where it came from.
Assume real conversations run 8 to 20 turns, based on actual week-three logs, not the 3-turn demo script the original spike used.
Naming the source of each assumption is what separates a defensible estimate from a guess.
U
Use a range. Low and high, not false precision.
6 cents on the clean spike sample versus 40 to 90 cents across real week-three traffic, settling near 80 cents before the fix.
A single number here would have hidden exactly how wide the real gap was.
N
Nail the sanity check. Does it survive comparison?
80 cents per conversation is more than double the 35-cent human agent cost it was meant to replace. That comparison is what caught the problem.
This is the step that actually catches the mistake: comparing the estimate to something known, not just trusting the arithmetic.
D
Direction. Which assumption swings it most?
Conversation length, driven by real customer messiness and escalation rate, moved the estimate far more than the model's own per-token price ever did.
Naming this is what a good estimator says, and what tells you where to watch after the spike ends.
The recap, one line per letter: break it down is naming the four real cost components before any number gets attached, own numbers is grounding conversation length in real logs, use a range is 6 cents against 40 to 90 cents, nail the sanity check is comparing against the 35-cent human baseline and finding it fails, and direction is that conversation length, not model price, is the lever that actually moves the number.
And if you want to be sure it really works, try it somewhere elseSame five letters, a bespoke tailoring shop's sizing chatbot instead of a ticketing platform. Different flip family entirely, the same cost trap.
Rosalind Vane runs Copperlatch Tailoring, piloting an AI chatbot that walks customers through self-measurement for made-to-order suits. Mapped onto BOUND: break it down is the same four-part cost equation, input tokens, output tokens, retries, human tailor escalation. Own numbers assumes real customers need 10 to 15 back-and-forth turns to get a shoulder measurement right, not the 4-turn ideal script used in the original spike. The flip here differs from Cue Row's over-trust flip: it's a scope flip, once cost per session started to sting, Rosalind's team quietly capped the chatbot's fitting session to fewer measurement points, undercutting the whole point of a full self-measurement tool without ever admitting cost was the real reason. Use a range shows the same wide gap between clean and real estimates. Nail the sanity check compares against the cost of a video call with a human tailor, the service this was meant to reduce. Direction is once again conversation length, this time driven by how many times a nervous first-time customer re-measures before trusting their own number.
The same four cost components, in a completely different business: the shape of the trap doesn't change with the industry.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "spike cost against real, messy conversations, and sanity-check it against what it's replacing," and stop.
Cost: there's no real traffic yet to sample from. Say so honestly, and pull the closest proxy, real historical support tickets read aloud into a conversation format, rather than trusting a team-written script.
The model gets cheaper per token, for real: even a genuine price drop is worth re-spiking against, since real conversation length can still grow faster than the per-token savings shrink the total.
Where people run it wrong.
They test cost on a small, clean sample the team wrote themselves, instead of real customer conversations.
They check quality carefully and treat cost as an afterthought, since a good answer feels like the whole win.
They stop watching cost once the first spike number looks settled, instead of keeping a running number as real traffic grows.
How to use it live. The moment an interviewer asks you to de-risk a cost assumption, buy two seconds by asking yourself: what would a real, frustrated customer actually type here, and is that anything like the sample the estimate was built on? That gap is usually the whole answer.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip: once the small, clean spike sample looked fine, the team stopped monitoring real cost per conversation and let traffic scale without a second check.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Devraj Anand, the AI PM at Cue Row, piloting a customer-support chatbot whose quality passed every check while its cost quietly tripled.
3 · THE HABIT
What did the team stop doing after the first spike looked settled?
Tap to flip
ANSWER
Watching cost per conversation. The early number looked fine, so nobody kept a running check as real, messier traffic scaled up.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Trusting a clean spike estimate completely versus tracking real cost continuously as traffic scales. No middle setting once monitoring stopped.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Keeping no ongoing record of cost per conversation after launch, since the first clean spike number looked settled and nobody expected it to move.
6 · THE NUMBER
Fill in the blank: the spike estimate was ___ cents per conversation, but real week-3 traffic cost about ___ cents.
Tap to flip
ANSWER
6 cents on the clean spike sample, about 80 cents on real week-3 traffic, against a 35-cent human agent baseline.
7 · THE REPLAY
Same real traffic, a six-turn length cap and handoff already shipped. What changes?
Tap to flip
ANSWER
Cost per resolved conversation settles at about 31 cents, just under the 35-cent human baseline, and finance gets a standing weekly cost number instead of a one-time spike result.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Copperlatch Tailoring's self-measurement chatbot. The flip there is a scope flip: the team quietly cut the number of measurement points to control cost, undercutting the tool's whole purpose.
Check yourself Score: 0 / 0
Multiple choice
1. Why did the clean 50-conversation spike sample estimate cost so far below the real number?
A. The model's price per token changed between the spike and real launch.
B. It used short, curated test conversations with no repeated questions or pasted context, unlike real, messy customer chats.
C. Cue Row switched providers after the spike.
D. The support agents started handling fewer tickets.
Show hint
Look at the comparison between the spike sample and real week-3 traffic.
Show answer
B. The team-written sample never included the length, repetition, and escalation that real customers actually produce.
True or false
2. True or false: the chatbot's answers were inaccurate, which is what drove the real cost overrun.
True
False
Show hint
Look at what quality checks found versus what the cost audit found.
Show answer
False. The chatbot passed every quality check. The cost overrun came from real conversation length and escalation, not from wrong answers.
Fill in the blank
3. Fill in the blank: real week-3 cost per conversation was more than double the ___ cent human agent baseline it was meant to replace.
Show hint
Look at the sanity check step, N, in the BOUND recap.
Show answer
35 cents. At roughly 80 cents per conversation, the chatbot was costing more than double what it was supposed to save.
Short answer, where it wouldn't matter
4. Name a situation on this same platform where this kind of continuous cost monitoring wouldn't be worth building.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A low-volume internal tool used by a handful of employees, where an undetected cost creep is tiny compared to a feature facing every customer on the platform.
Short answer, apply it yourself
5. Think of a subscription or tool you use where the actual cost surprised you once you used it in a real, messy way instead of how a demo showed it. What was different about your real usage?
Show hint
Think about what a demo or trial never shows you about how you'll actually use something.
Show answer
Model answer: A cloud storage plan that looked cheap in a demo with a few sample files, until real daily use, backups and shared folders included, blew past the free tier fast.
Short answer, work the number
6. If the length cap had brought real cost down to 45 cents per conversation instead of 31, would that still have been worth shipping?
Show hint
Think about what the sanity-check comparison value actually is.
Show answer
Model answer: No, not as is. At 45 cents, it's still above the 35-cent human baseline, meaning the fix wasn't enough yet, more escalation-avoidance would be needed before the business case works.
Before you close the answer
Why this works
Tests whether you'll treat cost as a real number to spike with the same rigor as quality, instead of assuming a good answer is automatically a cheap one.
Follow-up traps
"Isn't cost just an engineering or finance concern, not a product one?" Response: no, because the product decision, how long to let a conversation run before handing it to a human, is exactly what controls the cost, and only the product side has the context to make that call well.
"Why not just cap every conversation short from day one?" Response: that would have controlled cost but broken the actual value, some genuinely complex refund cases need the longer conversation; the real fix targets escalation-bound conversations specifically, not all long ones.
If pressed
The six-turn cap wasn't a flat rule, it triggered a handoff only when the model's own confidence in resolving the ticket had dropped below a set bar by turn six, so genuinely close-to-resolved conversations weren't cut off early.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.