Explain how routing between models affects blended cost.
Threadscribe's blended cost per ticket looked steady for a full quarter. A price cut on its expensive model tier was quietly canceling out a routing shift toward that same tier, and nobody had a number built to catch the difference.
- Watch the routing share to the expensive tier every week, on its own, never blended cost alone.Why: blended cost is a weighted average, and a price cut on the expensive tier can cancel a mix shift toward it in the very same number.
- Track each tier's own price separately too, not folded into the blend.Why: without that split you can't tell whether blended cost held flat because nothing changed, or because two real changes happened to cancel each other out.
- Set two real thresholds on the routing share and act at each one.Why: a metric nobody acts on past 25 percent, then past 40 percent, is a dashboard decoration, not a metric.
- Re-check the router's own complexity calls against a fixed set of labeled tickets on a schedule, not once at launch.Why: the small model deciding which tickets are complex enough for the expensive tier had never been checked again, so its line quietly moved as ticket wording changed.
- Reject a flat daily cap on expensive-tier calls per account as the fix.Why: it downgrades genuinely hard tickets right along with easy ones, and it doesn't touch what's actually pushing the share up.
- Gate any swap to a cheaper model within a tier behind an eval on that tier's own hardest tickets before rollout.Why: chasing a lower blended number by quietly downgrading the model itself is exactly how draft quality slips without anyone deciding it should.
How to answer this, stage by stage
Nobody is grading whether you know the word "routing." They're grading whether you can say, in one breath, why a calm average can hide two real changes at once.
Let's learn
Here's what happens when two real changes land in the same stretch and cancel each other out in the one number everyone is watching: nothing. The number just sits there. That's the whole trap.
Threadscribe is the tool Bexford built so a support agent gets a drafted reply the moment a ticket comes in, instead of writing every answer from scratch.
Before Threadscribe, an agent wrote every reply by hand. Anything past a one-line answer took about nine minutes, and a full shift closed out around 38 tickets. With Threadscribe, the same ticket arrives with a draft already sitting there to review and send. About three minutes a ticket, and a full shift closes out closer to 90.
The turn here isn't that Threadscribe's drafts got worse. Agents still accepted or lightly edited most of them, same as always. The real change was underneath: which pile each ticket landed in, Scout's pile or Vault's, had been quietly shifting for months, and the one cost number everyone watched never once showed it.
Here's why the number sat still. Bexford's model provider cut Vault's price by more than half in the same twelve weeks that Vault's own share of tickets climbed from 17 to 42 percent. Blended cost is a weighted average of both. A bigger share, at a smaller price, landed almost exactly where a smaller share, at the bigger price, had been.
At its worst, a cost number that holds still while the real mix keeps moving is worse than no number at all, because it tells finance the product is healthy right up until the quarter it isn't. Once Vault's price cut had been fully spent, and its share kept climbing anyway, blended cost jumped from $0.035 to $0.048 in eight weeks. Across roughly 2.1 million tickets a month, that's about $27,300 a month in AI spend nobody had budgeted for, found in a quarterly finance review, a full quarter after the mix had started moving.
What I'd leave alone: most of Bexford's customers run small, simple ticket volumes where Vault's share barely moves month to month. Building threshold checks and router audits around those accounts would spend engineering time on a group that was never the problem.
The lesson: a number that holds still isn't proof that nothing is changing. It can be proof that two things are changing in opposite directions and canceling out in the one place you're looking. A blended average can only ever tell you about the blend. If you want to know about the parts, you have to watch the parts.
Now here is the same thing as a story
Read the long version below when you want to feel why a canceling number is worse than a rising one, not just be told that it is.
Darragh Pellegrini could smell a pricing problem two weeks before finance could prove one. He'd built Bexford's original cost model himself, back when Threadscribe used one model for every single ticket and there was nothing to route.
Threadscribe had been live for a little over a year when Bexford shipped Vault, a second, bigger model built for the tickets Scout kept getting wrong: multi-item bundles, prorated refunds, anything a customer asked across two or three issues at once. For the first several months, the split held roughly steady. About one ticket in six went to Vault. Blended cost sat near $0.035, and Darragh checked the router's own split by hand most weeks, comparing it against what the dashboard showed.
By month four, the hand check had thinned to once every couple of weeks. By month seven, Darragh mostly just glanced at blended cost on the Monday dashboard and moved on if the number looked the same as last week, which it always did. By month ten, he'd stopped opening the router's own breakdown at all. The blended number had never once given him a reason to.
It came back from a question a new support-ops analyst asked in her second week: why did Vault's queue look busier than the ticket volume seemed to explain? Darragh didn't have an answer. He'd never actually looked at Vault's share on its own, only at what it cost blended into everything else.
He pulled the router logs going back a quarter. Vault's share of all tickets had climbed from 17 to 42 percent over twelve weeks. Blended cost hadn't moved, because Bexford's model provider had cut Vault's own price by more than half in that exact window, a routine update nobody thought to connect to the routing numbers. The two changes had landed on top of each other and canceled almost exactly.
The real cost wasn't in the blended number at all. It was in Corcannon, one of Bexford's larger customers, a home goods retailer that had launched a confusing new bundle-and-subscribe pricing plan around the same time. Corcannon's own Vault share had gone from 22 to 71 percent in those same twelve weeks, driven by real complexity in the tickets their own customers were now sending, not by anything Bexford's router did wrong.
The decision that opened the door went back to Threadscribe's very first pricing review, more than a year earlier, before Vault existed. The team picked blended cost per ticket as the number to watch, because at the time there was only one model and one price, and a blended number and a real number were the same thing. Nobody chose carelessly. It was the right number for the product that existed that day.
Run the same twelve weeks again with one change: Vault's share tracked on its own, every week, next to its own price. By week six, before the provider's price cut had even landed, the share crosses 25 percent, and Darragh's team goes looking, the way they eventually did, but six weeks sooner. Corcannon's launch gets flagged as the real driver behind its own account, instead of sitting lost inside a platform-wide average, and the router gets a stricter confidence bar for the gray-zone tickets before the quarter ends, not after it.
One design let a calm blended number speak for two tiers it was never built to describe together. The other watches the mix and the price on their own, and it would have rung six weeks sooner, before a single provider price cut ever had the chance to hide anything.
What I'd tell myself, back at that first pricing review: the blended number was never wrong, exactly. It just stopped being able to answer the question the day a second tier showed up, and nobody gave that day a name.
LEAD, the four letters behind the $0.048
This isn't a story wearing a metric's clothes. It's a metric question, and LEAD is what stops a calm blended number from standing in for two tiers it was never built to describe together.
Three things worth stating directly, since this is where the real judgment sits. The alternative Darragh's team considered first, and dropped, was a flat daily cap on how many tickets any one account could route to Vault. It lost because it downgrades genuinely hard tickets right along with easy ones, and it doesn't touch what's actually driving the share up, which is real ticket complexity plus a classifier nobody had re-checked. The AI-specific failure worth naming by name is silent classifier drift: the small model deciding which tickets are complex enough for Vault had never been checked again since launch, so its own line quietly moved as customers' wording and product features changed underneath it, with no code change and no decision behind it. The guardrail is a standing re-check of that classifier against a fixed set of labeled tickets on a schedule, plus an alert the moment routing share itself moves outside an expected band. That guardrail isn't free: a stricter confidence bar for what auto-routes to Vault means some genuinely hard tickets get downgraded to Scout first, and first-draft acceptance on those specific tickets drops from about 72 to 58 percent, a real quality-cost trade worth naming, not wishing away. And the bar Threadscribe holds itself to was never zero cost variance across 2.1 million tickets a month; no product with two model tiers can promise that. It's a threshold-specific bar, Vault's share held under a set line for the typical week, checked every week against the real mix, not one blended number standing in for two tiers pretending to be one.
And if you want to be sure it really works, try it somewhere else
Same four letters, an auto-insurance claims-triage tool instead of a support inbox, and this time a storm season drives the tier shift, not a pricing-plan launch.
Adjustly is a claims-triage assistant Quillane Insurance built for its adjusters. An incoming auto claim gets a drafted set of triage notes and next steps before a human ever opens the file. Anezka Osterlund runs product on it.
The build-up: Adjustly splits claims two ways too. Flint, the cheap tier, handles the clean ones: single vehicle, clear photos, an obvious at-fault party. Birch, the expensive tier, handles the rest: multiple vehicles, contested liability, an injury claim needing careful cross-referencing against the policy file. Blended cost per claim draft sat near $0.42 for most of a year.
Birch's share of claims climbed from 12 to 29 percent over the ten weeks after the storm, while blended cost barely moved, sliding from $0.42 to just $0.44, because Birch's own price per claim had been cut under the new contract in the very same stretch.
Same rank as before: track the tier's own share and its own price, on their own, and act before the blended number is forced to move. The fix is the same shape too: a stricter bar for what Adjustly auto-routes to Birch, a flag for the gray-zone claims instead of a guess, and a standing check on the router's own calls against a fixed set of claims nobody's relabeled in months.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: the two things that can cancel inside a blended number are the mix and the price, so track them apart, set a real threshold on the mix, and act on it before the bill does.
Cost: there's no budget this quarter for both a router audit and a pricing renegotiation. The router audit wins, since it's aimed at what's actually driving the mix, not just the number the mix produces.
The model got better, for real: say Vault's accuracy on complex tickets jumps for real. That's not proof the tail shrinks. Harder tickets keep arriving whether or not the model handling them got better at them, so a climbing share is a separate fact from the model's own quality.
Where people run it wrong.
They watch blended cost because it's the number finance already tracks, and never ask whether two tiers are hiding inside it.
They notice the mix moving and "fix" it by quietly swapping the expensive tier for a cheaper substitute model, without checking whether the harder tickets still get a usable draft.
They wait for the blended number to move before acting, when a price change on either tier can hold it still for months after the real shift already started.
How to use it live. Say the mechanism out loud before naming a metric: "blended cost is a weighted average, so it can hide a mix shift and a price shift at the same time." That buys a beat to think instead of reciting whatever the dashboard already shows.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just check blended cost weekly instead of monthly?" Response: because a price cut and a mix shift can land in the same week just as easily as the same quarter. Frequency doesn't fix a number that's built to hide the thing you need to see, decomposing it does.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Cost modeling and unit economics
- #1 Build the cost-per-interaction model for a feature with a 2,000-token prompt and a 500-token response.
- #2 What cost drivers exist for an AI feature beyond model tokens?
- #3 Explain how a RAG pipeline's cost structure differs from a single model call.
- #4 How does prompt caching change your unit economics, and when does it not help?
- #5 Model the monthly cost of a feature used by 50,000 users averaging 12 interactions each.
- #6 What is the cost impact of moving from a single call to a five-step agent?