CalculationAdvancedQuality, Cost & Token Economics / Cost modeling and unit economics / #15

Explain how routing between models affects blended cost.

LEAD · cost modeling and unit economics

Threadscribe's blended cost per ticket looked steady for a full quarter. A price cut on its expensive model tier was quietly canceling out a routing shift toward that same tier, and nobody had a number built to catch the difference.

The direct answer
Blended cost per ticket is a weighted average: the share of tickets going to the cheap tier times what that tier costs, plus the share going to the expensive tier times what that one costs. Move more tickets into the expensive tier and blended cost climbs, even with no price change anywhere, and a price cut on that same tier can cancel a mix shift out completely. At Bexford, blended cost held near $0.035 a ticket for twelve straight weeks while the expensive tier's share climbed from 17 to 42 percent, because a price cut on that tier landed in the same window and hid the whole shift. Track the routing share to the expensive tier, and each tier's own price, every week, on their own, not folded into one blended number, and set thresholds on the share itself.
Do this, in order
  1. Watch the routing share to the expensive tier every week, on its own, never blended cost alone.Why: blended cost is a weighted average, and a price cut on the expensive tier can cancel a mix shift toward it in the very same number.
  2. Track each tier's own price separately too, not folded into the blend.Why: without that split you can't tell whether blended cost held flat because nothing changed, or because two real changes happened to cancel each other out.
  3. Set two real thresholds on the routing share and act at each one.Why: a metric nobody acts on past 25 percent, then past 40 percent, is a dashboard decoration, not a metric.
  4. Re-check the router's own complexity calls against a fixed set of labeled tickets on a schedule, not once at launch.Why: the small model deciding which tickets are complex enough for the expensive tier had never been checked again, so its line quietly moved as ticket wording changed.
  5. Reject a flat daily cap on expensive-tier calls per account as the fix.Why: it downgrades genuinely hard tickets right along with easy ones, and it doesn't touch what's actually pushing the share up.
  6. Gate any swap to a cheaper model within a tier behind an eval on that tier's own hardest tickets before rollout.Why: chasing a lower blended number by quietly downgrading the model itself is exactly how draft quality slips without anyone deciding it should.

How to answer this, stage by stage

Nobody is grading whether you know the word "routing." They're grading whether you can say, in one breath, why a calm average can hide two real changes at once.

1
Scope it to one concrete product before answering in the abstract
Say it like this
"Let's ground this in one product. Threadscribe is a tool from Bexford. It reads an incoming support ticket and drafts a reply an agent can edit and send. Darragh Pellegrini owns cost and margin on it, across roughly two hundred and twenty customer accounts."
Why this works
An abstract "how does routing affect cost" question turns into a shrug fast. One product turns it into a number problem with real edges.
2
Say your structure out loud before touching a number
Say it like this
"I'm going to explain the mechanism first, in plain terms, then say what number actually catches a problem early, how that number gets gamed, and what I'd do at real thresholds."
Why this works
Tells the interviewer you have a method, not a guess, before you've said a single figure.
3
Answer the actual mechanism, before naming a metric
Say it like this
"Blended cost per ticket is a weighted average. Take the share of tickets going to the cheap tier, times what that tier costs. Add the share going to the expensive tier, times what that one costs. Move tickets toward the expensive tier and the average climbs, even with no price change anywhere. Cut the expensive tier's price at the same time, and the two can cancel out, so the average never moves at all."
Why this works
This is the actual question, answered directly. Everything after this is proof.
4
Give the one decision
Say it like this
"So I'd track the routing share to the expensive tier every week, and each tier's own price every week, on their own. Not folded into one blended number."
Why this works
This is the answer to the question. It's the thing a candidate can say and then defend.
5
Prove it with the failure, compressed
Say it like this
"At Bexford, blended cost sat at $0.035 a ticket for twelve weeks straight. Underneath it, the expensive tier's share climbed from 17 to 42 percent, because a price cut on that tier landed the same quarter and canceled the shift out in the total. Once the price cut was fully absorbed, blended cost jumped to $0.048 within two months, and nobody saw it coming."
Why this works
A real number with a real timeframe does more work than any adjective.
6
Say what you'd act on, at real thresholds
Say it like this
"Past 25 percent share, sustained two weeks, I'd go find out if it's a real jump in ticket difficulty or the router quietly drifting. Past 40 percent, I'd cap what the router can send automatically, downgrade the gray-zone tickets to a cheaper draft with a flag, and put the account driving it in front of a pricing conversation."
Why this works
A metric nobody acts on is decoration. Two thresholds turn it into a decision.
7
Close on the decision, not the arithmetic
Say it like this
"So: track the mix and the price separately, never trust one blended number to describe two tiers pretending to be one, and act on the mix before the bill forces you to."
Why this works
Ending on the rule, not the last number crunched, is what makes this sound like judgment instead of a spreadsheet read aloud.

Let's learn

Here's what happens when two real changes land in the same stretch and cancel each other out in the one number everyone is watching: nothing. The number just sits there. That's the whole trap.

Threadscribe is the tool Bexford built so a support agent gets a drafted reply the moment a ticket comes in, instead of writing every answer from scratch.

Before Threadscribe, an agent wrote every reply by hand. Anything past a one-line answer took about nine minutes, and a full shift closed out around 38 tickets. With Threadscribe, the same ticket arrives with a draft already sitting there to review and send. About three minutes a ticket, and a full shift closes out closer to 90.

Knowledge spark: what's a model tier? Threadscribe doesn't use one model for everything. Simple tickets, like a password reset or an order-status question, go to Scout, a small, cheap model. Harder tickets, like a multi-item billing dispute, go to Vault, a bigger, slower, more careful model that costs a lot more per ticket.

The turn here isn't that Threadscribe's drafts got worse. Agents still accepted or lightly edited most of them, same as always. The real change was underneath: which pile each ticket landed in, Scout's pile or Vault's, had been quietly shifting for months, and the one cost number everyone watched never once showed it.

The leading edge: Vault's share of all tickets, this half-year
20% 40% 60% 0 Vault price cut lands 17% 42% 64% Wk 0 Wk 12 Wk 24
Vault's share of all tickets platform-wide, tracked weekly. It climbed from 17 to 42 percent by week twelve and kept going, reaching 64 percent by week twenty-four, well before finance had any reason to look.
We didn't lose money on any single ticket. We lost track of which pile the tickets were landing in.
The lagging outcome: blended cost per ticket, platform-wide
$0.06 $0.03 0 $0.035 Wk 0 $0.035 Wk 12 $0.048 Wk 24
Held flatJumped, once the price cut was spent
Blended cost per ticket, checked monthly. It sat completely still for a full quarter while Vault's share nearly doubled underneath it, then jumped 37 percent in the eight weeks after the price cut had nothing left to absorb.

Here's why the number sat still. Bexford's model provider cut Vault's price by more than half in the same twelve weeks that Vault's own share of tickets climbed from 17 to 42 percent. Blended cost is a weighted average of both. A bigger share, at a smaller price, landed almost exactly where a smaller share, at the bigger price, had been.

The choice that mattered Bexford picked blended cost per ticket as the number to watch back when Threadscribe used a single model for every ticket. That made sense then, there was no mix to hide anything inside. It stopped making sense the week Vault shipped as a second tier, and nobody swapped the metric out.

At its worst, a cost number that holds still while the real mix keeps moving is worse than no number at all, because it tells finance the product is healthy right up until the quarter it isn't. Once Vault's price cut had been fully spent, and its share kept climbing anyway, blended cost jumped from $0.035 to $0.048 in eight weeks. Across roughly 2.1 million tickets a month, that's about $27,300 a month in AI spend nobody had budgeted for, found in a quarterly finance review, a full quarter after the mix had started moving.

What I'd leave alone: most of Bexford's customers run small, simple ticket volumes where Vault's share barely moves month to month. Building threshold checks and router audits around those accounts would spend engineering time on a group that was never the problem.

The lesson: a number that holds still isn't proof that nothing is changing. It can be proof that two things are changing in opposite directions and canceling out in the one place you're looking. A blended average can only ever tell you about the blend. If you want to know about the parts, you have to watch the parts.

Now here is the same thing as a story

Read the long version below when you want to feel why a canceling number is worse than a rising one, not just be told that it is.

Darragh Pellegrini could smell a pricing problem two weeks before finance could prove one. He'd built Bexford's original cost model himself, back when Threadscribe used one model for every single ticket and there was nothing to route.

Threadscribe had been live for a little over a year when Bexford shipped Vault, a second, bigger model built for the tickets Scout kept getting wrong: multi-item bundles, prorated refunds, anything a customer asked across two or three issues at once. For the first several months, the split held roughly steady. About one ticket in six went to Vault. Blended cost sat near $0.035, and Darragh checked the router's own split by hand most weeks, comparing it against what the dashboard showed.

By month four, the hand check had thinned to once every couple of weeks. By month seven, Darragh mostly just glanced at blended cost on the Monday dashboard and moved on if the number looked the same as last week, which it always did. By month ten, he'd stopped opening the router's own breakdown at all. The blended number had never once given him a reason to.

It came back from a question a new support-ops analyst asked in her second week: why did Vault's queue look busier than the ticket volume seemed to explain? Darragh didn't have an answer. He'd never actually looked at Vault's share on its own, only at what it cost blended into everything else.

He pulled the router logs going back a quarter. Vault's share of all tickets had climbed from 17 to 42 percent over twelve weeks. Blended cost hadn't moved, because Bexford's model provider had cut Vault's own price by more than half in that exact window, a routine update nobody thought to connect to the routing numbers. The two changes had landed on top of each other and canceled almost exactly.

It was never really about whether $0.035 was a healthy number. There was no single number that could describe two tiers pretending to be one.

The real cost wasn't in the blended number at all. It was in Corcannon, one of Bexford's larger customers, a home goods retailer that had launched a confusing new bundle-and-subscribe pricing plan around the same time. Corcannon's own Vault share had gone from 22 to 71 percent in those same twelve weeks, driven by real complexity in the tickets their own customers were now sending, not by anything Bexford's router did wrong.

The decision that opened the door went back to Threadscribe's very first pricing review, more than a year earlier, before Vault existed. The team picked blended cost per ticket as the number to watch, because at the time there was only one model and one price, and a blended number and a real number were the same thing. Nobody chose carelessly. It was the right number for the product that existed that day.

Run the same twelve weeks again with one change: Vault's share tracked on its own, every week, next to its own price. By week six, before the provider's price cut had even landed, the share crosses 25 percent, and Darragh's team goes looking, the way they eventually did, but six weeks sooner. Corcannon's launch gets flagged as the real driver behind its own account, instead of sitting lost inside a platform-wide average, and the router gets a stricter confidence bar for the gray-zone tickets before the quarter ends, not after it.

One design let a calm blended number speak for two tiers it was never built to describe together. The other watches the mix and the price on their own, and it would have rung six weeks sooner, before a single provider price cut ever had the chance to hide anything.

What I'd tell myself, back at that first pricing review: the blended number was never wrong, exactly. It just stopped being able to answer the question the day a second tier showed up, and nobody gave that day a name.

LEAD, the four letters behind the $0.048

This isn't a story wearing a metric's clothes. It's a metric question, and LEAD is what stops a calm blended number from standing in for two tiers it was never built to describe together.

LLink. What business outcome actually matters?
Gross margin on Threadscribe's contracts. The whole pricing model assumes AI cost per ticket stays a small, roughly steady share of what an account pays every month.
Not the model's own accuracy, and not total AI spend on its own. Margin is what actually breaks if that assumption stops holding.
EEarly signal. What moves weeks before the outcome does?
Vault's routing share, tracked weekly and kept separate from Vault's own price. It climbed from 17 to 42 percent over twelve weeks while blended cost sat still, because a price cut on Vault landed in the same window and absorbed the shift in the one number everyone watched.
This is the hardest step, and the one most answers skip. A number that looks perfectly calm right up until the morning it isn't, is exactly what blended cost was here.
AAbuse. How does this metric get gamed?
Loosen or tighten what counts as "complex" purely to move the routing-share number, without changing what agents actually deal with. Or keep swapping Vault for an even cheaper substitute model to hold blended cost down, without checking whether draft quality holds up on the hardest tickets.
A metric that can be hit without doing the real work isn't measuring the real work.
DDecision. What would you actually do at each threshold?
Past 25 percent Vault share, sustained two weeks: investigate whether it's a real jump in ticket difficulty or the router quietly drifting. Past 40 percent: cap what the router sends automatically, downgrade gray-zone tickets to Scout with a flag for a person, and start a pricing conversation with whichever account is driving it.
A metric nobody acts on is a dashboard. These two thresholds are what make it a decision instead of a chart.

Three things worth stating directly, since this is where the real judgment sits. The alternative Darragh's team considered first, and dropped, was a flat daily cap on how many tickets any one account could route to Vault. It lost because it downgrades genuinely hard tickets right along with easy ones, and it doesn't touch what's actually driving the share up, which is real ticket complexity plus a classifier nobody had re-checked. The AI-specific failure worth naming by name is silent classifier drift: the small model deciding which tickets are complex enough for Vault had never been checked again since launch, so its own line quietly moved as customers' wording and product features changed underneath it, with no code change and no decision behind it. The guardrail is a standing re-check of that classifier against a fixed set of labeled tickets on a schedule, plus an alert the moment routing share itself moves outside an expected band. That guardrail isn't free: a stricter confidence bar for what auto-routes to Vault means some genuinely hard tickets get downgraded to Scout first, and first-draft acceptance on those specific tickets drops from about 72 to 58 percent, a real quality-cost trade worth naming, not wishing away. And the bar Threadscribe holds itself to was never zero cost variance across 2.1 million tickets a month; no product with two model tiers can promise that. It's a threshold-specific bar, Vault's share held under a set line for the typical week, checked every week against the real mix, not one blended number standing in for two tiers pretending to be one.

And if you want to be sure it really works, try it somewhere else

Same four letters, an auto-insurance claims-triage tool instead of a support inbox, and this time a storm season drives the tier shift, not a pricing-plan launch.

Adjustly is a claims-triage assistant Quillane Insurance built for its adjusters. An incoming auto claim gets a drafted set of triage notes and next steps before a human ever opens the file. Anezka Osterlund runs product on it.

The build-up: Adjustly splits claims two ways too. Flint, the cheap tier, handles the clean ones: single vehicle, clear photos, an obvious at-fault party. Birch, the expensive tier, handles the rest: multiple vehicles, contested liability, an injury claim needing careful cross-referencing against the policy file. Blended cost per claim draft sat near $0.42 for most of a year.

The decision Anezka would take back Building the claims budget on the assumption that a busy storm season would show up in blended cost right away. Quillane renegotiated a volume discount on Birch pricing in the same contract cycle a regional storm pushed a wave of genuinely complex, multi-vehicle claims through the door, and the discount hid the whole spike.

Birch's share of claims climbed from 12 to 29 percent over the ten weeks after the storm, while blended cost barely moved, sliding from $0.42 to just $0.44, because Birch's own price per claim had been cut under the new contract in the very same stretch.

Hand sketched drawing on off-white paper. Left, a half circle gauge in green with its needle resting low, labeled blended cost, captioned sits flat, looks fine. A bold VS sits between the two halves in red. Right, a simple person figure in a red shirt, labeled hard claims, captioned keep walking to the costly tier.
The gauge looks calm. The hard claims keep moving to the expensive tier anyway. A blended number can only report what it's told to average, not what's actually shifting underneath it.

Same rank as before: track the tier's own share and its own price, on their own, and act before the blended number is forced to move. The fix is the same shape too: a stricter bar for what Adjustly auto-routes to Birch, a flag for the gray-zone claims instead of a guess, and a standing check on the router's own calls against a fixed set of claims nobody's relabeled in months.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: the two things that can cancel inside a blended number are the mix and the price, so track them apart, set a real threshold on the mix, and act on it before the bill does.
Cost: there's no budget this quarter for both a router audit and a pricing renegotiation. The router audit wins, since it's aimed at what's actually driving the mix, not just the number the mix produces.
The model got better, for real: say Vault's accuracy on complex tickets jumps for real. That's not proof the tail shrinks. Harder tickets keep arriving whether or not the model handling them got better at them, so a climbing share is a separate fact from the model's own quality.

Where people run it wrong.
They watch blended cost because it's the number finance already tracks, and never ask whether two tiers are hiding inside it.
They notice the mix moving and "fix" it by quietly swapping the expensive tier for a cheaper substitute model, without checking whether the harder tickets still get a usable draft.
They wait for the blended number to move before acting, when a price change on either tier can hold it still for months after the real shift already started.

How to use it live. Say the mechanism out loud before naming a metric: "blended cost is a weighted average, so it can hide a mix shift and a price shift at the same time." That buys a beat to think instead of reciting whatever the dashboard already shows.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
LEAD: find the signal that moves first. Built for metric questions, not a story about a single person's habit.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Darragh Pellegrini, who owns cost and margin for Threadscribe at Bexford. Built the product's original cost model himself.
3 · THE HABIT
What did Darragh stop doing because the number always looked fine?
Tap to flip
ANSWER
He stopped checking the router's own ticket split by hand, and started only reading blended cost off the Monday dashboard.
4 · THE CANCELING PAIR
What two numbers moved in opposite directions and canceled out in the total?
Tap to flip
ANSWER
Vault's own price fell from $0.16 to $0.07 a ticket while Vault's share of all tickets climbed from 17 to 42 percent, in the same twelve weeks. The two effects nearly canceled in blended cost.
5 · THE OLD DECISION
What decision would Darragh take back?
Tap to flip
ANSWER
Picking blended cost per ticket as the single cost metric at launch, back when Threadscribe had only one model, and never swapping it out once Vault shipped as a second tier.
6 · THE NUMBER
Fill in the blank: Vault's share of tickets climbed from 17 percent to ___ percent while blended cost never moved.
Tap to flip
ANSWER
42 percent. Blended cost held at $0.035 the whole time because Vault's own price had just been cut by more than half.
7 · THE REPLAY
Same quarter, new design, what changes?
Tap to flip
ANSWER
Vault's share crosses 25 percent by week six, before the price cut even lands. Darragh's team investigates six weeks sooner, Corcannon's launch gets flagged as the real driver, and the router gets a stricter bar before the quarter closes.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the matching blind spot?
Tap to flip
ANSWER
Adjustly, a claims-triage tool at Quillane Insurance. Same blind spot: a price renegotiation on the expensive tier hid a storm-driven jump in how many claims needed it.

Check yourself Score: 0 / 0

Multiple choice
1. Why did blended cost per ticket sit flat at $0.035 for twelve straight weeks even though Vault's share of tickets climbed from 17 to 42 percent?
  • A. Scout's price rose enough to offset the extra Vault tickets.
  • B. Vault's own price was cut by more than half in the same window, and the two changes canceled out in the total.
  • C. Bexford excludes Vault-tier tickets from the blended cost calculation.
  • D. The router stopped sending any tickets to Vault after week six.
Show hint
Look at the "why the number sat still" paragraph, right after the two charts in Section 1.
Show answer
B. A bigger share of tickets, at a much smaller price, landed almost exactly where a smaller share, at the bigger price, had been. That's what a weighted average lets happen.
True or false
2. True or false: once blended cost jumped to $0.048, the underlying shift in routing had just begun.
  • True
  • False
Show hint
Check the E step, Early signal, in the LEAD recap.
Show answer
False. The shift had been running for a full quarter already, invisible while the price cut absorbed it in the total. The jump was the moment the masking ran out, not the moment the drift started.
Fill in the blank
3. Vault's share of all tickets climbed from 17 percent to ___ percent over the same twelve weeks that its own price fell from $0.16 to $0.07.
Show hint
Check the direct answer and the leading-edge chart in Section 1.
Show answer
42 percent. A jump the blended number never showed a single hint of, because it was averaging the shift away against a price cut on the exact same tier.
Short answer, name the rejected alternative
4. What alternative did Darragh's team consider for controlling cost, and why did it lose?
Show hint
Look at the paragraph right after the four LEAD steps in the framework recap.
Show answer
Model answer: A flat daily cap on how many tickets any one account could route to Vault. It lost because it downgrades genuinely hard tickets right along with easy ones, and it doesn't touch what's actually driving the share up, which is real ticket complexity plus a classifier nobody had re-checked.
Short answer, apply it yourself
5. Pick an AI product you use that quietly hands different requests to different underlying models. Name one blended number it might report that a mix shift could hide, and how you'd check.
Show hint
Think of a product where an easy request and a hard request clearly cost different amounts to answer.
Show answer
Model answer: A writing assistant's "average cost per generation" might look steady, since most requests are short edits. But if more people started asking it to draft long documents, that mix shift could hide behind a provider price cut on its bigger model, the same way it did for Threadscribe. I'd ask for the share of requests going to each model size, tracked weekly, instead of trusting the one blended cost number.
Multiple choice
6. If Vault's price had never been cut and still cost $0.16 a ticket at week twelve, but its share had still climbed to 42 percent, what would blended cost have been instead of the $0.035 everyone saw?
  • A. $0.035, no change, since price and share always move together.
  • B. $0.049
  • C. About $0.072
  • D. $0.16, the full Vault price
Show hint
Work the weighted average from Stage 3 of the walkthrough: 58 percent at $0.009 plus 42 percent at $0.16.
Show answer
C. 0.58 times $0.009, plus 0.42 times $0.16, comes to about $0.072, more than double the $0.035 that actually showed up. The price cut is the entire reason the real number stayed hidden.
Before you close the answer
Why this works
Tests whether you can explain a mechanism, not just name a metric. Most candidates say "track cost per interaction" and stop, without ever asking what's hiding inside an average of two different prices.
Follow-up traps
"Isn't a 37 percent jump in one quarter just normal cost growth?" Response: blended cost sat completely flat for the twelve weeks before it, that's not growth, that's a masked shift finally showing up once the thing masking it ran out.

"Why not just check blended cost weekly instead of monthly?" Response: because a price cut and a mix shift can land in the same week just as easily as the same quarter. Frequency doesn't fix a number that's built to hide the thing you need to see, decomposing it does.
If pressed
The stricter confidence bar isn't the same for every account. Accounts with a track record of mostly simple tickets, like most of Bexford's small customers, keep the looser bar, since a false downgrade costs them little. Concentrated, complex accounts like Corcannon get the stricter bar first.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more