ConceptAdvancedQuality, Cost & Token Economics / Pricing AI products: seat, usage, outcome / #23

What is the relationship between pricing and your latency and quality tiers?

PICK · call transcription and QA scoring for a debt-collection call center

Verityline listens to a customer service call, writes it up, and scores it against a QA checklist: did the agent say the required words, confirm the problem got fixed, keep an even tone. Solworth Connect runs it across 40,000 calls a day, including a debt-collection line that has to open every call with one specific legal disclosure. Iwona Steenkamp owns Verityline's pricing at Duskharrow, the company that builds it. Dalibor Kucharski runs QA at Solworth. One quarter, both of them found out the hard way that Verityline's cheap tier and its careful tier were never really selling the same promise.

The direct answer
Price the model that scores the call, not the plan the customer bought. Every call defaults to a fast, cheap model. Any call that touches a required disclosure or an escalation gets pulled onto a slower, more careful model automatically, whatever tier the account is on. You charge more only for the calls where a wrong score costs more than the extra few cents ever could.
Do this, in order
  1. Route by what the call touches, not by which plan the customer bought.Why: a flat plan-wide setting either overspends on routine calls or underprotects the regulated ones, and it can't do both right at once.
  2. Build the routing rule off the call's content, before scoring starts, not after.Why: once a call is scored "compliant," nobody re-opens it. The routing decision has to happen before that label gets written.
  3. Spend the careful tier's accuracy on the hidden, expensive mistake, not the loud, cheap one.Why: overpaying shows up on an invoice and gets fixed next month. A missed disclosure hides inside a "clean" score for as long as nobody looks.
  4. Track the cheap tier's miss rate on a shared eval set, not just its overall accuracy.Why: one blended "97% agreement" number was exactly what let a budget review treat two very different tiers as the same tool at different speeds.
  5. Set a real number that would change the routing rule, and check it monthly.Why: a pick with no kill criteria is just a habit. This is what makes it a position instead of a preference.
  6. Don't put every call on the careful tier "to be safe."Why: three quarters of Solworth's calls are address changes and payment reminders. Paying careful-tier prices for those never changes a coaching decision.

How to answer this, stage by stage

Nobody is grading whether you can say "there's a tradeoff between speed and accuracy." They're grading whether you know which mistake costs more, in whose hands, and whether you'd actually price around that or just talk about it.

1
Ground it in one real case before answering in the abstract
Say it like this
"Let's ground this in one case. Verityline is Duskharrow's call transcription and QA scoring tool. Solworth Connect runs it across 40,000 calls a day, including a debt-collection line. Iwona Steenkamp owns pricing on it, and Dalibor Kucharski runs QA at Solworth."
Why this works
An abstract "speed versus accuracy" answer stays a slogan. One real account keeps it something you can actually walk through.
2
Say the plan out loud before naming a number
Say it like this
"I'm going to run PICK. Take a position on how price and the tiers actually relate, name who feels each kind of error, say which error is cheap and which is expensive, then say what evidence would change my mind."
Why this works
Signals a method already in motion, not four thoughts arriving in whatever order they occurred to you.
3
Give the position, plainly, before any reasoning
Say it like this
"Here's my position. Price should follow the model, not the plan. A cheap, fast model scores everything by default. A slower, pricier model with a second compliance pass takes over automatically the moment a call touches a required disclosure or an escalation, no matter what tier the customer bought."
Why this works
This is the direct answer, said in one breath, before the interviewer has to go looking for it.
4
Name the impact, in real units, for both sides
Say it like this
"If Solworth puts every call on the careful tier to be safe, that's about $4,800 a day, when three quarters of those calls are address changes nobody needs a second pass on. If they put every call on the cheap tier to save money, one cut-off disclosure on a debt-collection call sits inside a score marked 'compliant' for months, until a client audit opens it."
Why this works
Turns "there's a tradeoff" into two numbers a real person actually pays.
5
Name the cost asymmetry, and which mistake you're pricing against
Say it like this
"Overpaying is cheap and loud. It shows up on next month's invoice and somebody renegotiates it. Underpaying is quiet and expensive. It hides inside a score nobody re-checks, and it only shows up once a client's own compliance team pulls that exact call. I'd rather overpay on the boring calls than underpay on the ones that can turn into a finding."
Why this works
This is the hardest step in PICK. Naming which mistake actually costs more, out loud, is what makes it a real position instead of a shrug.
6
Give the kill criteria, so the pick isn't permanent by accident
Say it like this
"I'd change this the moment the cheap tier's miss rate on required disclosures closes to something like the careful tier's, under 2 percent on our own eval set. Until then, no compliance call runs on the cheap tier alone, no matter what the contract says."
Why this works
Shows the pick is tied to evidence, not to a preference you'd defend forever regardless of what the numbers do.
7
Prove it with the near miss, cut to four sentences
Say it like this
"Here's what happens without the routing rule. A Solworth agent got talked over mid-disclosure. The cheap tier still saw the words 'debt collector' show up and scored the call compliant. That score sat in Dalibor's coaching file for six months, until his client's own auditor pulled that exact call and found the required close was never actually said."
Why this works
Shows the real, countable cost of skipping the rule, not just "it could go wrong."
8
Close on the decision, not the story
Say it like this
"So: price the model, not the plan. Route by what the call touches, not by what the customer paid for. Overpaying is the mistake you can see and fix next month. Underpaying is the one that's already cost you before you find out it happened."
Why this works
Restates the direct answer in one breath, so the interviewer leaves with the decision, not just the story behind it.

Let's learn

What does it mean when a cheap AI tier and an expensive one both say a call is fine, but only one of them actually checked?

Verityline listens to a recorded customer service call and grades it against a checklist, so a QA team doesn't have to sit and listen to the whole thing themselves.

Hand sketched numbered icon list titled Solworth's QA desk before Verityline. Three rows: a headset icon, one reviewer samples 2 of every 100 calls. A document icon, each hand review takes about 12 minutes. A gauge icon, most agents go months without one checked call.
Before Verityline, Solworth's QA team could only ever see a sliver of what its own agents were saying.

Before Verityline, Solworth's QA team sampled 2 of every 100 calls by hand, about 12 minutes per review. Most agents went close to 90 days between one checked call. With Verityline, every single call gets scored. The Pulse tier, a small fast model, returns a score about 90 seconds after the call ends, for $0.045 a call. The Vault tier, a bigger and slower model, takes 20 to 30 minutes and adds a second pass built specifically to re-listen for required disclosure language, for $0.12 a call.

Hand sketched left to right flow diagram titled How Verityline routes a call. Five boxes connected by arrows: Call ends, Content check, this box outlined in amber to mark the routing decision, Pulse scores, Vault rechecks, Score saved.
One routing step, in the middle, decides which of the two models a call actually gets.
Knowledge spark: what's a false negative here? Verityline says a call passed the compliance check when it actually didn't. It missed a real problem instead of catching it. That's the kind of mistake that costs the most, because nobody goes looking for a problem the tool already told them isn't there.

Of Solworth's 40,000 calls a day, about 9,000 are debt-collection calls that must open with a specific legal disclosure: this is an attempt to collect a debt, anything said may be used for that purpose. The other 31,000 are routine, address changes, payment reminders, appointment confirmations, nothing a regulator has ever asked to see.

Here's the turn. Solworth's first instinct, when Vault tier launched, was to put every call on it, just to be safe. That's about $4,800 a day. Most of that money bought a second, careful pass on calls where the first pass was already going to be right. The extra accuracy on a payment reminder call never once changed what a supervisor did next.

Daily Verityline spend, three ways to route 40,000 calls
$5,000 $2,500 0 $1,800 All calls on Pulse $4,800 All calls on Vault $2,475 Risk-routed (recommended)
Cheapest, riskiestSafest, priciestOnly compliance calls pay the careful price
Routing everything to Pulse is cheapest, right up until one missed disclosure costs more than the difference ever saved. Routing everything to Vault burns about $2,325 a day, near $70,000 a month, paying careful-tier prices for calls that never needed them.
We did not lose six months to a worse model. We lost them to a score nobody was ever going to re-check.
Hand sketched quadrant diagram titled Two ways a pricing tier goes wrong. X axis how visible the mistake is, from hidden to obvious. Y axis how expensive it gets, from cheap to costly. Vault tier on routine calls plotted obvious and cheap. Pulse misses a disclosure plotted hidden and costly. Risk routed, right tier plotted in the middle, both low cost and mostly visible.
Overpaying sits out in the open, easy to see and easy to fix. A missed disclosure sits in the quiet corner, and that's exactly what makes it expensive.
The choice that mattered Duskharrow launched Verityline with one model and one blended price, backed by a single headline number: 97 percent agreement with human reviewers. That was fine when its first customers were small outbound sales teams with no regulated calls in the mix. It stopped being fine the day a debt-collection line signed up, because now some calls carried real legal weight and most didn't, and one blended number couldn't tell them apart.

What I would leave alone: appointment reminders, address updates, and every other zero-risk call type, whatever the account's budget looks like. Nobody needs a careful second pass on a call that was never going to change a coaching decision either way.

The lesson: a pricing tier is really a promise about which mistakes you're willing to make. That promise has to be attached to what the call actually is, not to which plan someone happened to buy.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a pricing page's single accuracy number was the actual bug, not the model underneath it.

Iwona Steenkamp has owned pricing at Duskharrow for three years, long enough to have built two products' worth of price sheets from scratch. She reads a usage chart the way some people read a weather map, past the total, straight to whichever slice is moving.

Solworth Connect signed up for Verityline eighteen months ago, back when it only ran customer service lines: billing questions, address changes, appointment scheduling. Pulse tier scored every one of those calls in about 90 seconds, for $0.045 a call, and Dalibor Kucharski's QA team went from checking 2 percent of calls to checking all of them overnight. The first two quarters were good ones. Dalibor could finally tell his own leadership, with real numbers, which agents needed coaching and which didn't. He called it having eyes everywhere for once.

Then Solworth won a debt-collection contract, its first regulated call type. Duskharrow had built Vault tier for exactly this: a slower model, 20 to 30 minutes instead of 90 seconds, with a second pass built to re-listen for the required opening disclosure word for word, not just check that the topic came up. Solworth turned it on for the whole account, all 40,000 calls a day, at $0.12 each. For a while, that felt like the responsible choice.

It thinned in three beats, and none of them looked careless. Beat one: a rough quarter company-wide sent Solworth's leadership looking for 15 percent out of every vendor line, and Verityline's bill was an obvious place to look. Beat two: someone pulled up Duskharrow's own pricing page, which still carried one accuracy number for the whole product, 97 percent agreement with human reviewers, no split by tier, no split by call type. Beat three: reading that single number, leadership decided Pulse and Vault were basically the same tool at different speeds, and moved the entire account, collections line included, back onto Pulse to save roughly $3,000 a day.

Hand sketched timeline titled The six months nobody reopened that call. Four milestones: Pulse scores it, marked compliant in 90 seconds. Score used, goes straight into coaching notes. Client audit, this milestone emphasized in red, reopens the same recording. Disclosure missing, the required close was never said.
Nothing looked wrong the week the account switched tiers. It took six months for anyone to open the one call that mattered.

The trigger wasn't a dashboard turning red. Six months later, one of Solworth's collection clients ran its own routine compliance check, the kind that happens on a schedule, not because anything looked off. Their auditor pulled 50 calls at random that Verityline had scored 100 percent compliant. Three of them had an incomplete disclosure. In each one, the customer had started talking before the agent finished the required line, and Pulse, hearing the words "debt collector" land somewhere in the transcript, marked the call clean anyway.

We did not lose the account to a mistake. We lost six months of trust to a mistake we had already told ourselves, in writing, wasn't there.

Dalibor's phone rang before nine that morning. The client wasn't threatening to leave over three calls. They were threatening to leave over what those three calls implied about the other 1.6 million collection calls Solworth had scored the same way across those six months. Duskharrow ran an overnight re-score of the whole stretch against Vault. It flagged about 97,000 likely misses, close to Pulse's own known miss rate on that model version, which was cold comfort: it meant the sample wasn't unlucky. It was exactly the rate the tool had been running at the whole time.

Hand sketched comparison diagram titled Trust in a QA score was never a dial. Left panel, a gauge icon labeled What we assumed, caption trust fades slowly as the tier gets cheaper. Right panel, a scale icon labeled What actually happens, caption fine, then it flips, the day one call gets reopened.
This is the whole answer to what pricing has to do with the tiers. Trust in a score doesn't fade a little as the price drops. It holds, then it flips, the day someone reopens one call.

The decision that opened the door went back to a fifteen-minute slide in a much earlier pricing review, the week Duskharrow first shipped Verityline. Someone had asked whether the pricing page needed a second accuracy number, split by call type, once regulated customers started showing up. The answer was no. At the time, every customer running Verityline was a small outbound sales team. One number was the whole truth for all of them.

Run the same six months again, with the routing rule in place instead of a plan-wide switch. Solworth's 9,000 daily collection calls get tagged the moment they enter the queue and always run on Vault, no matter what a budget review decides. The other 31,000 move to Pulse. New daily cost: $2,475, not $4,800, a real savings of $2,325 a day, close to $70,000 a month. It's less than the $3,000 a day leadership thought they were saving by cutting everything, but it comes without ever touching the calls that carried risk. An audit sample of 50 collection calls under the new rule finds zero incomplete disclosures.

One design let a budget spreadsheet decide which calls got the careful read. The other let the call's own content decide, whatever the spreadsheet said that quarter.

What Iwona would tell herself, back in that fifteen-minute slide: the one blended accuracy number wasn't a simplification, it was a promise about odds, and nobody had actually priced the odds. They'd only priced the speed.

PICK, or the four decisions hiding inside two price tags

Not a way to dress up "it depends" in four letters. PICK is what forces a real commitment before the reasoning, then makes you say, out loud, which mistake you'd rather live with.

PPosition. Your pick, in one sentence, before any reasoning.
Price follows the model that scores the call, not the plan the customer bought. Pulse, fast and cheap, scores every call by default. Vault, slower and pricier, takes over automatically the moment a call touches a required disclosure or an escalation.
Say the position before the reasoning, or the interviewer spends the next two minutes waiting to find out what you'd actually do.
IImpact. Who feels each kind of error, in what units.
Overpay, and every routine call on Solworth's floor costs real dollars, about $4,800 a day at full Vault pricing. Underpay, and a debt-collection agent's cut-off disclosure sits inside a "compliant" score for six months, until a client's own auditor opens that exact call.
Naming both people, the one paying too much and the one protected too little, keeps this from turning into a one-sided cost argument.
CCost asymmetry. The heart of it.
Overpaying is cheap and loud. It shows up on next month's invoice and gets renegotiated. Underpaying is quiet and expensive. It hides inside a score nobody re-checks until a client's compliance team goes looking. Spend the careful tier's accuracy against the second one, not the first.
This is the step that actually earns the pick. Anyone can say "it depends." Naming which mistake is worse, in real units, is what survives a follow-up question.
KKill criteria. What evidence would flip the pick.
If Pulse's miss rate on required disclosures falls under 2 percent on Duskharrow's own eval set, close to Vault's own 0.6 percent, the routing rule can relax. Until then, no compliance call runs on the cheap tier alone, whatever the contract says.
A pick with no kill criteria is just a preference you're defending. This is what makes it a position you'd actually change your mind about, on purpose.

And if you want to be sure it really works, try it somewhere else

Same four letters, a hospital's patient line instead of a call center, and this time the hidden mistake isn't a legal phrase. It's a symptom.

Voxbridge is a live AI interpretation tool. A patient calls a clinic in Spanish, and Voxbridge either translates the call fully by machine or hands part of it to a certified human interpreter who joins in near real time. Fennbridge Health runs it across roughly 6,000 patient calls a day, and about 900 of those involve a first visit or a new symptom being described for the first time, the calls where a wrong word actually changes what a nurse writes down.

Hand sketched comparison diagram titled Same four letters, a hospital instead of a call center. Left panel, a person icon labeled Dalibor, Solworth Connect, caption a missed disclosure, quiet for months on the cheap tier. Right panel, a person icon labeled Aksel, Fennbridge Health, caption a mistranslated symptom, caught inside one visit.
Same PICK, a different domain, a different hidden mistake. Both times, a flat plan-wide setting was the thing that let it hide.
The decision Fennbridge would take back Voxbridge was licensed per department, one flat annual price covering every clinic, with no split between a scheduling desk and a triage line. During a budget freeze, the hospital's IT department capped every department on the cheap, machine-only tier to hit a savings target, scheduling and triage together, because the license didn't distinguish between them.

Aksel Wieland runs Language Services at Fennbridge. He didn't choose the cap. It came down from a budget memo that never asked which calls actually needed the $1.10-a-minute human-reviewed tier versus the $0.30-a-minute machine-only one. On the machine-only tier, mistranslation of a specific medical term runs about 9 percent. On the human-reviewed tier, about 1.2 percent. A caller once described "dolor en el pecho que se irradia al brazo," pain in the chest radiating to the arm, a real warning sign. The machine-only tier flattened it into a vaguer note about chest discomfort, and the patient was booked for a routine visit. A bilingual nurse happened to review that recording two days later for an unrelated reason and caught it before the appointment date arrived.

Same rank, different lever: the fix isn't a bigger department budget or a better translation model. It's pricing and routing by what the call is about, a first-visit or new-symptom tag that forces the human-reviewed tier automatically, separate from whatever a department's annual license says it can afford that quarter.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: cap by what the call touches, not by department budget. Symptom calls always route to the human-reviewed tier.
Cost: there's no budget this quarter for a live content classifier. Ship the cheap version first, a short list of symptom and consent words checked before the call ends, not a trained model.
The model got better, for real: say the machine-only tier's accuracy on medical terms doubles overnight. The fix barely changes. You don't know it's doubled until your own eval set proves it under the kill line. Until then, the routing rule holds.

Where people run it wrong.
They price and cap access by department budget instead of by what a call is actually about.
They trust one blended accuracy number across every kind of sentence, when a handful of medical terms are the only ones worth a second look.
They treat a near miss as proof the process works, instead of proof it got lucky once.

How to use it live. Ask the split question before naming a fix: "Is this a price cap, or a content cap, because a flat department budget and a per-call-type routing rule protect two completely different things." That buys you the room to actually answer, instead of guessing at a discount.

Pulse's miss rate on required disclosures, six model versions
12% 6% 0% kill line: 2% Vault: 0.6% 11% 9.4% 8.2% 7.3% 6.6% 6.0% v1 v2 v3 v4 v5 v6
Pulse's own eval scoreKill line, 2%Vault's rate, 0.6%
Pulse has improved every version, 11 percent down to 6. It would need to fall under 2 percent, close to Vault's own six tenths of a percent, before the routing rule relaxes. It isn't there yet.

Three things worth stating directly, since this is where the real judgment sits. The alternative Duskharrow considered, and rejected, when fixing this was building one single best-of-both model and raising everyone's price to cover it, so every call always got the careful pass. It lost, because it prices out every customer whose calls carry no regulatory weight at all, and a small outbound sales team would rather leave for a cheaper competitor than fund a compliance pass it never needed. The AI-specific failure worth naming is confident wrongness: Pulse doesn't know it's unsure. It pattern-matches on a disclosure keyword appearing anywhere in the transcript, rather than checking the full required phrase was completed before the customer started talking, so it hands back "compliant: 100 percent" even when the phrase was cut off. The guardrail is Vault's dedicated second pass, built to check phrase completion rather than keyword presence, plus a monthly re-check of both tiers against a golden set of known partial disclosures, with the 2 percent line gating whether Pulse stays eligible for compliance calls at all. And the trade-off is real: Vault's extra accuracy costs both money, about 2.7 times the per-call price, and time, 20 to 30 minutes instead of 90 seconds, so a supervisor can't coach an agent live off a Vault score mid-shift. That delay is accepted on purpose, only for the calls where a wrong score is the expensive kind of wrong.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
PICK: commit to a position on a tradeoff, then show which of the two errors actually costs more, and to whom.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Iwona Steenkamp, who has owned pricing at Duskharrow for three years, and built Verityline's two-tier pricing after a client's audit exposed the gap in the flat, one-tier design.
3 · THE POSITION
What's the P step here, in one line?
Tap to flip
ANSWER
Price follows the model that scores the call, not the plan the customer bought. Routine calls default to Pulse. Any call touching a required disclosure or an escalation routes to Vault automatically.
4 · THE COST ASYMMETRY
Which error is cheap and visible, and which is hidden and expensive?
Tap to flip
ANSWER
Overpaying, putting every call on Vault, is cheap and visible: it shows up on the invoice. Underpaying, a missed disclosure on Pulse, is hidden and expensive: it hides inside a "compliant" score until an audit finds it.
5 · THE OLD DECISION
What decision would Iwona take back?
Tap to flip
ANSWER
Launching Verityline with one model and one blended accuracy number, 97 percent agreement with human reviewers, with no split by call type. It made sense when every customer ran only low-stakes calls.
6 · THE NUMBER
Fill in the blank: routing every call to Vault costs $___ a day. Risk-routing, Vault only for the 9,000 collection calls, costs $___ a day.
Tap to flip
ANSWER
$4,800 a day for all-Vault. $2,475 a day risk-routed, a savings of $2,325 a day, close to $70,000 a month, without touching the calls that actually carry risk.
7 · THE REPLAY
Same six months, new routing rule, what changes?
Tap to flip
ANSWER
Collection calls always run on Vault, whatever a budget review decides. Daily cost falls to $2,475 instead of $4,800, and an audit sample of 50 collection calls finds zero incomplete disclosures.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the hidden mistake there?
Tap to flip
ANSWER
Voxbridge, Fennbridge Health's live medical interpretation tool. The hidden mistake is a mistranslated symptom on the machine-only tier, caught by luck before it changed a patient's appointment.

Check yourself Score: 0 / 0

True or false
1. True or false: the missed disclosure that triggered the client audit happened because Verityline's model had gotten worse that quarter.
  • True
  • False
Show hint
Check what actually changed, the model, or which tier the calls were routed to.
Show answer
False. Pulse's model didn't change. Leadership moved the whole account, including collection calls, off Vault and onto Pulse to cut costs, without splitting which calls actually needed the careful pass.
Multiple choice
2. Why did Solworth's leadership originally move the debt-collection line onto the cheap tier?
  • A. Verityline's pricing page only showed one blended accuracy number, not a disclosure-specific rate.
  • B. Pulse tier had just received a major accuracy upgrade.
  • C. Dalibor personally recommended the switch to protect his own team's workload.
  • D. A regulator required every call to be scored in real time.
Show hint
Look at what leadership was actually looking at when they made the call, in the story's second habit-thinning beat.
Show answer
A. The 97 percent number looked the same for both tiers, so leadership read it as one tool at two speeds, not two very different odds on the same regulated mistake.
Fill in the blank
3. Pulse tier prices calls at $___ each. Vault tier prices calls at $___ each.
Show hint
It's stated where the two tiers are first introduced, in Let's learn.
Show answer
$0.045 and $0.12. A gap of $0.075 a call, which only matters at scale, and only where the calls actually carry risk.
Short answer, name the old decision
4. What old decision would Iwona take back, and why did it make sense when Duskharrow first built it?
Show hint
Look at the key point box titled "The choice that mattered," right after the first chart.
Show answer
Model answer: Launching Verityline with one model and one blended accuracy number instead of splitting it by call type. It made sense because every early customer ran small outbound sales calls, with no regulated call type in the mix.
Short answer, apply it yourself
5. Think of a service you use that has a fast, basic option and a slower, more thorough option. Name one kind of request you'd never want routed to the fast option, no matter the price.
Show hint
Think about which requests, if handled wrong, you wouldn't find out about for a long time.
Show answer
Model answer: A photo-printing app's automatic color correction is fine for everyday snapshots, but you'd want the slower, manually reviewed option for a printed portrait you're giving as a gift, because a bad print only gets noticed after it's already in someone's hands.
Short answer, work the number
6. If Pulse's disclosure miss rate ever fell to 1.5 percent, per the kill criteria in this answer, what should change?
Show hint
Check the K step in the framework recap, and the line chart's kill line.
Show answer
The routing rule should relax. 1.5 percent is under the stated 2 percent kill line, close to Vault's own 0.6 percent, so the evidence that justified forcing every compliance call onto Vault would no longer hold.
Before you close the answer
Why this works
Tests whether you treat a quality tier as a promise about one specific kind of mistake, not a general upgrade, and whether you'll route by what the call is rather than by which plan someone bought.
Follow-up traps
"Isn't it simpler to just charge one price and always use the good model?" Response: it prices out every customer whose calls carry no compliance risk, and it hides how much of what people pay for the careful model is protection they never actually used that day.

"What if a customer explicitly wants to save money and accept the risk?" Response: then it belongs in the contract in writing, but a customer who's only ever seen one blended number hasn't made that tradeoff on purpose, they've made it by accident.
If pressed
The content check that routes a call isn't a third AI model guessing at risk. It's a deterministic rule, this call belongs to the collections queue, or a flagged account code, so it never needs its own accuracy number, and the pricing stays exactly two tiers, not three.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more