What is the relationship between pricing and your latency and quality tiers?
Verityline listens to a customer service call, writes it up, and scores it against a QA checklist: did the agent say the required words, confirm the problem got fixed, keep an even tone. Solworth Connect runs it across 40,000 calls a day, including a debt-collection line that has to open every call with one specific legal disclosure. Iwona Steenkamp owns Verityline's pricing at Duskharrow, the company that builds it. Dalibor Kucharski runs QA at Solworth. One quarter, both of them found out the hard way that Verityline's cheap tier and its careful tier were never really selling the same promise.
- Route by what the call touches, not by which plan the customer bought.Why: a flat plan-wide setting either overspends on routine calls or underprotects the regulated ones, and it can't do both right at once.
- Build the routing rule off the call's content, before scoring starts, not after.Why: once a call is scored "compliant," nobody re-opens it. The routing decision has to happen before that label gets written.
- Spend the careful tier's accuracy on the hidden, expensive mistake, not the loud, cheap one.Why: overpaying shows up on an invoice and gets fixed next month. A missed disclosure hides inside a "clean" score for as long as nobody looks.
- Track the cheap tier's miss rate on a shared eval set, not just its overall accuracy.Why: one blended "97% agreement" number was exactly what let a budget review treat two very different tiers as the same tool at different speeds.
- Set a real number that would change the routing rule, and check it monthly.Why: a pick with no kill criteria is just a habit. This is what makes it a position instead of a preference.
- Don't put every call on the careful tier "to be safe."Why: three quarters of Solworth's calls are address changes and payment reminders. Paying careful-tier prices for those never changes a coaching decision.
How to answer this, stage by stage
Nobody is grading whether you can say "there's a tradeoff between speed and accuracy." They're grading whether you know which mistake costs more, in whose hands, and whether you'd actually price around that or just talk about it.
Let's learn
What does it mean when a cheap AI tier and an expensive one both say a call is fine, but only one of them actually checked?
Verityline listens to a recorded customer service call and grades it against a checklist, so a QA team doesn't have to sit and listen to the whole thing themselves.
Before Verityline, Solworth's QA team sampled 2 of every 100 calls by hand, about 12 minutes per review. Most agents went close to 90 days between one checked call. With Verityline, every single call gets scored. The Pulse tier, a small fast model, returns a score about 90 seconds after the call ends, for $0.045 a call. The Vault tier, a bigger and slower model, takes 20 to 30 minutes and adds a second pass built specifically to re-listen for required disclosure language, for $0.12 a call.
Of Solworth's 40,000 calls a day, about 9,000 are debt-collection calls that must open with a specific legal disclosure: this is an attempt to collect a debt, anything said may be used for that purpose. The other 31,000 are routine, address changes, payment reminders, appointment confirmations, nothing a regulator has ever asked to see.
Here's the turn. Solworth's first instinct, when Vault tier launched, was to put every call on it, just to be safe. That's about $4,800 a day. Most of that money bought a second, careful pass on calls where the first pass was already going to be right. The extra accuracy on a payment reminder call never once changed what a supervisor did next.
What I would leave alone: appointment reminders, address updates, and every other zero-risk call type, whatever the account's budget looks like. Nobody needs a careful second pass on a call that was never going to change a coaching decision either way.
The lesson: a pricing tier is really a promise about which mistakes you're willing to make. That promise has to be attached to what the call actually is, not to which plan someone happened to buy.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a pricing page's single accuracy number was the actual bug, not the model underneath it.
Iwona Steenkamp has owned pricing at Duskharrow for three years, long enough to have built two products' worth of price sheets from scratch. She reads a usage chart the way some people read a weather map, past the total, straight to whichever slice is moving.
Solworth Connect signed up for Verityline eighteen months ago, back when it only ran customer service lines: billing questions, address changes, appointment scheduling. Pulse tier scored every one of those calls in about 90 seconds, for $0.045 a call, and Dalibor Kucharski's QA team went from checking 2 percent of calls to checking all of them overnight. The first two quarters were good ones. Dalibor could finally tell his own leadership, with real numbers, which agents needed coaching and which didn't. He called it having eyes everywhere for once.
Then Solworth won a debt-collection contract, its first regulated call type. Duskharrow had built Vault tier for exactly this: a slower model, 20 to 30 minutes instead of 90 seconds, with a second pass built to re-listen for the required opening disclosure word for word, not just check that the topic came up. Solworth turned it on for the whole account, all 40,000 calls a day, at $0.12 each. For a while, that felt like the responsible choice.
It thinned in three beats, and none of them looked careless. Beat one: a rough quarter company-wide sent Solworth's leadership looking for 15 percent out of every vendor line, and Verityline's bill was an obvious place to look. Beat two: someone pulled up Duskharrow's own pricing page, which still carried one accuracy number for the whole product, 97 percent agreement with human reviewers, no split by tier, no split by call type. Beat three: reading that single number, leadership decided Pulse and Vault were basically the same tool at different speeds, and moved the entire account, collections line included, back onto Pulse to save roughly $3,000 a day.
The trigger wasn't a dashboard turning red. Six months later, one of Solworth's collection clients ran its own routine compliance check, the kind that happens on a schedule, not because anything looked off. Their auditor pulled 50 calls at random that Verityline had scored 100 percent compliant. Three of them had an incomplete disclosure. In each one, the customer had started talking before the agent finished the required line, and Pulse, hearing the words "debt collector" land somewhere in the transcript, marked the call clean anyway.
Dalibor's phone rang before nine that morning. The client wasn't threatening to leave over three calls. They were threatening to leave over what those three calls implied about the other 1.6 million collection calls Solworth had scored the same way across those six months. Duskharrow ran an overnight re-score of the whole stretch against Vault. It flagged about 97,000 likely misses, close to Pulse's own known miss rate on that model version, which was cold comfort: it meant the sample wasn't unlucky. It was exactly the rate the tool had been running at the whole time.
The decision that opened the door went back to a fifteen-minute slide in a much earlier pricing review, the week Duskharrow first shipped Verityline. Someone had asked whether the pricing page needed a second accuracy number, split by call type, once regulated customers started showing up. The answer was no. At the time, every customer running Verityline was a small outbound sales team. One number was the whole truth for all of them.
Run the same six months again, with the routing rule in place instead of a plan-wide switch. Solworth's 9,000 daily collection calls get tagged the moment they enter the queue and always run on Vault, no matter what a budget review decides. The other 31,000 move to Pulse. New daily cost: $2,475, not $4,800, a real savings of $2,325 a day, close to $70,000 a month. It's less than the $3,000 a day leadership thought they were saving by cutting everything, but it comes without ever touching the calls that carried risk. An audit sample of 50 collection calls under the new rule finds zero incomplete disclosures.
One design let a budget spreadsheet decide which calls got the careful read. The other let the call's own content decide, whatever the spreadsheet said that quarter.
What Iwona would tell herself, back in that fifteen-minute slide: the one blended accuracy number wasn't a simplification, it was a promise about odds, and nobody had actually priced the odds. They'd only priced the speed.
PICK, or the four decisions hiding inside two price tags
Not a way to dress up "it depends" in four letters. PICK is what forces a real commitment before the reasoning, then makes you say, out loud, which mistake you'd rather live with.
And if you want to be sure it really works, try it somewhere else
Same four letters, a hospital's patient line instead of a call center, and this time the hidden mistake isn't a legal phrase. It's a symptom.
Voxbridge is a live AI interpretation tool. A patient calls a clinic in Spanish, and Voxbridge either translates the call fully by machine or hands part of it to a certified human interpreter who joins in near real time. Fennbridge Health runs it across roughly 6,000 patient calls a day, and about 900 of those involve a first visit or a new symptom being described for the first time, the calls where a wrong word actually changes what a nurse writes down.
Aksel Wieland runs Language Services at Fennbridge. He didn't choose the cap. It came down from a budget memo that never asked which calls actually needed the $1.10-a-minute human-reviewed tier versus the $0.30-a-minute machine-only one. On the machine-only tier, mistranslation of a specific medical term runs about 9 percent. On the human-reviewed tier, about 1.2 percent. A caller once described "dolor en el pecho que se irradia al brazo," pain in the chest radiating to the arm, a real warning sign. The machine-only tier flattened it into a vaguer note about chest discomfort, and the patient was booked for a routine visit. A bilingual nurse happened to review that recording two days later for an unrelated reason and caught it before the appointment date arrived.
Same rank, different lever: the fix isn't a bigger department budget or a better translation model. It's pricing and routing by what the call is about, a first-visit or new-symptom tag that forces the human-reviewed tier automatically, separate from whatever a department's annual license says it can afford that quarter.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: cap by what the call touches, not by department budget. Symptom calls always route to the human-reviewed tier.
Cost: there's no budget this quarter for a live content classifier. Ship the cheap version first, a short list of symptom and consent words checked before the call ends, not a trained model.
The model got better, for real: say the machine-only tier's accuracy on medical terms doubles overnight. The fix barely changes. You don't know it's doubled until your own eval set proves it under the kill line. Until then, the routing rule holds.
Where people run it wrong.
They price and cap access by department budget instead of by what a call is actually about.
They trust one blended accuracy number across every kind of sentence, when a handful of medical terms are the only ones worth a second look.
They treat a near miss as proof the process works, instead of proof it got lucky once.
How to use it live. Ask the split question before naming a fix: "Is this a price cap, or a content cap, because a flat department budget and a per-call-type routing rule protect two completely different things." That buys you the room to actually answer, instead of guessing at a discount.
Three things worth stating directly, since this is where the real judgment sits. The alternative Duskharrow considered, and rejected, when fixing this was building one single best-of-both model and raising everyone's price to cover it, so every call always got the careful pass. It lost, because it prices out every customer whose calls carry no regulatory weight at all, and a small outbound sales team would rather leave for a cheaper competitor than fund a compliance pass it never needed. The AI-specific failure worth naming is confident wrongness: Pulse doesn't know it's unsure. It pattern-matches on a disclosure keyword appearing anywhere in the transcript, rather than checking the full required phrase was completed before the customer started talking, so it hands back "compliant: 100 percent" even when the phrase was cut off. The guardrail is Vault's dedicated second pass, built to check phrase completion rather than keyword presence, plus a monthly re-check of both tiers against a golden set of known partial disclosures, with the 2 percent line gating whether Pulse stays eligible for compliance calls at all. And the trade-off is real: Vault's extra accuracy costs both money, about 2.7 times the per-call price, and time, 20 to 30 minutes instead of 90 seconds, so a supervisor can't coach an agent live off a Vault score mid-shift. That delay is accepted on purpose, only for the calls where a wrong score is the expensive kind of wrong.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if a customer explicitly wants to save money and accept the risk?" Response: then it belongs in the contract in writing, but a customer who's only ever seen one blended number hasn't made that tradeoff on purpose, they've made it by accident.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pricing AI products: seat, usage, outcome
- #1 Compare seat-based, usage-based and outcome-based pricing for an AI product.
- #2 Why does seat-based pricing break when AI reduces the number of seats needed?
- #3 Design a pricing model for an AI feature with high variable cost and unpredictable usage.
- #4 What is the risk of usage-based pricing from the customer's point of view?
- #5 Explain how credits work as a pricing mechanism and their advantages.
- #6 How would you price an agent that completes a task rather than answers a question?