CalculationAdvancedAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #15

Explain how to negotiate around usage-based pricing uncertainty.

BOUND six calls a match on the demo, twenty-two calls a match on a bad Tuesday

LoadPilot AI sells a load-matching tool to freight brokers, billed per API call. Callum Bright runs vendor and procurement decisions at Haulwright, a regional freight brokerage moving about 1,200 loads a month.

The direct answer
Negotiate a price per completed load match, not per API call, and put a hard ceiling on how many calls the vendor's own model can burn chasing one match. That moves the pricing risk off the one variable neither side fully controls: how many tries the model needs before it lands a good pairing.
Do this, in order
  1. Push for pricing per completed match, not per raw call.Why: a completed match is the outcome you're paying for; a call count is a proxy that swings with the model's own uncertainty.
  2. Negotiate a hard retry ceiling per match, billed to the vendor past that point.Why: retries mean the model needed more tries to find a good answer, not that your load volume changed.
  3. Model your own worst case before signing, not just the sales demo's average case.Why: the average case is what gets sold; the worst case is what gets billed.
  4. Ask for a monthly cost cap or a true-up credit past an agreed ceiling.Why: it turns an open-ended bill into a bounded one you can actually plan a budget around.
  5. Track calls-per-match as its own number every month, not just the total invoice.Why: it's the number that moves before the bill does, so it tells you trouble is coming.

How to answer this, stage by stage

Nobody is scoring you on whether you can say "negotiate a discount." They're scoring whether you can show the arithmetic behind why the bill moved, and fix the actual variable.

Stage 1
Scope it to one vendor, one cost driver
Say it like this
"I'll answer this for Haulwright's contract with LoadPilot AI, and the one line item driving the pricing uncertainty: calls per load match."
Why this works
Keeps a broad "usage-based pricing is scary" question anchored to one real, calculable number.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down, own the numbers, use a range instead of one guess, nail the sanity check, then say what moves it most."
Why this works
Signals a real estimation method instead of a gut-feel guess dressed up with a dollar sign.
Stage 3
State the equation before the numbers
Say it like this
"Monthly cost equals loads processed, times calls per load match, times price per call. That's the whole equation, before I touch a single number."
Why this works
An estimate with no visible arithmetic is a guess wearing a confident tone.
Stage 4
Give the one decision
Say it like this
"Negotiate pricing per completed match, not per call, and put a hard retry ceiling on the vendor's own side of the equation."
Why this works
This is the direct answer, said plainly, before any numbers back it up.
Stage 5
Show the range, not a single number
Say it like this
"At six calls a match, we're near 1,300 dollars a month. At twenty-two calls a match, during a bad peak-season stretch, that same load volume runs closer to 2,300, without us processing one more load."
Why this works
A single point estimate implies a confidence nobody actually has about a probabilistic system.
Stage 6
Close on the one line
Say it like this
"Pay for the outcome, not for the vendor's uncertainty about how to reach it. Cap the retries, and the bill stops swinging."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

Let's learn

The laptop Callum Bright uses for vendor invoices has a cracked hinge and a spreadsheet tab that never closes. Say a freight brokerage buys an AI tool that matches loads to trucks and drivers automatically, billed by the API call it takes to find each match.

Before it, Haulwright's dispatchers matched about 1,200 loads a month by phone, roughly 20 minutes of calling per load. LoadPilot AI matches most loads in under a minute, billed at 18 cents a call, with about 6 calls per completed match under normal conditions: search, rank, confirm, recalculate the route, check for exceptions. The first few months, the bill ran close to 1,300 dollars a month, cheap against the dispatcher hours it replaced.

Hand sketched flow diagram titled How one load match actually runs. Five boxes in sequence: Search candidates, Rank matches, Confirm driver, Recalc ETA, Exception check highlighted.
Six steps, six calls, on a clean route with a driver who says yes the first time.

Here's the turn: the extra calls themselves were never the real problem. An extra call costs 18 cents. The real problem was what happened once peak season hit and drivers started cancelling more often: the model needed more attempts, more re-ranks, more re-matches, to land a good pairing. And nobody at Haulwright noticed the bill creeping, because no single month looked alarming on its own.

A peak-season month's invoice, broken into its two real parts
4,000 2,000 0 1,296 Base calls 2,424 Total: 3,720 base calls, 6 per match retry overage from re-matching
Almost two-thirds of a bad month's bill is retries. That's the model's own uncertainty, not Haulwright doing more business.

At its worst: by month six, the bill had tripled to over 4,100 dollars, on the same load volume, with nobody able to point to the day it happened, because there wasn't one.

The choice I would take back Haulwright signed a per-call price without ever modeling what happens to the call count on a bad month. That made sense at the time: the sales demo showed six calls a match, consistently, and modeling a worst case felt like borrowing trouble before there was any reason to. It stopped making sense the moment retries during peak season turned six calls into twenty-two.
Knowledge spark: why does a load-matching model retry at all? The model is searching for the best pairing of a load and a driver, and it isn't always sure on the first try. A cancelled driver, a route that no longer fits, or a low-confidence match sends it back to search again, the same way you'd redial if the first answer to your question didn't sound right.

What I would leave alone: on a normal week, at six calls a match, this pricing is honestly fine, and cheaper than the phone calls it replaced. Renegotiating the whole contract over one clean, uneventful month would waste time nobody has to spare.

The lesson: usage-based pricing on an AI tool isn't really priced on your volume. It's priced on how many tries the model needs to get it right, and that number is the one thing a vendor's demo will never show you going wrong.

Now here is the same thing as a story

The short version above is what you'd say defending this renegotiation to Haulwright's ownership. Read this one for how quietly it actually crept.

The dispatch floor at Haulwright gets loud around 4pm, when the next day's loads start posting. Callum Bright has run vendor decisions there for five years, and he reads an invoice the way some people read a weather report, looking for the line that doesn't match the season.

LoadPilot AI went live in January. The rep's demo, and the first three invoices, all landed the same way: about 1,200 loads, six calls each, 18 cents a call, a bill just north of 1,300 dollars. Callum stopped scrutinizing the invoice line by line around invoice four. It had been the same number three months running.

Hand sketched comparison titled The demo scenario vs the peak-season scenario. Left, a gauge icon labeled Demo, six calls a match, caption clean routes, willing drivers. Right, a scale icon labeled Peak season, twenty-two calls, caption cancellations force re-matching.
The vendor's demo was honest. It just only ever showed the left side.

Peak season started in April. Drivers cancelled more, routes got reassigned mid-morning, and the model started working harder to land each match, sometimes twenty tries instead of six. Nobody announced this. It just showed up, a few cents at a time, spread across thousands of calls Callum never individually saw.

By May the invoice was 1,890 dollars. By June it was 2,650. Neither jump, on its own, looked like an emergency. A vendor's usage bill goes up and down some months, everyone knows that. It was only in Haulwright's quarterly budget review, comparing six invoices side by side instead of one at a time, that anyone noticed the shape of the line.

We did not have a bad month. We had six ordinary-looking months that, laid end to end, added up to a bill that had tripled.
Hand sketched labeled parts diagram titled What's actually in the monthly invoice. A document icon at the center labeled Invoice, with four callouts: base calls, retry overage, rush surcharge, true-up credit.
LoadPilot's invoice never separated these out. Once Haulwright asked for that split, the real cost driver stopped hiding.

Callum went back through six months of invoices and split every bill into base calls against retry overage. The base number, the part tied to actual load volume, barely moved. The retry number was almost the entire increase.

Hand sketched timeline titled The range, three monthly cost scenarios. Three milestones: Best case, 1,296 dollars, highlighted. Expected, 1,890 dollars. Worst case, 2,333 dollars.
The original contract only ever priced the left end of this line. The bill lived on the right end for most of a season.

Here's what I'd take back. Haulwright signed a per-call rate based on the sales demo's six-calls-a-match number, without ever asking what a worse month would cost. That was a reasonable read of an honest demo. It stopped being reasonable the moment the season changed and nobody had a ceiling in place to catch it.

I would go back and negotiate a price per completed match, with the retries above a set ceiling billed to LoadPilot instead of Haulwright, before signing anything, not after six invoices quietly added up to three times the original number.

And the part I'd tell myself: we didn't get overcharged. We priced the wrong thing from the start, a call count instead of an outcome, and a call count was always going to move with the model's own confidence, not with our business.

BOUND, in one screenNot "guess a number and defend it." BOUND is what tells you which part of the equation you actually control, and which part you don't.

B
Break it down. State the equation first.
Monthly cost equals loads processed, times calls per load match, times price per call. Nothing gets estimated until this is on the table.
An estimate with no visible equation is a guess with a confident tone.
O
Own the numbers. State each assumption.
1,200 loads a month, 6 calls a match under normal conditions, 18 cents a call, from LoadPilot's own demo and the first three invoices.
Every number gets a source, even a rough one, instead of floating unattached.
U
Use a range, not one number.
Best case near 1,296 dollars at 6 calls a match. Worst case near 2,333 dollars at 22 calls a match during a bad peak-season stretch, same load volume both times.
A single point estimate hides exactly the uncertainty this whole question is about.
N
Nail the sanity check.
2,650 dollars a month against a route-planning department Haulwright used to staff at over 4,000 dollars a month in dispatcher hours. Even the worst case still beats the old process.
A number that fails this check means the arithmetic is wrong somewhere, before anyone signs off on it.

The recap, one line per letter: break it down is the equation itself, own the numbers is naming where 6 calls and 18 cents actually came from, use a range is best case against worst case instead of one guess, and nail the sanity check is comparing the worst case against the old dispatcher cost it replaced.

Hand sketched icon list titled What to put in a usage-based contract. Four items: price per outcome not per call, a hard retry ceiling per match, a monthly cost cap, calls-per-match tracked monthly.
This is the fix, in four lines, none of which requires the vendor to lower its base rate.

Direction. What moves the estimate most. Of every assumption in the equation, the retry rate swings the bill hardest, because it is the one number tied to the model's own uncertainty rather than Haulwright's business.

How much each assumption swings the monthly bill, at its realistic range
0 Retry rate, 6 to 22 calls +3,456 Peak season mix, 0 to 40% +1,382 Vendor price, +20% +259 Seasonal volume, +15% +194
The vendor's own price and Haulwright's own growth barely move the bill. The model's retry behavior moves it more than the other three combined.

And if you want to be sure it really works, try it somewhere elseSame four letters, an AI triage vendor billing a telehealth company per minute instead of per call. A different kind of retry.

Briarcliff Telehealth pays its AI triage vendor per minute of AI-handled call audio. Yusuf Ademola manages the vendor relationship there. Mapped onto BOUND: break it down is monthly cost equals calls handled, times AI-minutes per call, times price per minute. Own the numbers is 5,000 calls a month, 4 AI-handled minutes a call under normal conditions, 9 cents a minute, from the vendor's own reporting. Use a range is the honest split: simple, single-symptom calls average 4 minutes, but multi-symptom calls that need repeated clarifying questions can run to 14 minutes, and those make up roughly a fifth of monthly volume in a bad month. Nail the sanity check is comparing the worst-case bill, about 2,700 dollars, against the cost of a human triage nurse handling that same call volume, which would run several times higher.

The old decision here isn't an unmodeled peak season, it's a different reversal: Briarcliff signed a flat per-minute rate because the vendor's demo used only simple, single-symptom calls, and nobody modeled what happens when a caller describes three symptoms at once and the model loops through clarifying questions before it can triage confidently. That made sense when the demo was the only evidence anyone had. It stopped making sense once real multi-symptom calls turned out to be a fifth of monthly volume, not a rare exception.

Hand sketched flow diagram reused for Briarcliff Telehealth. Titled How one load match actually runs, repurposed to show a similar multi-step process with a highlighted exception check step standing in for a clarifying-question loop.
A different vendor, a different unit of work, the same shape of retry hiding inside it.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "price the outcome, cap the retries," and stop.
Cost: there's no leverage to renegotiate a signed contract mid-term. Say so honestly, and start by tracking calls-per-match monthly so the next renewal negotiation has real data behind it.
The model gets better, for real: if a vendor update genuinely cuts the average retry rate, that's the direction analysis doing its job, telling you the ceiling can loosen, not that the tracking can stop.

Where people run it wrong.
They accept the demo's average case as the number that will actually get billed.
They watch the total invoice instead of splitting it into base cost and retry overage.
They negotiate the price per call instead of the thing that actually varies, how many calls one outcome takes.

How to use it live. The moment someone says "usage-based pricing feels risky," ask back: which part of the usage is tied to volume, and which part is tied to the model's own uncertainty? Negotiate the second part separately, because it's the one that moves without your business changing at all.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits estimation and sizing questions, including pricing math?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, direction. It forces visible arithmetic instead of a confident-sounding guess.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Callum Bright, who runs vendor and procurement decisions at Haulwright, a freight brokerage moving about 1,200 loads a month.
3 · THE HABIT
What did Callum stop doing once the first three invoices matched the demo?
Tap to flip
ANSWER
Scrutinizing the invoice line by line. Three months of the same number made checking feel like wasted effort.
4 · THE EQUATION
What's the equation this whole answer is built on?
Tap to flip
ANSWER
Monthly cost equals loads processed, times calls per load match, times price per call.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Signing a per-call rate off the demo's six-calls-a-match number, without modeling a peak-season worst case first.
6 · THE NUMBER
Fill in the blank: the retry rate swinging from six to twenty-two calls per match changes the monthly bill by about ___ dollars, more than the other three assumptions combined.
Tap to flip
ANSWER
3,456 dollars, the single largest swing of any assumption in the equation.
7 · THE REPLAY
Same peak season, same driver cancellations, but the contract now caps retries at 10 calls with the rest billed to LoadPilot. What changes?
Tap to flip
ANSWER
Haulwright's bill stays close to the expected 1,890 dollars even during the bad stretch, since the overage past the ceiling lands on LoadPilot's side of the contract instead.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, using the same framework. Which product, and what's the reversal?
Tap to flip
ANSWER
Briarcliff Telehealth's per-minute AI triage billing. The reversal is signing a flat rate off a demo of only simple, single-symptom calls, never modeling multi-symptom clarifying-question loops.

Check yourself Score: 0 / 0

Short answer, state the equation
1. What is the equation behind Haulwright's monthly LoadPilot bill?
Show hint
Look at the B step in the BOUND recap.
Show answer
Model answer: Monthly cost equals loads processed, times calls per load match, times price per call.
Multiple choice
2. Why did the bill triple without any single alarming month?
  • A. LoadPilot secretly raised its per-call price.
  • B. Haulwright's load volume tripled.
  • C. The retry rate crept up gradually as peak season caused more re-matching.
  • D. Haulwright added a second vendor.
Show hint
Look at the sensitivity chart in Section 3.
Show answer
C. The retry rate swings the bill more than price or volume combined, and it moved gradually enough that no single month looked wrong.
Fill in the blank
3. Fill in the blank: the best-case monthly cost at 6 calls a match is about 1,296 dollars. The worst-case cost at 22 calls a match, same load volume, is about ___ dollars.
Show hint
Look at "The range" timeline diagram.
Show answer
2,333 dollars. Nearly double, from the same 1,200 loads, purely from the retry rate changing.
True or false
4. True or false: the fix Callum negotiated required LoadPilot to lower its base per-call price.
  • True
  • False
Show hint
Look at the direct answer and priority list.
Show answer
False. The fix was pricing per completed match and capping retries, not lowering the per-call rate.
Short answer, where it wouldn't matter
5. Name a situation where this pricing renegotiation genuinely wasn't worth pushing for.
Show hint
Look at "What I would leave alone."
Show answer
Model answer: A normal, non-peak-season month at six calls a match. The pricing is honestly fine there, and renegotiating over one clean month wastes everyone's time.
Short answer, apply it yourself
6. Think of a usage-based bill you've seen (cloud compute, an API, a subscription with overage fees). What's the one number in it that's actually tied to a model's uncertainty rather than your own usage?
Show hint
Look for the part of the bill that changes even when your own volume stays flat.
Show answer
Model answer: Anything like retries, re-generations, or extra reasoning steps, the part of the bill driven by how many tries the system needed, not by how much you actually used it.
Before you close the answer
Why this works
Tests whether you can show real arithmetic behind a pricing worry instead of just saying "negotiate a better rate," and whether you can separate the part of a usage-based bill tied to your own volume from the part tied to the model's own uncertainty.
Follow-up traps
"What if the vendor refuses to price per outcome instead of per call?" Response: negotiate a retry ceiling and a monthly cost cap instead, since both bound the same risk without requiring the vendor to change its billing model entirely.

"Isn't a retry ceiling just going to make the vendor cut corners on hard matches?" Response: no, because the ceiling is on calls billed to you, not calls the vendor is allowed to make; past the ceiling, the vendor keeps trying at its own cost.
If pressed
Haulwright's actual renegotiated ceiling landed at 12 calls per match, high enough to cover a normal bad day, low enough that a genuinely stuck match gets escalated to a human dispatcher instead of retried forever at LoadPilot's expense.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more