Describe how enterprise procurement changes AI pricing conversations.
A fraud-scoring tool can close its first real deal in one call, off one story about dollars saved. The moment that deal gets big enough to need procurement, security, and legal in the room, the story stops being enough, and somebody has to answer a much harder question live. This is about who that somebody is, and what you'd have needed to build months earlier so it isn't the founder, on every call, for six straight weeks.
- Build the evidence packet before procurement asks for it, not after.Why: by the time a security architect asks the question live, in front of the CFO and General Counsel, it's too late to go build the document you needed an hour earlier.
- Break the false positive and false negative rate out by channel, never ship one blended number as the whole story.Why: a blended number is exactly what a security review cracks open first, and by then it's a live seven-figure deal, not a slide.
- Pin the model version in the contract, with a required re-check before any swap touches live traffic.Why: this is the one term that turns "trust us" into something legal can actually put a signature under.
- Only join a deal personally for the final SLA line-item call, not the whole review.Why: reclaiming the whole deal from the champion is what actually costs you, their credibility and your calendar, not just hours.
- Don't force the packet-and-founder ritual onto every small pilot under the procurement line.Why: it would slow down champions who don't need it yet, for a threat that hasn't shown up in that room.
- Don't quote one blended "it's accurate" number to close faster.Why: an actuarial reviewer finds the segment gap eventually. Better to find it before signature than after, in production, off a support queue.
How to answer this, stage by stage
Nobody is grading whether you can list three things procurement checks, price, security, SLAs. They're grading whether you know exactly who has to walk into the room once a value story stops being enough, and what you'd have built beforehand so it isn't the founder, alone, for six weeks.
Let's learn
What happens the first time a founder gets asked to prove a number instead of just say it?
Ferrowatch is an API a payment processor calls on every transaction. Send it a purchase, and it sends back a score from 0 to 100 and a flag: approve, hold for a person, or decline.
Before Ferrowatch, Yarrowcross Pay ran its own rule book: if the amount was over a set limit and the shipping address didn't match the billing one, flag it. A team of six analysts reviewed about 3.4 percent of transactions by hand, roughly four minutes each, and caught about 61 percent of real fraud dollars that way.
With Ferrowatch running the pilot, only 1.1 percent of transactions needed a human at all, and chargeback dollar losses fell 38 percent in the first quarter. Tobenna Kanda closed that pilot off one 45-minute call with Julio Petracca, Yarrowcross's VP of Payments Risk: about 550,000 transactions a month across 40 stores, priced flat at six tenths of a cent per transaction scored, live within three weeks. No procurement, no security review. Julio had the budget authority to say yes himself.
Here's the turn. Nine months later, Yarrowcross wanted Ferrowatch across their full national volume, about 19 million transactions a month, 900 stores plus mobile and online ordering, worth close to $1.4 million a year at the same flat rate. That number crossed Yarrowcross's own procurement policy: any vendor contract over $250,000 a year goes through procurement, security, and legal, no exceptions. The extra scrutiny was never really the problem. What mattered was what Julio did once it arrived. He couldn't answer it. So Tobenna had to.
At its worst, every enterprise deal past that $250,000 line pulls the founder personally onto every call for weeks. Deals that used to close off one email start taking months. And the champion who brought you in stops being trusted to run the next one alone, which costs you something no dashboard tracks.
What I would leave alone: the flat $0.006-per-transaction sticker price itself, and the whole one-call pitch, for any deal still under Yarrowcross's own $250,000 line. Small pilots don't need a model-version-pinning clause. Building one for them anyway would slow down exactly the champions who don't have a procurement gate to clear.
The lesson: a pricing conversation was never really about the price. It's about who in the room can prove the number, and whether you built them anything to prove it with before they needed it.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a champion who'd never lost a deal in his life went quiet on a call in front of his own CFO.
Tobenna Kanda could read a term sheet the way a line cook reads a ticket rail: fast, in order, no second-guessing. Two years running pricing and packaging at Ferrowatch, and she'd never once needed to sit in on a deal Julio Petracca was closing. He didn't need the hand-holding. He closed the Yarrowcross pilot off a single 45-minute call, no deck, just the fraud numbers and a demo, and he told his own boss about it after the fact, not before.
For eight months, that was the whole relationship. Julio would call, ask a follow-up question, get an answer within the hour, and go run his own deal his own way. Tobenna barely thought about Yarrowcross between quarterly check-ins. The pilot worked, chargebacks fell, and Julio started quietly pitching the full national rollout to his own leadership, using nothing but the numbers from that first pilot deck.
It thinned in three beats, and none of them looked like trouble at the time. Weeks one through eight of the formal review, Julio fielded everything: the vendor risk questionnaire, the SOC 2 request, a round of questions about uptime. Week nine, he started forwarding a couple of the more technical questions to Tobenna's team, "just to be safe," but he still ran every call himself. By week eleven, a real security ticket had opened with Yarrowcross's own architect assigned to it, and Julio stopped answering anything at all. He just forwarded it, all of it, and waited.
The trigger wasn't two mistakes in a row. It was one written question, dropped into a shared review doc where Yarrowcross's CFO and General Counsel could both see it land: "Your dashboard shows a 2.1 percent false positive rate overall. Break that out by transaction channel, and tell us whether that number is guaranteed to hold if you retrain the underlying model mid-contract." Julio read it twice. He had never once seen Ferrowatch's accuracy broken out by channel. Nobody had ever shown it to him that way, because nobody had ever asked before.
Julio forwarded the thread to Tobenna at 6:40 that evening with one line: "I need you on the next call." That was the flip. Not a request for help on one question. A handover of the entire room. From that call forward, for five weeks, Tobenna sat in on every session personally, defending the model number by number: why in-store swipe ran a 0.9 percent false positive rate but mobile wallet ran 3.4, because mobile wallet was the newest payment rail Ferrowatch supported and had the thinnest training data of the three; why that number wasn't a fixed guarantee but a range, re-checked against a real eval set every quarter; and, line by line with Yarrowcross's own counsel, what a model-version-pinning clause would actually say.
Fourteen months earlier, when Ferrowatch's two-person go-to-market team built the pricing page and the one-slide ROI deck Julio had been using ever since, somebody asked, in a fifteen-minute meeting, whether they should also put together a fuller technical write-up, error rates, methodology, that kind of thing. The answer was no. Nobody had asked for one yet, and every deal so far had closed off the story alone. It was the sensible call. The company's biggest customer, at the time, was a single pilot doing 500,000 transactions a month.
Run the same nine weeks again, with the packet already built. Julio gets Ferrowatch's evidence packet the day the formal review opens: false positive and false negative rate by channel, the date of the last independent check, a plain-language paragraph on what a model-version-pinning clause is and why it exists. He answers the security architect's first two rounds of questions himself, straight off the document, no forwarding. He only pulls Tobenna in once, at week three, for a single clarifying call, about two hours. Tobenna joins one more time, at week seven, for the final SLA line-item negotiation, four hours across two calls. The deal signs by week eight. Tobenna's total time on it: six hours, not forty. Julio runs the whole review, keeps his own leadership's trust, and closes the deal he originally brought in the door, the way he closed the last five under his own name.
One design left a champion to defend a number nobody had ever shown him. The other handed him the number months before he needed it.
What I'd tell myself, back in that fifteen-minute meeting: skipping the write-up wasn't wrong for the product that existed then. It was wrong for the product everyone already knew they wanted to sell into accounts fourteen times the size of that first pilot. Ask how big a real deal gets before deciding a number never needs to be written down.
The five steps, if you want to remember it
Not a list of things a security review checks. FLIPS names the exact moment a champion who'd never lost a deal ran out of runway, and asks which old choice made that moment the only option on the table.
Three things worth stating directly, since this is where the real judgment sits. The alternative Ferrowatch's team considered, and rejected, was quoting one blended "it's over 98 percent accurate" number to close the expansion faster instead of breaking it out by channel. It lost because Yarrowcross's own actuarial reviewer would have asked for the channel split eventually, and finding the gap after signature, once mobile wallet's real number showed up in production, would have broken the SLA and the relationship, not just delayed the deal by a few weeks. The AI-specific failure worth naming by name is silent drift after a routine retrain: mobile wallet runs on the thinnest training data of the three channels, so any retrain moves that number first and moves it hardest, and Yarrowcross would only find out from a spike of angry calls about wrongly declined purchases. The guardrail is the model-version-pinning clause itself, no version swap touches Yarrowcross's live traffic without a shadow evaluation and a customer notice first. And the trade-off is real: pinning the version means Ferrowatch keeps serving a slightly older, slightly more expensive model on this one account for the length of the contract term, instead of rolling every account onto whatever's newest and cheapest. Yarrowcross's predictability costs Ferrowatch some of its own margin, accepted on purpose.
And if you want to be sure it really works, try it somewhere else
Same five letters, a county fire-rescue dispatch line instead of a checkout, and this time nobody argues about a false positive. They argue about which calls even get the model's attention.
Signalkeep is Anzhela Rioux's dispatch tool. It listens on a live 911 line, transcribes it in real time, and flags the words that mean somebody's in real trouble, chest pain, structure fire, child not breathing, so a dispatcher juggling four calls at once knows which one to answer first.
Rushcliffe County Fire-Rescue runs the whole county's 911 line out of one room, and its board can only approve a fixed annual figure. Procurement forced Signalkeep into a flat license: one price, one year, covering up to 240,000 calls. Once that cap was real, Rushcliffe's dispatchers didn't ration by day or by shift. They rationed by which calls felt like they mattered, running Signalkeep on structure fires and cardiac arrests, skipping it on routine calls and calm callers who could describe their own address. That meant spending their capped call-credits on exactly the audio Signalkeep transcribes worst: three people shouting over sirens is a different problem than one calm caller describing a fender bender. On calm, single-caller audio, Signalkeep's word error rate runs about 4 percent. On chaotic, multi-caller audio it runs closer to 22 percent. Before the cap, chaotic calls were about 18 percent of what Signalkeep processed. After it, that share rose to 54 percent, the same model, a different mix of calls being fed to it, and the flagged-miss rate rose from 6 percent to 15 percent without one line of the model changing.
Same rank, different lever: the fix here isn't a bigger cap or a better model. It's a flat price with a channel-mix warning built into the dashboard, one that fires the moment "urgent-only" rationing starts pulling the hardest audio toward the tool instead of away from it, before the missed-flag rate becomes a story a county commissioner reads about after the fact.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: cap the price, not the usage. If usage has to be capped at all, cap it with a channel-mix alarm, not a hard stop.
Cost: there's no engineering budget this quarter for a live channel-mix dashboard. Ship the cheap version first, a monthly manual check of what share of processed calls were chaotic audio, one query, not a live feed.
The model got better, for real: say Signalkeep's accuracy on chaotic audio doubles overnight. The fix barely changes. A capped, metered product still trains people to ration toward the hardest cases first. A better model changes how bad the rationed mix turns out, not whether the rationing happens at all.
Where people run it wrong.
They let a budget process pick the pricing shape instead of the usage pattern, and end up with a hard stop nobody actually designed for.
They measure accuracy in aggregate and never split it by which calls actually got run, so the drop reads as the model failing instead of the input changing under it.
They fix it by asking for a bigger cap, which buys a few more months before the same rationing comes right back.
How to use it live. Say the split out loud before answering: "is procurement really asking for a fixed price, or a fixed amount of usage, because those force two very different rationing problems." That buys a beat, and shows the interviewer you know a budget constraint and a usage constraint aren't the same fix.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a model-version-pinning clause just slowing down your own product improvements?" Response: only for that one account's contract term. Ferrowatch still ships new versions everywhere else, and re-validates against Yarrowcross's own traffic before the pin ever moves.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pricing AI products: seat, usage, outcome
- #1 Compare seat-based, usage-based and outcome-based pricing for an AI product.
- #2 Why does seat-based pricing break when AI reduces the number of seats needed?
- #3 Design a pricing model for an AI feature with high variable cost and unpredictable usage.
- #4 What is the risk of usage-based pricing from the customer's point of view?
- #5 Explain how credits work as a pricing mechanism and their advantages.
- #6 How would you price an agent that completes a task rather than answers a question?