CalculationIntermediateAI Opportunity & Model Strategy / Model selection from a PM lens / #4

How do you weigh a model that is 20 percent better and three times more expensive?

The direct answer
Split the work before you buy anything. Pay the 3x only where a mistake stays invisible, which here means clause extraction and risk flagging, and buy on price everywhere a lawyer reads the output anyway. Route by what an error costs, not by average quality.
Where the money goes, in order
  1. Split the work by what a mistake costs, then buy one model per job.Why: this product does two jobs with completely different stakes. One buying decision for both of them is the actual error, and it happens before anyone looks at a price.
  2. Get the absolute dollar gap before you argue about the multiple.Why: 3x here is 18,000 dollars a year. That is about 36 hours of partner time. A six-week bake-off costs more than the decision.
  3. Put the expensive model on clause extraction and risk flagging.Why: 5,400 dollars a year buys about 180 fewer missed clauses. Thirty dollars a miss, against a single miss that cost this firm 30,000 dollars in written-off partner time.
  4. Keep the cheap model on summaries and the first sort by contract type.Why: an associate opens the contract anyway before signing off. A nicer summary changes nothing they do next, so the extra 12,600 dollars buys a paragraph nobody keeps.
  5. Make the vendor re-measure the 20 percent on the four clause types that can cost real money.Why: an average across all clause types can hide a big gain on notice periods and no gain at all on indemnity. Averages are the wrong unit when one clause is worth 200 of the others.
  6. Run both models on extraction and keep every flag either one raises.Why: a junk flag costs 90 seconds and a miss costs years, so taking the union of both catches more than either alone, at about 1.3x, not 3x.

How to answer this, stage by stage

Six moves. This is a commit question, so the whole answer is one position with a price on it. Say the words, don't read them.

1
Split the question before you answer it
Say it like this
"Before I pick, I need two things the question doesn't give me. Twenty percent better at which job, and what's 3x in actual dollars a year. If it's 20 percent better at writing summaries I don't want it at any price, and if the gap is 18,000 a year, we're spending more on this meeting than on the decision."
Why this works
Most people answer yes or no. Both are wrong, because "20 percent better" is not a fact yet, it's a headline with the units cut off. Asking these two things in fifteen seconds says you have bought a model before.
2
Turn the multiple into money, out loud
Say it like this
"Let's price it. We run about 9,000 contracts a year. The cheap model is a dollar a contract all in, the good one is three. So that's 9,000 a year against 27,000. The whole argument is 18,000 dollars. That's about 36 hours of one partner's time."
Why this works
3x sounds huge and is usually a rounding error next to people cost. The absolute number ends the percentage argument in one line. It also flips the risk: being wrong now costs more than the model does.
3
Sort the work by who sees the mistake
Say it like this
"Now split the work. The tool does two jobs. It pulls out clauses and flags risk, and it writes a plain-English summary for the file. On the summary, an associate reads the contract anyway before they sign off, so a bad line gets caught in a minute. On extraction, if it misses an indemnity, nobody opens that file again for two years."
Why this works
This is what the question is really testing. Average quality is the wrong unit. Consequence of error is the right one, and once you say it, the routing answers itself in front of the interviewer.
4
Commit to a route, and put numbers on both halves
Say it like this
"So here's my call. Expensive model on extraction, cheap model on summaries. Extraction is the smaller job, so upgrading it runs about 5,400 a year and buys roughly 180 fewer missed clauses. That's 30 dollars a miss. Upgrading summaries costs 12,600 and changes nothing anyone does."
Why this works
PICK is testing whether you will commit. A route is a commitment with a price attached, and it is much harder to knock down than a yes or a no, because you have already costed both sides.
5
Price the worst miss you have actually had
Say it like this
"And here's why the extraction side isn't close. Last year we missed one auto-renewal clause on a facilities contract. The client got locked in for another year at 240,000, and we wrote off 60 hours of partner time cleaning it up. That's 30,000 dollars of our own money. That one clause cost us more than running the good model on every contract in the firm would cost for a year."
Why this works
One priced failure beats any amount of reasoning about risk. It also states the asymmetry in the only unit the room cares about, which is money, not model scores.
6
Say out loud what would change your mind
Say it like this
"Two things would flip me. If the 20 percent turns out to be an average across all clause types, and the gain on indemnity and change of control is basically zero, I buy the cheap one everywhere. Same if the gain is on precision instead, fewer junk flags. Junk flags cost us 90 seconds. Misses cost us years."
Why this works
A pick with no kill criteria reads as stubborn. Naming the exact evidence that would reverse you is what makes a commitment sound earned instead of lucky.
The move most people miss You are not limited to buying one. On extraction, run both models and keep every flag either one raises. Two models catch more than the better one alone, junk flags get deleted in 90 seconds, and the bill lands near 1.3x instead of 3x. Say this after you have committed, not instead of committing.

Let's learn

The tool is one button inside the document system a law firm already had open all day. It reads a commercial contract and does two separate jobs with it.

Contract in, cheap sort, costly extract, cheap summary, lawyer signs
One contract, three model calls, and only one of them needs the good model

Job one: pull out every clause with money or risk attached. Indemnity. Liability caps. Auto-renewal. Change of control. Job two: write a plain-English summary for the file, so a partner can read the shape of the deal in ninety seconds.

Before the tool, a second-year associate read each contract and built that list by hand. About 50 minutes a contract. Now the associate checks the tool's flags instead. About 15 minutes. The firm runs 9,000 contracts a year, so that is real time.

Then the vendor calls. There is a better model. Twenty percent better. Three times the price.

Per contract, both jobs
cheap
costly
Clause extraction and risk flags
$0.30
$0.90
Plain-English summary
$0.70
$2.10
Per contract, all in
$1.00
$3.00
x 9,000 contracts a year
$9,000
$27,000
The multiple is 3x. The gap is 18,000 dollars a year. Both statements are true, and only one of them tells you how hard to think about this.

Here is where most buying decisions go wrong. Someone asks "is 20 percent worth 3x" as one question, about one product, with one answer. It is not one question. It is two, because the tool does two jobs, and a mistake in each one is caught by a completely different person at a completely different price.

What each route costs a year
Clause extraction
Summaries
Cheap everywhere
$9,000
Split (my pick)
$14,400
Costly everywhere
$27,000
The middle row buys all of the safety for about a third of the extra money. The bottom row spends 12,600 more on summaries that a lawyer reads over anyway.

Take the two jobs one at a time.

Extraction. On a hundred contracts the cheap model misses about ten clauses that matter. The costly one misses eight. Two fewer per hundred, so across 9,000 contracts that is roughly 180 fewer misses a year. Upgrading only this job costs 5,400 dollars. Thirty dollars to stop one miss.

Summaries. The associate reads the contract regardless, because their name goes on the advice. A better summary saves them a few seconds and gets rewritten in their own words anyway. Upgrading this job costs 12,600 dollars and buys a habit change that never happens.

Twenty percent better is not a fact. It is a headline with the units cut off, and the units are the whole answer.
Knowledge spark: the two ways a model can be "better" It can catch more of the real stuff (fewer misses), or it can raise fewer false alarms (less junk). Vendors quote one number for both. They are not worth the same money. Fewer misses is worth a lot here, because a missed clause is invisible. Less junk is worth almost nothing, because junk is deleted in seconds.
A decision tree: who catches a mistake here, cheap model, pay the 3x, or test 200 files
One question decides the route
What I would leave alone The first sort, where the tool decides whether a document is a lease, an NDA or a supply agreement. It gets that right nearly all the time, and when it is wrong the associate sees it in the first line and fixes it. Twenty percent better at a job that is already visible and already right is 12,600 dollars for nothing. Do not upgrade a job whose mistakes have never cost you anything.

The week that made the decision easy

You do not need this to answer the question. Read it if you want to feel why the routing beats the average.

Ravi has run legal operations at a seventy-lawyer firm for nine years. He is not a lawyer. He is the person who knows the corporate team's contracts pile up on Thursdays, that the fee earners hate the intake form, and that Priya, second year, clears a stack faster than anyone without cutting a corner to do it.

The tool arrived in the spring and it was good. Not magic. Good. Review time went from about 50 minutes a contract to about 15. Priya stopped staying late on Thursdays. Nobody wrote a case study about it. It just quietly worked, month after month, which is the best thing software ever does.

Then two mistakes happened in the same week, and only one of them was ever noticed.

On the Tuesday, the tool threw eleven flags on a straightforward office lease. Nine of them were nothing. Priya read the nine, deleted them, and moved on. Under two minutes. She did not report it, because there was nothing to report. The tool was noisy and she handled it.

A junk flag seen at once versus a missed clause seen never
Same tool, same week, two very different bills

The other mistake was in March, on a facilities contract, and nobody found it until January. The auto-renewal sat on page 31, worded oddly, buried in a schedule. The tool did not flag it. Priya reviewed the flags it did raise and every one of them was correct. Nothing on her screen was wrong. There was no moment where a careful person could have caught it.

The client meant to leave that supplier. Instead they renewed automatically for another year at 240,000 dollars. The firm wrote off 60 hours of partner time sorting it out. At their rates, that is about 30,000 dollars of their own money.

The mistakes we could see cost us two minutes. The one we could not see cost sixty hours and a client's whole year.

So when the vendor rang about the better model, Ravi did not ask whether 20 percent was worth 3x. He asked which clause types the 20 percent was measured on. The vendor sent the sheet. Big gains on notice periods and governing law. Almost nothing on auto-renewal, which is the exact clause that had just cost them 60 hours.

He said no. Then he sent the vendor 200 of the firm's own contracts, already reviewed, answers known, and asked them to score only the four clause types that can actually cost money. On those four, the gain was real. He bought it that month, for extraction only, and left the summaries on the cheap model where they still are.

The part Ravi would take back is not the purchase. It is the month he spent comparing two models as if they were two products. They were never two products. They were one price list for four different jobs, and the last two years of write-offs already said which job deserved the money. He could have answered it in an afternoon, from his own files, without a single call.

PICK, and what each letter decided

This is a tradeoff question, so the framework is PICK. A "what if the error rate doubled" question would use FLIPS, and an "estimate the ROI" question would use BOUND. Different question shapes, different tools.

PICK: pick a side, name who pays, find the cost gap, say what would flip you
PICK, for tradeoff questions
P, position. Costly model on clause extraction, cheap model on summaries and the first sort. Said in one sentence, before any reasoning, because "it depends" fails this question.
I, impact. A junk flag is paid by the reviewer, in 90 seconds, and everyone can see it happened. A missed clause is paid by the client, two years later, at 240,000 dollars to them plus 30,000 in written-off partner time to us, and nobody sees anything at all.
C, cost asymmetry. Thirty dollars stops one miss on extraction. Twelve thousand six hundred a year makes a summary nicer that gets rewritten anyway. Those two are not close, and they point in opposite directions, which is what makes the split the answer instead of a compromise.
K, kill criteria. If the 20 percent is an average hiding no gain on the money clauses, buy cheap everywhere. If the gain is fewer junk flags rather than fewer misses, buy cheap everywhere. If contract volume goes up twenty times, redo the arithmetic, because 18,000 dollars stops being small.
Knowledge spark: the PICK test If both options in your answer cost about the same, you have not found the asymmetry yet, and there is no real tradeoff to argue about. Here one side costs 90 seconds and the other costs 60 partner hours. That gap is the answer.

Run PICK on a factory floor

A food manufacturer uses a model to read its own packaging artwork before print. Same question lands: a better checker, 20 percent better, 3x the price. Different industry, same shape.

P. Buy the expensive model for the allergen and ingredient lines only. Keep the cheap one on marketing copy, spelling and layout.
I. A wrong marketing word is caught by the brand team the same afternoon, and costs one email. A missed allergen goes on 400,000 packs, and costs a recall plus a phone call from the regulator.
C. The allergen check is a few lines of text per pack, so it is the cheapest thing on the page to upgrade. The marketing copy is most of the words, so it is most of the bill, and it is the part three humans read anyway.
K. If the packs go through a second allergen check downstream by a person whose only job is that, the invisible mistake stops being invisible, and the cheap model is fine on both.

Swap the trigger and it still runs

  • Speed: the better model takes 40 seconds instead of 4. Fine on extraction, which runs overnight. Not fine on the summary a partner is waiting for, so the route stays the same for a second reason.
  • Cost: the gap is 30x, not 3x. Now the absolute number is real money, so you route by contract value too: good model on anything over a million, cheap model under it.
  • The cheap model gets better: the gap closes to 4 percent. Drop back down, but only after re-scoring on the four money clauses. An average that moved 4 percent can still hide a big move on the clause that matters.

Where people run it wrong

  • Answering yes or no. Both lose. There is no single answer until you split the work, and the split is the thing being graded.
  • Arguing the multiple with no absolute number. 3x on 9,000 dollars and 3x on 9 million dollars are not the same question, and one of them does not deserve a meeting.
  • Routing by how hard a job looks instead of what a mistake costs. Summarising 90 pages looks harder than spotting one clause. It is not the one that can cost somebody a year.

If you are asked this cold

Say the split out loud before you know the answer. "Give me a second, I want to break this by job first." That buys you ten seconds and it is already stage 1, so you are not stalling, you are starting. Then name a concrete product, even if the interviewer did not give you one.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what is its hard step?
Tap to flip
ANSWER
PICK, for tradeoffs. The hard step is C, the cost asymmetry: one mistake is cheap and visible, the other is hidden and expensive. Optimise against the hidden one.
2 · THE MISSING WORDS
What two things does "20 percent better and 3x the price" leave out?
Tap to flip
ANSWER
Better at which job, and 3x of what absolute number. Without those, the question cannot be answered, only guessed at.
3 · THE PERSON
Who is this answer about, and what did he do that a buyer usually skips?
Tap to flip
ANSWER
Ravi, legal operations lead at a seventy-lawyer firm, nine years in. He asked which clause types the 20 percent was measured on, then re-scored it on 200 of the firm's own contracts.
4 · THE ASYMMETRY
Name the two kinds of mistake here and what each one costs.
Tap to flip
ANSWER
A junk flag: seen at once, 90 seconds to delete. A missed clause: seen never, 240,000 dollars to the client and 60 written-off partner hours, about 30,000, to the firm.
5 · THE POSITION
State the pick in one sentence, the way you would say it out loud.
Tap to flip
ANSWER
Expensive model on clause extraction, cheap model on summaries and the first sort. Route by what an error costs, not by average quality.
6 · THE NUMBER
Three times the price, on 9,000 contracts a year, is a gap of ______ dollars a year.
Tap to flip
ANSWER
18,000. That is 9,000 against 27,000, or about 36 hours of partner time. Small enough that the cost of deciding slowly is bigger than the decision.
7 · THE KILL CRITERIA
What evidence would flip this pick to "cheap model everywhere"?
Tap to flip
ANSWER
If the 20 percent is an average with no real gain on the money clauses, or if the gain is fewer junk flags rather than fewer misses. Junk is cheap. Misses are not.
8 · THE TRANSFER
Which product does the answer run PICK on second, and what is the route there?
Tap to flip
ANSWER
A food manufacturer's packaging checker. Expensive model on the allergen and ingredient lines, cheap model on marketing copy, spelling and layout.

Check yourself Score: 0 / 0

True or false
1. True or false: a model that is 20 percent better will cut missed clauses by about 20 percent on every clause type.
  • True
  • False
Show hint
Think about what an average is made of.
Show answer
False. The 20 percent is an average across all clause types. In Ravi's case the vendor's own sheet showed big gains on notice periods and governing law, and almost nothing on auto-renewal, which was the exact clause that had just cost the firm 60 partner hours. Ask which clauses moved before you pay for the average.
Multiple choice
2. A vendor tells you their new model is 20 percent better and 3x the price. What do you say first?
  • A. "Quality matters in legal work, so let's take it."
  • B. "Better at which job, and what's 3x in dollars a year?"
  • C. "Let's run a six-week bake-off against the current one."
  • D. "What does it score on the public benchmark?"
Show hint
One of these turns a headline into a decision. The others accept the headline.
Show answer
B. A commits before knowing anything. C spends more on deciding than the decision is worth, once you learn the gap is 18,000 dollars a year. D swaps one average for another. Only B gets the two facts the question is missing.
Multiple choice
3. Which job gets the expensive model, and why?
  • A. Summaries, because that is what the partner and the client actually read.
  • B. Clause extraction, because a miss there is invisible until it is expensive.
  • C. Both, because quality in one job feeds the other.
  • D. Neither, until the model is right more than 99 percent of the time.
Show hint
Ask who catches the mistake, and how long it takes them.
Show answer
B. A sounds right and is backwards: a lawyer reads the contract behind the summary anyway, so a weak summary is caught in a minute. C is the buying mistake the whole answer is about, and costs 12,600 dollars a year extra for nothing. D is a dial, not a decision, and waits forever.
Fill in the blank
4. The rule this answer runs on: route by ______, not by average quality.
Show hint
Two of the words are "a mistake" and "costs".
Show answer
What a mistake costs, and who catches it. Average quality treats every job in the product as if its errors were worth the same money. They never are. That single swap turns one unanswerable question into two easy ones.
Short answer
5. Redo it for a small firm doing 600 contracts a year instead of 9,000. Does the pick change?
Show hint
Run the same per-contract prices at the new volume, then ask what the split itself costs to build and run.
Show answer
Model answer: At 600 contracts, cheap everywhere is 600 dollars a year, the split is 960, and costly everywhere is 1,800. Splitting saves 840 dollars a year. That is less than a day of anyone's time to build and maintain the routing, so at this size you stop splitting and just buy the good model for everything. Same principle, opposite action: get the absolute number first, and let it decide how much cleverness the problem deserves.
Short answer, apply it yourself
6. Pick a tool you use yourself. Name one job inside it where a mistake is caught at once, and one where a mistake would never be seen. Which one deserves the better model?
Show hint
Look for a job where the output you never see is the one that matters.
Show answer
Model answer: "Take a spam filter. Junk that lands in my inbox is caught at once and costs me one second to delete. A real email sent to the spam folder is never seen at all, and I find out weeks later when someone asks why I ignored them. So every dollar of extra quality belongs on not losing real mail, and the junk side can stay cheap and noisy." Any answer works if you can name who catches each mistake and how long it takes them.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more