…
AI Product Case Questions

A model with 10x the capability at 10x the cost. What do you do?

A worked answer to a real AI PM interview question: a model has ten times the capability at ten times the cost. What do you do?

Transcript

Read the full transcript (1,433 words)

[INTERVIEWER] A model with 10x the capability at 10x the cost. What do you do? A vendor hands you a model with ten times the capability at ten times the cost. What do you do? Here's the move that wins the room. You refuse the framing. This isn't a yes or a no. You don't adopt or reject a model, you route it.

The premium model earns its cost only where ten times the capability actually changes the outcome, and everywhere else the cheap model already wins. The interviewer is testing whether you think about value and cost as flat, uniform things, or whether you break them down by task and by segment. That's the whole test. By the end of this, you'll be able to reframe the question onto value per task, lay out three real options, commit to a routing strategy, and back it with actual unit economics.

Start by framing the decision, because asking if we should use it is the wrong question. Cost and value aren't uniform across tasks or users. The right question is which tasks and which segments see ten times the capability produce more than ten times the value. Everywhere else, the cheaper model is the correct choice, and paying ten times for it is just burning margin.

So reframe onto value per task, not price. A model is worth its cost when the marginal quality lifts revenue, retention, or risk avoided by more than the marginal spend. Think about a low stakes autocomplete. Going from ninety five percent correct to ninety nine percent is barely worth anything, because nobody is harmed by a slightly worse suggestion. Now think about a legal contract review or a medical triage summary.

There, the last few points of accuracy are the entire value, and ten times the cost is cheap against the downside of one bad miss. Same capability jump, completely different worth, and that's the insight the interviewer wants. Now lay out the options honestly. Option A is the full switch. Everyone gets the premium model. It's simple, but you burn margin on the majority of traffic that never needed it, and your unit economics break at scale.

Option B is never adopt. You save money, but you hand the high value, high willingness to pay segment straight to a competitor who did adopt. Option C is tier and route. A cheap model handles the bulk of queries, and a router escalates to the premium model on the high stakes ones, or you package the premium model as a paid tier.

Three real options, each with a genuine downside, so the interviewer sees you're weighing, not guessing. Then recommend and commit, because fence sitting loses. Take Option C. Build a router that classifies each request by stakes and complexity, sends the easy traffic to the cheap model, and reserves the premium model for the hard or high value cases. In parallel, package the premium model as an enterprise tier aimed at buyers whose cost of error is high, because that's exactly the segment whose willingness to pay covers ten times the cost.

You're not choosing between margin and quality. You're aiming each at the traffic where it belongs. Finally, name the risks and the moat. The first risk is routing errors. You send a hard query to the cheap model and ship a bad answer, or you waste premium spend on something trivial. Mitigate it with a confidence check and a cascade.

Try the cheap model first, and escalate on low confidence. The second risk is that the cost gap closes fast. Inference prices fall every few months, so don't rebuild your entire product around today's ten times number. And the moat? It isn't the model. Anyone can rent the same model you can. The moat is the routing logic and the eval data that tells you which query needs which model, tuned to your own traffic.

That's the piece a competitor can't copy, because it's built on your data, not theirs. Let me put numbers on it. Picture a customer support assistant handling five million tickets a month. The cheap model resolves ninety percent of them, like returns, order status, and password resets, at a fraction of a cent each. A stakes classifier routes the other ten percent, the billing disputes, the cancellations, and the legal tone complaints, to the premium model.

Here, a better answer stops a churned account worth four hundred pounds in lifetime value. So your blended cost per ticket stays low, but the resolution quality on the cases that actually matter goes up. Then you sell the premium everything version to enterprise clients whose brand risk justifies it. The routing table, trained on your own ticket outcomes, is the part a competitor cannot copy.

Here's a second quick one to show the pattern holds. A coding assistant sends routine autocomplete to a small fast model, and only escalates a request to refactor a whole module to the expensive one, because that's where a wrong answer costs a developer an hour. Same logic, different product. Match the spend to the stakes. The interviewer will almost certainly press you on the router itself, because that's where the real design lives.

The first question is how the router decides and what it costs. You don't want a router that's itself a huge model, because then you're paying premium prices just to decide not to pay premium prices. So the classifier is small and cheap, a lightweight model or even a set of rules on features you already have, like query length, presence of certain keywords, the customer's tier, or the dollar value attached to the outcome.

Second, they'll ask what happens when the router is wrong. That's why you build a cascade, not a one shot decision. Try the cheap model first, have it return a confidence score or a self check, and escalate to the premium model only when confidence is low or the stakes are high. The cascade means a routing miss degrades gracefully into a second attempt, rather than shipping a bad answer with full confidence.

Third, and this is the one that separates a strong candidate, they'll ask how the economics actually pencil out. So do the sum out loud. If the cheap model is a tenth of a cent per call and the premium is a full cent, and you route ninety percent cheap, your blended cost is roughly nought point two of a cent, not the full cent you'd pay on a full switch.

That's a fivefold saving on inference, and you barely touched quality on the cases that matter. Then there's the pricing question they slip in at the end. Should the premium tier be a subscription or usage based? For an enterprise buyer with a high cost of error, a flat enterprise seat wins, because they want budget certainty and they'll happily overpay for the peace of mind.

That is exactly the willingness to pay the premium model was built to capture. Here's what makes them lean in. First, you rejected the binary and reasoned about value per segment and per task, which is the exact judgement the question is designed to probe. Second, you brought real unit economics into it, like blended cost, willingness to pay, and cost of error, rather than just vibes.

And third, you designed a concrete mechanism, a router with a confidence cascade, because a concrete mechanism always beats a platitude about using the right tool for the job. Now for the traps. The first is answering yes if it's better or no if it's too expensive, as if cost and value are flat across everything. That's the average answer and it dies immediately.

The second is forgetting that different segments have wildly different willingness to pay, so you miss the enterprise tier entirely. The third is betting the whole product on today's price gap, when inference costs drop every few months and your ten times assumption is stale by next year. So, to pull it together. Don't adopt or reject the model, route it.

Reframe onto value per task. Lay out full switch, never adopt, and tier and route, then commit to tier and route with a confidence cascade and an enterprise tier for buyers with a high cost of error. Mitigate routing errors, don't over fit to today's price, and remember the moat is your routing data, not the model. The one line to carry in is that ten times the cost is a bargain only where ten times the capability changes the outcome, so route it, don't rule on it.

Keep learning