ConceptFoundationalModel Fluency & the AI PM Role / The AI literacy baseline every PM needs / #12
Explain the tradeoff between a larger model and a smaller one in plain product terms.
PICK · Copperpot Kitchens lets its cheap model guess at a recipe conversion it was never built to solve, and the swap that looked like savings made the sandwich pricier
Copperpot Kitchens runs nineteen fast casual kitchens, and PriceThyme is the tool that costs out every one of its hundred and forty recipes against live ingredient prices. Zenzo Kalinowski runs menu engineering there, and PriceThyme actually does two very different jobs: repricing a known recipe against a new price, thousands of times a day, and working out a substitute when an ingredient's price jumps or disappears, maybe forty times a month. His team pointed both jobs at the same cheap model to save money. One of those jobs was never built for a model that cheap.
The direct answer
Send the easy, high volume jobs to the small model and the rare, hard ones to the large model, even though the large model costs more per call. The extra spent on a rare hard call is pocket change. Letting the small model guess on that same call is the mistake that quietly costs real money and trust.
Do this, in order
Send routine, high volume calls to the small model and rare, hard calls to the large one, even though it costs more per call.Why: the two jobs need different kinds of thinking, and the extra spent on a rare hard call is nothing next to what a wrong guess costs.
Never let the small model touch a ratio or unit conversion on its own.Why: that is exactly the kind of case it answers wrong with total confidence, no warning, no flag.
Set a dollar threshold that forces a check on any substitution above it.Why: not every swap needs the large model, only the ones where being wrong actually costs something.
Watch food cost by dish, every week, not just the whole menu once a month.Why: the drift was visible for weeks before anyone looked, twenty eight percent crept to thirty four percent in plain sight.
Do not overcorrect by routing every call to the large model.Why: that trades a rare, expensive mistake for a certain, everyday one. Here it would have cost about eighteen thousand seven hundred dollars a year for zero benefit on jobs a calculator could already do.
Ask one question before every model choice: does this call involve a ratio or a conversion a person would have to think through?Why: that is the kill test. Yes earns the expensive model. No means the cheap one is the right bar.
How to answer this, stage by stage
Nobody is grading whether you can define "parameters." They are grading whether you can point at one real call PriceThyme has to make and say, out loud, which size model should answer it and why.
1
Pin it to one real routing decision
Say it like this
"Let's make this concrete. Say there's a tool called PriceThyme. It costs out restaurant recipes against live ingredient prices, and every single time it runs, it has to pick between a small fast model and a large slow one."
Why this works
Keeps the interviewer from grading you on model theory in the abstract instead of a real product decision.
2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my actual pick and when each size earns its place. Impact, what breaks in each direction if you pick wrong. Cost asymmetry, which mistake is cheap and which one hides. Kill criteria, the one test that tells you which size a task actually needs."
Why this works
Two seconds of structure tells the interviewer you have a method, not just a hunch about which model sounds smarter.
3
Give the position, unhedged
Say it like this
"Here's my position. A bigger model reasons better on hard, unusual cases, but it costs more and answers slower. A smaller model is cheap and fast, and it's fine right up until it hits a case it can't actually do. Then it just answers wrong, with the exact same confidence as always."
Why this works
This is the direct answer, said plainly before any story gets a chance to soften it.
4
Split the traffic into what's actually easy and what's actually hard
Say it like this
"Most of what PriceThyme does is boring, on purpose. A supplier ticks a price, multiply it by the recipe quantity, done. That's the small model's job, about twenty nine hundred times a day. The rare job is different: an ingredient spikes or gets discontinued, and the tool has to suggest a substitute that keeps the same taste and yield. That's a conversion problem, not a lookup, and it needs the large model."
Why this works
Shows the tradeoff isn't abstract. It's a real split, in real volume, between two genuinely different kinds of thinking.
5
Show what broke, with a real number
Say it like this
"At Copperpot Kitchens, the substitution job got routed to the small model too, to save money. When black garlic paste spiked three hundred and forty percent, it suggested swapping in a milder purée at a straight one to one weight ratio, when the purée actually needed about one point seven times the weight to taste the same. It reported the sandwich getting thirty two cents cheaper. It actually got sixty nine cents more expensive, and that ran on prep cards for five weeks before anyone caught it."
Why this works
A precise number, over a stated period, is the thing this whole argument would fall apart without.
6
Point straight at the cost asymmetry
Say it like this
"Routing that one rare call to the large model instead would have cost about sixty cents a month, total, across every kitchen. Letting the cheap model guess on it cost close to four thousand dollars in five weeks, on one dish out of a hundred and forty, and it was invisible the whole time because nobody was watching that closely. One mistake is small and you can see it land. The other is hidden, and it keeps adding up until somebody happens to look."
Why this works
Naming which mistake is cheap and which one hides is the actual center of PICK.
7
Give the kill test, and close on one line
Say it like this
"Before I pick a size for any call, I ask one thing: does this involve a ratio, a conversion, or a substitute a person would have to think through? If yes, it earns the expensive model. If it's arithmetic on a known number, the cheap one is the right bar. So: small model for the thousands of routine repricing calls, large model for the rare substitution calls, and a dollar threshold that forces a check in between. That's the whole decision."
Why this works
Restates the direct answer in one breath and gives the test that makes it more than an opinion.
Let's learn
PriceThyme looks at a restaurant's recipes and today's ingredient prices, and tells the kitchen exactly what each dish costs to make right now.
One tool, two very different jobs. Most calls are the easy kind. The rare kind is where this whole answer lives.
Before anyone touched a keyboard, Zenzo Kalinowski did that job by hand. Once a week, he pulled the supplier price sheet for all nineteen Copperpot Kitchens locations and recalculated the cost of all one hundred and forty menu items on a spreadsheet. A full day, every week, and by Thursday the numbers were already three days stale.
PriceThyme did the routine part of that same job in under a second, thousands of times a day, every time a supplier's price ticked. For the everyday calls, that's arithmetic: take the new price, multiply by the recipe quantity, done. A small, cheap model handles that fine.
Knowledge spark: what makes a model "small" or "large" here?
Size is really about how much thinking a model does per answer. A small model is fast and cheap because it takes a shortcut on hard reasoning. A large model is slower and costs more because it works through more steps before it answers. Neither one is "smarter" in general. Each is built for a different kind of question.
The cost of the wrong default, same five weeks
Overcorrecting, large model everywhereUndercorrecting, one wrong substitution
Over the exact same five weeks, playing it safe with the large model on every call would have cost about eighteen hundred dollars for zero benefit. Letting the cheap model guess on the one hard call cost more than double that, on a single dish, and it was still climbing when it was caught.
Here is the turn. The extra mistakes were never the problem with the routine job. The small model has never once mispriced a straightforward reprice. The problem showed up somewhere nobody was measuring: the rare moment PriceThyme had to actually work something out, instead of just look it up.
We did not lose money because the model got worse at math. We lost money because we asked a lookup tool to do a reasoning job, and it never said it wasn't sure.
What that costs at its worst: if Zenzo had still been doing this by hand, he would have caught the ratio the same afternoon, the way any cook checks a jar of purée against a jar of paste before using it. Automating the swap did not just fail to catch the mistake. It hid a mistake a distracted line cook would have caught by taste, and it hid it behind a number that looked exact.
The choice I would take back
After finance flagged how much the large model cost the one time it ran on every call, the team quietly pointed the substitution job at the same cheap model as the routine reprice job, to keep the bill down. That made sense when nobody had priced out what a wrong substitution actually costs. It stopped making sense the moment a rare, hard call started running on a model built only for easy ones.
What I would leave alone: the routine repricing pipeline, still running on the small model, has never caused a single one of these problems. Thousands of correct calls a day is exactly the job it's suited to, and switching it to the large model would only add cost and lag to a live margin alert during a dinner rush, for no gain at all.
The lesson: one model size can't be the whole architecture. The moment you let a single size cover every job, you've quietly promised that the hardest case and the easiest case cost the same to get right. They never do.
Now here is the same thing as a story
The short version above is what you actually say out loud. Read this one for the Tuesday that got Zenzo here, and the one small question that finally cracked it open.
Every Sunday night, for six years before PriceThyme existed, Zenzo printed the week's supplier price sheet and circled anything that jumped more than ten percent in red pen. He knew Copperpot's hundred and forty recipes the way a mechanic knows an engine, which ingredient was load bearing to a dish's margin and which one barely mattered.
Nobody decided, on any single day, to stop checking. The checking just had nothing left to catch, because PriceThyme had never once been wrong on the routine job.
When PriceThyme arrived, the routine reprice work vanished into the background, correct and instant, thousands of times a day. Zenzo spot checked it for a month. It was always right. So he stopped, and put that hour back into menu planning, which is what the tool was supposed to free up in the first place.
The substitution job was different, and rarer, maybe forty times a month across all nineteen kitchens. Early on, Zenzo reviewed every one of those by hand before it hit a prep card. Then finance flagged how much the large model was costing after a month where every single PriceThyme call, easy and hard, ran on it by default. Someone pointed the whole pipeline, substitutions included, at the cheap model instead, to bring the bill down. Nobody in that meeting was being careless. On paper it looked like a clean fix to a real cost problem.
Black garlic paste spiked three hundred and forty percent that spring, a supplier shortage. PriceThyme flagged it inside a minute and suggested a substitute: a milder roasted garlic and shallot purée, swapped in at the same weight as the paste it replaced.
The paste is concentrated, used sparingly for its punch. The purée is milder and needed nearly twice the weight to taste the same. The small model treated the two as interchangeable by weight.
That is not how the two ingredients actually relate. The purée needed about one point seven times the weight of the paste to hit the same flavor. A straight one to one swap under-seasoned every batch, and the "cost savings" the model reported, thirty two cents off the Copperpot Smash, was really a cost the kitchen hadn't paid yet.
Food cost, Copperpot Smash, weeks 0 to 5
Weekly food cost, Copperpot SmashKill line crossed, week 2
The rate crossed the thirty percent kill line by week two. Nobody was watching dish level food cost that closely, only the monthly menu wide number, so it drifted three more weeks before a question forced a look.
The question came from a new sous chef, doing a manual recipe check for a health permit audit. Her paperwork said the Smash now used a cheaper garlic ingredient. The weekly numbers said the dish's food cost had climbed from twenty eight percent to thirty four percent. She turned to Zenzo with a fair, small question: "Why did the cheaper garlic make this more expensive?"
He pulled the batch math that afternoon. The purée needed one point seven times the weight the model had used. The reported savings, thirty two cents a sandwich, was actually a loss of sixty nine cents a sandwich once the correct amount went in. At around eleven hundred and fifty sandwiches a week, across five weeks, that was close to four thousand dollars gone, quietly, on one dish out of a hundred and forty.
The decision Zenzo would take back sits in that finance meeting, not in the kitchen. Routing the substitution job to the cheap model felt like a clean fix to a cost problem that was real. Nobody in that room asked the one question that actually mattered: does this job need arithmetic, or does it need a person's kind of reasoning about ratios? The routine reprice job needed arithmetic. The substitution job never did.
One alternative got seriously considered the week the mistake surfaced, and it's worth naming because it looked like the responsible fix: move every PriceThyme call, routine and rare alike, back onto the large model, for good. It lost. Twenty nine hundred routine calls a day on the large model would have cost about eighteen thousand seven hundred dollars a year for zero improvement on jobs the small model was already getting right, and it would have added real lag to a live margin alert that kitchen staff check mid service.
Run the same five weeks again, with a job type flag that routes substitution calls to the large model and leaves the routine reprice job exactly where it was. The large model correctly scales the purée to one point seven times the weight, on the first suggestion, at a cost of about sixty cents a month across all nineteen kitchens. The Copperpot Smash holds at twenty eight percent food cost. Nobody notices, because there's nothing to notice.
What Zenzo would tell himself, back in that finance meeting: a cheap model is not a smaller version of an expensive one. It's a different tool, built to skip exactly the kind of thinking that a ratio problem needs. We asked it to skip that thinking on the one job where skipping it cost real money.
PICK, routed by the job
Not a ranking of which model wins. PICK only works here if you refuse to crown one size, and instead say, out loud, which job belongs to which model.
PPosition. What actually trades off.
The large model reasons better on hard, unusual cases, the ones with a ratio, a conversion, or a judgment call in them. It costs more per call and answers slower.
The small model is cheap and fast, and it's the right tool for anything that's really just arithmetic on a known number. It's fine right up until it hits a case it can't do, and then it answers wrong with the exact same confidence as always.
Say both halves of the position before any numbers. A position that only shows up after the story looks reverse engineered from it.
IImpact. What breaks in each direction.
Default to the large model everywhere, and you burn budget and add real lag on the thousands of trivial requests that never needed it, about eighteen thousand seven hundred dollars a year here for zero gain.
Default to the small model everywhere, and it quietly produces worse answers on the rare hard cases where quality actually matters, invisibly, since nobody's watching for exactly that failure.
Naming what actually breaks, in both directions, keeps this from turning into a one sided caution story.
One mistake sits in the budget in plain view. The other hides inside a prep card that looks, on the surface, like it went fine.
CCost asymmetry. The heart of it.
Routing a request to a bigger model than it strictly needed just costs a little extra money, that one round, and it's visible on the bill the same day. Shipping the small model on a task it can't reliably do erodes trust in a way that's slow to win back, a chef stops believing the numbers, a sous chef starts double checking everything by hand again, and that habit doesn't come back cheap once it's gone. Start cheap by default. Spend on the expensive, slower option only where getting it wrong actually costs something.
KKill criteria. The one test that decides it.
Does the task's failure cost justify the large model's expense, or is good enough, fast, and cheap actually the right bar? Concretely here: does the call involve a ratio, a conversion, or a substitution a person would have to think through? Yes earns the large model. No means the small model is exactly the right tool, and asking for more would just be waste.
A position with no way to be checked is just an opinion. Naming the exact test is what makes this a real answer instead of a preference.
Three branches, one question asked every time a call comes in. The routine reprice answers known number. The garlic swap should have answered ratio to work out.
The trade off worth saying out loud: routing substitution calls to the large model adds about two seconds of extra time compared to the small model's instant answer, and it costs roughly sixty times more per call. Both were accepted on purpose, because the calls are rare enough, about forty a month, that the total extra cost barely clears sixty cents, while getting one of those calls wrong quietly cost close to four thousand dollars on a single dish.
And if you want to be sure it really works, try it somewhere else
Same four letters, a loading dock instead of a kitchen, and this time the thing that gets under quoted is a truckload instead of a sandwich.
Loadstone Freight brokers truckload shipments, and its quoting tool prices a load the moment a shipper posts one. Suleika Balint runs product there, and the tool splits into the same two kinds of job PriceThyme does: quoting a known, common lane in under a second, and pricing a load that needs real judgment, hazmat, multiple stops, or temperature control, where a straight per mile rate misses what the load actually costs to move.
Same split as Copperpot's kitchen, a completely different industry. The tool that fits each job doesn't change with the trade.
Loadstone pointed both jobs at its small model to keep quotes instant across the board. A temperature controlled load with two extra stops got quoted using the same flat per mile rate as a simple one stop lane, missing the reefer surcharge and the deadhead miles between stops entirely. The quote came in at two thousand one hundred fifty dollars. The load actually cost two thousand eight hundred ninety dollars to run. Loadstone ate the difference to keep the shipper, on a load that should never have needed a discount in the first place.
Mapped onto PICK: the position holds its shape, known lanes stay on the small model, hazmat and multi stop pricing needs the large one. The impact splits the same way, a stale simple quote is a minor annoyance, an under quoted reefer load is real money handed back on a single load. The cost asymmetry lands the same too, routing the rare hard quotes to the large model costs a small amount of extra compute; guessing on them cost seven hundred forty dollars on one load, invisible until a dispatcher noticed margins on multi stop reefer loads had been thin for weeks. The kill criteria transfers without changing a word: does this call involve a ratio or a conversion a person would have to think through.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: route anything with a ratio or a conversion in it to the larger model, before touching cost, no matter how tempting the cheap default sounds.
Cost: no budget this quarter for two model tiers. Fine, but get a written rule that nobody points a substitution or conversion job at the cheap model until the budget exists, full stop.
The model got better, for real: say the small model's overall accuracy improved on a fresh benchmark. Still doesn't matter much. A better score on ordinary requests says nothing about whether it can now handle a ratio conversion it was never trained to reason through.
Where people run it wrong.
They treat "the small model is cheaper" as the whole argument, and never ask what specific job it's cheaper at doing.
They assume a model that's accurate on thousands of easy calls will stay accurate on the one rare hard one, because the dashboard has looked fine for months.
They fix a cost overrun by moving everything to one tier, instead of asking which specific calls actually needed the move.
How to use it live. Before answering, ask yourself one thing: does this call need a person's kind of reasoning, a ratio, a conversion, a judgment, or is it really just arithmetic on a known number? That single question decides the model size. Everything else in the answer is just proving it.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question asking you to weigh a real tradeoff, not just describe two options?
Tap to flip
ANSWER
PICK: state a position, name the impact of each kind of error, find which mistake is cheap and which one hides, then name the one test that would flip your pick.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Zenzo Kalinowski, who runs menu engineering at Copperpot Kitchens, a nineteen location fast casual chain, using a costing tool called PriceThyme.
3 · THE POSITION
State the position in one line: what does a bigger model buy you, and what does it cost?
Tap to flip
ANSWER
A bigger model reasons better on hard, unusual cases, but costs more and answers slower. A smaller model is cheap and fast, fine until it hits a case it can't do, and then it answers wrong just as confidently as always.
4 · THE SPLIT
What's the real split in PriceThyme's calls, and which model handles which?
Tap to flip
ANSWER
Thousands of routine reprice calls a day, pure arithmetic, small model. A rare substitution call, maybe forty a month, that needs ratio reasoning, large model.
5 · THE COST ASYMMETRY
Which mistake is cheap and visible, and which one hides?
Tap to flip
ANSWER
Sending a rare hard call to the large model costs about sixty cents a month, visible on the bill. Letting the small model guess on it cost close to four thousand dollars in five weeks, on one dish, and stayed hidden the whole time.
6 · THE NUMBER
Fill in the blank: the paste spiked ___ percent. The model reported the sandwich ___ cheaper. It actually got ___ more expensive.
Tap to flip
ANSWER
Three hundred forty percent. Thirty two cents cheaper, reported. Sixty nine cents more expensive, actual.
7 · THE KILL CRITERIA
What is the one test that decides which size model a task needs?
Tap to flip
ANSWER
Does the task's failure cost justify the large model's expense, or is good enough, fast, and cheap the right bar? Concretely: does the call involve a ratio or conversion a person would have to think through?
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what plays the role of the garlic swap there?
Tap to flip
ANSWER
Loadstone Freight, a truckload quoting tool. A hazmat, multi stop, temperature controlled load quoted on a flat per mile rate, under quoted by seven hundred forty dollars.
Check yourself Score: 0 / 0
Multiple choice
1. Why did the substitution call need the large model instead of the small one?
A. The large model is always more accurate than the small one, on any task.
B. The swap involved a ratio conversion, weight for flavor strength, that the small model got wrong with full confidence.
C. The small model was offline that week.
D. The large model was actually cheaper for this specific call.
Show hint
Look at what the equivalence sketch and stage five of the walkthrough say the model actually got wrong.
Show answer
B. The small model treated a concentrated ingredient and a milder one as equal by weight. That's a reasoning gap, not a general accuracy problem.
True or false
2. True or false: the small model's everyday repricing work was the thing that caused this problem.
True
False
Show hint
Check the "what I would leave alone" line in Let's learn.
Show answer
False. The small model never mispriced a routine reprice, thousands of correct calls a day. The failure was the rare substitution job, which needed a different kind of reasoning entirely.
Fill in the blank
3. The garlic paste spiked ___ percent in price. Food cost on the Copperpot Smash drifted from ___ percent to ___ percent, crossing the thirty percent kill line by around week ___.
Show hint
Check the numbers in the story and the line chart under the K letter.
Show answer
Three hundred forty percent. Twenty eight percent to thirty four percent, crossing week two. Nobody was watching dish level food cost weekly, only the whole menu number once a month, so it drifted three more weeks past the kill line before anyone looked.
Short answer, name the rejected alternative
4. What alternative did the team seriously consider right after finding the mistake, and why did it lose?
Show hint
Look near the end of the story section, right before the replay.
Show answer
Model answer: Route every PriceThyme call, routine and rare alike, back onto the large model. It lost because the routine job alone runs about twenty nine hundred times a day, and paying large model prices and latency for jobs a calculator could already do would cost about eighteen thousand seven hundred dollars a year for zero benefit, plus real lag on a live kitchen alert.
Short answer, apply it yourself
5. Think of a tool you use that gives you a fast, cheap answer most of the time. Name one part of it that probably deserves a slower, more careful answer instead, and how you'd tell the difference.
Show hint
Look for the part of the tool that involves more than one changing thing at once, not just a lookup.
Show answer
Model answer: A delivery app that instantly estimates arrival time is fine for one simple order. A multi stop or cross border order probably deserves a slower calculation. Tell the difference by asking whether the case involves more than one variable changing at once, not just a single known number.
Short answer, work the number
6. If Copperpot's ingredient prices barely moved week to week instead of spiking suddenly, would routing substitution calls to the large model still matter as much? Why or why not?
Show hint
Think about what actually triggers a substitution call in the first place.
Show answer
Less, but not zero. Substitution calls only happen when a price spikes or an ingredient disappears. Fewer spikes means fewer hard calls, so the cost of skipping the large model shrinks. But the same wrong ratio risk is still there the moment a spike does happen, and it would be even more of a surprise, since nobody would be watching for it.
Before you close the answer
Why this works
Tests whether you can tell an easy call from a hard one and route accordingly, instead of treating model size as a single company wide setting. It also tests whether you'll name a real quality, latency, and cost trade instead of pretending the bigger model is a free upgrade.
Follow-up traps
"Why not just run the large model everywhere and stop worrying about it?" Response: because the routine job alone is thousands of calls a day, and paying large model prices and latency for a job a calculator could handle costs about eighteen thousand seven hundred dollars a year for nothing, plus lag on a live alert kitchen staff check mid service.
"How do you even know it's a hard call before you run it?" Response: check whether the request involves a unit or ratio conversion, or a price move past a set threshold, before picking the model. Or, since these calls are rare enough, just route anything tagged as a substitution to the large model by default, since the extra cost barely registers.
If pressed
Both jobs ran through the same API call with no field telling PriceThyme which kind of question it was answering, so there was no cheap way to route between model sizes until the team added a job type flag to the request itself.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.