Explain what a cost model in a portfolio signals to a hiring manager.
Meridian Legal Aid Network is a fictional nonprofit that connects volunteer lawyers to pro bono cases. Yusuf Demirci built Casebrief there: it reads a long case file and produces a short, structured summary a caseworker can read before a client meeting. Beatrix Coyle, a hiring manager, checks for a cost model before she reads anything else in an AI portfolio piece.
- Include a cost model with a real range, not a single clean number.Why: a single number implies a certainty about usage nobody building an AI feature actually has yet.
- Name the usage assumption behind the number, out loud.Why: a cost figure with no stated assumption can't be checked, and shouldn't be trusted either.
- Show the worst case, not just the typical case.Why: the worst case is exactly what a real production bill looks like on a bad month.
- Tie the cost to a real unit, per summary, per user, per month, not a lump sum.Why: a lump sum can't tell you what happens when usage doubles. A per-unit number can.
- Say when you'd re-check the model.Why: a cost model built once and never revisited is exactly how a real bill sneaks past everyone.
- Skip a full cost model only for a very early, throwaway prototype, and say so plainly if you do.Why: naming the gap yourself beats a reader discovering it and wondering what else got skipped.
How to answer this, stage by stage
Nobody is grading whether your cost estimate is exactly right. They're grading whether you understand that it needs a range at all.
Let's learn
A cost model is a short, honest account of what an AI feature actually costs to run, per use, not just once to build.
Beatrix Coyle reads dozens of AI portfolio pieces a year. For a long while, most candidates skipped a cost model entirely, or included one clean line: "runs for about 40 dollars a month." She had no way to tell if that number meant anything.
Lately, a small number of candidates include a real range instead: a low estimate for a short case file, a high estimate for a long one, with the actual token counts behind both numbers. Beatrix now spends real time on those pages, and treats the clean, unsourced numbers as a flag worth probing live.
Here's the turn: the interesting thing was never whether Yusuf's cost estimate was exactly right on day one. It was whether his model could see this exact drift coming, weeks before the actual monthly bill made it undeniable.
At its worst: a candidate ships a clean, confident-sounding number, a real bill arrives months later at three or four times that estimate, and nobody, including the candidate, saw it coming because nothing in the original model ever accounted for usage actually changing.
What I would leave alone: a rough, honestly-labeled early estimate for a genuine day-one prototype is fine, and doesn't need the full range treatment yet, as long as the candidate says plainly that it's a rough placeholder.
The lesson: a cost model isn't a math exercise you complete once. It's a habit of checking whether the assumption behind a number is still true, and most portfolios never show that habit at all.
Now here is the same thing as a story
The short version above is what you'd say defending your cost model out loud. Read this one for how Yusuf actually built his, and where it first went wrong.
Yusuf Demirci spent two years as a paralegal before he ever trained a model, and he can still tell, from the first page of a case file, roughly how long it'll take a lawyer to get through it.
His first version of Casebrief's cost model was one clean line at the bottom of his write-up: "Costs about 2 cents per summary to run." It looked tidy. It was also based only on the shortest sample case files he'd tested with.
Beatrix, reading an early draft, asked one plain question: "What happens to this number if a caseworker feeds it a genuinely long file, not your test sample?" Yusuf didn't know. He'd never actually tested one.
He went back and pulled real file-length data from his own pilot, run with five volunteer caseworkers over twelve weeks. The cost per summary had drifted from about 2 cents in week one to 11 cents by week twelve, as caseworkers realized they could paste in whole case files instead of the shorter excerpts he'd originally designed the tool around.
Rebuilding the model properly took Yusuf about three hours, mostly spent pulling his own usage logs and rerunning the math with real, not idealized, file lengths. The real cost was never those three hours. It was that his original clean number would have quietly justified never checking again.
Writing one confident, rounded number back in his first draft had felt like the responsible thing to do, tidy, and easy for a reader to skim. It stopped feeling responsible the moment Beatrix's one question showed it had never been tested against real usage at all.
The old model asked Beatrix to just trust a tidy number. The new one showed her exactly which usage pattern it was built on, and what happened outside that pattern.
I wrote one clean number because it looked responsible on the page. It took one direct question about a file length I'd never actually tested to see it was never really tested at all.
LEAD, what the cost model actually signalsNot a math check. LEAD is what shows a hiring manager whether this number would have moved before the real bill did.
The recap, one line per letter: link is what the candidate would actually catch on the job, early signal is the model surfacing that instinct in round one instead of month six, abuse is the cherry-picked or unsourced number, and decision is the real threshold a hiring manager acts on.
And if you want to be sure it really works, try it somewhere elseSame four letters, an HVAC dispatch tool instead of a case file. A different building, and the leading signal is a routing mistake, not a token count.
Thurlow HVAC Services is a fictional field-service company. Colette Marchand built a dispatch-routing assistant there, an AI tool that reads a repair request and matches it to the right technician. Grant Osei reads her cost model.
Mapped onto LEAD: the real link isn't whether Colette can price a model call, it's whether she'd notice the assistant quietly routing more jobs to a more expensive, faster on-call technician than necessary. The early signal is a per-dispatch cost broken out by technician tier, tracked weekly, since it would show the pattern drifting long before a quarterly labor-cost report ever would. The abuse: pricing the model using only the average dispatch, when the real cost swings heavily on rush and after-hours jobs specifically. The decision: a cost model that breaks costs out by job type and technician tier moves Colette's candidacy forward; one flat average number gets a direct follow-up question about rush jobs specifically.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a real range with a stated assumption is the signal, a suspicious flat number is the red flag," and stop.
Cost: there's no time to build a full sensitivity model before tomorrow's interview. A single honest sentence naming the assumption behind your number beats a polished number with none.
The model gets better, for real: if Casebrief's summarization gets cheaper per token next quarter, that's still worth modeling, since a falling cost curve is exactly the kind of change a real cost model should be built to notice.
Where people run it wrong.
They treat a cost model as a one-time calculation instead of a number that needs revisiting as real usage changes.
They price only the cheapest, cleanest scenario, and never say so.
They skip a cost model entirely for an AI feature, the same way nobody would skip a cost estimate for a physical product with a real bill of materials.
How to use it live. When someone asks what a cost model signals, ask yourself one question first: would this number have caught the bill drifting before anyone had to point at it. Say that, before you say anything about the arithmetic itself.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if usage is genuinely impossible to predict before launch?" Response: then say that plainly, give a wide range instead of a fake-precise point number, and name what you'd measure in week one to narrow it.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Building an AI PM portfolio
- #1 What does a hiring manager actually open first in an AI PM portfolio?
- #2 Describe the three artifacts that make the strongest AI PM portfolio.
- #3 How do you present a shipped AI project when you cannot share the internal metrics?
- #4 What does a portfolio project need to prove that a resume line cannot?
- #5 Critique a portfolio built entirely from case study write-ups with no build.
- #6 How do you build a credible AI PM portfolio with no AI job experience?