Describe how you would present unit economics to a CFO.
A one-page view that only shows the number when the number is calm is not honesty, it's a deck waiting to be caught out. Build the page so the ugly number already has a seat at the table.
- Build the page around one blended margin, shown as a range, with model retraining as its own permanent named line.Why: a range survives being wrong by a little; a point number that misses becomes the whole meeting.
- Give the retrain line a stated trigger, not a vague "AI costs vary."Why: a cost with a named cause is a budgeted risk. A cost with no cause is a red flag, even when the dollar amount is small.
- Show the retrain line every quarter, even the quarters it doesn't fire.Why: the first time a CFO sees a new line item is the worst possible time for it to also be a surprise bill.
- Hand the CFO exactly one lever, SKU count per store, not the raw cost breakdown.Why: a lever she can move herself builds trust. A spreadsheet she can't read just makes her trust the person, not the number.
- Set the retrain trigger as a stated quality-versus-cost choice, not an arbitrary schedule.Why: a tighter bar catches drift sooner and protects the retailer from stockouts, at a real cost. Naming that trade-off out loud is what makes the threshold defensible under a follow-up question.
- Skip the full one-pager for small pilot accounts still in free trial.Why: building a CFO-grade view for a five-store trial spends real time on a number no renewal decision is waiting on yet.
How to answer this, stage by stage
Nobody in the room is grading arithmetic. They're grading whether you know a probabilistic AI cost needs a visible seat before launch, not an apology after the invoice lands.
Let's learn
Fillwell looks at a retail chain's sales history and current stock, and predicts overnight, SKU by SKU, exactly what each store will need to reorder before it runs out.
Before this design existed, the CFO deck at Kestrahl, the company that sells Fillwell, was a printed export of the cloud bill: GPU inference cost broken out by model call type, a vector store fee, an orchestration line, added up into one blended total. It worked fine when Kestrahl had one customer and the CFO was also the technical co-founder, someone who could read a cloud invoice like a native language.
By the time Fillwell had 340 live stores across six retail chains, that same deck went to Faustina Merritt, a CFO with no reason to know what a vector store is. She read the total, trusted the person who built it, and moved on. Nothing in the deck was wrong. It just didn't map to a single question she actually had to answer: is this margin healthy, and what happens to it if we grow.
The turn is not that Fillwell's costs sometimes spike. Spikes on their own are survivable. The real problem shows up the moment a spike lands with no name attached to it, because a CFO who finds an unexplained line doesn't ask about that one line, she asks whether every other number in every other deck was ever explained either.
That was the cost at its worst: not the retrain bill itself, sixteen thousand four hundred dollars, small next to Fillwell's monthly revenue. It was that onboarding froze for six weeks across every product line while finance re-audited costs that had nothing to do with the actual incident, because one unexplained spike made every number look equally unexplained.
The choice Winnifred would take back is small and was completely sensible at the time: build the CFO deck by exporting the raw cloud dashboard and pasting the total into a slide. At one customer, with a technically fluent CFO reading it, that cost nothing and explained everything. At forty customers and a CFO who runs finance, not infrastructure, the same habit stopped being a shortcut and started being the thing that broke trust.
What I'd leave alone: a pilot account with under five stores, still in free trial, doesn't need this whole page. A single paragraph estimate is enough, since there's no renewal decision or board number resting on it yet. Building the full CFO-grade view there spends real design time on a number nobody is about to act on.
The lesson: a CFO doesn't distrust a number because it moved. She distrusts it because it moved somewhere she wasn't told to look. Give the moving part a name and a line before she ever has to go find it herself.
Now here is the same thing as a story
Read the story below when you want to feel why a named line survives a bad quarter that an unlabeled one never could, not just be told that it does.
Winnifred Cadwallader had presented Fillwell's numbers to finance for two years. She was good at it, careful with her slides, and Faustina Merritt, Kestrahl's CFO, had never once asked her to walk through the underlying spreadsheet. Winnifred's name on a deck was the check. That trust had been earned honestly, in the early days, when Fillwell had one customer and Faustina's predecessor had come up through engineering himself.
The habit that formed then never got questioned as the customer list grew. Every quarter, Winnifred opened the cloud billing console, exported the total, broke it into the categories the console gave her for free, GPU inference, storage, vector search, orchestration, and pasted the numbers into a slide. It took twenty minutes. It had never once been wrong, in the sense that the total always matched what the invoice said. Whether it was useful to the person reading it was a different question, one nobody had asked in over a year.
Forty customers in, the retrain schedule was no longer something anyone thought to mention out loud in a finance meeting, because it had genuinely never come up before. Fillwell's model retrained when its own forecast error, checked weekly against a holdout set of real SKU-week outcomes, crossed a twelve percent error bar. Most quarters it didn't fire at all. When it did, the retrain run cost around sixteen thousand dollars, and that cost landed inside the same "cloud compute" line as everything else, indistinguishable from a normal month with slightly higher usage.
The trigger wasn't a lawsuit or a board question. It was one regional grocery chain expanding its frozen-foods assortment by forty percent ahead of a seasonal push, a completely ordinary retail decision that Fillwell's model had never seen a version of before. Forecast error crept past twelve percent inside three weeks. The retrain fired automatically, the way it was built to. Nobody on the product team thought twice about it, because that's exactly what the system was supposed to do when demand patterns shifted.
Finance saw it differently. Cloud compute for that month came in at forty five thousand four hundred dollars against a normal run rate near twenty nine thousand, a fifty six percent jump, with nothing anywhere in Winnifred's deck that had ever mentioned retraining as a thing that happened, let alone a thing with a cost. Faustina didn't call Winnifred first. She called an emergency review of every AI product line Kestrahl sold, and froze new customer onboarding across the company for six weeks while a finance analyst tried to reconstruct, line by line, what every AI feature in the portfolio actually cost to run.
The real cost wasn't the retrain bill. It was that Winnifred's name stopped being the check. Every deck after that needed a second reviewer from finance before it reached Faustina's desk, which added most of a week to every quarterly close for the rest of the year.
The decision Winnifred would take back happened quietly, over a year earlier, the week she built the very first version of the CFO deck. Copying the cloud console's own categories straight into a slide felt like the honest, transparent choice, more detail, not less. Nobody in that first meeting flagged that "more detail" and "more legible to the person paying the bill" were two different things, and the shortcut sat there, unexamined, for two years.
Run the same incident the old way, and it repeats with a bigger number attached, since Fillwell now covers more stores than it did that quarter. Run it the new way: the page Winnifred builds afterward has three things on it. A blended margin range, sixty to sixty eight percent, wide enough to already include a retrain quarter inside it. A line called "model retrain," with the trigger, twelve percent forecast error against the holdout set, written next to it, present on the page even in quarters where it stays at zero. And one lever, SKU count per store, that Faustina can move herself to see what a bigger customer would do to the number, instead of asking engineering to explain a bill after the fact. The next time a chain expands its assortment and a retrain fires, cloud compute jumps the same way it did before. Faustina checks it against the range and the named line, sees it land inside both, and closes the deck in about two minutes. No emergency review. No freeze.
One design trusted a person because the total had always matched the invoice. The other trusted a page because every number on it, including the ones that move, had already been given a name and a reason before anyone had to go looking for one.
What I'd tell myself, back in that first meeting: more raw detail is not the same thing as more trust. A number the reader can't act on isn't transparency, it's a wall they have to climb before they can ask the question they actually came in with.
SPARK, in one page a CFO reads without an engineer in the room
Not a slide checklist. Each letter has to still hold up the day a chain doubles its SKU count without warning, the exact thing the story just walked through.
Three things worth stating directly, since this is where the real judgment sits. The alternative Winnifred considered and rejected was keeping the full cloud-console breakdown on the page, GPU cost by call type, vector store fee, orchestration overhead, in the name of transparency. It lost because more raw detail isn't the same thing as more trust, and it was the exact design that buried the retrain spike inside "cloud compute" in the first place. The AI specific failure worth naming by name is distribution drift: real-world demand patterns, a bigger frozen-foods assortment, a supply shock, a new product line, shifting enough that Fillwell's forecasts quietly get worse until the holdout-set check catches it. The guardrail is that same twelve percent MAPE threshold, now doing double duty as both the retrain trigger and the CFO's own named line item. And the quality, latency, and cost trade-off worth naming too: Winnifred set that threshold at twelve percent instead of a tighter eight or a looser eighteen. Tighter catches drift sooner and protects the retailer from stockouts and overstock, at the cost of more frequent, more expensive retrains. Looser saves money but lets forecasts drift further before anyone notices, and the retailer eats that cost in bad stock decisions long before Fillwell's own numbers show it.
And if you want to be sure it really works, try it somewhere else
Same five letters, a hospital's finance office instead of a retail chain's, and this time the closest comparison is a radiology worklist, not a grocery aisle.
ReadFirst flags which scans in a radiology worklist need a radiologist's eyes first, ranking studies by AI-estimated urgency across every imaging site in a hospital system. Galatea Marigny is the PM who owns its unit economics, presented once a quarter to the health system's finance office.
S, situation: before this method, the CFO's finance team got a raw GPU-cost export broken out by scan modality, MRI against CT against X-ray, a breakdown built for an engineer, handed to people who think in cost per scan and margin per site, not tensor operations per modality.
P, payoff: the habit worth building isn't "trust ReadFirst's vendor invoice." It's finance reading one blended number, margin per scan per month, and trusting it without re-deriving it from the raw cloud bill every quarter.
A, anchor: the page carries blended margin per scan as a range, a permanent line called "model recalibration," tied to a stated trigger, fires when triage accuracy against a radiologist-reviewed holdout set drops below a stated bar, and one lever, scans processed per site per month.
R, risk: a new imaging site comes online running a different scanner brand, its image quality distribution doesn't match what the model trained on, and the recalibration trigger fires sooner than the quarter budgeted for. Because it's a named line with a stated cause, finance reads it as an onboarding cost, not a mystery.
K, keep out: no per-modality cost breakdown on the CFO page, and no live dashboard tracking every scan. One report, one range, one named line, once a quarter.
Same method, a different weak spot: Fillwell's retrain trigger fires on a demand-pattern shift. ReadFirst's recalibration trigger fires on an image-quality shift from new hardware, a completely different cause, but the design answer is identical: give the moving cost a name and a stated trigger before finance ever has to ask what it is.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, a margin range, a named retrain line with its trigger, and one lever, and give the one concrete number, sixty to sixty eight percent blended margin, retraining budgeted at up to four times a year near sixteen thousand dollars each.
Cost: there's no budget this quarter for a fancier live dashboard. The one-page view costs almost nothing to build and it's the part that actually changes whether the CFO trusts the next number.
The model got better, for real: say the forecasting model gets thirty percent more accurate next year, needing fewer retrains. That's not a reason to quietly drop the retrain line from the page. A rarer retrain still needs its named slot, or the one time it does fire lands as a surprise again, in a page that had gone a year without mentioning it.
Where people run it wrong.
They keep every raw cost category from the cloud bill in the name of transparency, and bury the one line that actually needs a name inside a catch-all bucket.
They publish a single point margin to look confident, and quietly turn a real range into a promise the numbers can't always keep.
They only mention the retrain line the quarter it fires, which teaches the CFO that a new line item on the page always means bad news, instead of teaching her it's a normal, budgeted part of the product.
How to use it live. Say the real tension out loud before answering: "is this asking me how to format a slide, or how to make a moving number trustworthy." That buys a beat, and it's almost always the second one.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Doesn't showing a retrain line every quarter, even when nothing happened, just teach the CFO to ignore it?" Response: no, it's the one line she's trained to check against a stated trigger every time, the same way a budgeted maintenance reserve gets glanced at even in a quiet quarter. The risk of her tuning it out is smaller than the risk of it reappearing as a surprise.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Cost modeling and unit economics
- #1 Build the cost-per-interaction model for a feature with a 2,000-token prompt and a 500-token response.
- #2 What cost drivers exist for an AI feature beyond model tokens?
- #3 Explain how a RAG pipeline's cost structure differs from a single model call.
- #4 How does prompt caching change your unit economics, and when does it not help?
- #5 Model the monthly cost of a feature used by 50,000 users averaging 12 interactions each.
- #6 What is the cost impact of moving from a single call to a five-step agent?