CaseIntermediateAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #13
Your CFO wants predictable costs. How does that shape the build-buy-finetune choice?
PICKa flat bill that costs a bit more beats a cheap bill that occasionally explodes
Wrenhollow Hotels runs a small chain of boutique properties. GuestReply drafts a reply to a guest's WhatsApp message, a late checkout ask, a noise complaint, a room-service question, for the front desk to review and send. Isolde Marchbank is the AI PM who owns it, and Corrado Ashby is the CFO whose quarterly audit turned "predictable costs" from a slide-deck phrase into a real requirement.
The direct answer
Pick the fixed-cost path: a smaller, self-hosted fine-tuned model, or a vendor's flat-fee committed-throughput contract, over pay-per-token API calls, even though the API is cheaper in an average month. A predictable bill that costs a bit more every month beats a cheap bill that occasionally explodes on the exact week leadership is watching most closely.
Do this, in order
Pick the fixed-cost path over the cheaper-on-average one.Why: a surprise spike can get the whole AI budget frozen, a slightly higher flat bill never will.
Name who actually feels each kind of cost error.Why: finance feels the invoice spike in a single terrifying number, ops feels the flat plan as a quieter, ongoing overpay.
Say the trade-off out loud instead of pretending it's free.Why: the fixed path costs roughly 10 percent more a year. Hiding that number is how trust with the CFO gets spent early.
Set a real kill criteria for switching back.Why: if message volume ever becomes genuinely flat, the fixed plan stops earning its premium.
Build the overage alert either way.Why: predictable doesn't mean unmonitored, a flat plan can still get exceeded if volume grows past what it was sized for.
Revisit the pick every couple of quarters, not once and never again.Why: seasonality and volume both drift, and the pick was made against a specific shape of demand, not a permanent law.
How to answer this, stage by stage
Nobody is scoring whether you can name the cheapest option. They're scoring whether you can commit to a pick and defend the trade-off out loud, in front of the person who actually controls the budget.
Stage 1
Scope it to one real feature, with a real invoice behind it
Say it like this
"Let me ground this. Wrenhollow Hotels runs GuestReply, it drafts a reply to a guest's message for the front desk to review. Isolde owns it. Corrado, the CFO, just pulled last quarter's bill and found it spiked five times over during one holiday week."
Why this works
Keeps "the CFO wants predictable costs" from turning into a generic finance lecture with no real feature behind it.
Stage 2
Say your structure out loud before any content
Say it like this
"I'll run this as PICK. Position, my pick, stated first. Impact, who feels each kind of cost error and in what units. Cost asymmetry, which error is cheap and visible versus hidden and expensive. Kill criteria, what evidence would flip the pick."
Why this works
Signals a real, committed decision instead of a wishy-washy "it depends on a lot of factors."
Stage 3
Reframe the question: this isn't "which is cheapest"
Say it like this
"This isn't a question of which path costs less on paper. It's a question of which kind of cost error you can actually survive, one that shows up as a slightly bigger number every month, or one that shows up as a five-times spike on the one week finance is watching closest."
Why this works
This is where the answer separates from a spreadsheet comparison with no real judgment in it.
Stage 4
Give the one decision: the actual pick
Say it like this
"Here's what I'd pick. Move GuestReply off pay-per-token calls and onto a smaller, self-hosted fine-tuned model with a flat monthly infrastructure cost. It costs about ten percent more across a full year. It never spikes, and it never gives finance a reason to freeze the whole program."
Why this works
This is the direct answer, stated as a real pick with a real number attached, not a hedge.
Stage 5
Prove it with the compressed evidence
Say it like this
"On pay-per-token, quiet months ran about eighteen hundred dollars. One holiday month hit ninety four hundred, five times over, because guest message volume triples during peak weeks. The fixed path runs a flat thirty one hundred a month, every month, all year, no exceptions."
Why this works
Gives the interviewer real dollar figures instead of a vague "it was expensive."
Stage 6
Name the AI-specific reasoning and the trade-off being accepted
Say it like this
"The honest reason the spike is worse than it looks is that guest message volume isn't steady, it's driven by check-in surges and events, so a per-token bill scales with exactly the weeks a hotel is busiest and most exposed. We're accepting a real, roughly ten percent higher annual cost in exchange for a bill that can't blow a hole in a quarterly forecast."
Why this works
This is the load-bearing judgment, and the honest trade-off. It only makes sense because model usage cost scales with unpredictable message volume, not with a fixed software license.
Stage 7
Say what would flip the pick, then close on one line
Say it like this
"If Wrenhollow ever moves toward mostly corporate, contracted stays with genuinely flat message volume, the spike risk mostly disappears, and pay-per-token starts winning on pure cost. Today, with leisure-season swings, I'd pick the flat plan: predictable, a little more expensive, and safe from the one bill that could get this whole feature shut down."
Why this works
Closes with the kill criteria and restates the direct answer in one breath.
Let's learn
GuestReply is a feature inside Wrenhollow's front-desk software that reads a guest's incoming message and drafts a reply, which a staff member reviews and sends, instead of typing one from scratch.
A flat pass costs a little more most nights. A meter is cheaper most nights, and merciless on the one night it isn't.
For a year, GuestReply ran on pay-per-token calls to a general model. Quiet months, the bill sat around eighteen hundred dollars, and nobody thought twice about it. It looked like a solved cost problem.
Monthly bill, pay-per-token, across one year
Pay-per-token monthly bill
Eleven quiet months made the twelfth one look like a rounding error, until finance actually opened the invoice.
Corrado's team ran a routine quarterly audit and found the spike. Guest message volume roughly triples during peak holiday weeks, and since the API charged per message processed, the bill tripled and then some, once longer, more detailed guest messages during peak season were counted in.
Knowledge spark: why does a per-token bill spike faster than volume itself?
A busier week doesn't just mean more messages, it usually means longer ones too, more detail in a complaint, more back-and-forth in a request. Pay-per-token pricing charges for both the count and the length, so a 3x rise in messages can turn into a 5x rise in cost once longer messages are counted.
Both options cost money. Only one of them can get the whole feature's budget frozen in a single invoice cycle.
We weren't choosing the cheapest bill. We were choosing which kind of surprise we could actually survive.
Here's the turn: the spike itself wasn't really the damage. Ninety four hundred dollars for one month is real money, but Wrenhollow could have absorbed it once. What it couldn't absorb was what came after: Corrado's team froze new AI spending pending a review, which meant a planned expansion of GuestReply to two more properties sat on hold for six weeks.
Same year, fixed self-hosted plan versus pay-per-token
Pay-per-token, one spike monthFixed self-hosted plan
The flat plan costs about 10 percent more across the year. That's the honest price of never handing finance a 5x surprise.
The choice I would take back
Launching GuestReply on pay-per-token pricing without ever modeling what a peak week would cost. That made sense during the pilot, when volume was low and steady. It stopped making sense the moment GuestReply rolled out hotel-wide and inherited the same seasonal swings the front desk itself has always had.
What I would leave alone: Wrenhollow's back-office expense-report scanner, used by managers a handful of times a month with no seasonal pattern at all, stays on pay-per-token. There's no spike risk to protect against, so paying for unused flat capacity would be pure waste.
The lesson: "cheaper on average" and "predictable" are two different bars. When a CFO asks for predictable costs, they're not asking you to minimize the total, they're asking you to remove the one number that could blow up trust in the whole program.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what it actually felt like watching a routine audit turn into a real budget freeze.
Isolde Marchbank had shipped GuestReply eight months earlier, and the launch had gone quietly well. Front-desk staff liked it, guest response times were down, and the monthly bill, eighteen hundred dollars or so, barely came up in budget meetings.
The bill didn't creep upward. It sat quiet for months, then jumped all at once, on exactly the week nobody wanted a surprise.
Corrado Ashby's quarterly audit landed in Isolde's inbox with one line highlighted: a single month at $9,412, five times the usual run rate. "Nothing about this was flagged before it happened," Corrado said in the review meeting. "I need to know our AI costs before they show up on an invoice, not after."
"Predictable" turned out to mean four specific things, and Wrenhollow's setup had none of them yet.
Isolde's first instinct was to add usage caps and alerts on top of the existing pay-per-token setup. That would catch the next spike after it started, but it wouldn't stop message volume itself from tripling during the next holiday week. The real fix had to change what the cost curve looked like, not just watch it more closely.
A "flat" plan isn't one number, it's four commitments, and Wrenhollow's first contract only had one of them written down.
Isolde never had a fixed rule for exactly when a cost curve was risky enough to walk away from pay-per-token pricing. It came down to a feeling with two settings: either the busiest week costs roughly the same as the quietest one, or it costs meaningfully more, and the gap is wide enough that leadership would notice it as a surprise rather than a number they'd already budgeted for. A five-times spike, discovered by an audit instead of a forecast, was unmistakably the second setting.
The API wasn't a bad option, it was simply optimized for the wrong axis, once predictability became the actual requirement.
Back when GuestReply first launched, choosing pay-per-token pricing wasn't an unreasonable call. Volume was low, the pilot properties were quiet, and the team genuinely believed usage would stay roughly steady. It stopped being reasonable the moment GuestReply rolled out to every property and inherited the same seasonal swings the front desk had always lived with.
Here's the replay: Isolde's team stood up a smaller, self-hosted fine-tuned model sized for Wrenhollow's own guest-message patterns, with a flat infrastructure cost of $3,100 a month. The next holiday week arrived with the usual volume surge. The bill that month: $3,100. The same as the month before, and the month after.
One version of this story keeps a cheap-looking bill running for a year, then hands finance a 5x surprise at the worst possible moment and loses six weeks of momentum to a budget freeze. The other spends roughly ten percent more across the year and never gives anyone a reason to ask "what happened here" again.
What I'd tell myself, reading that audit line for the first time: an average cost that looks fine is not the same thing as a cost that's actually predictable. Somebody has to check the worst week, not the typical one, because the worst week is the one a CFO remembers.
PICK, run on a bill that looked fine until someone checked the worst weekNot a script for always choosing the cheapest option. PICK is what makes sure the option you choose survives the week that actually gets noticed.
P
Position. What's the pick, before any reasoning?
Move GuestReply off pay-per-token pricing onto a fixed-cost path, a self-hosted fine-tuned model or a flat-fee vendor contract, even knowing it costs more across a full year.
Stating the pick first is what separates a decision from a menu of options with no owner.
I
Impact. Who feels each kind of error, and in what units?
On pay-per-token, finance feels a $9,400 invoice in one line, right during peak season scrutiny. On the fixed plan, ops feels a flat $3,100 bill even in quiet months when usage would have cost less on the metered path.
Naming who feels each error in real units is what makes the trade-off concrete instead of theoretical.
C
Cost asymmetry. Which error is cheap and visible, which is hidden and expensive?
The flat plan's overpay is small, visible every month, and budgeted for. The metered plan's spike is hidden until the invoice lands, and it's the one that got the whole program's budget frozen for six weeks.
This is the actual reasoning behind the pick: optimize against the error that can kill the program, not the one that's merely annoying.
K
Kill criteria. What evidence would flip this pick?
If Wrenhollow's guest mix shifts toward mostly flat-volume corporate contracts, with little seasonal swing, the spike risk mostly disappears, and pay-per-token starts winning on pure cost with nothing left to protect against.
Naming this is what separates a confident pick from a stubborn one that never gets revisited.
The recap, one line per letter: position is choosing the fixed-cost path before any numbers get shown, impact is naming that finance feels the spike and ops feels the flat overpay, cost asymmetry is the reasoning that a program-freezing surprise beats a smaller, absorbable overpay, and kill criteria is the one condition, genuinely flat volume, that would flip the pick back.
And if you want to be sure it really works, try it somewhere elseSame four letters, a ride-share company instead of a hotel chain. This time the spike comes from a single viral weekend, not a holiday season.
Gideon Marchetti is the AI PM at Meridian Rides, which offers drivers an in-app chat assistant, FareGuard, that answers questions about surge pricing and payout timing. Mapped onto PICK: position is choosing a vendor's flat-fee committed-throughput contract over pure pay-per-token calls to a frontier model. Impact is that finance feels a metered spike as a single terrifying invoice line, while drivers feel a fixed plan as slightly slower rollout of new chat features, since the committed contract caps total monthly volume. Cost asymmetry holds the same shape: one weekend where FareGuard went viral on social media drove ten times the normal chat volume, and a metered bill for that weekend alone would have cost more than the fixed contract's entire month. Kill criteria is different here: if Meridian's driver base stops growing and chat volume flattens out for two straight quarters, the committed-throughput contract's unused headroom stops being worth its premium, and a lighter pay-per-token setup becomes the better pick again.
Same asymmetry, a different trigger. A viral weekend does to a chat feature what a holiday week does to a hotel's front desk.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to "pick the fixed-cost path, it costs more on average and never hands finance a surprise," and stop.
Cost: no budget yet to stand up self-hosted infrastructure. Say so honestly, and start with a vendor's flat-fee contract as the faster first step toward the same predictability.
The model got better, for real: say a future model cuts per-token pricing in half. The spike would shrink, but the ratio between quiet and peak weeks stays the same, so the pick doesn't change, only the exact dollar gap between the two options does.
Where people run it wrong.
They compare the two paths using an average month, and miss that the whole risk lives in the worst month, not the typical one.
They treat "predictable" and "cheapest" as the same request, and quietly pick the cheaper option without naming the trade-off out loud.
They fix the metered plan with usage alerts alone, which catches a spike after it starts instead of removing the spike's ability to happen at all.
How to use it live. The moment an interviewer says a CFO wants predictable costs, don't reach for the cheapest option. Ask yourself: which of these costs the same in its worst month as its best one? That's usually where the real pick is hiding.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a CFO's "predictable costs" request shaping a build-buy-finetune choice?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. This is a tradeoff question, not a perturbation, so a FLIPS flip family doesn't apply here.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Isolde Marchbank, the AI PM who owns GuestReply at Wrenhollow Hotels. Corrado Ashby is the CFO whose quarterly audit turns "predictable costs" into a real requirement.
3 · THE PICK
What's the actual position, stated before any reasoning?
Tap to flip
ANSWER
Move off pay-per-token pricing onto a fixed-cost path, a self-hosted fine-tuned model or a flat-fee vendor contract, even though it costs about 10 percent more across a full year.
4 · THE ASYMMETRY
Which cost error is cheap and visible, and which is hidden and expensive?
Tap to flip
ANSWER
Cheap and visible: the flat plan's modest monthly overpay in quiet months. Hidden and expensive: a metered spike that can freeze the whole program's budget, not just cost money once.
5 · THE OLD DECISION
What old decision does this answer take back?
Tap to flip
ANSWER
Launching GuestReply on pay-per-token pricing without ever modeling a peak week's cost. Reasonable during a quiet pilot. Wrong once the feature rolled out hotel-wide and inherited real seasonal swings.
6 · THE NUMBER
Fill in the blank: the quiet-month bill ran about $1,800. The holiday-month bill hit $___.
Tap to flip
ANSWER
$9,412, roughly 5 times the normal run rate, discovered by a routine quarterly audit rather than caught in advance.
7 · THE REPLAY
Same holiday-week volume surge, running on the fixed self-hosted plan instead. What changes?
Tap to flip
ANSWER
The bill stays at $3,100, the same as every other month. No spike, no audit finding, no six-week freeze on the planned expansion to two more properties.
8 · CROSS PRODUCT TRANSFER
Section 4 runs PICK again on a different product. Which one, and what triggers the spike there instead of a holiday season?
Tap to flip
ANSWER
Meridian Rides' driver chat assistant, FareGuard. There, the spike comes from a single viral social-media weekend, not a seasonal calendar, but the same fixed-cost pick still wins for the same reason.
Check yourself Score: 0 / 0
Multiple choice
1. Why did the fixed-cost path get picked even though it costs more across a full year?
A. The self-hosted model produced noticeably better replies than the API.
B. A predictable bill removes the risk of a surprise spike that could get the whole program's budget frozen.
C. Vendor contracts are always cheaper than pay-per-token pricing.
D. The CFO required all AI features to run on owned infrastructure.
Show hint
Look at the cost asymmetry step in the PICK recap.
Show answer
B. The pick optimizes against the error that can end the program, a budget-freezing spike, not just the error that costs a bit more on average.
True or false
2. True or false: the fixed self-hosted plan was actually cheaper than pay-per-token pricing across the full year.
True
False
Show hint
Look at the annual comparison bar chart.
Show answer
False. It cost about 10 percent more across the year, $34,000 versus $31,000. The trade-off was predictability, not a lower total.
Fill in the blank
3. Fill in the blank: the flat self-hosted plan cost $___ a month, every month, all year.
Show hint
Look at the annual comparison chart's note.
Show answer
$3,100 a month. The same figure in the quietest month and the busiest one, which is the entire point of the pick.
Short answer, where it wouldn't matter
4. Name a feature at Wrenhollow where this fixed-cost reasoning would NOT apply, and say why not.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The back-office expense-report scanner, used a handful of times a month with no seasonal pattern. With no spike risk to protect against, paying for unused flat capacity would be pure waste.
Short answer, apply it yourself
5. Think of a subscription or utility bill you personally have. Would you pick a flat plan or a metered one, and what's the "spike week" you'd be protecting against?
Show hint
Think about a bill that's usually low but occasionally jumps, like mobile data or electricity during a heat wave.
Show answer
Model answer: A flat mobile data plan over pay-per-gigabyte, specifically to protect against the one month of heavy travel or video calls that would otherwise turn into a shockingly large bill.
Short answer, work the number
6. If Wrenhollow's holiday spike had only reached $3,000 instead of $9,412, would the fixed-cost pick still make sense?
Show hint
Compare a much smaller spike against the $1,300 annual premium the fixed plan actually costs.
Show answer
Model answer: Less clearly. A $3,000 peak month is a much smaller surprise, closer to what a CFO might absorb without freezing the program, so the case for paying a real annual premium to avoid it gets weaker.
Before you close the answer
Why this works
Tests whether you can commit to a real pick under a business constraint, and whether you'll name the honest cost of that pick instead of pretending predictable and cheapest are the same thing.
Follow-up traps
"Couldn't you just cap the API spend instead of switching paths?" Response: a spend cap stops the bill, but it also stops GuestReply from working once the cap hits, mid-holiday-week, which trades a budget problem for an outage during the busiest, highest-visibility period.
"Isn't paying 10 percent more every year just wasteful?" Response: it is a real cost, named on purpose, not hidden, in exchange for removing a risk that already cost six weeks of stalled expansion once.
If pressed
The self-hosted model was sized against the 95th-percentile week, not the average one, so it has headroom for a typical holiday surge without needing to scale infrastructure up mid-week, which is what would have reintroduced cost unpredictability through the back door.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.