What happens to your pricing when model costs drop 60 percent in a year?
The model behind your product just got 60 percent cheaper over a year. Everyone in the room has an opinion about where that money should go. The interviewer is not grading whether you know the three places it could go. They are grading which one you protect first, and whether you can say why out loud before anyone has checked that the cheaper model is still any good.
- Bank the freed up margin first. Do not touch the price the day the new model's cost lands.Why: a price cut is the one move here nobody can quietly undo if the drop turns out to be a launch promotion, or the cheaper model turns out worse on a kind of job nobody has tested yet.
- Run the new model against your golden bid set and real shadow traffic, split by project type, before it touches a single customer's number.Why: a wrong bid does not get caught by a support ticket. It gets caught weeks later, when the real invoice does not match what the tool said it would cost.
- Reinvest part of the savings into a second, independent check pass, not more free bids.Why: contractors are not paying for volume. They are paying for a number they can put their name on, and that is what the second pass actually buys back.
- Split accuracy by project type before calling the eval passed.Why: a healthy average hid an 11.4 percent miss rate on commercial jobs over four floors, the 8 percent of bids carrying the most dollar risk.
- Only then phase the price down, in small steps across more than one pricing cycle, never as one headline cut.Why: raising a price back after contractors have already priced their own jobs against the lower number is the one apology nobody accepts.
- Do not match a competitor's same-week price cut on reflex.Why: a rival cut price the same week the model got cheaper. Matching it on the spot would have shipped the unverified model straight into commercial bids before anyone had looked at the segment split.
How to answer this, stage by stage
Nobody is grading whether you can name margin, price, and reinvestment as the three options. They are grading whether you can rank them live, with real numbers, and defend the order when someone points out that cutting price is the obvious way to pass along a win.
Let's learn
Craneline is the tool Rivercross Software sells to general contractors. A contractor uploads a plan set, and Craneline drafts a full line item cost estimate, materials, labor hours, subcontractor pricing, contingency, ready to become a real bid.
Before Craneline, an estimator on a mid size renovation spent about 14 hours hand building a bid, chasing quantities off paper plans with a calculator and a spreadsheet template. Most contractors could only realistically chase 3 to 4 full bids a week. Jobs with a tight deadline simply got skipped.
With Craneline, a full bid draft takes under 20 minutes. Contractors turned that into 10 to 12 bids a week instead of 3 to 4. By the start of the year this question is about, Craneline was live inside 1,400 contractor accounts, running about 12,600 bids a month.
Craneline's price was set by a simple rule: five times the trailing quarter's average compute cost, recalculated automatically every quarter. For the first year, a person, the company's finance partner, personally reviewed and signed off on each quarterly number by hand. Nothing ever needed to change. To move faster, the team folded that sign off into the formula itself, so the price would simply recalculate on its own from then on.
The year this question is really about starts right after that automation went live. Compute cost per bid opened at $36, price at $180. Two quiet quarters followed: cost drifted to $30 a bid, price to $150, then to $26 a bid, price to $130. Nobody double checked either move. Nothing had gone wrong, so nothing looked wrong.
Then the fourth quarter arrived. A new generation model shipped, and Rivercross's own benchmark showed compute cost falling all the way to $14.40 a bid, a full 60 percent below where the year began. A rival tool, selling into the same regional contractors, announced its own matching price cut the same week. Left alone, the formula's next recalculation would have set Craneline's price straight to $72, a 45 percent cut in a single quarter, with no human sign off at all.
Ahead of that quarterly recalculation, the engineer who owns Craneline's eval pipeline ran the new model against the golden bid set, 240 historical bids a professional estimator had already checked line by line. In aggregate, the new model actually looked slightly better than the one it replaced, 2.7 percent of line items flagged wrong, against a baseline of 3.1 percent. Split by project type, the picture changed. Commercial jobs over four floors, mostly mechanical and electrical rough in work, showed 11.4 percent of line items flagged wrong, more than three times baseline. Those jobs were only 8 percent of Craneline's volume by count, but carried far more dollar risk per bid than a small remodel.
One bid, for a six story mixed use build, had already been drafted with the new model during shadow testing, and was queued in the standard overnight batch to reach the contractor's account at 6am. It got pulled at 11:40pm, during a spot check that was not, at the time, a required step at all.
At its worst: had that bid gone out, and the contractor priced their own job on Craneline's number, an undercounted rough in on a multi story mechanical job could cost real money once the actual subcontractor invoice landed mid project. That is not a mistake a contractor quietly forgives, and it is the kind of story that spreads fast through a small regional trade network. It would have cost Craneline more than the year it took to build the price-savings feature in the first place.
What I'd leave alone: Craneline's flat monthly platform fee, account access, historical bid storage, team seats. It was never tied to per bid compute cost, so the 60 percent drop changes nothing about it. No review needed there.
The lesson: a cost drop is not a price signal by itself. It is a question about whether the thing behind the price still works the way it did. Answer that question first, with real evidence, before a formula or a competitor's move answers it for you.
Now here is the same thing as a story
Read the longer version below when you want to feel why an average accuracy number that looked like good news almost sent a stranger's plan set into a real bid, unchecked.
Every morning for six years, Endre Vantress has opened Rivercross's finance dashboard before the coffee is even done, and read it in the same order every time: cash first, then unit economics, then anything showing red. For Craneline's first year, that morning ritual included a quarterly ceremony nobody else in the company much noticed: comparing the new compute cost number against the model's own accuracy metrics, side by side, before signing off on the next quarter's price.
The good months were genuinely good. Craneline's price barely moved that first year, because compute cost barely moved, and every single quarterly sign off passed without a single incident. Watching the same review clear, quarter after quarter, it started to feel less like a decision and more like a formality. So the team automated it. Fold the review into a formula, let price recalculate itself from the live compute bill, free up Endre's mornings for something that actually needed a human. Nobody in that short meeting asked what the formula should do the one time cost dropped in a single big jump instead of a slow drift.
It faded in three quiet beats, and none of them looked like a mistake at the time. Beat one: the sign off simply stopped happening, replaced by a line item on a dashboard nobody was told to watch closely. Beat two: the second quarter's automatic recalculation landed, cost to $30, price to $150, and it was fine, so nobody checked the next one any harder. Beat three: the third quarter did the same, cost to $26, price to $130, fine again, and by now the formula had quietly become just how pricing worked.
The trigger was not two mistakes in a row. It was a new model generation shipping the same week a rival tool publicly slashed its own price to match the same industry wide cost drop. Left alone, the formula's Q4 recalculation was two days from setting Craneline's price at $72, a 45 percent cut, sight unseen, the same week everyone else in the market was also moving fast.
Dovid Palanca, the engineer who owns Craneline's eval pipeline, ran the new model against the golden bid set ahead of that recalculation, the way he always did before any model change. The aggregate number looked good, even slightly better than baseline. He almost signed off on the strength of that one number alone. Splitting it by project type, out of habit more than instruction, is what actually caught the 11.4 percent miss rate hiding inside commercial jobs over four floors. One bid built on that number, for a six story build, was already queued for the overnight batch. He pulled it at 11:40pm, off schedule, because he happened to look.
The decision that opened the door traced back to that short automation meeting a year earlier. Someone asked whether the quarterly sign off could just be folded into the formula, since it had cleared every single time. It was the fast, sensible sounding answer, and at the time, it was true. Nobody asked what the formula should do the one time cost fell all at once instead of drifting.
Run that meeting again with one change: price freezes at $130, the last verified level, for six weeks while a segment split shadow eval runs across every project type, not just the one that got caught. Part of the savings goes into a second, independent check, the model re run with a deliberately different quantity order, top down by floor instead of bottom up by trade, flagging anything the two runs disagree on. After four weeks, the commercial MEP segment's miss rate drops from 11.4 percent to 3.4 percent, back in line with baseline. Only then does price move, in two small steps across two cycles, $130 to $114, then $114 to $99, well short of the $72 the formula alone would have set. Contractor churn holds at 1.8 percent across both quarters of the fix.
One design let a single number decide when quality had been checked. The other made sure quality got checked before any number got to decide anything.
What I would tell myself, back in that short automation meeting: removing my own sign off was never really about moving faster. It was about assuming a year of zero incidents would keep holding on its own. That was never the formula's win to claim by itself.
ORDER, said out loud while the number is still moving
Not a story wearing a framework's clothes. This is a live ranking problem across three destinations for one saving, and ORDER is what stops "whichever move feels fastest today" from quietly standing in for "whichever move you can't take back tomorrow."
Three things worth stating directly, since the real judgment sits here. The alternative Rivercross's team actually considered, and rejected, was matching the rival's price cut the same week, letting the formula's automatic $72 recalculation ship on schedule. It lost because that move would have sent the unverified model straight into commercial bids before anyone had looked at the segment split, and because a contractor who had already priced a job off $72 could never be quietly told to plan for $99 instead. The AI specific failure worth naming by name is silent degradation hiding behind a healthy average: the new model's aggregate accuracy actually improved, which is exactly what made the commercial MEP segment's problem invisible without a deliberate split. The guardrail is making a segment split shadow eval a required gate for every future model swap, not a one time fix for the segment that happened to get caught this time. That guardrail is not free. Rivercross accepted higher compute spend, not lower, for six weeks, running the new model twice on the flagged segment to check its own work, and a slower rollout, in trade for a price move it would never have to walk back.
And if you want to be sure it really works, try it somewhere else
Same five letters, a radiology second read tool instead of a bid, and this time the lever is not which project type to trust more. It is which scan type actually carries the risk of being wrong.
Pellucid is Emberlyn Diagnostics' second read tool. A radiologist reads a scan, Pellucid reads it independently, and flags anything it thinks got missed, priced at $40 per scan reviewed. Corwen Ledford owns cost and quality on it.
The build up: a new vision language model cut compute cost per scan read from $9 to $3.60. Against Pellucid's own golden read set, the new model's overall discordance rate actually looked slightly better than baseline, 3.8 percent against 4.1 percent. Split by scan type, mammography screening reads alone showed 9.6 percent, more than double baseline for that scan type specifically, while routine chest and musculoskeletal reads held steady near baseline.
Same rank, different lever: the segment touching the highest stakes finding still needs the guarded slice here too, but the lever is not which archive to trust more, it is which uncertainty gets shown to a person before a price ever moves. Corwen's team banked price where it was, reinvested savings into a visible uncertainty flag on mammography reads specifically, shown to the radiologist directly rather than folded into one confidence score, and only phased price down for the scan types that had already cleared their own segment eval.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the order: bank first, verify by segment, reinvest in the thing that protects trust, cut price last and only where it is proven.
Cost: there is no budget this quarter for both a segment split shadow eval pipeline and a price cut. Fund the eval pipeline. A price cut you have to walk back later costs more trust than the eval infrastructure ever will.
The model got better, for real: say the new model is not cheaper, just more accurate across every segment, MEP and mammography alike. That changes what reinvestment buys, better numbers instead of just more of them, but the order barely moves. You would still want segment level evidence before making any promise a customer cannot un-hear, because better claimed by a vendor and better verified on your own golden set are not the same fact.
Where people run it wrong.
They let a pricing formula or a competitor's move make the timing decision instead of the eval.
They check accuracy in aggregate and miss the one segment quietly carrying all the risk.
They cut price first because it is the fastest way to feel like they passed the savings on, then discover the savings were never fully proven.
How to use it live. Say the real question out loud before naming a number: "before I touch price, has the thing generating the number actually been checked, split by the kind of job where being wrong costs the most." That buys a beat to actually rank instead of guessing at a percentage.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"The aggregate accuracy number looked fine, so why not ship it?" Response: an aggregate blends a segment that's fine with one that isn't. MEP bids were only 8 percent of volume but carried the real dollar risk, and the average hid that completely.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pricing AI products: seat, usage, outcome
- #1 Compare seat-based, usage-based and outcome-based pricing for an AI product.
- #2 Why does seat-based pricing break when AI reduces the number of seats needed?
- #3 Design a pricing model for an AI feature with high variable cost and unpredictable usage.
- #4 What is the risk of usage-based pricing from the customer's point of view?
- #5 Explain how credits work as a pricing mechanism and their advantages.
- #6 How would you price an agent that completes a task rather than answers a question?