ConceptAdvancedAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #20
How do you factor in that the buy option's capability improves without your effort?
BOUNDthe vendor was already closing most of the gap for free, nobody had charted it
Crestholm Mutual is an insurance company. ClaimLoop routes an incoming claim document to the right adjuster team and flags likely fraud cases. Selin Paskowitz is the AI PM weighing whether to build a custom classifier in house, and Graham Orenstein is the VP of Finance pushing to build, to lock in a four-point accuracy edge over the vendor Crestholm currently pays for.
The direct answer
Don't compare today's accuracy gap. Chart the vendor's own improvement over the last two years, project it forward, and compare the cost of closing the same gap at every future point, not just this one. If the vendor's own trend is on track to close most of your target gap within a year or two, building isn't buying a permanent edge, it's renting a temporary one at a much higher price.
Do this, in order
Chart the vendor's own accuracy over the last two years before comparing anything today.Why: a snapshot bake-off hides that one side of the comparison keeps moving for free.
Treat the vendor's improvement rate as a real, sourced assumption, not an afterthought.Why: it is the single number this whole decision is most sensitive to.
Price your own team's realistic improvement velocity, not just the cost to close today's gap.Why: a one-time sprint doesn't keep a lead open by itself, upkeep does.
Compare total cost over the same multi-year window for both options.Why: build's up-front cost and buy's ongoing fee only compare fairly over matching time.
Re-run the same eval set against the vendor's current version every quarter.Why: it turns "we think they're improving" into a real, checkable trend.
Name the one assumption that would flip your call, before anyone asks.Why: it's exactly what an interviewer, or a VP of Finance, will push on first.
How to answer this, stage by stage
Nobody is scoring whether you can recite build-vs-buy pros and cons. They're scoring whether you noticed that one side of the comparison is a moving target.
Stage 1
Scope it to one real, numbers-driven decision
Say it like this
"Let me ground this in one real case. Crestholm Mutual routes insurance claims with ClaimLoop. The vendor version hits 88 percent accuracy today. An in-house build could hit 92. That four-point gap is the actual number I'd stop treating as fixed."
Why this works
Keeps the answer from becoming a generic "build vs buy" lecture with no real gap behind it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as BOUND. Break down the real equation. Own each number and where it came from. Use a range instead of one figure. Nail a sanity check against something known. Name the one direction that would flip the answer."
Why this works
Signals a repeatable estimation method, not a gut call dressed up as financial rigor.
Stage 3
Reframe: it isn't "build vs buy today," it's "build vs a moving target"
Say it like this
"This isn't really a question of which option is better right now. It's a question of whether the gap you're planning to spend money closing will still exist by the time you've closed it."
Why this works
This is where a strong answer separates from someone who only ever compares a single snapshot.
Stage 4
Give the one decision: the actual math
Say it like this
"Cost of building equals the sprint to hit 92 percent now, plus the upkeep to keep re-earning that lead every year after. Cost of buying equals the subscription fee, minus the accuracy the vendor ships anyway. Over three years, building costs about 610 thousand dollars to hold a lead the vendor's own trend closes in under two years. Buying costs about 420 thousand and lands in roughly the same place."
Why this works
This is the direct answer, said as an actual number you could defend in the room, not a vague sense that "buy is usually safer."
Stage 5
Prove it with the compressed evidence
Say it like this
"Two years ago the vendor's accuracy was 82 percent. It's 88 now. That's about three points a year, for free, with zero engineering from Crestholm. Project that forward and the vendor crosses our 92 percent target in about a year and a half, right around when our own build would just be finishing its first full year of upkeep."
Why this works
Compresses the whole case into the one trend line that actually decides whether building is worth it.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The reason this isn't a generic vendor-negotiation question is that the vendor's model keeps improving on its own release schedule, something no ordinary software contract does. We considered running both models in parallel permanently as a hedge, and rejected it, since doubling infrastructure cost for a gap that's already closing isn't worth it. Building buys a temporary, real accuracy edge at real cost. Buying accepts convergence, at a fraction of the price, in exchange for never fully owning the differentiation."
Why this works
This is the load-bearing, AI-specific judgment. A normal vendor's product doesn't get better every quarter with zero action from either side.
Stage 7
Say what wouldn't change the calculus, then close
Say it like this
"I wouldn't apply this trend-based reasoning to a capability that's genuinely proprietary, like our own historical fraud patterns, since the vendor's general model never sees that data and won't improve on it for free. For the claims-routing gap specifically, the trend held: the vendor was already closing it, so we bought, and rechecked the eval every quarter to make sure the trend kept holding."
Why this works
Closes with real judgment about where the reasoning doesn't apply, and restates the direct answer in one breath.
Let's learn
Two years ago, a claim document at Crestholm Mutual got sorted correctly by a vendor model eighty-two times out of a hundred.
ClaimLoop routes an incoming claim document to the right adjuster team and flags likely fraud cases, instead of a mailroom clerk sorting by hand. Before ClaimLoop, a mailroom team sorted claims by reading the first page, about 30 seconds each, at 76 percent correct, since claims paperwork is genuinely ambiguous and some documents belong to two categories at once.
With the vendor version of ClaimLoop, routing hits 88 percent correct, in about a tenth of a second, for the same volume.
The vendor's own accuracy, charted, projected against the 92 percent in-house target
The dashed portion is projection, not measurement. Even at the low end of the range below, the vendor closes most of the gap within two to three years, for free.
Here's the turn: the extra four points of accuracy isn't really the decision anymore, since most of it may get won without Crestholm spending a dollar. The actual decision is whether paying real engineering money to grab those points now, instead of waiting, is worth it.
We weren't comparing build's cost against buy's cost today. We were comparing build's cost against how much of the gap the vendor was already going to close for free.
At its worst, Crestholm spends 610 thousand dollars over three years chasing a four-point lead that quietly disappears anyway, and ends up right back where a 420 thousand dollar subscription would have landed it, just with less money left over.
The five letters, held up as one page. The O step, owning the vendor's own trend as a real number, is the one this whole answer turns on.
The choice I would take back
Crestholm's procurement policy compared vendors with a single bake-off at contract renewal, never plotted as a trend. That made sense when models rarely changed year to year. It stopped making sense once providers started shipping meaningfully better versions every few months, for free, on their own schedule.
What I would leave alone: I wouldn't apply this same reasoning to a capability that's genuinely proprietary, like Crestholm's own historical fraud-pattern data. A vendor's general model never sees that data, so it can't improve on it for free, regardless of any trend line elsewhere.
The lesson: the fair fight isn't build's cost against buy's cost today. It's build's cost against how much of the gap the vendor was already going to close for you, for free, by the time you're done building.
Now here is the same thing as a story
Use the short version above in the room. Read this one for what it felt like the day a two-year-old chart quietly answered a six-month argument.
Selin Paskowitz had run the numbers on ClaimLoop three times before Graham Orenstein's team ever asked for a fourth.
Graham had spent four years pushing Crestholm's finance team to own more of its own technology instead of renting it, and he was good at his job: he could spot a recurring vendor fee compounding into real money faster than almost anyone in the building.
When the 92-percent in-house prototype numbers came back from a two-week internal test, Graham brought them straight to Selin's team with a clean case: four points of accuracy, and Crestholm would own every bit of it forever.
Knowledge spark: why would a vendor's model improve with zero effort from Crestholm?
ClaimLoop's vendor version runs on a general document-understanding model the provider keeps upgrading on its own schedule. Every time they ship a better version, Crestholm's accuracy can rise the next morning, with nobody at Crestholm changing a setting or paying for the improvement directly.
Nobody at Crestholm had ever charted the vendor's accuracy over time. It built up slowly: a small internal note here, a shrug there, a passing comment that "the vendor seems better than it used to be," never a real number, until Selin pulled two years of the same fixed eval set against every vendor release and actually plotted it.
Both options end up in nearly the same place. Only one of them costs 190 thousand dollars more to get there.
We weren't deciding whether Crestholm should own its accuracy. We were deciding whether to pay 190 thousand dollars extra to own it a little earlier than the vendor would have given it to us anyway.
Selin never had a fixed rule for when a gap was "worth building for" versus "worth waiting on." It came down to one real question: is the vendor's trend actually closing this gap, or standing still? For claims-routing, the trend was closing it. For Crestholm's own fraud-pattern data, there was no trend at all, because the vendor's model would never see that data no matter how many versions it shipped.
The sprint is only a third of the real cost. The other two thirds is upkeep, just to hold a lead the vendor's own trend was already closing.
Back when Crestholm's procurement policy was first written, comparing vendors once at renewal made sense, since the underlying models barely changed between one renewal and the next. It stopped making sense the moment providers started shipping real accuracy gains every few months on their own.
The base case is measured, not guessed. But the whole decision should still survive the low end of this range, not just the middle.
Rerun the argument with the range included: even at the low end, 1.5 points a year, the vendor still closes most of the gap within about three years, just more slowly. At the high end, it's closed within a year. Either way, building buys time, not a permanent edge, and the sanity check holds: 610 thousand dollars is real money to buy something the vendor's own roadmap was going to deliver regardless.
Three-year total cost: build versus buy
Building costs 190 thousand dollars more over three years, to hold a lead the vendor's own trend was already on pace to close.
One version spends 610 thousand dollars and ends up roughly where the vendor lands anyway. The other spends 420 thousand and gets there just as fast, since most of the gap was never going to require Crestholm's money to close.
What I'd tell myself, watching Graham's clean four-point case land on the table: a gap you haven't charted over time isn't a gap yet. It's just a number you haven't asked "compared to what, and by when" about.
BOUND, the five checks this decision actually neededNot a script for always choosing buy. BOUND is what makes sure build only wins when the gap is genuinely durable.
B
Break it down. State the real equation out loud.
Cost of build equals the sprint to close today's gap, plus the upkeep to keep re-closing it every year after. Cost of buy equals the subscription fee, minus whatever the vendor closes on its own.
Without stating both sides as equations, "build vs buy" stays a gut feeling wearing a spreadsheet's clothes.
O
Own numbers. Source the vendor's trend, don't assume it.
82 percent two years ago, 88 percent today, measured by re-running the same 3,000-document eval set against the vendor's own release history, not taken from their marketing page.
This is the hardest step and the one the whole decision actually turns on.
U
Use a range. Don't bet the decision on one point estimate.
1.5 to 4.5 points a year, with 3 as the measured base case. Even the low end still closes most of the gap within three years.
A single confident number hides how much of the conclusion depends on being right about the future.
N
Nail the sanity check. Does the number survive a smell test?
610 thousand dollars over three years to hold a lead the vendor closes in under two, versus 420 thousand to land in nearly the same place. If those numbers felt reversed, something in the math would be wrong.
A number that doesn't survive comparison to something known is a guess wearing a decimal point.
D
Direction. Which single assumption would flip the call?
The vendor's own improvement rate. If it's genuinely 1 point a year instead of 3, the gap stays open for years, and building starts to look like the right call instead of an expensive detour.
This is the answer to give when someone asks what would change your mind, before they have to ask twice.
The recap, one line per letter: break down build's real cost against buy's real cost as equations, own the vendor's measured 3-points-a-year trend instead of assuming it, use the 1.5-to-4.5 range instead of one number, nail the sanity check that 610k for a temporary lead against 420k for durable convergence, and name the vendor's own improvement rate as the one assumption that would flip the whole call.
And if you want to be sure it really works, try it somewhere elseSame five letters, a shipping terminal instead of an insurance office. This time the vendor's trend has gone flat, and that changes the answer.
Noor Zubrycki runs container-damage detection at Tallowreach Terminal, deciding whether to keep buying a vendor's computer-vision damage-detection API or build one in house. Mapped onto BOUND: break it down the same way, build's cost against buy's cost over three years. Own numbers: here the vendor's accuracy has moved only about 0.3 points a year over the last 18 months, essentially flat, because container-damage detection is a small market the provider hasn't prioritized, versus an in-house team that could realistically gain 2 points a year with dedicated investment. Use a range: there's a real, if small, chance the provider suddenly reprioritizes logistics and the trend jumps, so the range runs from flat to 1.5 points a year. Nail the sanity check: 180 thousand dollars to close a 6-point gap, plus 60 thousand a year to hold a 2-point pace, totals 360 thousand over three years, against a flat vendor that likely never closes the gap on its own, versus 270 thousand to keep buying and stay 6 points behind indefinitely. Direction: the one assumption that flips this is whether the vendor's parent company reprioritizes logistics, a real but low-probability risk, so the base case favors building, the opposite of Crestholm's call, because the trend line itself points the other way.
Same five checks, but the branch that fires depends entirely on which way the vendor's own trend line points.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "chart the vendor's trend, project it forward, and compare the cost of closing the same gap at every point, not just today," and stop.
Cost: no time to build a real two-year trend before deciding. Say so honestly, and use the widest defensible range for the vendor's improvement rate instead of guessing a single number.
The vendor got better, for real: if the vendor's trend accelerates past your projection, that's not a reason to panic-build, it's exactly the signal the range and the quarterly recheck exist to catch early.
Where people run it wrong.
They compare today's gap and stop, treating the vendor's number as if it were frozen.
They price the engineering sprint but forget the ongoing upkeep needed to keep any lead open.
They treat "the vendor might improve" as a vague risk instead of a real, chartable, sourced trend.
How to use it live. The moment an interviewer asks how you'd factor in a vendor's free improvement, ask yourself: what would this comparison look like if I ran it again in eighteen months? If the honest answer is "roughly the same," that's the whole argument, said out loud, in one breath.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits factoring a vendor's own improvement trend into a build-vs-buy call?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. It forces both options forward in time instead of comparing a single snapshot.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Selin Paskowitz, the AI PM at Crestholm Mutual, and Graham Orenstein, the VP of Finance who wanted to build in house to lock in a four-point accuracy edge.
3 · THE HABIT
What habit did Crestholm's procurement process have, that this decision broke?
Tap to flip
ANSWER
Comparing vendors with a single bake-off at contract renewal, never plotted as a trend over time.
4 · THE REAL COMPARISON
What comparison does this decision actually turn on?
Tap to flip
ANSWER
Today's four-point gap versus the cost of closing that same gap after the vendor's own trend has moved. The vendor's number keeps moving either way, so a static comparison is the wrong comparison.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Comparing vendors with a single point-in-time bake-off, never as a trend, back when models rarely changed year over year.
6 · THE NUMBER
Fill in the blank: the vendor's accuracy has risen about ___ points a year, projected to reach the 92 percent in-house target within about ___ months.
Tap to flip
ANSWER
3 points a year, about 16 to 18 months.
7 · THE REPLAY
Rerun the calculation with the vendor's trend included. What changes?
Tap to flip
ANSWER
Building costs $610k over three years to hold a lead the vendor's own trend closes in under a year and a half, versus $420k to buy and land in roughly the same place.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's different?
Tap to flip
ANSWER
Tallowreach Terminal's container-damage detection tool. There building wins instead, because the vendor's own improvement trend has gone flat, not because the math changed.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: the vendor's accuracy rose from ___ percent two years ago to ___ percent today, an average of about ___ points a year.
Show hint
Look at the line chart in "Let's learn."
Show answer
82 percent, 88 percent, 3 points a year. All three come from re-running the same eval set against the vendor's own release history, not a guess.
Multiple choice
2. Why does comparing today's four-point accuracy gap alone understate the real decision?
A. Four points is too small a gap to matter at all.
B. The vendor's own accuracy is already rising for free, so most of that gap may close without Crestholm spending anything.
C. In-house engineers always beat vendor accuracy eventually, no matter what.
D. Claims documents never change in structure over time.
Show hint
Look at the O step in the BOUND recap.
Show answer
B. The vendor's trend is a real, sourced number, and it was already closing most of the gap Crestholm was about to pay to close itself.
True or false
3. True or false: if the vendor's real improvement rate turns out to be only 1.5 points a year instead of 3, building in-house becomes a clearly worse decision.
True
False
Show hint
Look at the D step, direction, in the BOUND recap.
Show answer
False. A slower vendor trend means the gap stays open much longer, which makes building look better, not worse. The vendor's rate is the exact assumption that swings this call.
Short answer, where it wouldn't matter
4. Name a part of ClaimLoop where this vendor-trend reasoning would NOT apply, and say why not.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Crestholm's own historical fraud-pattern data. A vendor's general model never sees that proprietary data, so it can't improve on that specific capability for free, no matter its overall trend.
Short answer, apply it yourself
5. Think of a tool you use that's improved on its own, with no effort from you. What gap did that close, and would you have paid to close it yourself first?
Show hint
Think of an app or service whose accuracy or speed just got better one update, with no action from you.
Show answer
Model answer: A photo app's auto-tagging quietly went from mixing up two similar-looking pets to telling them apart, after a background model update. Paying to build a custom fix first would have been money spent on a gap that closed on its own within a year.
Short answer, work the number
6. If the engineering sprint cost $350k instead of $220k, would the three-year total for building still land below $700k?
Show hint
Add the new sprint cost to three years of $130k upkeep.
Show answer
Model answer: No. $350k plus $390k of upkeep is $740k, moving further above the $420k buy option and making the case for building even weaker, not stronger.
Before you close the answer
Why this works
Tests whether you project both options forward on their own trajectories, instead of comparing a snapshot and assuming it holds still.
Follow-up traps
"Isn't it worth building anyway, just to have full control?" Response: control has real value, but it should be priced as its own line item, not hidden inside an accuracy comparison the vendor was going to win on its own within eighteen months.
"What if the vendor's improvement rate is just a lucky streak?" Response: that's exactly why the range matters. Quarterly re-checks against a fixed eval set turn "a lucky streak" into a real, trackable trend within two or three quarters, instead of staying a guess either way.
If pressed
The historical trend was pulled from re-running the same 3,000-document eval set against the vendor's last four release notes, not from their marketing claims, which is what made the 3-points-a-year number defensible instead of assumed.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.