ConceptFoundationalAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #2
What conditions make buying a vendor solution the right call for an AI feature?
PICKthe expensive mistake was never the vendor's, it was the quarter of engineering time nobody priced against it
Hallowbrook Freight moves freight for other companies and handles the paperwork that comes with it. ManifestRead is a feature meant to read a scanned bill of lading and pull out the weight, the contents, and the destination, so a dispatcher doesn't retype it by hand. Odalys Ferran is the AI PM who has to decide whether to build that or buy it, and Grady Wemyss is the ops manager whose one question in a kickoff meeting reordered the whole decision.
The direct answer
Buy when three things are true at once: the capability is a problem other companies have already solved well, your own data would not teach a model something a stranger's documents couldn't, and the cost of being locked into a vendor's roadmap is smaller than the cost of the engineering quarter it would take to build it yourself. Build only when one of those flips, usually because your own data really is the edge.
Do this, in order
Buy when the capability is solved elsewhere and your own data isn't the edge.Why: this is the actual test, not a preference between "safe" and "innovative."
Test the vendor against your own messiest real documents, not their demo set.Why: a demo is chosen to look clean. Your worst 300 cases are the ones that actually decide it.
Price the real alternative: what your engineers would build instead if they weren't rebuilding a commodity feature.Why: that's the hidden, expensive mistake, not the vendor's occasional bad read.
Write a kill criteria down before you sign.Why: without one, "we already bought it" quietly becomes the whole argument, forever.
Check the switching cost, not just the sticker price.Why: lock-in is the part of the expensive mistake that doesn't show up until the day you try to leave.
Keep a simple manual override for the vendor's worst cases.Why: the cheap, visible mistake is fine to absorb. That's not the one worth designing against.
How to answer this, stage by stage
Nobody is scoring whether you can name the five build-versus-buy options. They're scoring whether you can actually commit to one, out loud, with a reason that survives a follow-up.
Stage 1
Scope it to one real feature
Say it like this
"Let me ground this in one real case. Hallowbrook Freight moves freight for other companies. ManifestRead is supposed to read a scanned bill of lading and pull out the weight, the contents, and the destination so a dispatcher doesn't retype it. Odalys Ferran, the AI PM, has to decide whether to build that or buy it."
Why this works
Keeps the answer from turning into a general build-versus-buy lecture with nothing real behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as PICK. Position, my actual pick, stated first. Impact, who feels each kind of mistake. Cost asymmetry, which mistake is cheap and which one is expensive in a way you don't see coming. Kill criteria, what evidence would change my mind."
Why this works
Signals a repeatable method for a buy decision instead of a gut call dressed up afterward as strategy.
Stage 3
Reframe the question
Say it like this
"This isn't really 'is the vendor good enough.' It's 'does owning this specific capability teach us anything a stranger's documents wouldn't.' That's the question that actually decides it."
Why this works
This is where a strong answer separates from a vendor-comparison spreadsheet with no real judgment in it.
Stage 4
Give the one decision
Say it like this
"Buy it. A bill of lading looks basically the same whether it's Hallowbrook's freight or somebody else's. Three vendors already read these well, and nothing about our own documents would make a model built just for us meaningfully better. Build only comes back if the vendor's ceiling turns out too low, or the contract gets too expensive to leave."
Why this works
This is the direct answer, said as an actual pick, not a list of factors to weigh someday.
Stage 5
Prove it with the compressed evidence
Say it like this
"We tested the vendor against 300 of our messiest real documents, water-damaged scans, handwritten totals, three different customs formats. It read the key fields correctly 91 percent of the time. An internal prototype, after six weeks of engineering, scored 89 percent on the same set. The gap wasn't in the accuracy. It was the six weeks."
Why this works
This is where the story lives, compressed to the one number pair that actually settled it.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this wasn't close is that a document-reading model gets most of its real skill from the sheer range of documents it's trained or grounded on, and a vendor's ten million documents beat our forty thousand every time, for a problem this common. We're accepting slower iteration on our own edge cases and dependence on someone else's roadmap, in exchange for not spending a quarter closing a gap that was mostly already closed."
Why this works
Names the load-bearing AI-specific judgment, data scale beats a small proprietary set here, and the trade-off actually being accepted.
Stage 7
Say what wouldn't apply, then close
Say it like this
"None of this holds for the part of ManifestRead that flags shipments that look like insurance fraud risks, based on patterns specific to Hallowbrook's own claims history. That's the opposite of a commodity problem, so that piece stays in-house. For plain field extraction, buy it, test it against your worst documents first, and write your kill criteria down before you sign."
Why this works
Closes with real judgment about where the pick stops applying, and restates the decision in one breath.
Let's learn
ManifestRead is a feature inside Hallowbrook Freight's dispatch software that reads a scanned bill of lading or customs form and pulls out the weight, the contents, and the destination, so a dispatcher doesn't have to retype it by hand.
Before anyone tested a vendor, the team assumed Hallowbrook's own documents were messy in some special way. Most of a bill of lading turned out to be four fixed things.
Before the test, the working assumption was that Hallowbrook's documents were too specific to freight brokerage for a generic vendor tool to handle well, water stains, handwriting in the margins, three different customs formats depending on the lane. It felt like exactly the kind of mess a stranger's model wouldn't have seen.
Field-extraction accuracy on 300 real documents, vendor versus a six-week internal prototype
Two points behind, and it took six weeks longer to get there. The accuracy question was never the one that decided this.
Here's the turn: the real test was never whether the vendor could read a bill of lading well enough. It was whether Hallowbrook's own 40,000 documents had anything left to teach a model that a vendor's own ten million documents hadn't already taught it.
We weren't testing whether the vendor could read a bill of lading. We were testing whether our own documents had anything left to teach it.
The actual tree Odalys ran, in the order the conditions get tested, not the order build and buy usually get argued in a meeting.
Knowledge spark: why does data volume matter this much here?
A document-reading model learns the shape of a form mostly from having seen thousands of versions of it before. A vendor selling to hundreds of freight companies has already seen more bill-of-lading formats than any one company will generate on its own. Hallowbrook's own 40,000 documents are a real number, just a small one next to that.
Monthly cost, vendor pricing versus in-house cost per document, by document volume
Vendor, flat 40 cents a documentIn-house, cost per document as volume grows
This is the kill criteria drawn as a number. Hallowbrook runs about 14,000 documents a month today. Below the crossover, buy wins outright.
At its worst, skipping this test costs exactly what almost happened here: real engineers spending a real quarter closing a gap that was mostly already closed, discovered only after a competitor shipped the same capability in three weeks using one of the same three vendors.
The choice I would take back
Assuming, without testing it, that Hallowbrook's own documents were too specific for a generic vendor tool to handle. That assumption felt reasonable before anyone had actually run the vendor against the worst 300 real cases. It stopped holding up the moment the numbers came back two points apart.
What I would leave alone: the part of ManifestRead that flags shipments matching Hallowbrook's own fraud patterns stays in-house. That's the opposite of a commodity problem, our claims history genuinely is the edge there, so the kill criteria never applies to it.
The lesson: "our documents are different" is an assumption, not a fact. Test it against your worst 300 cases before you let it decide a whole quarter of engineering time.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what it actually felt like to nearly commit a quarter before anyone ran the real test.
Odalys Ferran could price out an engineering quarter in her head before an estimate ever hit a slide. Three years running AI features at Hallowbrook Freight had taught her exactly what a rebuild costs in real weeks, not story points.
When ManifestRead came up, the plan leaned toward building it in-house from day one. Hallowbrook's bills of lading looked messy in a way that felt specific to freight brokerage: water stains, handwriting crammed into margins, three different customs formats depending on the lane. It felt like exactly the kind of mess a generic vendor tool wouldn't have seen before.
One of these mistakes gets noticed the same afternoon. The other one gets noticed a quarter later, usually from a competitor's launch announcement.
Then Grady Wemyss, the ops manager who'd actually live with whichever tool shipped, asked one question in the kickoff: "Have we run any of the three vendors against our own worst documents, or just watched their demo?" Nobody had.
Odalys asked for a week before committing engineering time to anything. The team pulled 300 of Hallowbrook's ugliest real scans, the ones clerks already dreaded, and ran them through the top vendor with no special setup. It read the key fields correctly 91 percent of the time. Meanwhile, an early internal prototype, six weeks into its own build, scored 89 percent on the identical set.
We were about to spend a quarter closing a two-point gap that a phone call could have closed in a week.
Odalys never had a fixed rule for exactly when "our documents are different" was true versus a story the team was telling itself. It came down to one real test: did the vendor's mistakes cluster somewhere Hallowbrook's own data would obviously fix, or were they scattered the same way the in-house model's mistakes were scattered. They were scattered the same way. That was the whole answer.
Nobody would weld their own forklift to avoid a purchase order. Odalys realized the team was about to do the document-reading version of exactly that.
Back in the kickoff, leaning toward building wasn't an unreasonable instinct. The team had genuinely never checked whether Hallowbrook's documents held anything a stranger's data hadn't already covered. It stopped being reasonable the moment a one-week test could answer the same question for almost nothing.
Here's the replay: the vendor got a capped, two-year contract with an exit clause, signed six weeks after that kickoff instead of a full quarter later. The two engineers who would have spent that quarter on ManifestRead went straight to the fraud-flagging model instead, the one place Hallowbrook's own claims history genuinely was the edge. That model shipped six weeks after the vendor deal closed.
One version of this story spends a full quarter building a document reader, ships two points worse than a vendor anyone could have called, and delays the one feature that actually needed Hallowbrook's own data. The other spends a week testing, signs a capped contract, and gets the real differentiator built on schedule.
What I'd tell myself, hearing Grady's question land in that kickoff: "our data is different" is the assumption every team reaches for first, and it's worth exactly nothing until you've actually tested it against your ugliest real cases.
PICK, run on a decision everyone assumed meant buildingNot a script for always buying. PICK is what stops "our data is different" from deciding a whole quarter by itself.
P
Position. The pick, stated before any reasoning.
Buy ManifestRead's field extraction. A bill of lading is a solved problem elsewhere, and Hallowbrook's own documents don't teach a model anything a stranger's ten million documents haven't already taught it.
Stating the pick first is what separates a decision from a list of factors still being weighed.
I
Impact. Who feels each kind of mistake, and how.
A clerk feels the vendor's rare miss, four minutes to re-key one document. Engineering time and the roadmap feel the cost of building a commodity feature instead of the real differentiator.
Naming both sides in real units is what makes the next step possible.
C
Cost asymmetry. Which mistake is cheap, which one is expensive and hidden.
The vendor's misses are cheap, visible, and absorbed the same afternoon. Building a solved problem in-house is expensive and hidden, it shows up a quarter later, often only when a rival ships first.
This is the reasoning the whole pick runs on, not a preference between "safe" and "innovative."
K
Kill criteria. What evidence would flip the pick.
If the vendor's accuracy plateaus below Hallowbrook's real bar, or in-house cost per document falls under the vendor's flat price at real volume, revisit building it.
A pick with no kill criteria isn't a decision, it's a habit nobody agreed to keep.
The recap, one line per letter: position is buy, stated first; impact is a clerk's four minutes against an engineering quarter; cost asymmetry is why the quarter is the real mistake, not the vendor's rare miss; and kill criteria is the accuracy floor or the volume crossover that would send Hallowbrook back to building.
And if you want to be sure it really works, try it somewhere elseSame four letters, a claims intake tool instead of a freight document reader. This time the capped-contract clause is what actually mattered.
Ansel Vance is the AI PM at Corrigan Claims, which processes insurance claims for regional carriers. A vendor pitches a tool that reads scanned claim intake forms and pulls out the policy number, the incident date, and the claimed amount. Mapped onto PICK: position is buy, the same logic applies, intake forms are a solved, common problem, and Corrigan's own forms don't teach a model anything a stranger's claims data hasn't already covered. Impact is the same shape, a processor re-keys the vendor's rare miss in under a minute, versus engineers spending a quarter on a commodity reader. Cost asymmetry holds too. Where this version differs is kill criteria: Corrigan's vendor contract had no exit clause and a three-year term, so Ansel's real fight wasn't build versus buy, it was capping the contract at one year with a real out, because the lock-in cost, not the accuracy, was the actual risk worth negotiating over.
Same four letters, same order of testing, but Corrigan's real negotiation was the exit clause, not the accuracy number.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "buy when it's a solved problem and your own data isn't the edge, test it against your worst cases first, cap the contract," and stop.
Cost: no time to run a real week-long test before a decision is due. Say so honestly, and commit to running it against ten worst-case documents by end of day, not a guess dressed up as confidence.
The model got better, for real: say the in-house prototype had scored 96 percent instead of 89. Rerun the pick anyway, a genuinely better in-house option just changes where the crossover sits, it doesn't change the method.
Where people run it wrong.
They assume "our data is different" without testing it against real worst cases.
They compare a vendor's demo accuracy to an in-house target instead of to the same ugly documents.
They negotiate the sticker price and never touch the exit clause, which is usually where the real cost is hiding.
How to use it live. The moment an interviewer asks a build-versus-buy question, ask yourself first: would our own data actually make this model better, or are we just assuming it would because it's ours? That question alone buys real thinking time, and it's usually where the honest answer starts.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits deciding when buying a vendor solution is the right call?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It commits to a pick first, then shows which mistake is the one worth designing against.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Odalys Ferran, the AI PM deciding ManifestRead's build-or-buy call at Hallowbrook Freight. Grady Wemyss is the ops manager whose question, "have we tested this against our own worst documents," reorders the decision.
3 · THE POSITION
What's the actual pick, stated in one line?
Tap to flip
ANSWER
Buy the field-extraction piece. It's a solved problem elsewhere and Hallowbrook's own documents don't teach a model anything a vendor's much larger document set hasn't already taught it.
4 · THE ASYMMETRY
Which mistake is cheap, and which one is expensive and hidden?
Tap to flip
ANSWER
The vendor missing a messy scan is cheap: a clerk re-keys it in minutes. Building a solved problem in-house is expensive and hidden, the cost shows up a quarter later, often only after a rival ships first.
5 · THE OLD DECISION
What old assumption does this answer take back?
Tap to flip
ANSWER
Assuming, without testing it, that Hallowbrook's own documents were too specific for a vendor tool. Reasonable before anyone ran the real test. Wrong the moment the numbers came back two points apart.
6 · THE NUMBER
Fill in the blank: the vendor scored ___ percent on the 300-document test, and the six-week internal prototype scored ___ percent on the same set.
Tap to flip
ANSWER
91 percent versus 89 percent. Two points apart, and one of them took six extra weeks to get there.
7 · THE REPLAY
Same decision, tested before committing instead of built on assumption. What changes?
Tap to flip
ANSWER
A capped, two-year vendor contract signed six weeks after kickoff, and two engineers redeployed onto the fraud-flagging model, the one place Hallowbrook's own data really was the edge, shipped on schedule.
8 · CROSS PRODUCT TRANSFER
Section 4 runs PICK again on a different product. Which one, and what mattered most there?
Tap to flip
ANSWER
Corrigan Claims' intake-form reader. Same buy decision, but the real fight there was the contract's exit clause, not the accuracy number, since the vendor's original three-year lock-in was the actual risk.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: the vendor scored ___ percent on the 300-document test, versus ___ percent for the six-week internal prototype.
Show hint
Look at the grouped bar chart in "Let's learn."
Show answer
91 percent and 89 percent. A two-point gap that took six extra weeks of engineering to close.
Multiple choice
2. Why is building a commodity document-reading capability the expensive mistake, per this answer?
A. Because vendor contracts are always cheaper than engineering salaries.
B. Because the cost is hidden and shows up a quarter later, often only after a competitor ships the same capability first.
C. Because in-house models are always less accurate than vendor models.
D. Because engineers refuse to build features vendors already sell.
Show hint
Look at the Cost asymmetry step in the PICK recap.
Show answer
B. The vendor's rare miss is cheap and visible the same afternoon. A wasted engineering quarter is expensive and invisible until it's too late to get the time back.
True or false
3. True or false: Hallowbrook decided to buy ManifestRead's fraud-flagging piece from the same vendor as the field extraction.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. Fraud-flagging stayed in-house, because Hallowbrook's own claims history genuinely is the edge there, the opposite of a commodity problem.
Short answer, name the reversal
4. What old assumption does this answer take back, and why did it make sense at the time?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Assuming Hallowbrook's documents were too specific for a vendor tool to handle. It made sense before anyone had actually tested it, and stopped making sense the moment a one-week test showed a two-point gap, not a real difference.
Short answer, apply it yourself
5. Think of a feature you've heard pitched as "we should build this ourselves." Would your own data actually make it better, or would it just feel more familiar?
Show hint
Ask whether the problem is common outside your company, and whether your own examples would teach a model something a stranger's examples couldn't.
Show answer
Model answer: A team considered building its own resume parser. Testing a vendor against 50 real resumes showed no real gap, resumes are a solved, common shape, so the honest call was buy, with the team's own data adding nothing a stranger's resumes hadn't already covered.
Short answer, work the number
6. If Hallowbrook's document volume grew from 14,000 to 50,000 a month, does the buy decision still hold, based on the crossover chart?
Show hint
Look at where the in-house cost line crosses under the vendor's flat 40-cent price in the second chart.
Show answer
Model answer: Not automatically. The crossover sits around 45,000 documents a month, so at 50,000, in-house cost per document may have genuinely fallen below the vendor's flat price, which is exactly the kind of change the kill criteria exists to catch before the contract auto-renews.
Before you close the answer
Why this works
Tests whether you can commit to a real buy-or-build pick and defend it with a genuine cost asymmetry, instead of listing pros and cons that never resolve into a decision.
Follow-up traps
"What if the vendor raises prices once you're dependent on them?" Response: that's exactly what the kill criteria and a capped contract term are for, the crossover chart gives you a real number to renegotiate against, not just a bad feeling.
"Isn't testing against worst-case documents unfair to the vendor?" Response: no, it's the only fair test, a demo is chosen to look clean, and the worst 300 real cases are what the tool will actually face in production.
If pressed
The vendor contract kept a data-export clause, so if the crossover volume is ever crossed for real, Hallowbrook can pull its own labeled corrections out and use them to bootstrap an in-house model, instead of starting the whole test from zero.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.