CaseAdvancedAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #8
A vendor offers 90 percent of what you need. How do you evaluate closing the last 10 percent yourself?
BOUND price the last 10 percent like a real budget line, not a feeling that it should be easy
Harborview Insurance processes auto claims. FieldPilot is the vendor tool that reads a submitted accident report and estimate, and pulls out the standard fields automatically. Renata Sooklal runs claims operations, and one near miss on a handwritten estimate is what forced her to actually price the gap instead of guessing at it.
The direct answer
Treat it as real arithmetic, not a feeling that the last 10 percent should be easy to mop up. Price what closing the gap yourself costs: engineering time, data curation, a new eval set, and ongoing maintenance. Price what living with the gap costs: the manual review time plus the real cost of the errors that slip through. Only build it yourself once the build cost is clearly, provably cheaper than the ongoing cost of the gap, and check that math again once a year, because the honest answer is often to keep routing it to a person.
Do this, in order
Write the equation before touching a single number.Why: cost to build versus cost to live with the gap, stated out loud, keeps the decision from becoming a gut call.
Pull a real sample of the missed 10 percent before estimating anything.Why: you can't price a gap you've never actually measured the shape of.
Give a range for the build cost, not one confident number.Why: engineering estimates for messy, rare cases are almost never as tight as they sound in a planning meeting.
Sanity-check the total against something you already know, like a headcount cost.Why: if closing the gap costs more than hiring a person to handle it, the arithmetic is telling you something.
Name which single assumption would flip the decision, and watch it.Why: the honest answer today can flip fast if the gap's volume grows or a missed case gets expensive.
How to answer this, stage by stage
Nobody is scoring whether you'd like to close the gap. They're scoring whether you can price it honestly enough that the answer might come back "don't."
Stage 1
Scope it to one real vendor gap
Say it like this
"I'll ground this in one real case. Harborview Insurance uses a vendor tool, FieldPilot, to pull data out of auto claims. It handles about 90 percent of claims cleanly. The other 10 percent are handwritten estimates and non-standard forms it can't read. That's the gap I'd actually be pricing."
Why this works
Stops "the last 10 percent" from staying an abstraction with no real shape to it.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as BOUND. Break it down, state the equation first. Own the numbers, and say where each one came from. Use a range, not one falsely precise figure. Nail a sanity check against something I already know. Direction, which single assumption would swing this the most."
Why this works
Shows a repeatable estimation method instead of a guess dressed up with confidence.
Stage 3
Reframe: this isn't "can we build it"
Say it like this
"The question isn't whether engineering could close this gap. Almost any gap can be closed with enough time. The real question is whether it's cheaper to close than to just keep routing it to a person, forever."
Why this works
This is where a strong answer separates from "yes, we could build that," which answers a different question.
Stage 4
Give the one decision, with the equation behind it
Say it like this
"Cost to close it equals engineering time, plus data curation, plus a new eval set, plus a year of maintenance. Cost to live with it equals manual review hours times the loaded cost per hour, plus whatever real losses slip through. Here, that came out to about 120,000 dollars to build against roughly 24,000 dollars a year to live with the gap, so the honest call is: don't build it yet, keep routing it to a person, and revisit the math if the gap grows."
Why this works
This is the direct answer, with the actual numbers behind it instead of a confident-sounding guess.
Stage 5
Prove it with the compressed near miss
Say it like this
"A handwritten estimate came through with a smudged digit. FieldPilot's fallback OCR misread it, and the claim nearly auto-approved 8,400 dollars over the real estimate. A reviewer caught it by chance, not by design. That's the real cost of living with the gap: it's not zero, it's just rare enough that it didn't change the math on its own."
Why this works
Grounds the abstract "cost of errors" term in the equation with one real, countable near miss.
Stage 6
Name the AI-specific reasoning, the trade-off, and close
Say it like this
"The honest reason this gap is expensive to close isn't the coding, it's that handwriting recognition on messy, real-world estimates fails in ways that are genuinely hard to eval against, so the new eval set and ongoing tuning cost more than people expect. We accepted living with a small, known error rate in exchange for not sinking six figures into a fix that might not even fully work. If the gap doubled in volume, or one of these errors actually shipped instead of getting caught, I'd rerun this math today."
Why this works
Names the model-specific reason the estimate is hard (handwriting, not a generic coding task) and states the trade-off being accepted out loud.
Let's learn
Every month, Renata Sooklal's claims team spent roughly 560 hours simply retyping paperwork a vendor tool was supposed to have already handled.
FieldPilot reads a submitted accident report and estimate and pulls out the standard fields, date, VIN, adjuster notes, automatically. On about 2,000 claims a month, it gets roughly 1,800 of them clean. The remaining 200, mostly handwritten damage estimates from small independent repair shops and a handful of non-standard state forms, still need a person to type them in by hand, about 14 minutes each.
The last 10 percent is rarely one thing. It's usually a handful of different messy cases, each with its own real cost to close.
Here's the turn: the manual typing itself was never really the problem, the team had already built that into their staffing. The real question showed up the first time someone asked whether Harborview should just build its own handwriting extractor to close the gap for good, and nobody had actually priced what that would cost against what the gap was already costing.
First-year cost to close the gap yourself, by line item
EngineeringData curationEval setMaintenance
Against a gap costing roughly 24,000 dollars a year to live with, 120,000 dollars to build doesn't pay for itself for five years, before counting any maintenance after year one.
We weren't asking whether the gap could be closed. We were asking whether closing it was cheaper than the person already closing it by hand.
At its worst, skipping this arithmetic means sinking six figures into a custom handwriting extractor for a slice of claims that was already being handled, reliably enough, by one person's typing, and still not fully closing the gap, because messy handwriting stays messy no matter how much engineering you throw at it.
The choice I would take back
Pricing FieldPilot per claim processed. That made the team reluctant to pull a large enough sample of the missed 10 percent to know its real shape, so early guesses about the gap were based on a handful of complaints, not a real count. It made sense when the vendor bill was the only cost anyone was watching. It stopped making sense once a real decision about building a patch needed a real number to stand on.
What I would leave alone: a genuinely rare edge case, like claims involving classic or vintage cars, under half a percent of volume, doesn't earn its own custom build no matter how annoying it is when it shows up. Route it to a person and move on.
The lesson: "the last 10 percent" isn't one number, it's usually several different gaps with different real costs. Price each one against what it already costs to live with, and let the arithmetic say no when it says no.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what actually pushed Renata to run the real numbers instead of trusting a feeling.
Renata Sooklal could tell a real coverage dispute from a simple paperwork error before she'd finished her coffee, four years running claims operations at Harborview will do that. FieldPilot had been live for a year, and her team had settled into a rhythm: 1,800 claims a month moved through clean, and the 200 that didn't got quietly typed up by a rotating pair of processors.
The mistake wasn't ignoring the gap. It was treating every part of it as equally worth building for.
A new claims director had floated the idea of building a custom handwriting model to close the gap for good, and it sounded reasonable, ten percent felt small enough to mop up quickly. Nobody had actually priced it. Then, on a Tuesday afternoon, a handwritten estimate came in with a smudged final digit.
Nobody designed a catch for this. It worked because a reviewer happened to double back on the file.
We didn't almost lose 8,400 dollars to a hard problem. We almost lost it to a gap nobody had ever actually priced.
FieldPilot's fallback OCR read the smudged estimate as 2,400 dollars instead of 8,400. The claim moved into the auto-approval queue, and a reviewer only caught it because she happened to reopen the file to check an unrelated detail. It rattled the team enough that Renata finally asked for real numbers instead of a feeling.
A two-week sample of 200 flagged claims showed the gap breaking into a few real shapes: handwritten estimates, a couple of out-of-state forms, and a small number of second-language submissions. Engineering estimated six weeks to build a rules-plus-model patch for the handwriting piece alone, roughly 75,000 dollars, plus 15,000 for curating enough real examples, 8,000 for a proper eval set, and 22,000 a year to keep it maintained as handwriting styles and forms kept drifting.
Four real line items, not one round number. Each one had to be priced on its own.
Knowledge spark: why is handwriting harder to close than it sounds?
A standard form has fields in the same place every time. Handwriting varies person to person, and a small crimp or smudge changes a digit's shape entirely. Closing this gap needs enough real examples to teach a model the actual variety, not just a handful of clean samples, which is what pushes the cost past what a typical coding task would run.
Against 200 manual claims a month at 14 minutes each, roughly 24,000 dollars a year in loaded labor cost, the 120,000 dollar first-year build didn't come close to paying for itself, even before counting the 22,000 dollars a year it would keep costing after that. Building it made sense as an idea. It stopped making sense the moment real numbers sat next to each other.
Harborview's gap was common but not costly enough yet, and the vendor had no roadmap fix in sight, which is what kept the honest answer at "don't build."
Here's the replay: with the real numbers in hand, Renata kept the gap routed to her team, added a second reviewer check specifically for estimates over 5,000 dollars, and set a standing quarterly recheck of the math. Six months later, the gap's volume hadn't grown enough to change the answer, and the near miss never repeated, because the new check would have caught it even without a model.
What I'd tell myself, standing in that room when the handwriting-model idea first came up sounding so reasonable: ten percent is not a size, it's a placeholder for several different gaps, and none of them deserve a build decision until you've actually priced what they cost to live with.
BOUND, run on Harborview's vendor gapOne line per letter, if you want to say it fast under pressure.
B
Break it down. State the equation before any numbers.
Cost to build equals engineering plus data curation plus a new eval set plus maintenance. Cost to live with it equals manual review hours times loaded cost, plus the real cost of errors that slip through.
Without the equation stated first, every number that follows is just a guess wearing a decimal point.
O
Own numbers. State each assumption and where it came from.
200 affected claims a month, from a real two-week sample. 14 minutes per manual entry, from timing the actual process. Six weeks of engineering, from a real estimate against a scoped patch.
A number nobody can trace back to a source is a number nobody should trust in a decision this size.
U
Use a range, not one falsely precise figure.
Build cost: 95,000 to 140,000 dollars in year one, depending on how much real handwriting data curation actually turns up. Cost to live with it: 20,000 to 30,000 dollars a year, depending on claim volume that quarter.
A single number here would claim a confidence the estimate hasn't earned.
N
Nail the sanity check. Does it survive a smell test?
120,000 dollars is more than a full year of two extra claims processors. Unless the gap keeps growing or a missed case gets genuinely expensive, hiring beats building here, and living with it beats both.
If the number surprises you compared to something familiar, that's the moment to double check it, not ship it.
D
Direction. Which single assumption swings this most?
The gap's volume. At 10 percent it doesn't clear the bar. At 25 percent of claims, the same 120,000 dollar build starts paying for itself inside two years, and the decision flips.
Naming the swing assumption is what a real estimator says out loud, and a rushed one skips.
What moves the estimate most, by assumption
Gap volumeError costEstimate accuracyVendor timing
Gap volume swings the decision more than any other assumption. Watch that number, not the engineering estimate, if you want to know when to revisit the call.
The recap, one line per letter: break it down is the equation, build cost against living-with-it cost, own numbers is the real 200-claim sample behind every figure, use a range is 95 to 140 thousand dollars instead of one falsely precise number, nail the sanity check is comparing it to a headcount cost, and direction is the gap's volume, the one thing that would flip this decision if it changed.
And if you want to be sure it really works, try it somewhere elseSame five letters, a permit reviewer's vendor gap instead of a claims processor's.
Efe Adaramola reviews building-permit applications at Millbrace Permits, which sells software cities use to process them. A vendor tool pre-checks applications for missing required documents and clears 92 percent of submissions cleanly. The remaining 8 percent are applications bundled as single scanned PDFs instead of separate files, which the vendor tool can't split apart. Mapped onto BOUND: break it down is engineering time to build a PDF-splitting step, plus a small eval set of misfiled bundles, plus light ongoing tuning as new scanner software changes bundle formats. Own numbers: a one-month sample shows 340 affected applications, each costing a reviewer 9 minutes to manually split and refile. Use a range: build cost lands between 30,000 and 45,000 dollars, far lower than Harborview's gap, because splitting a PDF is a much narrower problem than reading handwriting. Nail the sanity check: 340 applications a month at 9 minutes is only 51 hours, about 15,000 dollars a year in labor, so even the low end of the build estimate takes two years to pay back. Direction: the swing factor here is whether the city's own scanning vendor fixes bundling on their end first, which would erase the gap for free. Efe's team decides to wait and recheck in two quarters rather than build, since the vendor fix is plausible and cheap to wait for.
Same method, a different shape of picture: Millbrace's decision plotted as a timeline instead of a tree, since the real story here is when to recheck, not which branch to take.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "price the build cost against the cost of living with the gap, in real numbers, and only build if the build cost clearly wins," and stop.
Cost: no time to pull a real sample before a decision is due. Say so honestly, and make pulling that sample the very next step, not a guess dressed as an estimate.
The model got better, for real: a new vendor release closes half the gap on its own. Rerun the numbers, since the remaining gap is now smaller and the build case gets weaker, not stronger.
Where people run it wrong.
They estimate the build cost and never price the cost of living with the gap, so there's nothing to compare it against.
They treat "10 percent" as a size instead of asking what specific cases make up that ten percent.
They skip the sanity check, so a 120,000 dollar build against a 24,000 dollar problem ships without anyone noticing the mismatch.
How to use it live. The moment someone says "the vendor gets us most of the way there," ask what the missing piece actually costs to live with, before asking what it would cost to build. That number alone usually settles the argument.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits evaluating whether to close the last 10 percent of a vendor gap yourself?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. It prices the decision as real arithmetic instead of a gut feeling.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Renata Sooklal, who runs claims operations at Harborview Insurance and is the one who finally priced FieldPilot's 10 percent gap after a near miss.
3 · THE EQUATION
What's the equation this whole answer runs on?
Tap to flip
ANSWER
Cost to build (engineering, data curation, eval set, maintenance) versus cost to live with the gap (manual review hours plus the cost of errors that slip through).
4 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pricing FieldPilot per claim processed, which made the team reluctant to sample the missed 10 percent enough to know its real shape, so early guesses were built on a handful of complaints instead of a real count.
5 · THE NUMBER
Fill in the blank: closing the gap was estimated at ___ dollars in year one, against ___ dollars a year to keep living with it.
Tap to flip
ANSWER
120,000 dollars, against 24,000 dollars a year. The build wouldn't pay for itself for five years, before counting any maintenance after year one.
6 · THE REPLAY
Same near miss, but after the real numbers got run. What changes?
Tap to flip
ANSWER
Renata keeps the gap routed to her team, adds a second reviewer check for estimates over 5,000 dollars, and the near miss doesn't repeat, no custom model required.
7 · WHERE IT WOULDN'T MATTER
Name a gap that would NOT be worth pricing out this carefully.
Tap to flip
ANSWER
A genuinely rare edge case, like claims involving classic or vintage cars, under half a percent of volume. Route it to a person and move on, no build decision needed.
8 · CROSS PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which one, and what's the call?
Tap to flip
ANSWER
Millbrace Permits' bundled-PDF gap. The call is wait, not build, since the build cost is small but so is the gap, and the vendor might fix it for free first.
Check yourself Score: 0 / 0
True or false
1. True or false: Harborview's real numbers showed that building a custom handwriting patch was clearly worth it in year one.
True
False
Show hint
Compare the build-cost chart against the yearly cost of living with the gap.
Show answer
False. The 120,000 dollar build cost didn't come close to paying for itself against a roughly 24,000 dollar a year cost of living with the gap, so the honest call was to keep routing it to a person.
Multiple choice
2. Why did closing Harborview's handwriting gap cost more than a typical coding task would?
A. The vendor refused to share their own extraction code.
B. Handwriting varies enough from person to person that closing the gap needs real, varied examples and a genuine eval set, not just a quick rule.
C. Harborview's engineering team had never built an extraction tool before.
D. Insurance regulations required a third-party audit of any new model.
Show hint
Look at the knowledge spark on why handwriting is harder to close than it sounds.
Show answer
B. Handwriting's real variety, not the coding itself, is what pushes the cost of data curation and evaluation higher than people expect going in.
Fill in the blank
3. Fill in the blank: the near miss involved an estimate that FieldPilot misread as ___ dollars instead of the real ___ dollars.
Show hint
Look at "here's the turn" in the story section.
Show answer
2,400 dollars instead of 8,400 dollars. A reviewer caught it by chance, not by any built-in safeguard, which is what finally pushed Renata to price the gap properly.
Short answer, where it wouldn't matter
4. Name a vendor gap at Harborview where this careful pricing exercise would NOT be worth running, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Claims involving classic or vintage cars, under half a percent of volume. It's rare enough that no build decision is worth pricing out, just route it to a person.
Short answer, apply it yourself
5. Think of a tool you use that mostly works but has a known gap. What would it actually cost, in real hours or dollars, to live with that gap forever instead of trying to fix it?
Show hint
Think about how often the gap shows up and what you or someone else does instead when it does.
Show answer
Model answer: A calendar app that can't parse handwritten meeting notes means someone retypes them, maybe 10 minutes a week. That's under 9 hours a year, cheap enough that building a fix would rarely be worth it.
Short answer, work the number
6. If the gap's volume had been 500 claims a month instead of 200, would the same "don't build" call still hold?
Show hint
Recompute the yearly cost of living with the gap at the higher volume and compare it to the 120,000 dollar build estimate.
Show answer
Model answer: No, likely not. At 500 claims a month, the yearly cost of living with the gap climbs toward 60,000 dollars, which starts to make the 120,000 dollar build pay for itself in about two years, closer to a real case for building.
Before you close the answer
Why this works
Tests whether you can price a build decision with real arithmetic, including the honest possibility that the answer is no. Most candidates assume closing any gap is automatically worth it.
Follow-up traps
"But doesn't the near miss prove you need to build it?" Response: the near miss proved the gap has real risk, not that building beats a cheaper fix like a targeted second review on high-value estimates.
"What if engineering underestimated the build cost?" Response: that's exactly why a range was used instead of one number, and why the sanity check against a headcount cost was run before committing.
If pressed
The added second-review rule specifically targets estimates over 5,000 dollars, since that threshold covers every near miss the team has seen so far while adding review time to only about 30 claims a month, not the full 200.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.