CaseAdvancedAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #18
What would make you reverse a build decision six months in?
FLIPSthe tell wasn't the near miss, it was the spreadsheet a lawyer built to stop trusting the tool
Ledgerwick sells contract review software to corporate legal teams. Clausevane is the in-house-built feature that reads a vendor contract and flags the clauses worth a second look. Aurelius Radulescu is the AI PM who chose to build it instead of buying a vendor's clause-detection API, and Petra Ulvestad is the senior counsel whose private spreadsheet became the real signal that the build needed reversing.
The direct answer
I would reverse the build the moment the review team starts building its own manual system to double-check it, not the moment a single error happens. That is the real signal: the people using the tool no longer trust its coverage enough to skip checking, which means the in-house model's upkeep has quietly gotten more expensive than paying a vendor whose whole job is tracking this kind of drift every day.
Do this, in order
Reverse the moment the team starts a private manual workaround, not at the first single error.Why: a workaround is proof the model's coverage no longer holds, whether or not anyone has said so out loud yet.
Watch a fixed weekly eval set's catch rate, not the raw count of complaints.Why: it is the number that moves for weeks before any single contract actually goes wrong.
Trace any drop back to the date of the last underlying model update before blaming your own data.Why: a foundation-model version bump can quietly change behavior nobody at your company touched.
Price the hours the team spends working around the tool against a vendor's quote.Why: in-house upkeep is a real cost even when it never shows up on an invoice.
Keep the eval history and version log you didn't think you needed at launch.Why: without a history, drift looks like one bad day instead of a pattern worth acting on.
Don't reverse on one scare alone; confirm the workaround is real and still growing first.Why: a single near miss can be bad luck, but a spreadsheet that keeps growing every week is a trend.
How to answer this, stage by stage
Nobody is scoring whether you can describe a postmortem. They're scoring whether you can name the tell that shows up while there's still time to act on it.
Move 1
Scope it to one real build decision
Say it like this
"Let me pick one real case. Ledgerwick sells contract review software to corporate legal teams. Six months ago we built Clausevane in house, a feature that flags risky clauses, instead of buying a vendor's clause-detection API. That's the decision I'd actually watch for reversal signals."
Why this works
Keeps the answer from turning into a generic "know when to pivot" speech with nothing real underneath it.
Move 2
Say the structure out loud before any content
Say it like this
"I'll run this as FLIPS. Find the person it actually affects. Locate the habit the tool built in them. Identify the flip, the exact moment their behavior snaps. Pinpoint the old decision that only made sense before. Show the replay with the decision reversed."
Why this works
Signals a repeatable way to think about "when do you reverse," not a one-off gut call dressed up as instinct.
Move 3
Reframe: it isn't "was building wrong," it's "what's the tell"
Say it like this
"This isn't really asking whether building was a mistake. It's asking what specific, visible thing would tell me that's happening, before the postmortem, while there's still time to actually reverse it."
Why this works
Separates a strong answer from someone who can only describe reversing after the fact, once the damage is already done.
Move 4
Give the one decision
Say it like this
"I'd reverse the moment the legal team starts keeping their own private list to double-check Clausevane instead of trusting its flags. That's not a mood, it's an action I can see: a spreadsheet, a shared doc, a second read nobody asked them to do. Once that exists, the model's real coverage has already dropped below what the team needs, whether or not anyone's told me yet."
Why this works
This is the direct answer, stated as something you could actually watch for, not a vague sense that "the vibes are off."
Move 5
Prove it with the compressed story
Say it like this
"Clausevane launched catching 97 percent of our clause eval set. Nobody touched the model, but the provider pushed an update in week seven, and by week ten we were down to 81, and one contract nearly got signed with the wrong renewal window. Petra caught it by luck, and by the next week she'd started her own spreadsheet, checking every contract by hand."
Why this works
Compresses the whole case into the two real numbers and the one behavior change that actually mattered.
Move 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't about one bad clause is that the model's owner, the provider, can change its behavior under us without us shipping anything. We could keep patching, retraining a small model of our own every time a clause type slips, but I'd reject that, since it just means we're permanently chasing a provider's release schedule, a race two people can't win. Buying instead means a recurring fee, in exchange for a vendor whose whole team owns watching for drift, not two of ours."
Why this works
This is the load-bearing, AI-specific judgment. A normal feature doesn't quietly get worse because someone else updated a dependency's brain.
Move 7
Say what wouldn't make you reverse, then close
Say it like this
"I wouldn't reverse over one near miss alone, or because a competitor bought a vendor tool instead. I'd reverse when the workaround is real and still growing, week over week, on a metric I already track. For Clausevane, that line got crossed by week ten, and that's the one thing that actually changed my decision."
Why this works
Closes with judgment about false alarms, and restates the direct answer in one breath.
Let's learn
Here is what happens when a tool works fine for months, then, for one week in the middle of a normal quarter, quietly stops.
Clausevane is a feature inside Ledgerwick's contract review software. It reads a vendor contract and flags the clauses worth a second look: indemnification, liability caps, auto-renewal windows.
Before Clausevane, in-house counsel read every full vendor contract start to finish, about 45 minutes each, across roughly 30 contracts a month. With Clausevane, a lawyer reads the 3 or 4 flagged clauses instead of the whole document, about 8 minutes a contract, for the same 30 contracts.
Clausevane's clause catch rate, weekly, on a fixed 240-contract eval set
Nobody at Ledgerwick changed Clausevane's code in week seven. The provider updated the general model underneath it, and the catch rate started sliding the same week, three weeks before anyone noticed.
Here's the turn: the occasional missed clause was never the real problem. The real problem was that nobody had a way to see the miss rate creeping up, so the team's trust broke before anyone knew the model had.
We did not lose one flagged clause. We lost the eight minutes Clausevane was supposed to save Petra every morning, and then some.
The five moves, held up as one page. The I step is the hard one, and the only one drawn in a different color on purpose.
At its worst, the review team ends up spending nearly as much time double-checking Clausevane by hand as they'd have spent just reading the contract themselves, while the tool still looks trusted on paper, since nobody has said otherwise out loud.
The choice I would take back
Ledgerwick never kept a running eval history, a way to compare this month's catch rate against last month's. That made sense at launch, when there was only one version of the model to test. It stopped making sense the moment an outside provider could change that model without anyone at Ledgerwick touching a line of code.
What I would leave alone: I would not touch how Clausevane handles ordinary, boilerplate indemnification language in standard vendor paper. That behavior hasn't drifted at all, and rebuilding it would spend real engineering time solving a problem that doesn't exist.
The lesson: six months isn't a deadline, it's just how long silent drift takes to add up into something someone finally builds a workaround for. Watch for the workaround. Don't wait for the calendar.
Now here is the same thing as a story
Use the short version above out loud in the room. Read this one when you want to feel exactly how a spreadsheet quietly became Ledgerwick's real quality process.
The eval set lived in a spreadsheet nobody at Ledgerwick had opened in four months.
Petra Ulvestad had spent eleven years reading vendor paper by the time Clausevane launched. She could spot a hidden auto-renewal clause on the second page of a contract before she'd finished her coffee.
When Clausevane arrived, Petra's mornings changed. She'd open a contract, glance at the three flagged lines, sign off by 9:15, instead of reading start to finish until noon. For six weeks, that held.
Six months, laid flat. The gap between the update and the near miss is the part worth staring at.
The habit thinned in three beats. At first she still skimmed the rest of the contract for ten seconds out of old habit. By week four she'd stopped that. By week seven she wasn't rereading the flagged clauses twice either, just the model's one-line summary of each.
Knowledge spark: why would a model get worse with no one touching it?
Clausevane runs on a general-purpose model Ledgerwick doesn't own. When the company behind that model ships a new version, the way it reads and summarizes text can shift, even though nobody at Ledgerwick changed a setting. That's a real risk with any bought or API-based model, not a bug in Clausevane specifically.
In week nine, a vendor renewal contract came through. Clausevane flagged the auto-renewal clause as standard, 30-day notice. It wasn't. This one had a 90-day window, buried in a cross-reference two sections later. Petra almost signed off. A junior paralegal, checking something unrelated, mentioned the 90-day line out loud in the hallway.
The number moved slowly for ten weeks. Petra's behavior moved once, on one afternoon.
Petra didn't file a ticket that week, or tell Aurelius. She opened a blank spreadsheet that night and started logging, contract by contract, what Clausevane flagged against what she found reading the full document herself. Within a month it was eating six, then nine hours a week, on top of her regular caseload.
We didn't take one clause from her. We took the eight minutes Clausevane was supposed to give her back, and then some.
Petra never had a fixed number for when she stopped trusting the flags. She had a switch: trust the model's read of a contract, or read the whole thing herself and log it. There was no setting in between where she "checked a bit more carefully."
Thirteen more misses out of the same 240. Small enough on paper. Large enough for Petra to stop trusting the sample entirely.
Back when Clausevane was first scoped, nobody suggested tracking catch rate over time. "It's a static model, we test it once at launch," someone said in that meeting, and it sounded reasonable, since nothing was going to change unless Ledgerwick decided it should.
The one full-page idea of this whole answer: Petra never had a dial to turn. She had a switch, and it flipped once.
Rerun with a version log in place: a weekly eval set catches the drop at week eight, 88 percent, before the near miss ever happens in week nine. Aurelius reruns the eval against the new provider model, confirms the drop, and greenlights a vendor pilot the same week, three weeks before Petra would have opened her spreadsheet at all.
Hours per month spent on Petra's private workaround spreadsheet
The near miss lands in month three. The workaround was already 9 hours a month before that, and it never stopped climbing once it started.
One version of this story needs a lucky hallway comment to catch a 90-day clause. The other has already caught the drift on a Tuesday afternoon dashboard, two mornings before anyone signs anything.
What I'd tell myself, sitting in that scoping meeting: the test you skip because "it's static" is exactly the one you'll wish you'd kept the day it stops being static without telling you.
The five moves, in the order I'd actually check themFLIPS, run backward from the spreadsheet Petra built, not forward from the launch date.
F
Find the person. Whose morning is this?
Petra Ulvestad, senior contract counsel, eleven years reading vendor paper, the one whose sign-off actually carries the legal risk.
Without a real person, "reverse the decision" stays an abstract policy nobody can actually watch for.
L
Locate the habit. What did she stop doing because it worked?
Rereading the full contract after the flagged clauses. First a ten-second skim, then nothing at all by week seven.
That habit, not the time saved, is the actual thing Clausevane shipped.
I
Identify the flip. What verb snaps?
Trusting the flags outright, or quietly rebuilding her own manual check in a spreadsheet next to them. No middle setting, and it doesn't flip back on its own.
This is the hard step, and the one that gives the whole answer its shape.
P
Pinpoint the old decision. Which choice only made sense before?
No running eval history was kept, since the model was assumed to be static after launch, before an outside provider could change it out from under Ledgerwick.
A reasonable call in month one becomes a blind spot the day the model stops being static.
S
Show the replay. Same day, new design.
A weekly eval set catches the same drop at week eight, three weeks before the near miss, and Aurelius reverses to a vendor pilot before anyone has to sign anything on faith.
This is the direct answer, proven with a real Tuesday instead of asserted as a policy.
The recap, one line per letter: find the person who actually carries the risk, locate the habit the tool quietly builds in them, identify the two-setting flip with no middle, pinpoint the old decision that assumed the model would hold still, and show a replay where a kept eval history catches the drop three weeks earlier than a hallway comment did.
And if you want to be sure it really works, try it somewhere elseSame five letters, a grain co-op instead of a law firm. This time the flip runs the other way: the model got better, and that's what broke the habit.
Isbrand Kloss is the agronomist at Rowancroft, a grain cooperative, who spent ten years walking fields before every planting decision. Yieldbrace is the crop-yield tool Rowancroft built in-house instead of buying a specialized agronomy vendor model. Mapped onto FLIPS: find the person is Isbrand. Locate the habit is his weekly walk through a sample of fields to sanity-check Yieldbrace's number by eye. The flip here is an over-trust flip, not a workaround one: after a real model upgrade pushed overall forecast accuracy from 82 to 95 percent, Isbrand stopped walking the sample fields at all, not more carefully, not less often, just not at all. Pinpoint the old decision: Yieldbrace only ever showed one confidence number on screen, never a range, so once that single number "felt right," nothing prompted a spot-check. Show the replay: adding a visible band on days when the model's own uncertainty runs high brings back exactly one field walk a month, on exactly the days it matters, and it catches one bad frost-damage read before a planting decision gets made on it.
Same five letters, a different flip, and a decision that branches on how far the workaround has actually spread.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "reverse when the team quietly builds its own manual check around the tool, not at the first bad output," and stop there.
Cost: no time to build a real eval set before deciding. Say so honestly, and start counting workaround hours as your interim signal instead.
The model got better, for real: if the team just relaxed and stopped checking entirely once the news was good, that's still a flip worth watching, an over-trust one, and it still deserves a visible signal instead of silence just because the news is good.
Where people run it wrong.
They wait for one dramatic failure instead of watching for a trend building over weeks.
They blame the review team for "not trusting the tool enough" instead of treating the workaround itself as data.
They treat six months as a deadline instead of realizing it's just how long silent drift usually takes to add up into something visible.
How to use it live. The moment an interviewer asks what would make you reverse a build, ask yourself: what would the team be doing with their hands the week before I found out? If the honest answer is "quietly building their own version of the tool by hand," that's your signal, and it buys you real thinking time to say it out loud.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
The workaround flip: the person keeps using the tool, but builds a private manual process around it because they no longer trust its coverage.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Petra Ulvestad, senior contract counsel at Ledgerwick, eleven years reading vendor paper, could once spot a hidden clause before finishing her coffee.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Rereading the full contract after Clausevane's flagged clauses. First a ten-second skim, then nothing at all by week seven.
4 · THE FLIP
What's the two-setting switch here?
Tap to flip
ANSWER
Trusting Clausevane's flags outright, or quietly rebuilding her own manual check in a spreadsheet next to them. No middle setting.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never keeping a running eval history to compare this month's catch rate against last month's, since the model was assumed to be static after launch.
6 · THE NUMBER
Fill in the blank: Clausevane's catch rate started at ___ percent and had dropped to ___ percent by week ten.
Tap to flip
ANSWER
97 percent, then 81 percent, after the provider's underlying model update in week seven.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
A weekly eval set catches the drop at week eight, 88 percent, three weeks before the near miss, and the team greenlights a vendor pilot before anyone signs anything on faith.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Rowancroft's Yieldbrace crop-yield tool, the over-trust flip: an agronomist stops field-checking the model entirely once an upgrade made it look better, not worse.
Check yourself Score: 0 / 0
Multiple choice
1. Why does this answer treat Petra's spreadsheet as the reversal signal, instead of the near miss itself?
A. A spreadsheet is easier for management to notice than a single contract error.
B. The near miss could be one-off bad luck, but a growing private workaround shows the team's trust has actually broken.
C. Aurelius asked Petra to keep it as part of her job.
D. The model's accuracy score had already proven the tool had failed.
Show hint
Look at the difference between a single scare and a growing, weekly habit.
Show answer
B. One scare can be luck. A workaround that keeps growing week over week is a real, measurable trend in trust, not a mood.
True or false
2. True or false: Clausevane's own code changed in week seven, and that's what caused the drift.
True
False
Show hint
Look at the knowledge spark about the provider's own model update.
Show answer
False. The underlying provider updated its general model. Nobody at Ledgerwick touched Clausevane's own code that week, which is exactly why the drift went unnoticed without a rerun eval set.
Fill in the blank
3. Fill in the blank: the catch rate dropped from ___ percent at launch to ___ percent by week ten, after the provider's update in week ___.
Show hint
Look at the line chart in "Let's learn."
Show answer
97 percent, 81 percent, week seven. All three real, measured points, and the update date is what turns "a bad week" into "a traceable cause."
Short answer, where it wouldn't matter
4. Name a place in Clausevane where this same kind of drift would NOT be worth reversing the build over, and say why not.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: How Clausevane handles ordinary, boilerplate indemnification language in standard vendor paper. That behavior hasn't drifted, and it isn't tied to the provider update, so rebuilding there fixes nothing while spending real engineering time.
Short answer, apply it yourself
5. Think of a product you use yourself. What's a habit it built in you that you'd stop doing if it got a little worse, and what would you start doing instead, by hand?
Show hint
Think of something you no longer double-check because it's usually right.
Show answer
Model answer: A map app's estimated arrival time. If it started running consistently late, the habit I'd stop is trusting the single number, and I'd start padding it myself with an extra ten minutes I calculate by hand.
Short answer, work the number
6. If the catch rate had only dropped to 90 percent by week ten instead of 81, would the same reversal decision still make sense?
Show hint
Think about whether the workaround, not the raw number, is the real trigger.
Show answer
Model answer: Not necessarily on its own. 90 percent might still clear the bar the team needs. The real test is whether the workaround keeps growing regardless. If Petra's hours were still climbing even at 90 percent, that's still the tell worth acting on.
Before you close the answer
Why this works
Tests whether you watch for a trend building in the team's actual behavior, or wait for a single dramatic failure to tell you what already happened weeks earlier.
Follow-up traps
"Isn't a spreadsheet just someone being extra careful, not a real signal?" Response: a one-off double-check is caution. A spreadsheet that keeps growing and eats nine hours a week is a private QA process, and it means the real cost of the build now includes paying for that hidden process too.
"What if the workaround exists because the person is just resistant to new tools?" Response: check whether it started before or after a measured drop in the eval set. Here it started two weeks after a real ten-point drop, not on day one, which rules out plain resistance.
If pressed
The vendor Ledgerwick piloted reruns its own model weekly against a shared clause eval set and publishes version notes with every change, the exact audit trail the in-house build never had. That transparency, not just a fresh model, is what actually closed the gap.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.