CaseAdvancedAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #21
Your competitor built; you bought. How do you assess whether that was a mistake?
TRACEthe seven-point gap was real, just not where the conference slide said it was
Norvale Residential manages apartment listings. TenantVane is the vendor chat assistant, built by Delcastro, that answers renter questions and pre-qualifies leads before a human agent calls. Wren Loewenthal is the AI PM who has to answer a hard question: Anchorbrace Properties, a competitor, built their own assistant in house and claims a much higher conversion rate. Did Norvale make the wrong call by buying?
The direct answer
Don't compare a competitor's headline number to yours directly. Rebuild the timeline first, recut your own number by segment, rule out a measurement mismatch, then name real cause candidates and run the one test that separates them. Only if the gap survives every one of those checks, same metric, same segment, same window, was buying instead of building actually the wrong call.
Do this, in order
Rebuild the timeline before comparing anything: when did each number actually start moving.Why: a competitor who launched later can still look ahead sooner if their curve is steeper, not just higher.
Recut your own number by segment before trusting a single average.Why: a plateau can hide one segment doing fine and another dragging it down.
Rule out a measurement mismatch before trusting a competitor's public number.Why: a conference-talk statistic might not even be the same metric as your own dashboard.
Name real cause candidates, not just "their AI is better."Why: inventory mix, pricing, and retraining cadence can all explain a gap that has nothing to do with model quality.
Run the one evidence test that actually separates the candidates.Why: a real test beats another round of guessing which story sounds most plausible.
Only call it a mistake once the gap survives every filter above.Why: otherwise you've diagnosed a bad quarter, or a bad slide, as a bad decision.
How to answer this, stage by stage
Nobody is scoring whether you can admit fault gracefully. They're scoring whether you'd actually go check before agreeing anything went wrong.
Stage 1
Scope it to one real comparison
Say it like this
"Let's ground this in one real comparison. Norvale bought a vendor leasing assistant, TenantVane, fourteen months ago. Our competitor, Anchorbrace, built their own in house, four months later. Leadership just found Anchorbrace's own conference talk claiming a 34 percent tour-to-lease rate against our 27. That's the actual number I'd investigate before agreeing anything went wrong."
Why this works
Keeps the answer from turning into a generic "how do you handle failure" speech with no real gap behind it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as TRACE. Timeline, when did each number actually start moving. Recut, sliced by segment instead of one average. Assume nothing, rule out a measurement mismatch first. Cause candidates, name real competing explanations. Evidence test, the one check that tells them apart."
Why this works
Signals a repeatable diagnostic method, not a defensive reaction to an uncomfortable slide.
Stage 3
Reframe: it isn't "were we wrong," it's "what would tell us"
Say it like this
"This isn't really asking whether we made the wrong call fourteen months ago. It's asking what evidence would actually prove that, versus just noticing a competitor's number looks better in a conference slide."
Why this works
This is where a strong answer separates from someone who either defends the decision or apologizes for it, neither backed by evidence.
Stage 4
Give the one decision
Say it like this
"Here's what I'd actually conclude, once the checks are done: TenantVane isn't a mistake everywhere, but it is genuinely behind on family-sized units. And the real cause isn't the vendor's base model, it's that we've never once retrained it on our own fourteen months of conversation logs, while Anchorbrace retrains theirs monthly. That's fixable without switching vendors."
Why this works
This is the direct answer, stated as an actual, checkable conclusion, not a vague sense that "we should look into it."
Stage 5
Prove it with the compressed evidence
Say it like this
"TenantVane launched, and conversion rose from 22 to 27 percent in two months, then went flat for a year. Anchorbrace launched four months later and has climbed from 24 to 34 and is still climbing. Recut by unit type, our 27 percent average is actually 32 on studios and 19 on family units, and family units are exactly where Anchorbrace's portfolio, and their in-house model, both skew."
Why this works
Compresses the whole case into the timeline and the one segment split that actually explains the gap.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't just 'their AI is better' is that TenantVane runs on a static prompt template we've never retrained on our own conversation history, while Anchorbrace fine-tunes theirs monthly on their own transcripts. We accepted a lower up-front cost and a faster launch by buying a generic assistant, in exchange for a model that doesn't get sharper on our own objections the way a retrained one would."
Why this works
This is the load-bearing, AI-specific judgment. A generic vendor comparison doesn't hinge on whether a model gets retrained on your own conversation data.
Stage 7
Say what wouldn't change your mind, then close
Say it like this
"I wouldn't call the original buy decision a mistake for the whole product, since it still beats our old manual process by a wide margin on most units. I'd call the missing retraining pipeline the actual mistake, and that's the one thing I'd change, not the vendor."
Why this works
Closes with real judgment about what's actually wrong versus what merely looks bad, and restates the direct answer in one breath.
Let's learn
What happens when a competitor's number looks better than yours, and you don't yet know why?
TenantVane is a chat assistant on Norvale's listing pages that answers renter questions and pre-qualifies leads before a human agent ever calls. Before any assistant, a human leasing agent answered listing questions by phone, at whatever hours they worked, converting about 22 percent of tours into signed leases.
With TenantVane answering instantly, day or night, conversion rose to 27 percent in the first two months, then held flat there for a year.
Tour-to-lease conversion, month by month: Norvale vs Anchorbrace
Norvale, TenantVaneAnchorbrace, in-house
Anchorbrace launched four months later and started lower. The gap isn't that they were ahead from day one, it's that their curve never stopped climbing.
Here's the turn: the plateau at 27 percent was never really the problem on its own. The problem showed up the day a competitor's number, on the surface the exact same metric, came in seven points higher, and nobody at Norvale had a fast way to say whether that gap was real, or a trick of segments and self-reported statistics.
A competitor's better headline number isn't proof you chose wrong. It's a reason to check, on the same segment, the same metric, before you decide anything.
At its worst, leadership overreacts to one comparison, rips out a working vendor tool, and spends real months rebuilding in-house before ever confirming the actual cause of the gap.
The five checks, held up as one page. Assume nothing is the one most people skip on the way to a conclusion.
The choice I would take back
Norvale's contract with the vendor never included a retraining clause, an agreement to periodically fine-tune the assistant on our own leasing conversations. That made sense at signing, when a fast, generic launch mattered more than a slow, customized one. It stopped making sense once a full year of conversation logs sat unused while a competitor's in-house model kept sharpening on exactly that kind of data.
What I would leave alone: I wouldn't touch how TenantVane handles studio-unit conversations, since that segment is already converting at 32 percent, ahead of even Anchorbrace's own studio number.
The lesson: a competitor's better number isn't a verdict. It's a question with a real answer sitting in your own data the whole time, if you're willing to go recut it before you accept the story on the slide.
Now here is the same thing as a story
The short version above is what you'd say in the room. Read this one for what it actually felt like the day a VP's slide put a number next to Wren's without asking first.
The leasing office at Norvale's downtown property goes quiet around 9 pm, right when TenantVane starts answering on its own.
Wren Loewenthal had been Norvale's AI PM for two years, good at turning messy support logs into a clear read on what was actually broken, not just what looked broken. When TenantVane launched, the early bump felt like a clean win: 22 to 27 percent in two months, minimal fanfare, agents relieved to stop fielding 11 pm questions about pet policies.
The habit thinned in three beats. At first Wren checked the weekly conversion dashboard every Monday. After the number held steady at 27 for three months, she checked it monthly. By month nine, she barely opened it at all, since 27 percent had become the number everyone already expected to see.
Fourteen months laid flat. The gap between "plateaus" and "still climbing" is the part worth staring at.
In month fourteen, a regional VP, prepping for a board meeting, pulled up Anchorbrace's public conference talk, a slide claiming "34 percent tour to lease," and put it side by side with Norvale's own dashboard, then asked Wren directly whether buying instead of building had been the wrong call.
Knowledge spark: why would two companies' "same" metric not actually match?
A conference talk usually reports whichever number looks best, and "conversion" can mean different things: lead-to-tour, tour-to-lease, or something blended across both. Two companies can both say "conversion" and be measuring completely different steps of the funnel.
Wren didn't answer that afternoon. She pulled the full fourteen months of TenantVane's dashboard, recut it by unit type, and found studios at 32 percent, family units at 19, a gap nobody had looked at all year.
Before comparing the two numbers at all, Wren had to confirm they were even measuring the same thing.
We weren't behind by seven points. We were behind on exactly the units where our own portfolio, and Anchorbrace's, actually overlapped.
Wren never had a fixed number for "good enough." She had a real question: is this gap the same metric, the same segment, the same time window as theirs, or not? Once she checked, the honest gap was five points on family units specifically, not seven points across the board, and it was explainable.
Three honest explanations, and one real check, TenantVane's own version history, that told them apart.
When the vendor contract was signed, nobody asked for a retraining clause. "Let's just launch and see," someone said in that meeting, since speed mattered more than customization for a first version, and at the time, that was a reasonable trade to make.
Conversion by unit type, month 14: Norvale vs Anchorbrace
NorvaleAnchorbrace
Norvale actually leads on studios. The whole seven-point gap lives in one segment, family units, and it's a 17-point gap there, not seven.
Rerun the same audit with a retraining pipeline live for three months: family-unit conversion rises from 19 to 26 percent, closing most of the real gap, verified on the same segment, the same metric, month over month.
One company treats a competitor's headline number as a verdict. The other treats it as a question with a real answer sitting in their own data the whole time.
What I'd tell myself, watching that VP's slide land on the table: don't defend the seven points. Go find out what they actually are.
TRACE, the five checks before I'd call it a mistakeNot a script for defending every buy decision. TRACE is what keeps you from confusing a bad quarter, or a bad slide, with a bad call.
T
Timeline. When did each number actually start moving?
TenantVane launched and plateaued at 27 percent by month 3. Anchorbrace launched four months later and has climbed continuously through month 14, not yet plateaued.
A steeper, later curve can pass a flatter, earlier one without either side doing anything dramatic.
R
Recut. Slice by segment before trusting one average.
Norvale's 27 percent splits into 32 on studios and 19 on family units. The blended number hid a real, specific gap.
An average is only as honest as the segments it's hiding.
A
Assume nothing. Confirm the metric before trusting the number.
Confirm Anchorbrace's "34 percent" is actually tour-to-lease, not lead-to-tour or some blended figure that would flatter a conference slide.
This is the hardest step, and the one most people skip on the way to a conclusion.
C
Cause candidates. Name real competing explanations.
A genuine product gap on family-unit objection handling, a measurement mismatch in Anchorbrace's reported number, or a non-AI cause like a better inventory mix.
Three named hypotheses beat one assumed story every time.
E
Evidence test. The one check that tells them apart.
Check TenantVane's version history for retraining events. Find none in fourteen months, while Anchorbrace retrains monthly, confirming the real product gap, not the other two candidates.
This is the direct answer, proven with a real audit trail instead of asserted as a hunch.
The recap, one line per letter: timeline shows Anchorbrace's later, steeper climb overtaking Norvale's early plateau, recut shows the whole gap concentrated in family units, assume nothing catches that the two "conversion" numbers needed confirming before comparing, cause candidates names three honest explanations, and the evidence test, TenantVane's own version history, confirms the real one: a missing retraining pipeline, not the vendor's base model.
And if you want to be sure it really works, try it somewhere elseSame five letters, a waste-hauling company instead of a leasing office. This time the real cause isn't the model at all.
Idris Gauthier runs operations at Grovewick Waste Services, which bought a vendor route-optimization tool for its collection trucks. A competitor built their own routing engine in house and claims 18 percent lower fuel costs. Mapped onto TRACE: timeline shows the competitor's savings appearing almost immediately after launch, not climbing slowly like Norvale's case. Recut by route type shows the gap concentrated on Grovewick's older, rural routes, not the dense urban ones. Assume nothing confirms both companies are measuring the same thing, fuel cost per stop, so the numbers are genuinely comparable this time. Cause candidates: a worse routing model, a measurement issue, or bad underlying address data. The evidence test settles it: Grovewick's own address database has 12 percent stale entries on rural routes, entries the vendor's routing model can't fix no matter how good it is, while the competitor rebuilt their address list before ever building their router. The real cause here isn't build versus buy at all. It's data quality, sitting one layer below the model either company chose.
Same five checks, a different real cause entirely. Sometimes the model was never the problem.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "rebuild the timeline, recut by segment, confirm the metric, then name and test real causes, before calling anything a mistake," and stop.
Cost: no time to run a full audit before a meeting. Say so honestly, and name the one check you'd run first, the metric confirmation, since it's the cheapest and most likely to change the whole conversation.
The competitor's number turns out to be wrong, or inflated: that's not a reason to stop checking, it's exactly what the "assume nothing" step exists to catch, and it's often the actual answer.
Where people run it wrong.
They compare a competitor's public number to their own private one without confirming it's the same metric.
They treat one blended average as the whole truth instead of recutting it by segment first.
They jump straight to "we should have built it ourselves" without naming or testing any other cause.
How to use it live. The moment an interviewer asks you to assess whether a competitor's approach beat yours, ask yourself: what would I need to check before I even believed their number? Say that check out loud first, and it buys you real thinking time before naming a cause.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits assessing whether a competitor's build beat your buy?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. It stops you from treating a competitor's headline number as a verdict before checking it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Wren Loewenthal, AI PM at Norvale Residential, who had to answer whether buying TenantVane instead of building was a mistake.
3 · THE HABIT
What habit did Wren fall into, that the audit broke?
Tap to flip
ANSWER
Checking the conversion dashboard less and less often once it held steady at 27 percent, down to barely checking it at all by month nine.
4 · THE REAL QUESTION
What's the actual comparison being tested here?
Tap to flip
ANSWER
Whether the seven-point gap versus Anchorbrace was a real product failure, or an unchecked mix of segment differences and a mismatched metric.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Not including a retraining clause in the vendor contract, so fourteen months of Norvale's own conversation logs sat unused while a competitor's in-house model sharpened on exactly that kind of data.
6 · THE NUMBER
Fill in the blank: TenantVane converts ___ percent on studios and ___ percent on family units, against a blended average of ___ percent.
Tap to flip
ANSWER
32 percent, 19 percent, 27 percent blended.
7 · THE REPLAY
Same audit, retraining pipeline live for three months. What changes?
Tap to flip
ANSWER
Family-unit conversion rises from 19 to 26 percent, closing most of the real gap, verified on the same segment and metric, month over month.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this again for a different product. Which one, and what's the real cause there?
Tap to flip
ANSWER
Grovewick Waste Services' route-optimization tool. There the real cause turns out to be stale address data, not the vendor's model or the build-vs-buy choice at all.
Check yourself Score: 0 / 0
Short answer, first check
1. Before agreeing a competitor's number proves a mistake, what's the very first thing you should check?
Show hint
Look at the "assume nothing" step in the TRACE recap.
Show answer
Model answer: Whether their reported number is even the same metric, on the same segment, over a comparable time window, since a conference-talk statistic can quietly use a friendlier definition than it looks.
Multiple choice
2. Why does the segment recut matter here?
A. It doesn't; the blended average already tells the full story.
B. It reveals that the real gap is concentrated on family units, not spread evenly across all listings.
C. It proves Anchorbrace's number was fabricated.
D. It shows studio units are actually underperforming.
Show hint
Look at the grouped bar chart, by unit type.
Show answer
B. Studios actually lead. The entire gap lives in family units, a fact the blended 27 percent number completely hid.
True or false
3. True or false: Anchorbrace's 34 percent figure was confirmed to be the exact same metric as Norvale's 27 percent before any conclusion was drawn.
True
False
Show hint
Look at the knowledge spark about matching metrics.
Show answer
False. The "assume nothing" step exists exactly because that confirmation hadn't happened yet, and conference-talk numbers often use a friendlier metric.
Fill in the blank
4. Fill in the blank: TenantVane has never been ___ on Norvale's own conversation logs in fourteen months, while Anchorbrace's in-house model is ___ every month.
Show hint
Look at the evidence test step.
Show answer
Retrained; retrained. The version history check is what confirmed this, not a guess about which vendor tries harder.
Short answer, apply it yourself
5. Think of a time you compared your own results to someone else's public claim. What check would have told you if the comparison was fair?
Show hint
Think about whether the two numbers were measuring the exact same thing, over the same period.
Show answer
Model answer: Comparing a personal running pace to a friend's app-reported pace. The check would have been confirming both apps measured pace the same way, since one rounded up on hills and the other didn't.
Short answer, work the number
6. If Anchorbrace's 34 percent turned out to be lead-to-tour, not tour-to-lease, would the seven-point gap still be meaningful?
Show hint
Think about what part of the funnel each metric actually measures.
Show answer
Model answer: No, not directly. Lead-to-tour and tour-to-lease measure different steps of the funnel, so the two numbers wouldn't be comparable, and the entire seven-point gap could be an artifact of comparing two different metrics.
Before you close the answer
Why this works
Tests whether you treat a competitor's number as an investigation to run, not a verdict to accept, before deciding your own past decision was wrong.
Follow-up traps
"What if the gap is real even after all the checks?" Response: then it's a real, specific, fixable gap, in this case a missing retraining pipeline, not proof that buying instead of building was categorically the wrong call.
"Isn't recutting by segment just looking for excuses?" Response: no, the recut is what actually located the fixable cause. Ignoring it would have meant tearing out a vendor tool that outperforms on most of the portfolio to chase a gap that only existed in one segment.
If pressed
The version-history check that confirmed zero retraining events came from the vendor's own admin console logs, not a guess, the kind of audit trail every AI vendor contract should guarantee access to from day one.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.