InterviewAdvancedResponsible AI & Advanced Practice / AI product case study teardowns / #20
Tear down a product I name, live, in ten minutes.
BOUND the named product is Kestrel Ledger, an AI churn-prediction and support-triage tool for a telecom carrier
Kestrel Ledger flags customers likely to cancel their phone plan, so a support rep can call before they leave. Idris Kavanjian runs support operations at Basalt Meridian Communications and was handed exactly this product, live, in a ten-minute stakeholder review, no prep time.
The direct answer
The 89% figure is a flag-accuracy number, not a save rate, and those are not the same thing. Ask one question before anything else: of the customers Kestrel Ledger correctly flags, how many actually stay because a rep called them. That single number, not the 89%, decides whether this product is worth its cost.
Do this, in order
Ask for the save rate, not the flag accuracy, in the first thirty seconds.Why: 89% flag accuracy says nothing about how many flagged customers actually stay.
Break the value into an equation before touching a single number.Why: ten minutes isn't enough time to think in numbers and structure at once.
Give a range, not a single net-value figure.Why: a save rate this uncertain can't honestly produce one clean number.
Sanity-check the range against the cost of the support team itself.Why: if the tool's value range is smaller than a single rep's salary, that's the real finding.
Name the one assumption that swings the estimate most, out loud, before time runs out.Why: that's the one number worth actually going and measuring after the meeting ends.
Don't let the demo's flashy interface distract from the missing save-rate number.Why: a polished screen can't answer a question the vendor never measured.
How to answer this, live, minute by minute
Six moves, budgeted against an actual ten-minute clock. This is the one skill that matters most when someone hands you a product cold.
0:00 - 1:00
Scope it out loud, fast
Say it like this
"Kestrel Ledger, one product, one job: flag customers likely to churn so a rep can call first. Let's tear that down, not AI churn tools in general."
Why this works
Ten seconds spent scoping saves two minutes of rambling later.
1:00 - 2:30
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down, the equation. Own the numbers. Use a range. Nail the sanity check. Direction, what would move this the most."
Why this works
A named structure, said fast, buys you room to think without looking lost.
2:30 - 4:30
Break down the equation, then flag the gap
Say it like this
"Net value equals saved customers times what a saved customer's worth, minus outreach cost on every flag, minus flagged customers who leave anyway. The 89% they gave us fits nowhere in that equation. It's flag accuracy, not save rate."
Why this works
States the real equation, then names exactly which number is missing from it.
4:30 - 6:30
Own the numbers and give a range
Say it like this
"I'd assume a save rate somewhere between 10% and 30% of flagged customers, since industry proactive-retention calls usually land there. That swings the net value from a genuine loss to a solid win, and we can't tell which without the real number."
Why this works
A range under time pressure beats a confident, made-up single figure every time.
6:30 - 8:00
Sanity-check it against something known
Say it like this
"At the low end of that range, Kestrel Ledger's net value is smaller than the fully-loaded cost of one support rep. If that's true, we shouldn't sign until we know which end of the range we're actually on."
Why this works
Comparing to something the room already understands makes an abstract range feel real.
8:00 - 10:00
Name the direction and close on the clock
Say it like this
"The save rate is the one number that swings this most, more than accuracy, more than cost. Before we sign anything, run a thirty-day pilot on one region and measure exactly that number."
Why this works
Ends on a concrete next step, not just a list of doubts, with time to spare.
Let's learn
Kestrel Ledger scans a customer's usage pattern, billing history, and support-ticket history, and flags the ones it thinks are about to cancel, so a rep can call before they do.
Before Kestrel Ledger, Basalt Meridian's retention team called customers only after they'd already started a cancellation request, catching roughly 15% of them with a last-minute offer.
With Kestrel Ledger, reps now call proactively, before any cancellation request starts, based on a churn-risk flag alone.
Knowledge spark: what's a save rate?
Of the customers a tool correctly flags as likely to leave, the share who actually stay because someone reached out. It's a completely different number from how often the flag itself is correct, and vendors rarely lead with it, because it's usually the weaker number.
The turn: the 89% figure sounds like the whole story, but it only measures whether the flag matched a customer who later left. It says nothing about whether the phone call changed anything.
Estimated net value of Kestrel Ledger, low vs. high save-rate scenario
The same product swings from an eighteen-thousand-dollar loss to a sixty-two-thousand-dollar gain, entirely on one unmeasured number.
At its worst: Basalt Meridian signs a two-year contract on the strength of "89% accuracy," and the actual save rate turns out low enough that the tool costs more in outreach and licensing than it returns in kept customers, discovered only at renewal.
The one question that mattered most
Ask for the save rate before anything else. A flag-accuracy number and a save-rate number can both be true at once and mean completely different things about whether this product is worth buying.
What I would leave alone: Kestrel Ledger's underlying flagging model itself is probably fine. The 89% figure, whatever it actually measures, likely reflects real work. The gap is entirely in what got reported, not necessarily in what got built.
Ten minutes was never enough time to know if this product works. It was enough time to find the one question that would actually tell us.
The lesson: a live teardown isn't about producing a verdict in ten minutes. It's about finding the one unmeasured number the whole decision quietly depends on, and saying so plainly before the meeting ends.
Now here is the same thing as a story
The minute-by-minute version above is what you'd actually say in the room. Read this one for how the real ten minutes played out.
Idris has run support operations for six years and has sat through more vendor pitches than he can count, most of them longer than ten minutes and worth less than five.
The flagging step gets all the marketing attention. The step after it, whether the call actually works, gets none.
The meeting opened with a slide: "89% flag accuracy, validated across 40,000 customer records." Everyone in the room, including Idris for the first thirty seconds, read that as "this product works."
Five checkpoints across ten real minutes, not a leisurely hour-long teardown.
The trigger was one word in the slide's own footnote: "accuracy measured against churn outcome, independent of outreach." Idris caught it on the second read and asked the vendor directly: "so this number doesn't include whether the call worked?"
Only one of these four branches is a genuine save. The vendor's 89% doesn't distinguish between any of them.
The vendor's answer confirmed it: no, the 89% was purely about whether the flagged customer actually churned, not whether the rep's call changed the outcome. Nobody at Basalt Meridian had asked for the save rate before that day.
Not every flagged customer is worth calling, and not every call is worth the same amount, a distinction the 89% number can't make.
With the range worked out live, from an 18,000-dollar loss at a conservative save rate to a 62,000-dollar gain at an optimistic one, the room agreed to a thirty-day, single-region pilot instead of signing the full contract on the spot.
The vendor's slide answered zero of these four questions. All four had to be asked live, in the room.
The old habit was to read a headline accuracy number as the whole verdict on a product. The new one asks, in the first minute, what that number is actually measuring, before any range gets built at all.
Idris used to let a clean-sounding percentage carry more weight than it earned, since pushing back live, in front of a room, felt like it cost more time than it was worth. It took one footnote, caught on a second read with the clock running, to see that ten minutes spent finding the right question beats an hour spent trusting the wrong answer.
BOUND, in one screenNot a leisurely estimate. BOUND is the shape a good teardown takes when the clock is actually running.
Five letters, and under a real clock, direction is the one that has to land before time runs out.
B
Break it down.
Net value equals saved customers times customer worth, minus outreach cost on every flag, minus flagged customers lost anyway.
States the equation before a single vendor number gets accepted at face value.
O
Own numbers.
Save rate assumed at 10-30% based on industry proactive-retention norms, since Kestrel Ledger never reported its own.
Every assumption stated with where it came from, live, out loud.
U
Use a range.
Net value spans an 18,000-dollar loss to a 62,000-dollar gain, depending entirely on the unmeasured save rate.
A single confident number here would have been a guess wearing a suit.
N
Nail the sanity check.
At the low end, the tool's net value is smaller than one support rep's fully-loaded salary.
Compares an abstract range to a number the whole room already understands.
D
Direction.
The save rate swings the estimate more than anything else, and it's the one number a thirty-day pilot could actually measure.
The hardest step under a clock, and the one that turns ten minutes of doubt into a concrete next step.
What would move the estimate most, if wrong
The save rate alone swings the estimate almost four times more than the other two assumptions combined.
The recap, one line per letter: break it down is the equation with the save rate named as missing, own numbers is the 10-30% industry-range assumption, use a range is the eighteen-thousand-loss-to-sixty-two-thousand-gain spread, nail the sanity check is comparing it to one rep's salary, and direction is the save rate as the number worth a real pilot.
And if you want to be sure it really works, try it somewhere elseSame five letters, a school district's enrollment-forecasting tool, ten minutes, a different clock entirely.
Tindersong forecasts which students are likely to withdraw mid-year, so a school counselor can intervene early. Marguerite Osafo, a district administrator, gets ten minutes to tear it down live at a budget review.
Mapped onto BOUND: break it down is net value equals students retained times per-student funding, minus counselor time spent on flagged students, minus students who withdraw despite intervention. Own numbers: Tindersong's vendor reports "84% withdrawal-prediction accuracy" but no retention rate after counselor contact, so Marguerite assumes a 15-25% retention range based on comparable early-intervention programs. Use a range: net value spans a modest loss to a genuine gain, same shape as Kestrel Ledger's spread. Nail the sanity check: at the low end, the tool costs more in counselor hours than the per-student funding it protects. Direction: the retention-after-contact rate, not the prediction accuracy, is what swings this the most, and it's the one thing a single semester's pilot could measure directly.
Swap "save rate" for "retention-after-contact rate" and the exact same four missing pieces show up in a school district's pitch too.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds, not ten minutes. Say "89% is flag accuracy, not save rate, ask for the save rate before anything else," and stop.
Cost: there's no time in the actual meeting to build the full range live. Say so honestly, and commit to one number to go verify afterward, rather than pretending to finish an estimate you don't have inputs for.
The model gets better, for real: if Kestrel Ledger's flag accuracy improves to 95% next quarter, that's still not an answer to the save-rate question, a more accurate flag on a customer nobody successfully calls back is still zero saved revenue.
Where people run it wrong.
They accept the vendor's headline number as the whole answer instead of asking what it actually measures.
They try to build one precise final number under time pressure instead of giving an honest range.
They spend all ten minutes debating the model's accuracy instead of finding the one input nobody has measured yet.
How to use it live. When someone hands you a product to tear down on the clock, spend your first sixty seconds finding the one number the whole pitch quietly assumes but never states. Everything else you say afterward should point back at that number.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "tear down a product I name, live, in ten minutes"?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. A live teardown is really a timed estimation problem.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Idris Kavanjian, who runs support operations at Basalt Meridian Communications and caught the gap in a vendor's slide, live.
3 · THE MISSING NUMBER
What number does Kestrel Ledger's pitch never actually provide?
Tap to flip
ANSWER
The save rate: of customers correctly flagged, how many actually stay because a rep called them.
4 · THE RANGE
What's the actual net-value range worked out live?
Tap to flip
ANSWER
An 18,000-dollar loss at a 10% save rate, up to a 62,000-dollar gain at a 30% save rate.
5 · THE DIRECT ANSWER
What's the one question that matters more than the 89% figure?
Tap to flip
ANSWER
Of the customers correctly flagged, how many actually stay because a rep called them, the save rate.
6 · THE NUMBER
Fill in the blank: Kestrel Ledger's own slide claimed ___% flag accuracy.
Tap to flip
ANSWER
89%. A real number, but one that measures the wrong thing for this decision.
7 · THE DIRECTION
Which single assumption swings the estimate the most?
Tap to flip
ANSWER
The save rate, an 80,000-dollar swing, almost four times more than outreach cost or customer value combined.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's its equivalent of "save rate"?
Tap to flip
ANSWER
Tindersong, a school-district enrollment-forecasting tool. Its equivalent is the retention-after-contact rate, not its prediction accuracy.
Check yourself Score: 0 / 0
True or false
1. True or false: an 89% flag-accuracy number and a high save rate necessarily mean the same thing about whether the product is worth buying.
True
False
Show hint
Look at the "knowledge spark" on save rate.
Show answer
False. Flag accuracy measures whether the prediction matched who left. Save rate measures whether the outreach actually changed the outcome. They can diverge completely.
Multiple choice
2. Why does this answer say the 89% figure "fits nowhere" in the net-value equation?
A. It's mathematically impossible for a percentage to be part of an equation.
B. It measures flag accuracy against churn outcome, not whether outreach on a flagged customer actually worked.
C. The vendor refused to explain how it was calculated.
D. It only applies to customers who never got called.
Show hint
Look at the "break it down" step.
Show answer
B. The equation needs a save rate, and flag accuracy is a different measurement entirely.
Fill in the blank
3. Fill in the blank: at the low end of the range, Kestrel Ledger's estimated net value is an ___,000-dollar loss.
Show hint
Look at the stacked bar chart of the low vs. high scenario.
Show answer
18,000. Against a 62,000-dollar gain at the high end of the same range, entirely dependent on the unmeasured save rate.
Short answer, apply it yourself
4. Pick an AI product's marketing claim you've seen. What's the one unstated number that claim is quietly assuming?
Show hint
Think about "accuracy," "engagement," or "satisfaction" claims that skip a step.
Show answer
Model answer: Many people point to a resume-screening tool claiming "top-candidate accuracy," which quietly assumes but never states how many of those top candidates actually got hired and succeeded.
Short answer, where it wouldn't matter
5. Name a part of Kestrel Ledger's design where the missing save-rate number genuinely wouldn't change the verdict.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The underlying flagging model itself. Whatever the 89% actually measures, the flagging logic is probably sound, the gap is in what got reported, not necessarily what got built.
Short answer, the number question
6. If the vendor revealed the true save rate was only 5%, well below the assumed 10-30% range, what would you recommend doing next?
Show hint
Think about the sanity check against a support rep's cost.
Show answer
Model answer: At 5%, the tool likely costs more than it returns, well below even the low-end scenario. The honest recommendation would be not to sign, or to renegotiate pricing against the real number.
Before you close the answer
Why this works
Tests whether you can produce real structure and a real range under genuine time pressure, instead of either freezing or bluffing a confident-sounding but baseless verdict.
Follow-up traps
"Isn't ten minutes too little time to say anything meaningful?" Response: ten minutes is enough to find the one missing number the whole decision depends on, that's a genuinely useful outcome, even without a final verdict.
"What if the vendor just makes up a save-rate number on the spot to answer your question?" Response: that's exactly why the recommendation is a real pilot, not just asking the vendor a second question, a made-up number under pressure needs field verification either way.
If pressed
Kestrel Ledger's flag model actually recalculates risk daily rather than at contract signing, which means even the 89% figure itself is a moving target quarter to quarter, a detail worth knowing but secondary to the save-rate gap.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.