CaseAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #19

What is the ROI framing for a defensive AI feature that only prevents churn?

PICK · customer-review summarization and aggregation for retailers

Fenwrix reads every review Southwold Home's shoppers leave, on their own site and three marketplaces, and turns them into a plain digest merchandising checks each morning. Shorecliff is the one part of it built to do nothing, on purpose, when it's working. Katsuro Fentiman owns what that's actually worth. Quang Trenholm, Southwold's finance partner, wanted one honest number for it. Loreto Fitzhugh, Loamstead's own VP of Product, nearly gave Shorecliff's two engineers to a feature that had something to show.

The direct answer
Frame Shorecliff's ROI as retained-revenue-at-risk avoided, never as growth. Build the number from one matched-cohort baseline, what an uncaught batch defect actually costs in quiet repeat-purchase loss, not from crediting every flagged cluster automatically. Then defend the framing against two different mistakes: don't let it get inflated with saves that were never really at risk, and don't let "nothing happened" talk you into reporting nothing at all.
Do this, in order
  1. Frame Shorecliff's ROI as retained-revenue-at-risk avoided, not growth.Why: it's a purely defensive feature. Scoring it against a growth yardstick makes it look weak next to a feature that was never trying to do the same job.
  2. Build the number from one matched-cohort baseline, not from every flagged cluster.Why: crediting every catch equally is the overclaim that gets checked and corrected in the same meeting, fast and cheap.
  3. Put the number in front of Quang before he has to go looking for it.Why: a finance partner who has to hunt for a defensive feature's value already assumes it doesn't have one.
  4. Don't let "no visible incident" read as "nothing to report" on the roadmap.Why: that exact read is what nearly cost Shorecliff its two engineers, right when a second cluster needed them.
  5. Set a kill line that can move the framing in either direction.Why: without one, "modeled" turns into a permanent hedge instead of a number that can actually update.
  6. Keep the 15-review, 10-day confidence floor, even though it's slower.Why: a faster trigger catches noise too, and merchandising stops trusting the alert the first time it's wrong.

How to answer this, stage by stage

Nobody is grading whether you know the dictionary definition of cost avoidance. They're grading whether you can defend a feature whose whole job is to leave nothing behind, and still put a real number on it.

1
Put one feature, one contract, and one name on it
Say it like this
"Let's ground this in one feature. Shorecliff is the part of Loamstead's Fenwrix platform that watches Southwold Home's reviews for an emerging complaint cluster, early enough that nobody's noticed it yet. Katsuro Fentiman owns what it's actually worth."
Why this works
Keeps every claim after this checkable against a real number, not a category.
2
Say the method out loud before naming a framing
Say it like this
"I'll run this as PICK. Take a position on the framing, say who feels it, name which mistake actually costs more, then say what would change my mind."
Why this works
Signals a plan is already running, so the next few minutes read as a plan, not a ramble.
3
Give the position, committed, in one breath
Say it like this
"Here's my position. Frame Shorecliff's ROI as retained-revenue-at-risk avoided. Not growth, not 'complaints resolved.' The number is what Southwold would have quietly lost in repeat purchases if that batch defect had taken 47 days to surface instead of 9."
Why this works
This is the direct answer, said before the interviewer has to dig for it.
4
Name who's standing in the room when the framing lands
Say it like this
"Two people feel this choice. Quang Trenholm, Southwold's finance partner, wants one defensible number for the renewal budget. Loreto Fitzhugh, our own VP of Product, is deciding whether Shorecliff keeps its two engineers or loses them to a feature that shows up in every demo."
Why this works
Turns an abstract framing debate into two real people who each pay a price for the wrong call.
5
Name which mistake is the expensive one
Say it like this
"Overclaiming is cheap. If I credit Shorecliff for a save that was never really at risk, Quang's team checks it against the real repeat-purchase data and corrects me in that same meeting. Underselling is expensive. If 'nothing happened' reads as 'nothing to report,' Shorecliff quietly loses its engineers, its eval coverage thins, and the next real cluster gets caught late instead of early."
Why this works
The hardest part of PICK. Naming which direction actually costs someone something is what makes this a real stance.
6
Prove it with the near miss that almost happened
Say it like this
"Here's what almost went wrong. While we were debating moving Shorecliff's engineers to Grayling, a second cluster started forming on the Ansley kettle, right at the model's confidence floor, sixteen reviews in ten days. With fewer people on eval review, that confirmation would have slipped two or three weeks, right into the window where it shows up as returns instead of an alert."
Why this works
Makes the hidden-expensive mistake real and countable, not hypothetical.
7
Give the number, built from a real baseline
Say it like this
"The Cotterell dutch oven batch hit 4,100 buyers. Two years ago, an uncaught defect on a different line took 47 days to surface and dropped that cohort's repeat-purchase rate from 38 percent to 11. Apply that same drop here, and catching it on day 9 instead of day 47 is worth somewhere between 610,000 and 860,000 dollars in retained purchases over the next year."
Why this works
Shows the number is built from a real comparison, not asserted, and gives it a range instead of false precision.
8
Say what would change the pick, then close
Say it like this
"If a matched-cohort check ever shows the repeat-purchase gap doesn't hold up, I'd shrink this to just the avoided replacement cost. If it holds for two years running, I'd call it measured, not modeled, a stronger number, not a weaker one. But right now: frame it as retained revenue at risk, avoided, put a real range on it, and don't let 'nothing happened' mean 'nothing to report.'"
Why this works
Closes on the decision and shows the position is falsifiable, not stubborn.

Let's learn

What does it look like when a feature works so well that nothing happens? For Shorecliff, it looks like an ordinary Tuesday. No return spike. No angry review thread. No story anyone tells at standup. Just a batch of dutch ovens that shipped, sold, and got used, the same as every other batch.

Fenwrix reads every review Southwold Home's shoppers leave, on their own site and three marketplaces, about 2.1 million reviews a year across 340 SKUs, and turns them into a plain digest merchandising checks each morning. Shorecliff is the one part of it built to only ever prevent something.

Hand sketched icon list titled Before Shorecliff existed. Three rows: a document icon, 47 days before a defect showed up as a return-rate spike. A gauge icon, that cohort's repeat-purchase rate fell from 38 percent to 11 percent. A question mark icon, no alert, no owner, no name attached to the loss.
Two years ago, before Shorecliff existed, this is what an uncaught batch defect actually cost Southwold, and nobody had a name for it at the time.

Two years ago, a saucepan handle started loosening after about three months of normal use. Nobody caught the pattern in the reviews. It took 47 days to show up as a spike on Southwold's return-rate dashboard, and by then the damage wasn't in the returns. It was in the buyers who never returned anything at all. They just stopped buying Southwold cookware. That cohort's 12-month repeat-purchase rate fell from a category baseline of 38 percent down to 11.

With Shorecliff, a batch of 4,100 Cotterell dutch ovens shipped with a lid rim about a millimeter out of spec, just enough to rattle and vent a little steam. Shorecliff's clustering model watches every SKU's reviews against a rolling 90-day baseline, and flags a theme once it clears 3.5 standard deviations above normal and shows up in at least 15 reviews inside a 10-day window, a bar tuned against 60 real past batch incidents and 400 cases of ordinary noise. On the Cotterell batch, it crossed that line at 41 reviews, on day 9.

Hand sketched flow diagram titled How Shorecliff catches a cluster, five boxes connected by arrows: reviews arrive, model clusters, confidence check, this box outlined in red-orange, person confirms, alert sent.
Five steps, and one of them is a person, on purpose, right where the model is least sure.
Knowledge spark: what's an emerging complaint cluster? A group of reviews about the same real problem, arriving faster and tighter than normal chatter ever does. Southwold's cookware SKUs always get a little background grumbling. A cluster is when that grumbling stops looking random.

Here's the turn. The interesting part isn't that Shorecliff worked. It caught the Cotterell batch four days faster than Southwold's own quality team would have managed on their best week. The interesting part is that when it worked, there was nothing left to show anyone. No return chart moved, because the returns never happened.

The model caught it four days faster than any person would have. Then there was nothing left to show anyone.
Cost, by the numbers: the fix, the whole contract, and the save
$900k $450k 0 $58,000 Cost of the fix $240,000 Whole Fenwrix contract $735,000 Retained revenue, modeled
Cost of the fixWhole Fenwrix contractRetained revenue preserved (range $610,000 to $860,000)
The modeled save is about three times the whole yearly contract, and about thirteen times what it cost Southwold to fix the batch. None of that shows up unless someone builds the middle number on purpose.

At its worst, that invisibility gets read as absence. For 14 months, Shorecliff never once appeared on Loamstead's product roadmap deck as a win, while Grayling, the feature that turns trending review quotes into marketing copy, showed a clean 1.8 percent conversion lift on every page it touched. Loreto Fitzhugh, weighing where to put two open engineering seats, nearly moved both of Shorecliff's off it and onto Grayling, because Grayling had something to point at.

The choice that mattered When Shorecliff launched, the dashboard gave it exactly one number: incidents flagged, this month. That was fair, then, there was no baseline yet to build a dollar figure from. Nobody ever came back and gave a defensive feature a revenue line once real incidents existed to model one from.

What I would leave alone: Grayling itself. A growth feature should keep being measured as growth, a clean conversion lift, not forced into "retained-revenue-at-risk avoided" framing just to match Shorecliff. That's the wrong tool for a feature built to add, not protect.

The lesson: a defensive feature's ROI isn't a number you compute once you're forced to defend it. It's a habit you build the same week you ship the feature, because the exact quality that makes it valuable, nothing bad happening, is the same quality that makes it invisible if nobody wrote down what almost happened instead.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a save with no bruise on it is the one that almost got cut.

For fourteen months, Shorecliff's biggest win never showed up anywhere Loreto Fitzhugh looked.

Katsuro Fentiman built Shorecliff to do one thing: watch Southwold Home's reviews for the shape of a real problem, tight and fast, before it turned into a quiet wave of buyers who just stopped. Southwold's own quality team was good, genuinely good, but they worked off returns data, and by the time a return spike shows up, the buyers who never bothered returning anything are already gone.

Shorecliff shipped without a revenue line. At launch, there was nothing yet to put a dollar figure on, no incident, no baseline, no way to say what a catch was worth. So the launch dashboard gave it exactly one number: incidents flagged, this month. Zero, most months. That was fine, then. The whole feature was too new to owe anyone a bigger story.

For a year, that felt like enough. Katsuro's quarterly update to Loamstead's leadership had one line for Shorecliff: "No incidents this quarter, coverage holding." Nobody minded. Grayling, the feature that lifts trending review quotes into marketing copy, had a real number every single quarter: conversion up 1.8 percent on any product page that used it. It was easy to love. It came with a receipt.

Hand sketched comparison diagram titled Two people, two different reasons this framing matters. Left, a person icon labeled Quang Trenholm, finance partner, wants one defensible number for the renewal budget. Right, a person icon labeled Loreto Fitzhugh, VP of Product, deciding whether Shorecliff keeps its two engineers.
Neither of them was being unreasonable. They were both just missing the same number.

Shorecliff didn't come with a receipt. It came with an absence, and an absence is hard to put in a deck.

Then, in March, the Cotterell dutch oven batch went out with a lid rim about a millimeter too tight in the wrong direction. Nothing dramatic. Just enough that the lid rattled in transit and let a little steam vent where it shouldn't. Shorecliff caught the pattern on day 9, 41 reviews in, well inside its calibrated bar. Southwold's team pulled the remaining 2,300 units off the floor, mailed every one of the 4,100 buyers a free replacement lid, and updated the supplier's tolerance spec before a single customer had to escalate anything.

It went about as well as a batch defect can go. Which is exactly the problem. Because it went well, it produced nothing anyone could point at in a budget meeting. No spike. No headline. Just 4,100 people who kept buying Southwold cookware instead of quietly not.

Its whole job was to make sure nothing happened. The renewal deck had no line for nothing.

Three weeks later, in the ordinary course of a headcount review, Loreto proposed moving both of Shorecliff's engineers onto Grayling's next phase. She wasn't being careless. She had a real number to grow and a feature with fourteen straight months of showing nothing on a slide. "I need the team somewhere I can point at what it did," she said, in the flat tone of someone who has already half-decided and is explaining, not asking.

Katsuro didn't have a number ready either. That was the real problem, not Loreto's math.

Hand sketched timeline titled Fourteen months with nothing to point at. Five milestones: Shorecliff launches, no revenue line. Good quarters, quiet, no deck line. Headcount review, engineers up for reassignment. Ansley near miss, this milestone emphasized in red-orange, right at the confidence floor. The fix, a real number with a range.
No single bad month. A number that stayed at zero for fourteen of them, until a headcount review and a near miss landed three weeks apart.

It almost went through. What stopped it was a second cluster, already forming, on the Ansley kettle: a run of reviews saying the handle got hot faster than it should. Sixteen reviews in ten days, right at Shorecliff's confidence floor, the exact kind of case that needs a person, not just the model, to say whether it's a real pattern or an ordinary spike from a regional heat wave pushing more people to use their kettles outside. Nazrin Belfour, who does that check, confirmed it was real on a Thursday morning. With two fewer engineers already reassigned, that same confirmation would likely have slipped two or three weeks, past the point where it shows up as returns instead of an alert.

Katsuro took both numbers, the Cotterell catch and the Ansley near miss, and built the sheet that should have existed since launch: Shorecliff's real value, framed as retained-revenue-at-risk avoided, not growth, not incidents caught. Using the saucepan incident from two years earlier as the matched baseline, a comparable batch defect that went uncaught and cost that cohort 27 points of repeat-purchase rate, the Cotterell catch alone is worth somewhere between $610,000 and $860,000 in purchases Southwold didn't quietly lose.

Quang Trenholm, sitting down to build Southwold's renewal budget that same week, had been about to flag the whole Fenwrix contract, not just Shorecliff, for a hard look. "Everything else I fund shows up in a chart somewhere," he said. "What does $240,000 a year actually buy me here?" Katsuro walked him through the matched-cohort number instead of the old incident count. Quang didn't need convincing that Shorecliff mattered. He needed a number in the same units as everything else on the page.

The decision that opened this door traced back to a fifteen-minute stretch in Shorecliff's original launch meeting. Someone asked whether the dashboard needed a rough dollar line for defensive features from day one, even before there was real data. The room said no, there wasn't anything to model yet. That was true, then. Nobody wrote down when to come back and check.

Hand sketched full page metaphor scene titled The whole answer in one picture. Left panel, a plain document icon labeled a mark, caption shows on a return chart, easy to defend. Right panel, a filled box with a large question mark labeled no mark, caption the churn wave that never happened, easy to cut.
The whole answer to this question in one picture. A number that leaves a mark gets defended by itself. A number that leaves nothing needs someone to go build it.

Run that meeting again, with one more line added: give every defensive feature a modeled retained-revenue baseline from its first real incident, however rough. Same Cotterell batch, same Ansley near miss. This time, Loreto's headcount review opens with Katsuro's number already on the slide, not built under pressure three weeks later trying to save two jobs and a contract at once.

One design let "nothing happened" mean nothing worth writing down. The other lets it mean exactly what it's worth, in dollars, the same as everything else competing for the same two engineers.

What Katsuro would tell their past self, back in that fifteen-minute meeting: a feature that's built to prevent something isn't finished at launch. It's finished the day someone gives its absence a number.

PICK, or why a save with no bruise is the one that gets cut

Not a way to dress up "defensive work matters" in four letters. PICK forces a real commitment about which framing this feature earns, then makes you say, out loud, which kind of mistake actually costs someone something.

PPosition. Your pick, in one sentence, before any reasoning.
Frame Shorecliff's ROI as retained-revenue-at-risk avoided. Never growth, never a raw count of clusters caught. The number is what Southwold would have quietly lost in repeat purchases if the Cotterell batch had taken 47 days to surface instead of 9.
Say the framing before the dollar figure, or the room spends the next five minutes guessing what kind of win this even is.
IImpact. Who feels each kind of error, and in what units.
Book it wrong toward growth, and Quang Trenholm can't reconcile it against anything else on the renewal invoice, every other line has a chart behind it. Book it wrong toward silence, and Loreto Fitzhugh keeps reading fourteen quiet months as fourteen months of nothing, and nearly moves the two engineers who make the catches possible.
Naming both people, the one who wants a number and the one who's deciding without one, keeps this from turning into a one-sided caution story.
CCost asymmetry. The heart of it.
Overclaiming is cheap. Credit Shorecliff for a save that was never really at risk, and Quang's team checks it against the real repeat-purchase data and corrects it in the same meeting. Underselling is expensive. Let "nothing happened" read as "nothing to report," and it sits quiet for over a year, until it's two engineers and a near miss away from losing the coverage that makes the next catch possible.
This is the step that earns the pick. Anyone can say defensive work is undervalued. Naming which direction actually costs someone something is what survives a follow-up.
Hand sketched comparison diagram titled Two ways to get a quiet save wrong. Left, a plain document icon labeled overclaim, caption credit Shorecliff for a save that was never at risk, checked and fixed in the same meeting. Right, a filled box with a large question mark labeled undersell, caption let nothing happened mean nothing to report, sits quiet for months, then costs the feature its team.
One of these mistakes gets caught in the room. The other one gets caught fourteen months later, if it gets caught at all.
KKill criteria. What evidence would flip the framing.
Right now about 54 percent of Shorecliff's flagged clusters have a validated matched-cohort outcome behind them. If that validated share stalls, or the repeat-purchase gap it's supposed to predict doesn't hold up under a real check, shrink the claim down to just the avoided replacement cost, a smaller, safer number. If it crosses 75 percent with the gap holding, call the number measured, not modeled, and that's a stronger claim, not a weaker one.
A pick with no kill criteria is a caution you're defending forever. This makes it a live position, one that can get stronger, not just smaller.
Share of Shorecliff's flagged clusters with a validated outcome, against the kill line
100% 50% 0% kill line: 75% validated 9% 17% 29% 40% 54% Q1 Q2 Q3 Q4 Now
Share of flagged clusters validated against a matched cohortKill line, 75% validated
19 of 35 flagged clusters have a confirmed matched-cohort outcome so far, a little over half. The framing stays modeled until that share crosses 75 percent with the repeat-purchase gap holding.

And if you want to be sure it really works, try it somewhere else

Same four letters, a city permitting office instead of a kitchenware brand, and this time the thing that almost went uncounted wasn't a repeat customer. It was a permit applicant who just stopped calling.

Underbridge is Ashfold Civic Systems' permitting workflow platform. Sandover is its one purely defensive feature: it reads inspector notes and applicant complaint tickets, and flags an emerging stall pattern in a permit category before the average processing time balloons enough to trigger a city council motion to cancel the contract. Fairholt runs Underbridge across roughly 41,000 permits a year.

What Sandover's catch was worth to Fairholt, built up from three parts
$650k $325k 0 $210,000 Underbridge contract $615,000 total Sandover, avoided cost
Underbridge annual contractAvoided temp-inspector costWaived rush feesAvoided renegotiation discount
Same shape as Shorecliff's chart: the avoided cost is about three times the contract that pays for it, and none of it shows up unless someone builds the middle bar on purpose.

Callisto Corriveau, who owns Fairholt's IT-services budget, was preparing the same kind of renewal hearing Quang was. Sandover caught a stall in electrical inspections, one inspector's medical leave quietly backing up a whole category, on day 6 instead of the historic 34-day lag before it would have shown up as a council complaint. Modeled against Fairholt's own past incident, that catch is worth about $615,000: avoided temp-inspector cost, waived rush fees, and the contract discount Fairholt would have demanded after a public backlog story, against an annual contract of $210,000.

The decision Ashfold's own team would take back Sandover's rollout deck had one number, "incidents flagged," the same shape of gap Shorecliff started with. Nobody came back and gave it a dollar line once real incidents existed to model one from, until a renewal hearing forced the question.

Same rank, different lever: the fix isn't a smarter model or a bigger inspection team. It's the same framing, run against Fairholt's own numbers: retained processing capacity avoided, built from one matched past incident, defended against overclaiming every minor stall and against underselling the one that actually mattered.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: frame it as loss avoided, build the number from one matched incident, name which mistake costs more, keep a kill line.
Cost: no budget this quarter for a real matched-cohort study. Ship the cheap version first: one rough range from the closest past incident, labeled "modeled," revisited next quarter with better data.
The model got better, for real: say Shorecliff's confidence threshold tightens and it catches clusters two days earlier on average. The framing doesn't change. A faster catch just means a bigger number inside the same "retained revenue avoided" frame, not a reason to call it growth.

Where people run it wrong.
They credit the feature for every flagged cluster equally, instead of building one matched-cohort baseline and applying it honestly.
They let "no incident this quarter" sit on a dashboard with no dollar line, until someone reads the blank as zero value instead of a save with no bruise.
They set no kill line, so "modeled" becomes a permanent hedge nobody ever revisits either direction.

How to use it live. Ask the framing question before naming a number: "Is this feature supposed to add something, or only supposed to stop something from being lost, because those two get defended with completely different evidence." That buys real thinking time, and it reframes the whole question before you have to guess at a figure.

Three things worth stating directly, since this is where the real judgment sits. The alternative Loamstead considered, and rejected, was scoring Shorecliff by a raw count of clusters caught instead of a dollar figure, since a count sidesteps the whole modeling problem. It lost, because a count can't be checked against anything else on the invoice, and it treats a $735,000 catch the same as a trivial one, which would have made the underselling problem worse, not better. The AI-specific failure worth naming is a false confirmation at the model's confidence floor: a cluster that clears the review-count bar by a hair, like the Ansley kettle at sixteen reviews in ten days, can just as easily be a real batch pattern or an ordinary seasonal spike, and the model alone can't always tell the difference. The guardrail is the mandatory human check on any cluster that clears the bar but sits close to it, exactly where Nazrin Belfour's review caught the Ansley pattern before it could go either way unconfirmed. And the trade-off is real: Shorecliff waits for 15 reviews across 10 days before it alerts anyone, slower than it could technically fire, and Loamstead accepted that delay on purpose, because a lower bar catches more noise, and merchandising stops trusting the alert the first time it sends them chasing a false one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
PICK: commit to a position on a tradeoff, then show which of two mistakes actually costs more, and to whom. Used here because the question is a framing tradeoff, not a perturbation or a metric question.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Katsuro Fentiman, who owns Shorecliff's ROI story at Loamstead Analytics, and has to defend a feature whose whole job is to leave nothing behind.
3 · THE POSITION
What's the P step here, in one line?
Tap to flip
ANSWER
Frame Shorecliff's ROI as retained-revenue-at-risk avoided, never growth or a raw count of clusters caught.
4 · THE COST ASYMMETRY
Which mistake is cheap and visible, and which is hidden and expensive?
Tap to flip
ANSWER
Overclaiming a save is cheap: someone checks it against real repeat-purchase data and corrects it in the room. Underselling real defensive work is expensive: it reads as "nothing to report" and can cost the feature its own engineers.
5 · THE OLD DECISION
What decision would Katsuro take back?
Tap to flip
ANSWER
Launching Shorecliff with only an incident count on the dashboard, no dollar line, because at launch there was nothing yet to model. Nobody came back to add one once real incidents existed.
6 · THE NUMBER
Fill in the blank: catching the Cotterell batch on day 9 instead of day 47 preserved somewhere between $___ and $___ in retained purchases.
Tap to flip
ANSWER
$610,000 and $860,000, modeled off a matched batch-defect incident from two years earlier.
7 · THE REPLAY
Same headcount review, new framing, what changes?
Tap to flip
ANSWER
Katsuro's retained-revenue number is already on the slide before Loreto has to ask for it, so the review starts from a real figure instead of fourteen quiet months and a near miss.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the modeled number there?
Tap to flip
ANSWER
Sandover, Ashfold Civic Systems' backlog-detection feature inside Underbridge. The modeled number is about $615,000 in avoided cost for Fairholt, caught on day 6 instead of a historic 34-day lag.

Check yourself Score: 0 / 0

True or false
1. True or false: Shorecliff's ROI should be framed the same way as Grayling's, since both features come out of the same Fenwrix platform.
  • True
  • False
Show hint
Check "what I would leave alone" in Let's learn.
Show answer
False. Shorecliff is purely defensive. Framing it as growth, the way Grayling's conversion lift gets framed, makes it look weak against a feature that was designed to add something, not protect something. They need different framings, not a shared one.
Multiple choice
2. Why did Quang Trenholm nearly flag the whole Fenwrix contract for a hard look?
  • A. Shorecliff's model accuracy had dropped that quarter.
  • B. Every other line he funded showed up in a chart somewhere, and Shorecliff's value never had one.
  • C. Southwold's return rate had spiked that quarter.
  • D. Loamstead raised the contract price mid-year.
Show hint
Check the story section, the paragraph where Quang is building the renewal budget.
Show answer
B. The model hadn't gotten worse. Nobody had ever built the matched-cohort number that would let Shorecliff's value be checked the same way everything else on the invoice could be.
Fill in the blank
3. Shorecliff flags a cluster once a complaint theme clears ___ standard deviations above baseline and shows up in at least ___ reviews within a ___-day window.
Show hint
Look at the paragraph describing the Cotterell batch catch, right before the flow diagram.
Show answer
3.5 standard deviations, 15 reviews, a 10-day window. That's the calibrated bar, tuned against 60 real incidents and 400 false alarms, that keeps Shorecliff fast without flooding merchandising with noise.
Short answer, name the rejected alternative
4. What alternative did Loamstead consider instead of the matched-cohort dollar framing, and why was it turned down?
Show hint
Look at the "three things worth stating directly" paragraph at the end of Section 4.
Show answer
Model answer: Scoring Shorecliff by a raw count of clusters caught, since it would have sidestepped the whole modeling problem. Turned down because a count can't be compared against the invoice cost, and it treats a $735,000 catch the same as a trivial one.
Short answer, apply it yourself
5. Think of a feature at your own job, or in a product you use, whose whole job is to prevent something bad. What would you have to measure to prove it's working, given that success looks like nothing happening?
Show hint
Think about what you'd need to compare against, not just what the feature caught.
Show answer
Model answer: A spam filter's real value isn't the junk it blocks, it's whether people who'd have quit the inbox over spam kept using it. You'd need a matched comparison, a group without the filter, to see what churn would have looked like without it.
Short answer, work the number
6. If the historic repeat-purchase drop had been 15 points instead of 27, would the modeled retained-revenue range for the Cotterell batch still clear the $58,000 cost of the fix?
Show hint
Use the same arithmetic as the number chart: 4,100 buyers times the point drop, times the per-buyer repeat-purchase value.
Show answer
Yes, easily. 15 points of 4,100 buyers is about 615 buyers. At $550 to $780 each, that's roughly $338,000 to $480,000, still many times the $58,000 fix cost and comfortably above the whole $240,000 Fenwrix contract.
Before you close the answer
Why this works
Tests whether you'll fight for a defensive feature's real value or let "nothing happened" talk you into reporting nothing. Most candidates either inflate the number to compete with growth features, or undersell it out of caution.
Follow-up traps
"Isn't $610,000 to $860,000 just a made-up number, since nothing actually happened?" Response: it's modeled off a real, matched incident from two years earlier, not invented. That's exactly why it's a labeled range, not a hard figure.

"Why not just count how many clusters Shorecliff catches, simpler for everyone?" Response: a count can't be checked against the invoice, and it treats a $735,000 catch the same as a trivial one, exactly the undercounting this framing is meant to fix.
If pressed
The matched-cohort baseline only pulls comparison incidents from a similar price band, a $130 to $180 item against another $130 to $180 item, never against a $12 accessory, because repeat-purchase elasticity moves very differently by price point, and mixing bands would make the range meaningless.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more