CalculationAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #16
How do you separate the effect of the AI feature from general product growth?
The direct answer
Do not compare Warmline's numbers to before Warmline. Hold back a slice of new accounts from Warmline for a few weeks and compare the group that got it against that group. The gap between the two, not the before-and-after gap, is what Warmline actually did. Report it as a range, because the holdback group is small and its own number moves around on its own.
Do this, in order
Hold back a slice of new accounts from Warmline and compare the two groups over the same weeks.Why: a before-and-after number cannot tell Warmline's effect apart from the quarter just getting better on its own.
Write the equation out first: attributable lift equals the Warmline group's move minus the holdback group's move.Why: naming the subtraction stops anyone from quietly reporting the Warmline group's raw move as the whole story.
Give a range, not one number, sized off how small the holdback group actually is.Why: a 100-account holdback group's own number is noisy, and pretending it isn't invents confidence nobody earned.
Check the range against a number finance already believes, built without any reference to Warmline.Why: if the two don't roughly agree, one of the two estimates is wrong, and it's worth knowing which before anyone repeats the number in a deck.
Name which single assumption would swing the answer most, and say so out loud.Why: it's usually the holdback group's size, and a reader deserves to know how fragile the range is.
Never let self-selected adopters stand in for the holdback group.Why: accounts that turn a new AI feature on themselves were already growing faster, so comparing them to everyone else double-counts who they already were.
How to answer this, stage by stage
Nobody is grading whether you can say the word "counterfactual." They are grading whether you can turn a plausible-sounding before-and-after chart into an honest number, with the arithmetic shown. Seven moves get you there.
Stage 1
Ground it in one real feature, one real number
Say it like this
"Let's ground this. Mailtide is a marketing-automation platform. Warmline is the feature inside it: it writes the subject line and picks the send time for each recipient on its own, instead of a marketer setting one subject line and one send time for the whole list. Marceline Iyer runs growth analytics there, and the number in question is campaign conversion rate, the share of people who get an email and end up buying."
Why this works
Names the product, the feature, and the exact metric before any argument starts, so nothing later has to be taken on faith.
Stage 2
Say the trap out loud before falling into it
Say it like this
"Here's the trap. Warmline launched right into Mailtide's Q4, when every customer's conversion rate rises anyway because shoppers buy more in the run-up to the holidays. If Marceline just compares accounts before Warmline to the same accounts after Warmline, she will credit Warmline for the season."
Why this works
Naming the specific AI-adjacent trap, crediting a rising tide to a feature that launched into it, shows the interviewer you know why before-and-after fails here specifically, not just in general.
Stage 3
Write the equation before touching a single number
Say it like this
"So here's the equation. Warmline's true effect equals the outcome for the group that got Warmline, minus what that same kind of account would have done anyway. And the only honest way to get the second half of that equation is a group that looks just like the treated group but never got Warmline, over the same weeks."
Why this works
This is the B step. Saying the subtraction out loud, before any figures, is what separates a real estimate from a number that happens to sound right.
Stage 4
Own the real numbers, plainly
Say it like this
"Here's what I'd assume, and where each number comes from. Mailtide onboarded about 1,000 new accounts in a four-week window. Marceline held back 10 percent, 100 accounts, from Warmline for those four weeks, purely at random. Conversion rate for both groups started at 2.10 percent. The Warmline group ended the four weeks at 2.60 percent. The holdback group ended at 2.35 percent. One more assumption worth stating: Warmline's own writing model got retrained partway through, in week three. Marceline pinned the comparison to that single model version on both sides, so a mid-window model update couldn't sneak into the number disguised as growth."
Why this works
This is the O step. Every number gets a source, and pinning the model version is what makes this an AI-specific estimate rather than a generic before-and-after, a model retrain is a second moving confound a non-AI feature would never have.
Stage 5
Do the subtraction, then widen it into a range
Say it like this
"The naive number is 2.10 to 2.60, a 0.50 point lift, and that's what a before-and-after slide would claim as Warmline's win. But 100 accounts is a small group, so its own 2.35 percent number moves around, it could honestly be anywhere from 2.20 to 2.50 given that sample size. Subtract that whole band from the Warmline group's 0.50 point move, and Warmline's real, attributable lift is somewhere between 0.10 and 0.40 points, not a flat 0.50."
Why this works
This is the U step, and it matches the direct answer. A single point estimate here would claim precision nobody has.
Stage 6
Run the sanity check against something already believed
Say it like this
"Before Marceline reports this anywhere, she checks it against finance's own Q4 model, built off past years' holiday seasonality and never touching Warmline at all. That model expects roughly 0.20 to 0.30 points of pure seasonal lift this quarter. That's almost exactly what the holdback group showed on its own. When an outside number and your holdback group agree, that's a sign the holdback number isn't a fluke."
Why this works
This is the N step. Checking the number against an independent source is what stops a confident-sounding estimate from being a wrong one nobody caught.
Stage 7
Name what would swing the answer most, and close on the trade
Say it like this
"The single assumption that moves this most is the holdback group's size. At 100 accounts the band is already 0.10 to 0.40. Cut the holdback to 20 accounts and the band gets so wide it swallows the whole result, you couldn't say anything useful. We rejected the easier route, just comparing accounts that turned Warmline on themselves against everyone else, because the accounts that opt into a new AI feature are usually the ones already growing fastest. That would have overstated Warmline's effect, not understated it. The real cost we accepted instead: 100 real accounts got a slightly worse product for four weeks, and Mailtide waited a month for a number it could defend instead of reporting one on day one."
Why this works
Naming both the rejected shortcut and the real cost is what turns a estimation exercise into a defensible decision, not just an arithmetic trick.
If you remember one thing
A number that moved is not proof your feature moved it. Two things can rise in the same quarter for two different reasons, and only a holdback group can tell you how much belongs to each.
Let's learn
What happens when a company launches an AI feature right into a season where the underlying number was already about to climb on its own?
Mailtide sells marketing-automation software to mid-market companies that run email campaigns. Warmline is the AI feature inside it: instead of a marketer writing one subject line and picking one send time for an entire list, Warmline writes a slightly different subject line for each recipient and decides, recipient by recipient, when to actually send it.
Before Warmline, the average Mailtide customer's campaign conversion rate, the share of recipients who opened an email and went on to buy something, sat around 2.10 percent. That number had been roughly flat for a year.
Four weeks, two groups, one question: how much of this is Warmline?
The Warmline group climbed 0.50 points. The holdback group, which never touched Warmline, still climbed 0.25 points on its own. Only the piece above the holdback's own climb belongs to Warmline.
Ten weeks after Warmline shipped, Marceline Iyer pulled the first read. Accounts using Warmline were converting at 2.60 percent, a jump that looked, on a simple before-and-after chart, like a 24 percent relative lift. Her VP of growth, Freida Lund, wanted that number in the board deck by Friday.
The chart said Warmline moved the number. It never said how much of the number was already moving before Warmline touched it.
Here is the turn. When Marceline pulled up the 100 new accounts she had, on a hunch, held back from Warmline for those same four weeks, their conversion rate had also risen, from 2.10 percent to 2.35 percent. Nobody had given those accounts a single AI-written subject line. Q4 was doing that on its own.
The naive number was 0.50 points. Once the holdback group's own noisy climb gets subtracted out, the honest answer is a band, and finance's own Q4 model, built with no reference to Warmline at all, lands right inside it.
Knowledge spark: what is a holdback group?
A slice of accounts, chosen at random, that get everything except the new AI feature, for a set stretch of time. It is the only honest stand-in for "what would have happened anyway." Without one, a rising number and a rising season look identical on a chart.
At its worst, this costs more than an overstated slide. If Mailtide had reported the full 0.50 point lift as Warmline's number, and then priced Warmline as a paid add-on based on that claim, customers who bought it expecting a 24 percent lift would have gotten closer to half of that, and Mailtide's own renewal conversations would have opened with a number it could not defend.
The choice I would take back is not the holdback group, it's that Mailtide almost skipped one. Early on, someone suggested a faster read: just compare accounts that turned Warmline on themselves against accounts that hadn't gotten around to it yet. That felt reasonable, it used real data already sitting there, and it would have shipped a number two weeks sooner. It made sense until someone asked who turns on a new AI feature the week it ships. Usually the more engaged, faster-growing accounts already. Comparing them to everyone else would have credited Warmline with growth those accounts already had coming.
What I would leave alone: Warmline's send-time picking, the half of the feature that decides when an email goes out rather than what it says. That part barely needs this whole exercise, its effect shows up as a clean shift in open-time distribution within days, no four-week holdback required, because timing has almost no seasonal confound to separate out.
The lesson: a number that climbed during the same quarter you shipped a feature is not evidence the feature caused the climb. It is evidence you now have two candidate causes and no way yet to tell them apart.
Now here is the same thing as a story
The short version sits above. Read on for how ordinary the week looked from inside Mailtide, right up until the board deck almost went out wrong.
Marceline Iyer has run growth analytics at Mailtide for three years. She is the person who, months earlier, caught a reporting bug that had been quietly double-counting free-trial conversions for two full quarters, just by refusing to trust a number that looked a little too good. That habit, distrust anything suspiciously clean, is why she is the one who gets handed Warmline's first performance read.
Warmline launched in early October. By the second week, the growth team's shared dashboard showed conversion climbing, steadily, account after account. Freida Lund, who runs growth for Mailtide, started referencing it in Monday standups as "the Warmline lift" before the team had actually measured a lift at all.
Marceline didn't argue in the moment. She just quietly added one line to the account provisioning script: any new signup in the next four weeks had a 10 percent chance of being flagged "Warmline: hold," a random pick, nothing about the account itself decided it. Nobody outside her team noticed. It cost those 100 accounts a feature everyone else was getting.
She didn't slow the launch down. She just made sure someone, somewhere, wasn't getting it.
By the third week of November, Freida had a Friday board deck due and wanted the number: "Warmline drove a 24 percent lift in conversion, quarter to date." Marceline pulled the Warmline group's own numbers and they matched, 2.10 to 2.60 percent. For about ten minutes, before she pulled up anything else, that was going to be the number in the deck.
Then she checked the 100 held-back accounts, mostly out of habit. They had moved too, from 2.10 to 2.35 percent, with nothing but the ordinary Q4 shopping season touching them. She sat with that for a minute before opening a new tab.
Two years earlier, when Mailtide's growth team first built its feature-launch reporting process, the design decision had been simple: measure before and after, on the accounts using the new thing. It was fast, it matched what every dashboard already tracked, and building a randomized holdback into every launch felt like slowing down a team that was already stretched thin. Nobody in that meeting was picturing a specific number that would turn out to be double what it should have been. Why would they. Nothing had gone wrong yet.
What Marceline actually did that Thursday: she pulled Mailtide finance's own Q4 seasonality model, built a year earlier off holiday-shopping history alone, with no line in it about Warmline at all. It predicted 0.20 to 0.30 points of pure seasonal lift for Q4, across the whole customer base. That range sat almost exactly on top of what her 100 held-back accounts had shown on their own. Two numbers, built by two teams who'd never compared notes, landing in the same place. That was the sanity check passing.
Friday's deck went out with a different number: "Warmline's attributable lift is 0.10 to 0.40 points, roughly 5 to 19 percent relative, with 0.25 points as our best estimate, confirmed against finance's independent seasonal model." Slower to write. Half the size of the number Freida had wanted. It was also the number nobody on the board could later prove wrong.
What I would tell myself, before any of this: the week a number starts climbing is the week to ask what else is climbing at the same time, not the week to claim it.
BOUND, the five moves for an estimate you can defend
This is an estimation question, isolate one cause's share of a number that moved for several reasons, so BOUND fits. Not a habit with two settings, not a ranked list of what to build.
B
Break it down. State the equation before touching a number.
Warmline's true effect equals the Warmline group's move, minus what that same kind of account would have done with no Warmline at all.
In this story: 2.60% minus (what the holdback group's own number says the group would have reached anyway).
O
Own the numbers. Say where each assumption came from.
1,000 new signups, 10 percent (100 accounts) held back at random for four weeks. Both groups started at 2.10%. Warmline group ended at 2.60%. Holdback group ended at 2.35%.
Every figure traces back to the same four-week window, nothing borrowed from a different period.
U
Use a range. A 100-account group's own number is noisy.
The holdback group's 2.35% could honestly sit anywhere from 2.20% to 2.50% given its size. Attributable lift: 0.50 minus that band, so 0.10 to 0.40 points, not a flat 0.50.
A single point estimate here would be false precision.
N
Nail the sanity check. Compare it to something already believed.
Finance's Q4 seasonality model, built with zero reference to Warmline, independently predicted 0.20 to 0.30 points of general lift. That's almost exactly the holdback group's own move.
Two independent numbers agreeing is what makes the range trustworthy, not just plausible.
D
Direction. Which one assumption would swing this most.
Holdback group size. Shrink it from 100 accounts to 20 and the noise band widens so far the whole exercise stops meaning anything.
This is the number Marceline would defend a bigger holdback for, next launch.
Two things worth naming directly, since this is where the AI-specific judgment actually lives. First, the alternative most teams reach for is comparing self-selected adopters, accounts that turned Warmline on themselves, against everyone else. That got rejected on purpose: the accounts most likely to opt into a new AI feature early are usually the ones already growing fastest, so that comparison overstates the feature's own effect, it does not just add noise, it adds bias in one direction. Second, an AI-specific failure worth guarding against sits inside Warmline itself: its subject-line writer can drift toward clickbait-style phrasing that spikes opens without spiking actual purchases, or worse, nudges up spam complaints. The guardrail is watching unsubscribe and spam-complaint rate inside the same holdback comparison, not just the topline conversion number, so a subject line that wins on opens but loses on trust gets caught before it scales. The trade accepted here is real: a slower, more defensible number costs Mailtide a month of reporting silence and costs 100 real accounts a slightly worse product for four weeks, against shipping a headline number that might be roughly double the truth.
And if you want to be sure it really works, try it somewhere else
Same five letters, ticket sales instead of email marketing, so the method proves itself instead of repeating a story I happened to prepare.
TicketFathom runs PriceScout, an AI feature for independent concert venues that recommends a dynamic ticket price per show based on demand signals. Bastien Dumont handles venue analytics there. Live music demand across the whole industry was already climbing that quarter, a well-known post-summer bump that has nothing to do with any one venue's software.
B, break it down. PriceScout's true effect equals the sell-through rate for venues using it, minus what those same venues would have sold anyway, in a season where demand was already rising industry-wide. O, own the numbers. TicketFathom held back 15 percent of new venues, 45 of them, from PriceScout for six weeks. Both groups started the season around 71 percent sell-through. Venues with PriceScout reached 79 percent. Holdback venues reached 75 percent. U, use a range. With only 45 venues in the holdback group, its own 75 percent could plausibly be 73 to 77 percent. Attributable lift: 8 points minus that band, so roughly 2 to 6 points, not a flat 8. N, nail the sanity check. An industry ticketing trade report, built off nationwide sales with no mention of TicketFathom, put the seasonal sell-through bump at 3 to 5 points. That lands right inside the holdback group's own range. D, direction. The assumption that would swing this most is the length of the measurement window. Six weeks was long enough for the holdback group to show a stable number; at two weeks, a single big show in either group could have swung the whole comparison.
Both groups of venues sold more tickets this season. The holdback group's own 4-point rise is the season. Only the gap past that belongs to PriceScout.
Same shape, different stakes
At Mailtide, the shared rising tide was the holiday shopping season. At TicketFathom, it's the post-summer concert bump. The method doesn't change: hold back a real slice, measure both groups over the same window, and subtract before you claim a single point of it.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the subtraction: never report a treated group's before-and-after move as the feature's effect, always subtract what a real holdback group did over the same window.
Cost: engineering says a proper randomized holdback can't ship this launch, only next quarter's. Don't report the naive before-and-after number as a placeholder in the meantime, say plainly that the number is unattributed until the holdback exists.
The model got better, for real: say Warmline's writing model gets meaningfully sharper at subject lines. That's still not proof the attributable lift grew, a better model chasing the same rising season needs the same subtraction as a worse one.
Where people run it wrong.
They report the treated group's own before-and-after number and call it the feature's lift, with no holdback in sight.
They build a holdback group so small its own number is meaningless, then report a single point estimate off it anyway.
They skip the sanity check, so nobody outside the analytics team ever finds out the number disagreed with something the company already believed.
How to use it live. Open with the subtraction, before any numbers: "I'd never trust a before-and-after number here, because I can't tell a feature's effect apart from the quarter just getting better on its own, so the first thing I'd build is a real holdback group." That buys you room to walk through real arithmetic instead of reciting "we'd A/B test it" on reflex.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
What framework fits separating an AI feature's effect from general growth, and why?
Tap to flip
ANSWER
BOUND. It's an estimation question, isolating one cause's share of a number that moved for several reasons, not a habit with two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marceline Iyer, who runs growth analytics at Mailtide, the marketing-automation company behind the AI feature Warmline.
3 · THE EQUATION
What is the equation behind the B step here?
Tap to flip
ANSWER
Attributable effect equals the treated group's move, minus what a matched holdback group did over the same window with no access to the feature.
4 · THE TRAP
What is the AI-specific trap this answer names?
Tap to flip
ANSWER
Crediting a feature for a general seasonal or growth trend that was already rising in the same quarter the feature happened to launch.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Almost measuring feature launches by before-and-after alone, no holdback. It made sense because it was fast and matched every existing dashboard, before anyone had a specific overstated number to point to.
6 · THE NUMBER
Fill in the blank: the naive before-and-after lift was 0.50 points. The honest, holdback-adjusted range was ___ to ___ points.
Tap to flip
ANSWER
0.10 to 0.40 points. The naive number was up to five times the low end of the honest range.
7 · THE SANITY CHECK
How did Marceline confirm the range wasn't just a fluke of one small group?
Tap to flip
ANSWER
She checked it against finance's independent Q4 seasonality model, built with no reference to Warmline, which predicted 0.20 to 0.30 points of pure seasonal lift, almost exactly matching the holdback group's own move.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what plays the holdback group's role?
Tap to flip
ANSWER
TicketFathom's PriceScout, for concert venues. The 45 venues held back from PriceScout for six weeks play the holdback group's role, against the post-summer demand bump.
Check yourself Score: 0 / 0
Multiple choice
1. What is Warmline's actual attributable lift in conversion rate, once the holdback group's move is subtracted out?
A. A flat 0.50 percentage points, exactly what the before-and-after chart showed.
B. A range of roughly 0.10 to 0.40 percentage points, not a single number.
C. Zero, because the holdback group also grew, so Warmline did nothing.
D. 0.25 percentage points exactly, with no uncertainty stated.
Show hint
Remember the U step: a small holdback group's own number is noisy, so the answer isn't one point.
Show answer
B. Subtracting the holdback group's noisy 0.10-to-0.40-point band from Warmline's 0.50-point move leaves a range, with 0.25 points as the best single guess inside it.
True or false
2. True or false: comparing accounts that turned Warmline on themselves against accounts that hadn't yet is just as trustworthy as a random holdback group.
True
False
Show hint
Ask who chooses to turn on a brand-new AI feature the week it launches.
Show answer
False. Accounts that opt into a new feature early tend to already be growing faster than average, so that comparison overstates the feature's effect. A random holdback avoids that bias entirely.
Fill in the blank
3. Finance's own Q4 seasonality model, built with no reference to Warmline, predicted a general lift of ___ to ___ percentage points, which almost exactly matched the holdback group's own move.
Show hint
Check the sanity-check stage in the walkthrough.
Show answer
0.20 to 0.30 percentage points. An independent number landing inside the holdback group's own range is what made the range trustworthy, not just plausible.
Short answer
4. Which single assumption would swing this estimate the most if it were wrong, and why?
Show hint
Look at the D step of BOUND.
Show answer
Model answer: The size of the holdback group. At 100 accounts, the noise band is already wide (0.10 to 0.40 points). Shrink the holdback to something like 20 accounts and the band widens so far the whole subtraction stops meaning anything, because you'd be subtracting a number you can't actually trust.
Short answer, apply it yourself
5. Think of an AI feature you've used in an app that launched around a time when usage was already rising for another reason (a new season, a viral moment, a broader trend). How would you build a holdback group to find out how much of that rise was really the feature?
Show hint
Look for a slice of users you could randomly withhold the feature from, without breaking the product for them.
Show answer
Model answer: A fitness app that added an AI-generated workout plan feature right as New Year's resolutions were driving sign-ups anyway. Randomly hold the AI plan back from 10 percent of new January sign-ups for four weeks, compare their 30-day retention to the group that got it, and subtract the holdback group's own retention rise (which is just New Year's motivation) from the treated group's rise.
True or false
6. True or false: because Warmline's send-time picking has almost no seasonal confound to worry about, it needed the same four-week holdback treatment as the subject-line writer.
True
False
Show hint
Check "what I would leave alone" in "Let's learn."
Show answer
False. Send-time picking shows its effect as a clean shift in open-time distribution within days, with little seasonal confound to separate out, so it didn't need the same holdback effort as the harder-to-isolate conversion number.
Before you close the answer
Why this works
Tests whether you'll subtract a real counterfactual before claiming a number, or report whatever a before-and-after chart hands you because it happens to look good in a deck. Most candidates stop at "we'd A/B test it" without ever doing the subtraction out loud.
Follow-up traps
"Isn't holding back 100 accounts from a feature just leaving money on the table?" Response: yes, deliberately, for four weeks, because that small real cost is what buys a number Mailtide can defend, instead of a number that's roughly double the truth.
"What if the holdback group happened to have an unusually bad month for unrelated reasons?" Response: that's exactly why the sanity check against finance's independent seasonality model matters. If the holdback number had disagreed sharply with that outside model, that would have been the signal to distrust the holdback read, not just accept it.
If pressed
Mailtide doesn't run the holdback once and stop. It rotates a fresh 10 percent of each new signup cohort into holdback for every major Warmline update, because a model that gets meaningfully better six months later needs its own attributable-lift number, not a number borrowed from launch day.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.