CalculationAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #4

How do you attribute a revenue change to an AI feature specifically?

TRACE · attributing a revenue jump to one legal-drafting feature, before Tuesday's board deck

Every quarter, before anyone else reads Solventra AI's revenue dashboard, a printed ledger lands on Nadine Studholme's desk, and she reads it before she reads anyone's slide about it. This quarter the slide already existed. Grier Halstrom, who runs sales, dropped one line into Tuesday's board deck on a Friday afternoon: expansion revenue up twenty two percent, credited to Argument Draft, the feature inside Motionforge that drafts a legal brief's argument section instead of just finding the cases for it. Nadine had until Monday night to find out if that line was true.

The direct answer
Don't credit Argument Draft from the topline number. Check whether the money lands on each firm's own access date to the feature, not on the price change or the model swap that shipped in the same window, using firms that got the feature later, or never, as your control. If it lands there, the feature earns the credit. If it lands somewhere else, it doesn't, no matter how good the story reads in a board deck.
Do this, in order
  1. Test whether the money lands on each firm's own access date to the feature, not on anything else that shipped nearby.Why: it's the one check that tells the feature apart from a price change or a model swap sitting in the same month.
  2. Rebuild the full release timeline before you touch the revenue number.Why: three other things shipped inside the same five weeks, and any one of them could explain a lift on its own.
  3. Check how the number itself got counted before you credit anyone.Why: a billing change quietly moved six real signings out of last quarter and into this one, before anyone asked why revenue was up.
  4. Split the blended total into the lines that actually make it up.Why: one line moved 31 percent, the other moved 9, and averaged together they hid both real numbers.
  5. Name every rival cause a skeptic would raise, before they raise it.Why: a cause list with only the flattering answer on it hasn't been tested, it's been decided.
  6. Don't run this whole process on a change with nothing else shipping near it.Why: a feature that launches alone, with no price change and no model swap nearby, doesn't need five steps to explain a bump.

How to answer this, stage by stage

Nobody is grading whether the twenty two percent turns out to be true. They're grading whether you check it before you repeat it, out loud, to the people who sign your budget.

1
Scope it to one product, one number, one quarter
Say it like this
"Let's say I'm a PM at an AI company that sells legal drafting software to law firms. We just shipped a feature that drafts the argument section of a brief, not just finds the cases. Expansion revenue is up twenty two percent this quarter, and sales wants to credit the feature. I need to find out if that's actually true before it goes in the board deck."
Why this works
Naming the product, the number, and the deadline turns a vague prompt into something with real stakes, and it tells the interviewer exactly what's being tested.
2
Say your plan out loud
Say it like this
"I'm going to run TRACE. Timeline, recut, assume nothing, cause candidates, evidence test. Five steps, and the last one is the one that actually decides it."
Why this works
Naming the method up front buys thinking time and tells the interviewer a real process is coming, not a guess dressed up as confidence.
3
Reframe what's actually being tested
Say it like this
"This isn't really asking whether the feature is good. It's asking whether I'll believe a story the moment it's flattering, or whether I'll go check it before it becomes a fact everyone repeats."
Why this works
This is where a shallow answer, one that just praises the feature and moves on, gets separated from a real one.
4
Give the one decision, before any reasoning
Say it like this
"Here's what I'd do: check whether the extra revenue lands on each firm's own access date to the feature, not on the price change or the model swap that shipped the same month. If it does, the feature gets the credit. If it doesn't, I say so, even in the board deck."
Why this works
This is the direct answer, said in the room, before the interviewer has to wait through five steps to find out what you'd actually do.
5
Walk the timeline, name every rival cause
Say it like this
"Four things shipped inside five weeks. The feature went out to half our sales reps on April third and the other half on April twenty fourth. A price increase was announced April tenth. A new model went live April seventeenth. Any one of those three could explain a revenue bump that shows up by June thirtieth."
Why this works
Naming the other suspects out loud, unprompted, is what tells the interviewer you're not building a case for the answer you already wanted.
6
Check the plumbing and recut the number before you believe any of it
Say it like this
"Before I trust the twenty two percent, I check whether the number itself changed. It had. Six of the firms counted this quarter actually signed in the last two weeks of last quarter. Once I correct that, it's fifteen firms and sixteen percent, not twenty one firms and twenty two."
Why this works
A tracking or billing change looks exactly like a real lift on a dashboard, and checking it first is what stops you explaining something that never happened.
7
Run the one test that decides it, and close on the line
Say it like this
"Eleven of the fifteen real upgrades signed within about ten days of their own rep pitching them the feature, wave one or wave two. Only two signed anywhere near the price date. So: the feature gets the credit for the upgrade line, a real but smaller number than sales wanted, and I can defend every point of it."
Why this works
Ending on a number you can defend, not a bigger number you can't, is what survives the follow-up question.

Let's learn

Motionforge is Solventra AI's tool for litigation teams. It reads the facts of a case, searches real case law for the arguments that fit, and drafts the pieces of a legal brief around them. Two tiers exist. Research tier finds the cases and the citations. Pro tier adds Argument Draft, a feature that writes the argument section itself, in full sentences, ready for an attorney to edit instead of write from a blank page.

Hand sketched labeled parts diagram titled Who checks the number before the board sees it. A person icon in the center labeled Nadine Studholme, with four callouts around her: Owns Solventra's revenue numbers, Rebuilds it from raw events, Never credits a story cold, Marks up a printed sheet by hand.
Most quarters, this check takes an hour and confirms what everyone already believed.

Nadine Studholme owns Solventra's revenue numbers, and every quarter she pulls the raw billing events herself before she trusts a chart built from them. This quarter, Grier Halstrom, who runs sales, had already written a line into Tuesday's board deck: expansion revenue up twenty two percent, credited to Argument Draft. He sent it to her Friday afternoon, mostly as a courtesy.

Knowledge spark: why does a drafting feature need its own guardrail? A model that drafts an argument can also invent a case that sounds real but isn't, a well known real world problem for AI legal tools. Argument Draft checks every citation it writes against Solventra's real case law database before the sentence ever reaches the page. Anything that doesn't match a real, on point case gets flagged red, not quietly inserted. That check is slower than a plain citation search, which is part of why Argument Draft only ships on the slower, costlier tier.

For the last four quarters, expansion revenue held close to $186,000 a quarter: about $61,000 from firms upgrading Research to Pro, and about $125,000 from existing firms adding attorney seats. About twelve firms upgraded in a typical quarter. This quarter the dashboard read $227,000, twenty two percent higher, and it credited twenty one new upgrades to Pro.

Hand sketched horizontal timeline titled Four things shipped, one number moved. Five milestones left to right: Argument Draft Wave 1, caption reps start pitching it Apr 3. Price change announced, caption Research tier list price Apr 10. New citation model live, caption every tier Apr 17. Argument Draft Wave 2, caption reps start pitching it Apr 24. Quarter closes, this milestone highlighted in red, caption revenue read up 22 percent Jun 30.
The dashboard only shows the one number that moved, at the very end of four separate clocks.
Nobody asked whether Argument Draft was good. They asked whether twenty two percent was true, and those turned out to be two different questions.

Before Nadine believed any of it, she checked how the number itself gets counted. Billing had changed how a signed contract gets dated, starting April first: revenue now posts by invoice date instead of contract date. Six of the twenty one credited upgrades had actually signed in the last two weeks of March, before some of them had even seen Argument Draft. Push those six back to the quarter they actually happened in, and twenty one becomes fifteen. Twenty two percent becomes sixteen.

Hand sketched comparison diagram titled The line that used to be two lines. Left panel, a document icon labeled Before April, captioned tier upgrades and seat growth, two separate lines. Right panel, a document icon in red labeled After April, captioned both folded into one line, expansion revenue.
A billing change on April 1st quietly turned two tracked lines into one, before anyone asked why the one line moved.
The correction that mattered $227,000 was never real, not because anyone lied, but because the ruler moved on April first and nobody told the team reading it. The honest number, once the six firms land in the right quarter, is $216,000: still a real increase, still worth explaining, just not the one already sitting in the board deck.

That $216,000 is two different lines added together. Tier upgrades: $80,000, up from $61,000, a 31 percent jump. Seat growth on existing accounts: $136,000, up from $125,000, a 9 percent jump. Averaged together they read like one modest, believable story. Split apart, they're two different stories, and only one of them has anything to do with Argument Draft.

At its worst, this doesn't cost Solventra two days of Nadine's time. It costs them a story that gets repeated to the board, then to next quarter's sales targets, then to a pricing strategy built on a number that was never going to happen twice: a one time billing correction and a one time price deadline don't repeat every quarter. The feature would still be good. The story about it would just be wrong, and wrong stories get expensive exactly when nobody's checking anymore.

The choice I would take back Solventra shipped a price change, a model migration, and a two wave feature rollout inside the same five weeks, and nobody flagged that a billing definition was changing in that same window too. Nobody decided on purpose to make this hard to trace. The release calendar just never asked what would happen if three causes and one metric change all landed in the same month. I would hold billing changes to their own release, on their own date, every time, specifically so a number like this only ever has one explanation to check.

What I'd leave alone: if Motionforge shipped a single feature with nothing else changing that month, no price move, no model swap, I wouldn't run any of this. A number that moves right after one clean, isolated change earns the easy explanation. This quarter wasn't that quarter.

The lesson: a number that moves the same month your favorite feature ships is not proof the feature moved it. If more than one thing changed in that window, the topline can only tell you that something happened. Only the date on each customer's own decision can tell you which something it was.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a number that was never going to repeat itself almost went out under Solventra's name.

Nadine Studholme has closed Solventra's books for six years, and she has exactly one rule she's never broken: nothing she didn't rebuild herself goes in front of the board. She built that rule the ordinary way, by watching two people above her get burned by a chart neither of them had checked.

Most quarters reward the rule by being boring. The dashboard closes, Nadine spot checks a sample of the biggest accounts against the raw ledger, the numbers match, and she signs off by Wednesday lunch. She's done this close to fifty times. Fifty times it held.

It thinned in three quiet beats. Beat one: in February, the billing team changed how a signed contract gets dated, folding it into the same release as an unrelated invoicing fix, because it was a two line change and nobody thought a dating rule needed its own announcement. Beat two: in March, marketing started using the phrase "AI driven growth" in an internal channel before Argument Draft had even reached half the sales floor, because the early numbers already looked good and a good story travels faster than a caveat. Beat three: by the first week of July, the phrase had made it into a board deck draft as a stated fact, twenty two percent, credited, with nobody between that channel and the slide who had actually pulled a single raw event.

Grier didn't ask Nadine to check anything. He sent her the slide as a courtesy, Friday afternoon, the way you'd forward someone a nice thing said about their work. "Going in Tuesday's deck, thought you'd want to see it," the message said. Nothing about it looked like a request.

Here's what Nadine did next, and it wasn't sign off from her phone the way the message invited her to. She pulled the raw billing events for every account that touched Pro tier that quarter, printed the ledger, and spent most of Saturday and half of Sunday with a red pen, marking every signed date against every rollout date by hand. By Sunday night she'd found the six firms. By Monday afternoon she'd built the wave by wave timing test. Call it sixteen hours, across a weekend that was supposed to be hers.

We didn't nearly lose two days of Nadine's weekend. We nearly told the board a story that was never going to happen again, because a billing date and a price deadline don't repeat every July.

It was never really about whether the answer was twenty two percent or sixteen. It was about whether one single person, somewhere between an internal message and a board slide, was going to check before the number left the building. This time there was. Most Fridays, there might not have been.

Months earlier, in a routine release planning meeting, someone had proposed bundling the price change with Argument Draft's launch, to make one announcement instead of two. The billing team, separately, slotted their invoicing fix into the same two week release window because the calendar had room. Nobody in either meeting asked what would happen if a customer's decision, a company's pricing letter, and a ledger's own definition of "this quarter" all landed in the same five weeks. It wasn't a bad meeting. It was a normal one. Nobody in it was trying to make anything hard to explain.

Run it again with one change: billing changes ship on their own dated release, always, flagged to analytics the day they go out, no exceptions for a two line fix. Same price change, same feature, same model swap. The board still gets a real number in July. It's sixteen percent instead of twenty two, broken into two lines instead of one, sourced to the exact date on each firm's own contract. Nadine builds that slide in about two hours instead of sixteen, because nothing about how the quarter closed needs reconstructing by hand.

One version of this quarter needed a weekend and a red pen to become true. The other version was already true the moment it closed.

What I'd tell myself, back in that release planning meeting: convenient timing for a launch calendar is not the same thing as safe timing for a number somebody will have to defend later. I would have asked one question nobody in the room thought to ask: what else is shipping this month, and who's going to be able to tell them apart in July?

TRACE, the five checks a revenue story has to survive

Not a checklist for finding the true cause. TRACE is what stops you from crediting the first cause that happens to flatter the team that built it.

TTimeline. When did the number actually move, and what shipped nearby?
Four things shipped in the same three and a half weeks: Argument Draft went out to the first half of Nadine's sales reps on April 3rd, a Research tier price rise was announced April 10th with a ninety day grace period, a new citation retrieval model went live for every tier on April 17th, and Argument Draft reached the second half of the reps on April 24th. The dashboard only shows one number moving, on June 30th. It never shows that four separate clocks started ticking first.
Start further back than the number moved. The causes here all sit inside one month; the effect gets read three months later.
RRecut. Split the blended total before trusting the average.
"Expansion revenue" is not one thing. It's tier upgrades, a Research firm paying to become a Pro firm, plus seat growth, an existing firm adding more attorneys, added onto the same line since April. Split back apart: the upgrade line moved 31 percent. The seat line moved 9. A single 16 percent average hides both numbers and explains neither.
A blended average is where the interesting story hides. Always ask what's inside the number before asking why it moved.
The recut: two lines hiding inside one 16 percent average
$150k $75k $0 $61k $80k Tier-upgrade revenue $125k $136k Seat-expansion revenue
Baseline, trailing quarter averageTier-upgrade line, correctedSeat-expansion line, corrected
Tier upgrades moved 31 percent. Seat growth moved 9. The blended 16 percent headline is the average of two very different stories.
AAssume nothing. Rule out the plumbing before you rule on the person.
Six of the twenty one firms the dashboard credited to this quarter had actually signed their upgrade contracts in the last two weeks of March, before some of them even had Argument Draft. A billing change, effective April first, started posting revenue by invoice date instead of contract date, and it quietly pulled six real signings out of the previous quarter and into this one. Fix that first, and twenty one becomes fifteen, and twenty two percent becomes sixteen.
Check whether the ruler moved before you check whether the thing you're measuring did. A tracking change looks exactly like a real lift on a dashboard.
CCause candidates. Three real stories, not a list of everything possible.
One, Argument Draft is genuinely worth upgrading for. Two, firms are paying to lock in today's price before the grace period ends. Three, the new citation model made the whole product better, Research tier included, so more attorneys everywhere started trusting it enough to add seats, whether or not they ever touched Argument Draft.
Name the ones a skeptic would actually raise. A cause list with only the flattering explanation on it hasn't been tested, it's been decided.
Hand sketched numbered icon list titled Three stories that would explain the same number. Item 1, green, Argument Draft, firms want the drafting feature. Item 2, grey, Price change, firms upgrade before the grace period ends. Item 3, amber, New citation model, the whole product got better.
Only one of these three turns out to explain the tier-upgrade line specifically.
EEvidence test. The one check that tells the three apart.
Look at where each real upgrade landed in time, against each firm's own pitch date, not the calendar date everyone's staring at. Eleven of fifteen signed within ten to twelve days of the date their own rep started pitching Argument Draft to them, wave one or wave two. Only two signed anywhere near the price announcement. And Research tier firms, the ones who never got Argument Draft at all, still grew their seat-expansion revenue by 9 percent, right on schedule with the new citation model going live. Three different clocks, three different signatures, and only one of them lines up with Argument Draft.
This is the whole method in one move. Whichever cause is real leaves a fingerprint at the level of one customer, one date. An average can't show you that. Only the raw events can.
Hand sketched scatter diagram titled Where the upgrades actually cluster. Vertical axis, days after the price announcement, from same week at the bottom to eight or more weeks later at the top. Horizontal axis, days after this firm's own pitch, from same week at the left to eight or more weeks later at the right. Eight labeled firm points cluster tightly on the left side of the chart, close to their own pitch date, while spreading widely up and down relative to the price announcement date.
Eight sample firms, plotted against both dates. They line up on one axis and scatter on the other.
When the money actually moved, week by week
4 2 0 Wave 1 Price, no bump Model live Wave 2 Apr 3 Apr 24 Jun
Upgrades signed that week
Two clusters, one right after each wave's own pitch date. No third cluster around the price announcement, which is the whole evidence test in one line.

One thing worth stating directly, since this is where the real judgment sits. There was a cleaner way to run this test: pick a matched set of Research tier firms on purpose and hold Argument Draft back from them for a quarter, purely to measure the effect. Nadine didn't ask for that, and Solventra didn't build it. Denying a paying-eligible law firm a feature every similar firm gets, just so an internal team can measure something, is a bad trade against a real firm's real business, and a customer would have felt that decision before any dataset did. The staggered training rollout worked as a stand in because it already existed for an ordinary reason: the enablement team could only run so many rep trainings a week. A found natural experiment beats a built one you'd have to justify to a customer.

Knowledge spark: what makes a staggered rollout a real test? Wave one and wave two got Argument Draft on different dates for a boring reason: only so many reps could be trained per week. Nobody picked which firms went first based on size, region, or how likely they were to upgrade. That's what makes the split usable as a natural experiment instead of a coincidence: the timing was decided by a training calendar, not by anything about the firms themselves.
Hand sketched decision tree titled The one test that tells the three stories apart. Root box, Where did the extra revenue land, in time? Three branches: clusters near each firm's own pitch date leads to Argument Draft gets the credit, outlined in red. Clusters near the price announcement instead leads to Price change gets the credit, outlined in red. Spread evenly, even firms with no Pro access, leads to Model upgrade gets the credit, outlined in red.
The whole page compressed into one test. Solventra's real quarter landed on the leftmost branch.

And if you want to be sure it really works, try it somewhere else

Same five letters, a farm cooperative instead of a law firm, and this time the feature is not the one that wins.

FieldSentry is Pallwick AI's tool for spotting crop disease from drone and satellite images, sold to farm cooperatives. Katalin Meissner owns its growth numbers. Subscriptions to the Early Warning tier jumped 19 percent the same month FieldSentry shipped a new Outbreak Confidence Score, the same month the region had its wettest spring in six years, and the same month a distributor ran a subscription push through its usual referral code.

Hand sketched left to right flow diagram titled Same tangle, a wet spring instead of a price change. Four steps connected by arrows: Score ships, this step outlined in green, Wet season begins, Distributor runs a push, Subs jump 19 percent.
Three plausible causes again. This time the feature is not the one that survives the test.

T, timeline: three things landed inside the same five weeks, the Confidence Score feature rolled out to co-ops in two batches, the wettest spring in six years began, and a distributor's push email went out in week three.

R, recut: split subscription growth by co-op region. The three wettest regions accounted for most of the lift, not the two rollout batches evenly.

A, assume nothing: Pallwick's own signup tool had started crediting organic signups to the distributor's referral code that same month, a tracking default, not a real lift in referred signups.

C, cause candidates: the Confidence Score feature, real fear driven demand from a wet season, and the distributor's push.

E, evidence test: co-ops in dry regions that got the same two rollout batches of Confidence Score show almost no lift at all, while wet-region co-ops jumped regardless of which batch they were in. That rules out the feature as the primary driver and points at the weather.

Same rank, different winner: at Solventra the feature earned its credit. At Pallwick, the weather did. Different climate, same asymmetry: whatever a topline number can't see, it can't warn anyone which cause actually deserves the line in the deck.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the decision: "test whether the money lands on each customer's own access date; if it does, the feature earns the credit."
Cost: no time to pull raw events this quarter. Compare the two waves' upgrade rates over the same number of days since each wave started, using whatever the CRM already logs, cheaper than a full event pull.
The model got better, for real: say the citation model migration alone drove a 10 percent lift in retention, no feature, no price change involved. Credit the model migration specifically, by name, not the product broadly. A better model is still a nameable cause, not "things got better somehow."

Where people run it wrong.
They credit the topline number because it moved the same quarter a favorite feature shipped, without checking what else moved too.
They treat "up 22 percent" as one fact instead of asking whether the meter itself changed.
They average across firms who could and couldn't have used the feature, instead of splitting them apart first.

How to use it live. Ask one buying-time question before answering: "before I credit anything, can I ask what else shipped in that same window?" That question alone buys a breath, and it signals the right instinct before you've said a single number.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
TRACE: rule out the fake causes before crediting the real one. Built for diagnosis questions, when a number moved and several things could explain it. (Swapped in for "flip family," which TRACE doesn't have.)
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Nadine Studholme, senior product manager at Solventra AI, who owns Motionforge's revenue numbers and rebuilds every one from raw billing events before it goes in front of the board.
3 · THE BLENDED NUMBER
What single number did the dashboard show, and what was it hiding?
Tap to flip
ANSWER
Expansion revenue, reported at $227,000, up 22 percent. It blended two different lines, tier upgrades and seat growth, that had actually moved by very different amounts, 31 percent and 9 percent.
4 · THE THREE SUSPECTS
Name the three rival causes for the revenue jump.
Tap to flip
ANSWER
Argument Draft itself (feature value), the price change's grace period (pull-forward upgrades), and the new citation model migration (a broad quality lift showing up as seat growth, even on Research tier).
5 · THE OLD DECISION
What decision would Nadine's team take back?
Tap to flip
ANSWER
Bundling the price change, the model migration, and Argument Draft's two waves into the same five-week release window, with no flag to analytics that a billing definition was changing too.
6 · THE NUMBER
Fill in the blank: the dashboard first read expansion revenue up ___ percent. After Nadine corrected the billing-timing error, the real figure was ___ percent.
Tap to flip
ANSWER
22 percent, then 16 percent, once six misattributed firms were moved back to the quarter they actually signed in.
7 · THE REPLAY
Same board deck, corrected analysis, what changes?
Tap to flip
ANSWER
The slide reads 16 percent, not 22, broken into two lines instead of one, sourced to the exact signing date on each firm's own contract, finished by Monday evening instead of stretching the whole weekend.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what does the test find this time?
Tap to flip
ANSWER
FieldSentry, Pallwick AI's crop-disease tool, at farm cooperatives. This time the feature is NOT the primary cause; a wet growing season is, confirmed because dry-region co-ops with the same feature rollout showed almost no lift at all.

Check yourself Score: 0 / 0

Fill in the blank
1. The dashboard first reported expansion revenue up ___ percent. Once Nadine corrected the billing-timing error, the real figure was ___ percent.
Show hint
Check the Assume Nothing step, and the key block titled "The correction that mattered."
Show answer
22 percent, then 16 percent. Six firms that signed in the last two weeks of March were counted in the wrong quarter, inflating the headline number.
Multiple choice
2. Which single check actually separated "Argument Draft caused this" from "the price change caused this"?
  • A. Comparing total revenue across the last four quarters
  • B. Checking whether each firm's upgrade date clustered near its own access date to the feature, rather than the price date
  • C. Asking the sales team which cause they believed
  • D. Waiting until next quarter to see if revenue kept climbing
Show hint
This is the E step, and the direct answer at the top of the page.
Show answer
B. Eleven of fifteen real upgrades signed within ten to twelve days of their own wave's pitch date; only two signed anywhere near the price announcement.
True or false
3. True or false: the new citation model migration explains most of the tier-upgrade revenue increase.
  • True
  • False
Show hint
Look at which revenue line grew among Research-tier firms, who never had Argument Draft at all.
Show answer
False. The model migration shows up in the seat-expansion line, including among Research-tier firms with no Pro access. The tier-upgrade line tracks Argument Draft's own two wave rollout instead.
Multiple choice
4. Which of these changes would NOT need a full TRACE investigation before crediting it?
  • A. A single, isolated feature ships, with no price change or model update anywhere nearby
  • B. Three unrelated changes ship in the same week as a metric move
  • C. A metric's own tracking definition changes in the same month it moves
  • D. A blended total combines two very different customer segments
Show hint
Check "What I'd leave alone" in Let's learn.
Show answer
A. A number that moves right after one clean, isolated change earns the easy explanation. TRACE is for when more than one plausible cause is sitting in the same window.
Short answer, apply it yourself
5. Think of a product you use whose usage or spend changed around the same time as two other things (a price change, a redesign, a season). What single event-level check would tell you which one actually caused it?
Show hint
Think about what you could check for individual users or accounts, not the average, that would look different under each explanation.
Show answer
Model answer: A grocery delivery app raised its free-delivery threshold the same month it added a loyalty badge feature, and orders per week went up. Checking whether order size grew mostly among shoppers who were already just under the new threshold, versus shoppers who engaged with the badge feature specifically, would separate a price-driven basket-padding effect from a genuine engagement effect.
Short answer, work the logic
6. If Wave 2 had gone live on the exact same day as Wave 1, instead of three weeks later, would the access-date test still be able to separate "the feature caused this" from "the price change caused this"?
Show hint
Think about what made the staggered rollout useful as a natural experiment in the first place.
Show answer
No, not on its own. With both waves live on the same day, and the price announcement only a week later, every candidate cause would predict upgrades clustering in roughly the same window. The test only works because the two waves' access dates are far enough apart, and far enough from the price date, that each cause predicts a different pattern.
Once the answer's out of your mouth
Why this works
Tests whether you'll trust a number the moment it flatters the team that built it, or check it against the raw events before it becomes a fact the whole company repeats. Revenue-attribution questions reward the candidate who slows down exactly when everyone else in the room wants to move fast.
Follow-up traps
"Isn't the staggered rollout basically luck? What if the two waves hadn't been three weeks apart?" Response: then the test loses its power, and you'd need a different control, like comparing upgrade rates over equal windows since each firm's own access date instead of raw counts.

"The tier-upgrade line only moved 31 percent on top of $61,000. Does a gap that size even matter?" Response: yes, because it's the line the whole board-deck claim was actually about; a smaller, defensible number beats a bigger, indefensible one the moment anyone asks a follow-up question.
If pressed
Argument Draft's citation check has to clear a strict match threshold against the real case database, not just "looks plausible," and any citation below that bar gets flagged red rather than silently smoothed over, which is also why the feature runs on a slower pass than plain citation search.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more