Describe an experiment design that would isolate an AI feature's business impact.
Jorvik sits on a busy executive's inbox. It ranks what's urgent, drafts short replies to the routine stuff, and flags anything that actually needs a decision. Reinier Casswell, Managing Partner at Faircroft Capital, has run it for five months. Fiora Mears, who owns product analytics at Vespermark, the company that builds it, has to prove, before Faircroft renews, that Jorvik is actually why his inbox got faster, and not something else that happened to land in the same quarter.
- Split new users at random into two groups the moment they onboard, not after.Why: a random split inside the same company, in the same weeks, is the only way to know the gap in the numbers came from the feature and not from the quarter.
- Keep the silent group truly silent, not just switched off.Why: if it still scores and logs every item without showing anything, you can grade what the model would have flagged against what actually mattered, with no risk of the group knowing it's being watched.
- Measure on the thing that happens often, not the thing that happens rarely.Why: a dozen people on each side won't produce enough closed deals in one quarter to prove anything. Flagged threads happen dozens of times a week per person, and that's where the real signal lives.
- Watch for the two groups talking to each other.Why: at a firm small enough that everyone sits in the same partner meeting, one shared conversation about a flag can quietly erase the gap you're trying to measure.
- Leave the big, rare outcome out of the claim for now.Why: a dozen people over one quarter isn't enough weight to put a dollar figure on deals closed, and a number built on that little data won't survive one hard question.
- Never tell the silent group they're the control.Why: a person who knows they're being watched stops behaving normally, and normal behavior is the whole point of the comparison.
How to answer this, stage by stage
Nobody is grading whether you know the word "randomize." They're grading whether you can spot that a before and after chart has three plausible causes hiding under it, and whether you'd actually build the thing that tells them apart.
Let's learn
What does it actually take to know that an AI feature caused a number, instead of just showing up next to it?
Jorvik reads a busy executive's inbox. It sorts what needs them from what doesn't, and it drafts short replies to the routine stuff so the executive only has to skim and send.
Before Jorvik, Reinier's own habit was to work through his inbox top to bottom, no screen telling him what mattered. On average it took him about fourteen hours to get to a flagged deal thread buried somewhere in the pile. Five months after all thirteen original partners went live on Jorvik, that number is three hours. On paper, urgent mail gets answered eleven hours sooner.
Here's the turn. That eleven-hour drop isn't proof of anything on its own. In that same window, Faircroft hired its first dedicated deal-flow analyst, whose whole job is pre-screening the partners' mail before they see it. The wider industrial buyout market also sped up that quarter, so more people everywhere were answering faster. The eleven hours might be Jorvik. It might be the analyst. It might be the market. Right now, nobody at Vespermark or Faircroft can tell you which.
At its worst, that costs Vespermark the renewal. Not because Jorvik doesn't work. Because nobody can prove it does, and a finance team asked to keep paying without proof will happily credit the good quarter to something cheaper.
Faircroft is onboarding twenty-four more partners this quarter for a second, industrials-focused fund. Fiora randomizes them into two arms of twelve at the moment each one signs on. The full-Jorvik arm sees drafts and flags exactly like Reinier does now. The silent arm gets Jorvik scoring and logging every email in the background, with nothing shown, so those twelve partners triage their own inbox by hand, same as everyone did before Jorvik existed. Over the thirteen-week quarter, that's roughly 1,080 flagged threads logged in each arm, plenty of volume even though only twelve people sit on each side.
The gap didn't appear on day one. It built over the quarter, as partners in the full arm learned which flags to trust, while the silent arm stayed roughly where everyone starts.
Shadow mode does one more thing worth naming directly. Jorvik's flag has its own cut off point for how sure the model needs to be before it raises its hand. Set that cut off too cautiously and it stays quiet on a real, rare problem, a false negative, exactly the kind of miss nobody notices because the tool never claimed to have caught it. Because the silent arm is scored and logged even though nobody sees it, a reviewer can go back later and check every miss against what actually happened.
What Fiora would leave alone: Jorvik's routine drafts, the ones that just confirm a meeting time or forward a document. Nobody's renewal depends on whether a calendar reply saved four seconds, and this level of proof isn't worth building for it.
The lesson: a before and after number only tells you that something changed. It never tells you which of the three things that changed that month actually caused it. If a design can't tell those apart, the number was never really a measurement. It was a guess with a date stamp on it.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why eleven hours saved was never a number Fiora could actually defend.
Fiora Mears has run product analytics at Vespermark for four years. She built the usage dashboard every account team on the sales side leans on, and she can open a report and land on the one row worth arguing about before anyone else has finished their coffee.
Reinier and the other twelve partners at Faircroft Capital onboarded Jorvik in one week last spring. The account became the one Vespermark's sales team pointed to on every call. Every quarterly review carried the same slide, fourteen hours to three, and Fiora built that slide herself. For a while she was proud of it.
There wasn't one bad meeting that changed things. It built slowly, over about six weeks, ahead of the renewal. Faircroft's CFO kept asking, in slightly different words each time, the same plain question: did Jorvik do this, or would it have happened anyway. Fiora had a good answer the first time. By the fourth time, she noticed she was saying the same sentence with less confidence, and she couldn't put a finger on the exact call where it stopped sounding like a compliment and started sounding like a challenge.
She'd been treating a chart with one line on it as an argument. It was actually a coincidence with a fourteen and a three written on it.
In the meeting where she first tried to fix it, her first idea was to stagger the next twenty-four partners instead of splitting them: four onboard this week, four more in two weeks, and so on, so she could compare the early group to the late group. She liked it for about a day. Then she said it out loud to her own team and heard the hole in it before anyone else pointed it out: if the deal market kept speeding up mid-quarter, the way it just had, the early group and the late group would differ by time as well as by the tool. She'd have built the exact same tangle again, just with a calendar instead of a chart.
What she built instead was the coin flip. Twelve new partners see Jorvik in full. Twelve run it silently, logged, invisible, for the same thirteen weeks. She kept the silent arm off a shared distribution list and asked the account team not to discuss flags in the joint partner meeting until the quarter closed, so nobody in either arm would know which side of the flip they'd landed on.
At the end of the quarter, the full arm's median time to answer a flagged thread was 2.6 hours. The silent arm's was 11.4 hours. Same firm. Same thirteen weeks. Same deal market, moving the same way for both of them. The only real difference was whether Jorvik's output ever reached the screen, and that difference was 8.8 hours, measured across roughly 1,080 flagged threads on each side, not guessed at from one inbox.
What Fiora would tell herself, back when she built that first slide: fourteen to three was never a lie. It just was never an answer either. She'd built a picture that flattered the product before she built one that could survive being questioned.
SPARK, or the difference between a chart and a comparison
Not a way to dress up a before and after number in five letters. SPARK is what forces you to name the exact design decision an experiment turns on, and to say plainly what it still can't prove.
- S: one inbox, three plausible causes, no way to tell them apart.
- P: a renewal number both sides can actually trust.
- A: randomize new partners at onboarding, full versus silent, same firm, same weeks.
- R: spillover between arms, a control that notices, and too few people for a rare outcome.
- K: no revenue claim yet. Only the response time gap, until the design has run more than once.
And if you want to be sure it really works, try it somewhere else
Same five letters, a school district instead of a private equity firm, and this time the rare outcome isn't a closed deal. It's an escalation to the board.
Loksworth triages a school superintendent's inbox: parent complaints, board member emails, state compliance notices, and it drafts short acknowledgements for the routine ones. Rowanmere Unified School District rolled it out to Superintendent Pryderi Coombeshire, and a few months later is deciding whether to buy it for the district's other four office heads too, off the strength of one number: Coombeshire's time to respond to an urgent parent complaint, which fell from two full school days to about four hours.
Same problem sits underneath it. That number lands right in the middle of a school year that also saw a new assistant superintendent hired, and a new state reporting deadline that made everyone answer faster in general, not just Coombeshire.
S: one office head's before and after number, with a new hire and a new deadline both sitting under it. P: a decision the district's board can trust without re-litigating it every budget cycle. A: when Loksworth rolls out to the other four office heads next term, randomize which two go live with the tool visible and which two run it silently logged, for the same eight weeks. R: four office heads is an even smaller group than Faircroft's twelve, and they already meet every Wednesday, so contact between the two arms is a bigger risk here, not a smaller one. Guard it by running the silent pair off a separate list and keeping flags out of that Wednesday meeting until the term ends. K: don't claim Loksworth reduced complaints escalating to the board. Four office heads, one term, isn't enough to say that yet. Trust the response time gap on flagged threads only.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: randomize the next onboarding wave, full tool versus silent logged tool, same weeks, same place. Measure the thing that happens often, not the thing that happens rarely.
Cost: no budget this quarter for a full fifty-fifty split. Ship the cheap version: hold back just a handful of the next cohort instead of half. A small silent arm still beats no comparison at all.
The model got better, for real: say Jorvik's flagging accuracy doubles overnight. The design barely changes. You still don't trust that claim until the shadow log proves it against real outcomes, not the vendor's own number.
Where people run it wrong.
They compare one company to a different company instead of holding the company constant and randomizing inside it.
They let the "silent" group learn it's the control, and its behavior quietly stops being normal.
They try to prove a rare, high-stakes outcome, revenue, an escalation, with a sample sized for a common one.
How to use it live. Ask before you answer: "Is the number we have a comparison, or is it a coincidence with a date stamp on it?" That buys you the room to actually design something, instead of grabbing the nearest flattering before and after chart.
Three things worth stating directly, since this is where the real judgment sits. The alternative Fiora considered and rejected was a staggered rollout, early partners first, late partners later, so she could compare the two groups. It lost, because if the deal market keeps moving mid-quarter, an early group and a late group differ by time as well as by the tool, which reintroduces the exact tangle the design was supposed to remove. The AI-specific failure worth naming is a model that's tuned to stay quiet unless it's sure, which means its cut off point can let a real, rare problem through unflagged without anyone knowing it happened. The guardrail is shadow mode's own silent log: because it scores and records every item in the holdout arm too, a neutral reviewer can check months later whether a miss like that actually mattered, and whether the cut off needs to move. And the trade off is real and accepted on purpose: running Jorvik silently for twelve partners still costs Vespermark the same inference spend as running it for the twelve who can see it, with nothing to show a customer for that spend all quarter, and those twelve partners get none of Jorvik's time savings for thirteen weeks. That cost buys a number worth defending instead of one worth apologizing for.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the silent group figures out they're being logged?" Response: that's exactly why shadow mode has to be invisible in the product itself, no banner, no different screen, and nobody told which arm they landed in.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #7 What ROI argument works for an internal AI tool with no revenue line?