CaseAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #6

Describe an experiment design that would isolate an AI feature's business impact.

SPARK · AI inbox triage and prioritization for a private equity managing partner

Jorvik sits on a busy executive's inbox. It ranks what's urgent, drafts short replies to the routine stuff, and flags anything that actually needs a decision. Reinier Casswell, Managing Partner at Faircroft Capital, has run it for five months. Fiora Mears, who owns product analytics at Vespermark, the company that builds it, has to prove, before Faircroft renews, that Jorvik is actually why his inbox got faster, and not something else that happened to land in the same quarter.

The direct answer
Split the next wave of new users at random the moment they sign on. Half get the AI feature in full. Half get it running silently in the background, scoring and logging every item, showing them nothing. Measure the gap between the two groups over the same weeks, inside the same company. That gap, not a before and after chart, is what the feature actually did.
Do this, in order
  1. Split new users at random into two groups the moment they onboard, not after.Why: a random split inside the same company, in the same weeks, is the only way to know the gap in the numbers came from the feature and not from the quarter.
  2. Keep the silent group truly silent, not just switched off.Why: if it still scores and logs every item without showing anything, you can grade what the model would have flagged against what actually mattered, with no risk of the group knowing it's being watched.
  3. Measure on the thing that happens often, not the thing that happens rarely.Why: a dozen people on each side won't produce enough closed deals in one quarter to prove anything. Flagged threads happen dozens of times a week per person, and that's where the real signal lives.
  4. Watch for the two groups talking to each other.Why: at a firm small enough that everyone sits in the same partner meeting, one shared conversation about a flag can quietly erase the gap you're trying to measure.
  5. Leave the big, rare outcome out of the claim for now.Why: a dozen people over one quarter isn't enough weight to put a dollar figure on deals closed, and a number built on that little data won't survive one hard question.
  6. Never tell the silent group they're the control.Why: a person who knows they're being watched stops behaving normally, and normal behavior is the whole point of the comparison.

How to answer this, stage by stage

Nobody is grading whether you know the word "randomize." They're grading whether you can spot that a before and after chart has three plausible causes hiding under it, and whether you'd actually build the thing that tells them apart.

1
Scope it to one real account before answering in general
Say it like this
"Let's ground this in one case. Jorvik is Vespermark's inbox tool for busy executives. It ranks the mail, drafts routine replies, and flags anything that needs a real decision. Faircroft Capital has run it with thirteen partners for five months, and their finance team wants proof before they renew."
Why this works
An abstract "how do you measure AI ROI" answer stays a slogan. One real account keeps it something you can actually walk through.
2
Say the plan out loud before naming a design
Say it like this
"I'm going to run SPARK here. Say where things actually stand today, name the habit I want to build, give the one design decision, say what breaks if I'm wrong about it, then say what I'm not going to try to prove yet."
Why this works
Tells the interviewer a real method is in motion, not five thoughts arriving in whatever order they occur to you.
3
Name the real problem before naming a fix
Say it like this
"Right now all Fiora has is one inbox, before and after one date. Response time on key threads went from fourteen hours to three. That looks great, but Faircroft also hired a new junior analyst that same month, and the whole deal market sped up that quarter. Nobody can say which of those three things actually did it."
Why this works
Shows you understand the actual failure, mixed-up causes, not just that "we should measure something."
4
Give the one decision, plainly
Say it like this
"Here's what I'd do. Faircroft is onboarding twenty-four more partners this quarter for a second fund. The moment each one joins, flip a coin. Twelve get Jorvik in full, drafts and flags on screen. Twelve get Jorvik running silently, scoring every email in the background, showing them nothing. Same firm, same weeks, same market. The only real difference between the two groups is whether Jorvik's output ever reached the screen."
Why this works
Matches the direct answer, said in one breath, concrete enough for someone to argue with.
5
Name what breaks it before someone else does
Say it like this
"Three things could ruin this. One, partners talk. If someone in the full group mentions a flag at a partner meeting, someone in the silent group starts asking their assistant to watch for the same thing, and the gap disappears. Two, if the silent group ever learns it's the control, it stops behaving normally. Three, twelve partners a side, one quarter, isn't close to enough closed deals to prove anything about revenue. I'd measure flagged threads instead, which happen dozens of times a week per person, not deals, which happen two or three times a quarter."
Why this works
Shows you're pressure-testing your own design, not just proposing a holdout and hoping.
6
Say what you'd leave out, on purpose
Say it like this
"I would not put a dollar figure on closed deals from this one test. Twelve or even twenty-four partners over one quarter isn't enough weight for that claim to survive a hard question. What I'd trust, right now, is the gap in response time on flagged threads. The revenue question waits until this same design has run across a few more accounts."
Why this works
Shows judgment, not a wish list. A strong candidate says what they're not claiming yet, as clearly as what they are.
7
Close on the decision, not the story
Say it like this
"So: randomize new partners into full Jorvik versus silent Jorvik, inside the same firm, the same weeks. Measure the gap in thread response time, not deals. That gap is the real number, the one that survives the renewal conversation, because it isn't a story anymore. It's a comparison."
Why this works
Leaves the interviewer with the decision, not just the walk-through that got there.

Let's learn

What does it actually take to know that an AI feature caused a number, instead of just showing up next to it?

Jorvik reads a busy executive's inbox. It sorts what needs them from what doesn't, and it drafts short replies to the routine stuff so the executive only has to skim and send.

Hand sketched labeled parts diagram titled Reinier's inbox, before Jorvik. A center figure labeled Reinier, 7am, with four radiating labels: 200 plus overnight emails, no ranking. Skims subject lines for keywords. Answers by gut, not urgency. Quiet, important ones get buried.
Before Jorvik, Reinier's mornings had no ranking at all. Just a long list and a guess at what mattered.

Before Jorvik, Reinier's own habit was to work through his inbox top to bottom, no screen telling him what mattered. On average it took him about fourteen hours to get to a flagged deal thread buried somewhere in the pile. Five months after all thirteen original partners went live on Jorvik, that number is three hours. On paper, urgent mail gets answered eleven hours sooner.

Knowledge spark: what's a confound? Something else that changed at the same time as the thing you're trying to measure, and could just as easily explain the result. If two things move together, you can't tell which one, if either, actually caused the other.
The eleven hours were real. Proving Jorvik caused them was not.

Here's the turn. That eleven-hour drop isn't proof of anything on its own. In that same window, Faircroft hired its first dedicated deal-flow analyst, whose whole job is pre-screening the partners' mail before they see it. The wider industrial buyout market also sped up that quarter, so more people everywhere were answering faster. The eleven hours might be Jorvik. It might be the analyst. It might be the market. Right now, nobody at Vespermark or Faircroft can tell you which.

At its worst, that costs Vespermark the renewal. Not because Jorvik doesn't work. Because nobody can prove it does, and a finance team asked to keep paying without proof will happily credit the good quarter to something cheaper.

The choice Fiora would take back When Jorvik rolled out at Faircroft, all thirteen partners got it on the same day. That felt like the fast, generous thing to do. It also meant there was never a group inside Faircroft that didn't have it, so there's nothing real left to compare the fourteen-to-three number against. The fix: build the missing comparison into the next rollout wave, on purpose, instead of handing everyone the tool again.
Hand sketched decision tree titled The anchor, a coin flip at onboarding. Root box: New partner onboards. Two branches labeled random draw heads and random draw tails, leading to two outcomes: Full Jorvik, drafts and flags shown, and Silent Jorvik, scored, logged, shown to no one.
Every new partner lands in one of two arms by a coin flip, inside the same firm, the same weeks.

Faircroft is onboarding twenty-four more partners this quarter for a second, industrials-focused fund. Fiora randomizes them into two arms of twelve at the moment each one signs on. The full-Jorvik arm sees drafts and flags exactly like Reinier does now. The silent arm gets Jorvik scoring and logging every email in the background, with nothing shown, so those twelve partners triage their own inbox by hand, same as everyone did before Jorvik existed. Over the thirteen-week quarter, that's roughly 1,080 flagged threads logged in each arm, plenty of volume even though only twelve people sit on each side.

Median time to answer a flagged thread, after one randomized quarter
12 hrs 6 hrs 0 2.6 hrs Full Jorvik arm 11.4 hrs Silent Jorvik arm
Saw the flags and draftsSame firm, same weeks, saw nothing
Twelve partners a side, same account, same quarter. The only thing different between them is whether Jorvik's output ever reached the screen. That gap, 8.8 hours, is the number that survives a hard question.

The gap didn't appear on day one. It built over the quarter, as partners in the full arm learned which flags to trust, while the silent arm stayed roughly where everyone starts.

Median response time by week, both arms, thirteen-week quarter
12 hrs 6 hrs 0 wk 1 wk 3 wk 5 wk 7 wk 9 wk 11 wk 13
Full Jorvik armSilent Jorvik arm
The silent arm stays noisy but flat all quarter, business as usual. The full arm falls steadily from week one. That slow separation, not a jump, is what rules out a fluke week.

Shadow mode does one more thing worth naming directly. Jorvik's flag has its own cut off point for how sure the model needs to be before it raises its hand. Set that cut off too cautiously and it stays quiet on a real, rare problem, a false negative, exactly the kind of miss nobody notices because the tool never claimed to have caught it. Because the silent arm is scored and logged even though nobody sees it, a reviewer can go back later and check every miss against what actually happened.

Hand sketched five step flow diagram titled The day the model is wrong, and shadow mode still catches it. Steps: Email arrives, Scored low, this step outlined in red to mark the routing decision, Logged anyway, Reviewer opens it, It mattered.
Jorvik scored one thread too low to flag it. Because shadow mode logs everything anyway, that miss didn't disappear.

What Fiora would leave alone: Jorvik's routine drafts, the ones that just confirm a meeting time or forward a document. Nobody's renewal depends on whether a calendar reply saved four seconds, and this level of proof isn't worth building for it.

Hand sketched comparison diagram titled What this test can prove now, and what it cannot yet. Left panel, a gauge icon, labeled Measured now, caption response time on flagged threads. Right panel, a question mark icon, labeled Left for later, caption a dollar figure on closed deals.
The response time gap is real and defensible today. The revenue number is a guess dressed as a chart until it's been pooled across more accounts and more quarters.

The lesson: a before and after number only tells you that something changed. It never tells you which of the three things that changed that month actually caused it. If a design can't tell those apart, the number was never really a measurement. It was a guess with a date stamp on it.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why eleven hours saved was never a number Fiora could actually defend.

Fiora Mears has run product analytics at Vespermark for four years. She built the usage dashboard every account team on the sales side leans on, and she can open a report and land on the one row worth arguing about before anyone else has finished their coffee.

Reinier and the other twelve partners at Faircroft Capital onboarded Jorvik in one week last spring. The account became the one Vespermark's sales team pointed to on every call. Every quarterly review carried the same slide, fourteen hours to three, and Fiora built that slide herself. For a while she was proud of it.

There wasn't one bad meeting that changed things. It built slowly, over about six weeks, ahead of the renewal. Faircroft's CFO kept asking, in slightly different words each time, the same plain question: did Jorvik do this, or would it have happened anyway. Fiora had a good answer the first time. By the fourth time, she noticed she was saying the same sentence with less confidence, and she couldn't put a finger on the exact call where it stopped sounding like a compliment and started sounding like a challenge.

She'd been treating a chart with one line on it as an argument. It was actually a coincidence with a fourteen and a three written on it.

In the meeting where she first tried to fix it, her first idea was to stagger the next twenty-four partners instead of splitting them: four onboard this week, four more in two weeks, and so on, so she could compare the early group to the late group. She liked it for about a day. Then she said it out loud to her own team and heard the hole in it before anyone else pointed it out: if the deal market kept speeding up mid-quarter, the way it just had, the early group and the late group would differ by time as well as by the tool. She'd have built the exact same tangle again, just with a calendar instead of a chart.

Hand sketched quadrant diagram titled How much contact between the two groups. X axis how much the two groups talk, from rarely to constantly. Y axis how much the gap shrinks, from barely to a lot. Separate deal teams plotted low contact low shrink. Same partner meeting plotted high contact high shrink. Different office entirely plotted lowest contact lowest shrink.
The closer the two groups sit, the more a stray remark at a partner meeting can quietly erase the gap Fiora needed to measure.

What she built instead was the coin flip. Twelve new partners see Jorvik in full. Twelve run it silently, logged, invisible, for the same thirteen weeks. She kept the silent arm off a shared distribution list and asked the account team not to discuss flags in the joint partner meeting until the quarter closed, so nobody in either arm would know which side of the flip they'd landed on.

She didn't need a better chart. She needed a group that never saw the tool at all.

At the end of the quarter, the full arm's median time to answer a flagged thread was 2.6 hours. The silent arm's was 11.4 hours. Same firm. Same thirteen weeks. Same deal market, moving the same way for both of them. The only real difference was whether Jorvik's output ever reached the screen, and that difference was 8.8 hours, measured across roughly 1,080 flagged threads on each side, not guessed at from one inbox.

What Fiora would tell herself, back when she built that first slide: fourteen to three was never a lie. It just was never an answer either. She'd built a picture that flattered the product before she built one that could survive being questioned.

SPARK, or the difference between a chart and a comparison

Not a way to dress up a before and after number in five letters. SPARK is what forces you to name the exact design decision an experiment turns on, and to say plainly what it still can't prove.

SSituation. What's actually true today, before you touch anything.
Faircroft has one number, from one inbox, with at least three plausible causes sitting underneath it: Jorvik, the new analyst, and a faster deal market that quarter. Nobody at Vespermark can say, in one sentence, why they believe Jorvik did this.
This is why the renewal conversation felt shaky. Not because the number was bad, because nobody could defend where it came from.
PPayoff. The habit you actually want to build.
A team, and a customer, that trusts the number on the renewal slide because it came from a real comparison, not because it's the only number in the room.
This is what lets Vespermark say "here's what we measured it against" instead of "look how much faster it got."
AAnchor. The one design decision everything else hangs on.
Random split at onboarding. New partners get either full Jorvik or a silent, logged version of it, inside the same firm, the same weeks.
This is the answer to the question. Everything else in this recap defends it.
RRisk. What breaks the first time you're wrong.
Partners talk to each other, a silent group that suspects it's a control stops acting normal, and a dozen partners a side isn't enough weight to prove anything about revenue in one quarter.
Each risk gets its own guard: keep the arms apart in day to day contact, keep the silent arm genuinely invisible, and measure on threads instead of deals.
KKeep out. What you won't try to prove yet.
A dollar figure on closed deals, from this one test, at this one firm, in one quarter.
Claiming that now would be a guess wearing a chart's clothes. The response time gap is real and defensible. The revenue number isn't, not yet.
  • S: one inbox, three plausible causes, no way to tell them apart.
  • P: a renewal number both sides can actually trust.
  • A: randomize new partners at onboarding, full versus silent, same firm, same weeks.
  • R: spillover between arms, a control that notices, and too few people for a rare outcome.
  • K: no revenue claim yet. Only the response time gap, until the design has run more than once.

And if you want to be sure it really works, try it somewhere else

Same five letters, a school district instead of a private equity firm, and this time the rare outcome isn't a closed deal. It's an escalation to the board.

Loksworth triages a school superintendent's inbox: parent complaints, board member emails, state compliance notices, and it drafts short acknowledgements for the routine ones. Rowanmere Unified School District rolled it out to Superintendent Pryderi Coombeshire, and a few months later is deciding whether to buy it for the district's other four office heads too, off the strength of one number: Coombeshire's time to respond to an urgent parent complaint, which fell from two full school days to about four hours.

Hand sketched comparison diagram titled Same design, a school district instead of a fund. Left panel, a person icon labeled Faircroft Capital, caption 24 new partners split by coin flip. Right panel, a person icon labeled Rowanmere Unified, caption 4 new office heads split in pairs.
Same anchor, a much smaller population. Four office heads instead of two dozen partners, which changes what the design can safely promise.

Same problem sits underneath it. That number lands right in the middle of a school year that also saw a new assistant superintendent hired, and a new state reporting deadline that made everyone answer faster in general, not just Coombeshire.

S: one office head's before and after number, with a new hire and a new deadline both sitting under it. P: a decision the district's board can trust without re-litigating it every budget cycle. A: when Loksworth rolls out to the other four office heads next term, randomize which two go live with the tool visible and which two run it silently logged, for the same eight weeks. R: four office heads is an even smaller group than Faircroft's twelve, and they already meet every Wednesday, so contact between the two arms is a bigger risk here, not a smaller one. Guard it by running the silent pair off a separate list and keeping flags out of that Wednesday meeting until the term ends. K: don't claim Loksworth reduced complaints escalating to the board. Four office heads, one term, isn't enough to say that yet. Trust the response time gap on flagged threads only.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: randomize the next onboarding wave, full tool versus silent logged tool, same weeks, same place. Measure the thing that happens often, not the thing that happens rarely.
Cost: no budget this quarter for a full fifty-fifty split. Ship the cheap version: hold back just a handful of the next cohort instead of half. A small silent arm still beats no comparison at all.
The model got better, for real: say Jorvik's flagging accuracy doubles overnight. The design barely changes. You still don't trust that claim until the shadow log proves it against real outcomes, not the vendor's own number.

Where people run it wrong.
They compare one company to a different company instead of holding the company constant and randomizing inside it.
They let the "silent" group learn it's the control, and its behavior quietly stops being normal.
They try to prove a rare, high-stakes outcome, revenue, an escalation, with a sample sized for a common one.

How to use it live. Ask before you answer: "Is the number we have a comparison, or is it a coincidence with a date stamp on it?" That buys you the room to actually design something, instead of grabbing the nearest flattering before and after chart.

Three things worth stating directly, since this is where the real judgment sits. The alternative Fiora considered and rejected was a staggered rollout, early partners first, late partners later, so she could compare the two groups. It lost, because if the deal market keeps moving mid-quarter, an early group and a late group differ by time as well as by the tool, which reintroduces the exact tangle the design was supposed to remove. The AI-specific failure worth naming is a model that's tuned to stay quiet unless it's sure, which means its cut off point can let a real, rare problem through unflagged without anyone knowing it happened. The guardrail is shadow mode's own silent log: because it scores and records every item in the holdout arm too, a neutral reviewer can check months later whether a miss like that actually mattered, and whether the cut off needs to move. And the trade off is real and accepted on purpose: running Jorvik silently for twelve partners still costs Vespermark the same inference spend as running it for the twelve who can see it, with nothing to show a customer for that spend all quarter, and those twelve partners get none of Jorvik's time savings for thirteen weeks. That cost buys a number worth defending instead of one worth apologizing for.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
SPARK: design against the failure before you build. Used here to design an experiment that survives being questioned, not just to design the tool itself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Fiora Mears, who has run product analytics at Vespermark for four years, and built the usage dashboard every account team on the sales side relies on.
3 · THE PAYOFF
What's the P step here, the habit Fiora actually wants?
Tap to flip
ANSWER
A team, and a customer, that trust the ROI number on the renewal slide because it came from a real comparison, not because it's the only number anyone has.
4 · THE ANCHOR
What's the two-setting switch in Fiora's design?
Tap to flip
ANSWER
Full Jorvik, drafts and flags shown, versus silent Jorvik, scored and logged but shown to no one. Every new partner lands in one or the other by a coin flip.
5 · THE OLD DECISION
What decision would Fiora take back?
Tap to flip
ANSWER
Rolling Jorvik out to all thirteen original partners on the same day, leaving no group inside Faircroft that never had it, so there was nothing real to compare the fourteen-to-three number against.
6 · THE NUMBER
Fill in the blank: full Jorvik arm median response time was ___ hours. Silent arm was ___ hours.
Tap to flip
ANSWER
2.6 hours for the full arm, 11.4 hours for the silent arm, across about 1,080 flagged threads on each side over the thirteen-week quarter.
7 · THE REPLAY
Same renewal conversation, new design, what changes?
Tap to flip
ANSWER
Fiora shows the CFO a real gap, 2.6 hours versus 11.4, measured inside one firm in the same weeks. Faircroft renews without the finance team re-litigating whether the market caused it.
8 · CROSS-PRODUCT TRANSFER
The "try it elsewhere" section answers this same question for a different product. Which product, and what's the design there?
Tap to flip
ANSWER
Loksworth, used by Rowanmere Unified School District's superintendent Pryderi Coombeshire. The district randomizes which two of four new office heads see the tool and which two run it silently, for the same eight weeks.

Check yourself Score: 0 / 0

Multiple choice
1. Which design actually isolates Jorvik's own effect from everything else that happened that quarter?
  • A. Compare Faircroft's response time this quarter to last quarter.
  • B. Compare Faircroft's partners to a similar firm that doesn't use Jorvik.
  • C. Randomly split new partners into full Jorvik and a silent, logged version of it, in the same weeks.
  • D. Ask partners in a survey whether Jorvik made them faster.
Show hint
Think about what stays exactly the same between the two groups in each option, and what doesn't.
Show answer
C. Only a random split inside the same firm, in the same weeks, holds the market, the season, and the new hire constant for both groups.
Multiple choice
2. What old decision does this answer take back?
  • A. Choosing shadow mode instead of a full holdout with no Jorvik at all.
  • B. Giving Jorvik to all thirteen original partners on the same day, with no group inside the firm that never had it.
  • C. Hiring a dedicated deal-flow analyst.
  • D. Pricing Jorvik per partner instead of per firm.
Show hint
Look at the key point box titled "The choice Fiora would take back," right after the turn.
Show answer
B. Rolling out to everyone at once felt generous, but it meant there was never a comparison group inside Faircroft to check the number against.
True or false
3. True or false: the eleven-hour drop in Reinier's response time is solid proof that Jorvik caused it.
  • True
  • False
Show hint
Check what else changed at Faircroft in that same window besides Jorvik.
Show answer
False. Faircroft also hired a new deal-flow analyst and the industrial buyout market sped up that same quarter, so a single before and after number can't separate the three causes.
Fill in the blank
4. In the randomized quarter, the full-Jorvik arm's median response time on a flagged thread was ___ hours. The silent arm's was ___ hours.
Show hint
It's stated right under the first bar chart in Let's learn.
Show answer
2.6 and 11.4. A gap of 8.8 hours, measured across roughly 1,080 flagged threads on each side, not guessed at from one inbox.
Short answer, where it wouldn't matter
5. Name one part of Jorvik where this level of proof would not be worth the trouble.
Show hint
Look at "What Fiora would leave alone," right before the lesson.
Show answer
Model answer: Jorvik's routine drafts, the ones that just confirm a meeting time or forward a document. Nobody's renewal depends on whether a calendar reply saved four seconds.
Short answer, apply it yourself
6. Think of a product you use yourself that changed recently. What's one way you could check whether the product caused the change, rather than something else happening at the same time?
Show hint
Think about whether you could hold something constant and only change the one thing you're curious about.
Show answer
Model answer: A fitness app added AI workout suggestions and your weekly step count went up. Before crediting the app, check whether the season changed too, better weather alone could explain more walking, and you'd want a friend on the old version of the app, in the same city, in the same weeks, to compare against.
Before you close the answer
Why this works
Tests whether you can tell a real controlled comparison from a flattering before and after chart, and whether you know an AI-specific way to check the model's own judgment, scoring a holdout silently, without ever changing what it saw.
Follow-up traps
"Isn't a coin flip inside one small firm still too small to trust?" Response: measure on threads, not partners. About 1,080 threads a side over the quarter carries real weight even though the partner count is small, and the revenue question stays out of scope until it's pooled across accounts.

"What if the silent group figures out they're being logged?" Response: that's exactly why shadow mode has to be invisible in the product itself, no banner, no different screen, and nobody told which arm they landed in.
If pressed
Each partner generates about 30 flagged threads a month, so twelve partners a side works out to roughly 360 threads a month per arm, more than enough to detect even a modest shift with real confidence. The same design could never do that if it tried to power itself off closed deals instead, which might only happen two or three times a quarter per partner.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more