CaseIntermediateQuality, Cost & Token Economics / Measuring ROI and business impact / #18

How would you present ROI when the honest answer is that it is too early to tell?

SPARK · real-time odds-setting and risk management for a sportsbook

Vigwatch is Sableridge Sportsbook's engine for pricing in-play bets and managing how much risk the book carries, updated live on every market. Tamzin Barrowclough owns its product story. Eleven weeks ago Vigwatch took on a new job: pricing thin, rarely-watched markets like darts and county cricket on its own. Benedikt Pemberton-Hale, the Chief Risk Officer, wants the ROI for the board deck. Tamzin doesn't have one yet, and has to say so without it sounding like the project failed.

The direct answer
Split the report by how much graded data actually backs each number. Marquee sports already have enough settled markets for a real Vigwatch ROI, so report that number now. For the new micro-market pricing, say plainly it is too early to tell, in the same breath as two things: the leading indicators already tracked, and the exact point, about 32,000 graded events, roughly week 50, when a real hold-percentage read becomes possible. Never blend the two into one figure, and never let a single loud week decide the story before the data can.
Do this, in order
  1. Split the ROI report by which number actually has enough graded volume behind it, marquee real now, micro-markets not yet.Why: mixing them lets a solid marquee win quietly vouch for a micro-market number nobody has earned.
  2. Hand over the leading indicators already known, in the same breath as the "not yet."Why: override rate, repricing speed, and pricing sharpness are real signals right now; staying silent about them is what makes "too early" sound like an excuse.
  3. Name the exact threshold and the week it resolves.Why: about 32,000 graded events, around week 50, gives leadership a date to hold you to instead of a mood to worry about.
  4. Never blend marquee and micro-market numbers into one figure.Why: a real shortcut that got rejected, because it would let this quarter's marquee win mask a micro-market loss nobody's found yet.
  5. Treat one loud week, good or bad, as noise until the threshold clears.Why: the week 6 stale-line loss swung the number more than nine points in seven days; that's what thin samples do, not proof of anything.
  6. Fix the actual gap a near miss reveals, not just the metric that flagged it.Why: watching the number harder wouldn't stop the next stale line; a kill switch that pulls a market during a surprise event does.

How to answer this, stage by stage

Nobody is grading whether you can say "we'd be transparent." They're grading whether you can hand someone a real page instead of a shrug, and whether you know which numbers a small sample can and can't support.

1
Anchor it to one product, one number, one real person
Say it like this
"Let's ground this in one case. Vigwatch is Sableridge Sportsbook's engine for setting in-play odds and managing risk exposure. Tamzin Barrowclough owns its product story. Benedikt Pemberton-Hale, the Chief Risk Officer, is the one who needs an honest ROI answer for the board."
Why this works
An abstract "how do you handle uncertainty" answer stays a platitude. One real report and one real person keeps every claim checkable.
2
Say the structure out loud
Say it like this
"I'll run this as SPARK. Ground it in what happens today without a good answer, name the habit I want leadership to build instead, give the actual thing I'd hand them, say what breaks if I'm wrong, then say what I'd deliberately leave out."
Why this works
Two seconds of structure tells the interviewer a plan is running, not four thoughts arriving as they occur.
3
Name what "too early" gets replaced with today
Say it like this
"Right now, when a number genuinely isn't ready, it gets replaced with silence, or a soft 'still ramping.' Both read the same way to a risk chief prepping a board deck: something's being hidden. He doesn't wait for the real number. He assumes the worst one and starts planning around it."
Why this works
This is the reframe the whole answer turns on. Skip it and the rest sounds like "have patience," which no boardroom accepts.
4
Give the anchor, the actual thing you'd hand over
Say it like this
"Here's what I'd actually hand him. Two parts, one page. Part one, what we know now: override rate's down from 38 percent to 6 in eleven weeks, and the model's opening price already lands within 1.9 percent of where each market closes. Part two, what we can't know yet: hold percentage, because eleven weeks of a thin market isn't enough data, and here's the number, 32,000 graded events, and the week, around week 50, when it will be."
Why this works
This is the direct answer, and it's concrete enough that Benedikt could hold Tamzin to it later.
5
Show the risk, and prove the anchor survives it
Say it like this
"Here's what almost broke it. Week 6, a rain delay in a county cricket match, our price didn't update fast enough, and a sharp bettor hit the stale line for $86,400 in one afternoon. Old habit, that week gets reported alone and reads as proof the whole thing failed. With the report already saying 'expect noise this size until week 50,' it's one point on a range we'd already flagged, not a crisis."
Why this works
A risk you only describe is a warning. A risk you show surviving is a design decision.
6
Say what's deliberately left out
Say it like this
"Three things I won't do to make this easier. I won't blend the micro-market number into the marquee number to get one clean board slide. I won't borrow marquee's accuracy and assume it carries over. And I won't quietly tighten Vigwatch's risk limits on micro-markets just to smooth the hold rate before week 50."
Why this works
Naming the shortcuts you're refusing is what makes "too early to tell" sound like discipline, not an excuse.
7
Close on the decision, in one breath
Say it like this
"So: split the report by what the data can actually support, hand over the leading indicators you do have, name the threshold and the date, and never let one week, good or bad, stand in for the real answer before it's earned."
Why this works
Restates the direct answer plainly, so the interviewer leaves with the decision, not just the story.

Let's learn

Here's what happens when eleven weeks of promising numbers and one very bad week land on the same desk, before anyone can say which one is real.

Vigwatch is Sableridge Sportsbook's engine for pricing in-play bets and managing how much risk the book is carrying, on every sport, updated live as each match moves.

Before Vigwatch's newest job, micro-markets, darts, table tennis, county cricket, the bottom tier of English football, got priced by hand. A trader set a rough multiplier off the pre-match line and checked it once a day, if that, because the volume never justified more attention than that.

Hand sketched comparison diagram titled Two markets, two sample sizes. Left panel, a document icon labeled Marquee sports, caption 380,000 graded events over 14 months. Right panel, a document icon labeled Micro-markets, caption 7,040 graded events over 11 weeks.
Marquee sports have years of settled markets behind them. Micro-markets have eleven weeks. That gap is the whole reason this question is hard.

With Vigwatch's new job, every micro-market reprices itself within a couple of seconds of any real move, automatically, all day. Traders stopped babysitting markets nobody was watching closely enough. Their override rate on these markets fell from 38 percent to 6 in eleven weeks.

Knowledge spark: what is closing line value? The closing line is where a market settles right before an event starts, shaped by thousands of sharp bettors and other books all pricing the same thing. If a book's own price beats where the line closes, more often than not, that's real skill, not luck. That's closing line value, or CLV, and it's the number marquee sports already have enough history to prove.

On marquee sports, football, basketball, the tennis majors, Vigwatch has been live for 14 months and graded about 380,000 settled markets. CLV moved from -0.3 percent before Vigwatch to +1.1 percent after. Hold percentage climbed from 6.8 to 7.4. Those numbers are solid enough to defend in front of anyone, and Tamzin reports them every quarter without a second thought.

Micro-markets are eleven weeks old. Here's the turn. The headline number, hold percentage up from a 4.1 percent baseline to about 4.4, isn't really the story yet, and neither is the one bad week that dragged it to negative 3.1 for seven days. The real risk sits three weeks later, in a boardroom, when Benedikt asks Tamzin for the ROI on micro-markets and gets either silence or a shaky guess instead of a plan for finding out.

Micro-market hold percentage, week by week
7% 0% -4% legacy baseline, 4.1% W6 W1 W2 W3 W4 W5 W7 W8 W9 W10 W11 -3.1%, $86,400 lost
Weekly micro-market holdWeek 6, the stale-line loss
Eleven weeks average out to about 4.4 percent, a little above the 4.1 percent baseline. One week alone swung more than nine points. That is the size of noise a real signal still has to climb out of.
The bad week wasn't proof the model failed. The silence about it was the actual risk.

What it costs at its worst: silence reads as failure. If Tamzin doesn't answer, Benedikt fills the gap with his own worst guess, and proposes handing micro-market pricing back to manual trader control at the very meeting the number was supposed to inform, before the real answer has had a single extra week to arrive. That would also end the only thing capable of ever producing the honest number: the data still being collected.

The decision that mattered Sableridge built its reporting habit around one rule: every AI capability gets a quarterly ROI number starting the first quarter after launch, no exceptions. That made sense for marquee sports, backed by years of trader-priced history where a month of live data was already plenty. It broke down the day a capability launched into a market with no deep history behind it and sixty settled events a day, not sixty thousand.

What I would leave alone: marquee sports' pricing report doesn't need any of this. Football, basketball, and the tennis majors already clear 380,000 graded markets over 14 months, more than enough for a real number every quarter. Building a leading-indicator dashboard for markets that already have a trustworthy answer would just add paperwork.

The lesson: an honest "not yet" is not a weaker answer than a number. It's a different kind of answer, and it only holds up if it always ships with what's already known and a real date attached. Without those two things, nobody can tell it apart from stalling.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a stale price, not a bad model, is what almost cost Tamzin the board's trust.

For fourteen months, the best part of Tamzin Barrowclough's job was a number that always arrived early.

Tamzin has led product for Vigwatch since before it went live, back when the whole idea was still a slide deck: reprice every in-play market in real time, on every sport Sableridge carries, and manage how much risk the book takes on while it does. Marquee sports went first, sports with years of graded history behind them and enough weekly volume that a real answer showed up within a month of any change. Every quarter, Tamzin walked into Benedikt's office with a number, and the number held up.

Eleven weeks ago, Vigwatch took on something harder: pricing the markets nobody had bothered to watch closely. A trader used to set these by hand, a rough multiplier off the pre-match line, glanced at once a day if that. It made sense. Nobody was going to give a real trader's full attention to a market that settled sixty times a week.

Hand sketched flow diagram titled How silence becomes an assumption. Four connected boxes reading left to right: Asks for the number, No update no date this box emphasized in red, Fills the gap himself, Hears it is failing.
This is the chain Tamzin was walking toward, without meaning to, every week she didn't say anything at all.

The first weeks looked almost boring, the good kind of boring. Overrides fell from 38 percent to under 10 within a month. The model's opening price crept closer and closer to where each market actually closed. Tamzin watched the hold rate too, out of habit more than anything: 4.0 percent, 5.3, 4.7, a touch above the old 4.1 baseline, nothing dramatic yet.

Then, on a Thursday in week six, rain stopped a county cricket match for forty minutes.

Nothing about that should have mattered. Rain delays happen. But Vigwatch's live price didn't know the delay had changed the shape of the match, not for twenty-six seconds, well past its usual two, on a sport where Vigwatch had barely a season of history to lean on. A professional bettor, watching for exactly that kind of gap, hit the stale line across three connected markets before it moved. Sableridge lost $86,400 that afternoon. Chibundu Kowalik, who runs the numbers behind the micro-market rollout, pulled the trade log before lunch and found it in minutes. There was nothing subtle about it.

Hand sketched decision tree titled Week 6 goes wrong, does the report survive it. Root box reads A stale line loses $86,400 in one week, branching into two outcomes: no report exists yet leading to reads as proof it failed, and report already flags this as noise leading to reads as expected nobody panics.
The same $86,400 loss reads completely differently depending on whether a report existed the week before it happened.

Here's what almost got lost in that one number. Week six's hold rate cratered to negative 3.1 percent, for that week alone. On its own, dropped into a slide with no context, it reads as proof the whole rollout had failed. It wasn't proof of that. It was one professional bettor, one delay, one gap in the model's reaction time, on a sport where eleven weeks of data barely covers a season.

We did not lose $86,400 to a bad model. We lost it to a price that hadn't heard the weather had changed.

Three weeks later, prepping the quarterly board deck, Benedikt asked Tamzin the question he'd asked every quarter for fourteen months: "What's the ROI?" This time she didn't have one. Eleven weeks of a genuinely thin, genuinely volatile market isn't fourteen months of football. She could say the hold rate looked a little better, 4.4 percent against the 4.1 baseline. She could not honestly say that gap was real and not just the same kind of noise that had swung the number nine points in a single week already.

Hand sketched labeled parts diagram titled The two-part ROI report. A central document icon labeled Micro-market update, with five labeled callouts around it: What we know now, Override rate and latency, What we cant know yet, Threshold 32,000 events, and Date it resolves.
Not a bigger number. A different shape of page: what's already known, sitting right next to what isn't and exactly when it will be.

What Tamzin didn't do was stay quiet and hope the next few weeks looked better. She built the report Benedikt actually needed: what Vigwatch already knows for certain, override rate down to 6 percent, opening prices landing within 1.9 percent of the close, and what it still can't know, the real hold-percentage lift, alongside the number that would make it knowable. Chibundu ran the arithmetic: with the variance micro-markets carry, you need about 32,000 graded events before a real signal separates from noise this size. At 640 graded a week, that's roughly week 50. Eleven weeks in, thirty-nine to go.

Hand sketched horizontal timeline titled The eleven weeks before the board meeting. Five milestones: Vigwatch launches, marquee sports 14 months ago. Micro markets go live, 11 weeks ago. Rain delay stale line, week 6, $86,400 lost, this milestone emphasized in red. Baptiste asks for the number, week 11 board prep. Threshold reached, around week 50.
Fourteen months of marquee history sits behind this. Only eleven weeks sit behind micro-markets, and the board meeting landed right in the middle of them.

The decision Tamzin would take back traced to how Sableridge built its own reporting habit in the first place: every AI capability gets a quarterly ROI number starting the first quarter after launch, no exceptions, no waiting for enough data to say something true. That rule made sense when every capability launching was another marquee sport, backed by years of trader-priced history that made a month of live data plenty. It stopped making sense the day a capability launched into a market with no deep history behind it and sixty settled events a day, not sixty thousand.

Run the same board meeting again, with that rule fixed. Benedikt still asks for the ROI. Tamzin still doesn't have one for micro-markets. But now she has a page: what's known, what isn't, and a date, week 50, that Benedikt can hold her to. He doesn't propose pulling micro-market pricing back to manual trading. He circles the date on his own calendar and asks what happens if week 50 comes and the number still isn't good.

One version of that meeting reads a bad week as a verdict. The other reads it as one point on a range that was labeled ahead of time.

What Tamzin would tell herself, back when Sableridge first wrote the quarterly-ROI-no-exceptions rule into how it reports on Vigwatch: a rule built for a market with fourteen months of trader history behind it was never going to survive being handed to a market with eleven weeks of anyone's history at all, and nobody had written down what the rule should do when that happened.

SPARK: the report that turns "not yet" into a real answer

Not a way to dress up "we don't know" so it sounds better. SPARK forces the actual page you'd hand over, and makes you prove it still holds up the week a real number goes wrong.

SSituation. What happens today, without this design.
Right now, when a real number genuinely isn't ready, it gets replaced by silence or a soft "still ramping." Leadership doesn't wait for the real answer. It fills the gap with its own worst guess and starts planning around that instead.
Say what actually happens today before naming the fix, or the anchor sounds like a nice-to-have instead of a repair.
PPayoff. The habit worth building.
Not "better communication." Leadership trusts an honest "not yet" because it always arrives packaged with what's already known and a real date for the rest, so nobody ever has to fill the gap themselves again.
Name the habit, not the mood. A habit is something you can check for on the next report; a mood isn't.
AAnchor. The actual thing you'd hand over.
A two-part report. What we know now: override rate, repricing speed, and how close the model's opening price lands to the eventual closing line, all real signals that don't need a big settled sample. What we can't know yet: hold percentage, with the exact threshold, about 32,000 graded events, and the week it resolves, around week 50.
This is the concrete answer to the question. Everything else in the framework exists to protect it.
RRisk. What breaks the first time it's tested.
The first loud, bad week arrives before the threshold clears, and the temptation is to read it as the real answer. The report has to survive that moment already labeled as expected noise, or the first bad week undoes the whole design.
Design the anchor against this specific risk, not against a generic "someone might doubt me."
KKeep out. What doesn't ship on day one.
No blended marquee-plus-micro ROI figure. No borrowing marquee's accuracy for micro-markets. No quietly tightening Vigwatch's risk limits on micro-markets just to smooth the number before week 50.
Naming the shortcuts you're refusing is what makes "too early to tell" sound like discipline instead of stalling.
Hand sketched numbered icon list titled What stays out of the report. Three rows: a scale icon, no blended marquee plus micro ROI number. A gauge icon, no borrowing marquee accuracy for micro markets. A funnel icon, no quietly tightening limits to smooth the number.
Three things that would each, quietly, make this quarter's slide look cleaner and next quarter's honesty harder.
Micro-market progress toward a trustworthy read
32,000 16,000 0 7,040 Graded so far (week 11) 32,000 Threshold needed (around week 50)
Graded events so farThreshold for a trustworthy read
At 640 graded events a week, week 11 has cleared about a fifth of the distance to a number anyone should trust.
Why the anchor survives the risk Check it against the near miss. Does the two-part report still hold up the week a stale line loses $86,400? Yes, because the range was already labeled "expect noise this size until week 50" before the loss happened, not after. Does it avoid quietly hiding the loss? Yes, because K refuses to blend or borrow numbers that would smooth it out of sight.

And if you want to be sure it really works, try it somewhere else

Same five letters, a drought insurer instead of a sportsbook, and this time the rare event isn't a rain delay. It's the severe drought year the model hasn't seen yet.

Netherwood Underwriters sells parametric drought insurance to smallholder farmers through Rainmark, a satellite index that reads vegetation and rainfall signals and decides when a payout triggers, without anyone filing a claim or an adjuster visiting a field. Otoniel Ochieng, who owns Rainmark's model performance, onboarded a new farming region eight months ago. The region has been through exactly one growing season since. A reinsurance partner wants to know whether Rainmark's new drought trigger, tuned for that region's soil and rainfall pattern, actually cuts basis risk, the gap between what farmers really lost and what the index paid, better than the old, blunter regional index it replaced.

Hand sketched comparison diagram titled Same wait, a different rare event. Left panel, a gauge icon labeled Sableridge Sportsbook, caption a stale odds line in a rain delay. Right panel, a document icon labeled Netherwood Underwriters, caption a drought trigger with one growing season of data.
Same wait for a real number. A sportsbook waits on graded bets; an insurer waits on a growing season severe enough to actually test the trigger.
The decision Otoniel would take back Netherwood's own reporting rule called for a basis-risk number by the first renewal after any regional model update, no matter how many seasons had actually passed. That made sense in regions with a decade of matched claims and satellite data behind them. It didn't survive being applied to a region with one season on record, since a single mild season can't say whether a trigger tuned for that region's rare, severe droughts is actually well calibrated.

Same rank, different lever, mapped straight onto SPARK: the situation is a reinsurance partner asking for a basis-risk number the same way Benedikt asked for a hold-percentage number, before enough seasons exist to answer honestly. The payoff is the same trust: the partner accepts "not yet" because it comes with real numbers already in hand, how closely Rainmark's rainfall readings track ground-truth gauges this season, how fast payouts reached farmers. The anchor is the same shape: what's known now, and the number of full growing seasons, probably three, spanning at least one genuinely severe drought year, before the basis-risk claim is trustworthy. The risk is the same shape too: one unusually mild season could get mistaken for proof the trigger works. And what stays out is the same discipline: no borrowing the old, decade-proven regional index's track record to vouch for the new one early.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: report what you can measure now, name the number of full cycles still needed, and never let one loud result, good or bad, stand in for the real answer.
Cost: no budget this quarter to wait three full seasons before saying anything. Ship the honest interim version: leading indicators only, with the same "not proof yet" label attached, updated every season instead of staying silent until the end.
The model got better, for real: say Rainmark's satellite resolution doubles overnight. The two-part shape doesn't change. A sharper sensor narrows how many seasons are needed, it doesn't excuse skipping straight to a claim before a real severe-drought year has actually been observed.

Where people run it wrong.
They wait for total certainty before saying anything at all, and leadership reads the silence as a worse verdict than the honest one would have been.
They blend a new, thin result into an old, proven number to get one clean line for a deck.
They let one loud period, good or bad, decide the story before enough of them have actually happened.

How to use it live. Ask the coverage question before naming a number: "How much of the real variance has this actually been tested against yet, and what's the smallest bad case it hasn't seen?" That question alone usually separates a real answer from a guess dressed up as one.

Three things worth stating directly, since this is where the real judgment sits. The alternative Tamzin's team considered, and rejected, was blending marquee and micro-market performance into one "Vigwatch Impact" figure for the board deck, the same shortcut Netherwood's team was tempted by too. It lost, because a statistically solid marquee win would quietly vouch for a micro-market number nobody had actually earned yet, and if micro-markets turn out net-negative once real data arrives, that blended number will have already told the board it was a win. The AI-specific failure worth naming is distribution shift under a rare event: Vigwatch's live price kept pricing off the pre-delay distribution for twenty-six seconds because nothing had told it the match had genuinely changed shape, and a model that doesn't know to say "I'm less sure right now" keeps serving a confident, stale number. The guardrail is a stale-price kill switch: any flagged surprise, a rain delay, a red card, an injury, pulls that specific market off the board for a few seconds until Vigwatch recalculates, instead of continuing to quote the old price. And the trade-off is real: pulling a market costs Sableridge a few seconds of betting handle on it, every time the switch fires, even on the vast majority of surprises that turn out harmless. That's accepted on purpose, because the alternative is finding out about the next stale line from an $86,400 trade log instead of a switch that fired quietly and cost nothing but a few seconds of handle.

Hand sketched full page metaphor scene titled Silence lets someone else write the ending. Left panel, a question mark icon labeled SILENCE, caption leadership fills the gap themselves. Right panel, a gauge icon labeled A DATE, caption leadership waits for the real number.
The whole answer to this question, in one picture. Silence hands the ending to someone else. A date keeps it in your hands.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits this question, and what's its job?
Tap to flip
ANSWER
SPARK: design the actual thing you'd hand someone against the risk it has to survive, before you build it. A forward-running method: situation, payoff, anchor, risk, keep out.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tamzin Barrowclough, who has led product for Vigwatch at Sableridge Sportsbook since before it launched, and owns the honest answer when a number genuinely isn't ready.
3 · THE SITUATION
What happens today, without this design, when a real number isn't ready?
Tap to flip
ANSWER
It gets replaced by silence or a soft "still ramping," and leadership fills the gap with its own worst guess instead of waiting for the real one.
4 · THE ANCHOR
What's the concrete anchor, the actual thing Tamzin hands over?
Tap to flip
ANSWER
A two-part report: what's known now (override rate, repricing speed, opening-to-close price gap), and what isn't yet (hold percentage), with the exact threshold and the week it resolves.
5 · THE REJECTED ALTERNATIVE
What alternative did Tamzin's team reject, and why?
Tap to flip
ANSWER
Blending marquee and micro-market performance into one ROI figure. Rejected because a solid marquee win would quietly vouch for a micro-market number nobody had actually earned yet.
6 · THE NUMBER
Micro-markets had graded ___ events by week 11. A trustworthy hold-percentage read needs about ___.
Tap to flip
ANSWER
7,040 graded by week 11. About 32,000 needed, around week 50 at the current pace of 640 a week.
7 · THE REPLAY
Same board meeting, new report, what changes?
Tap to flip
ANSWER
Benedikt still gets no ROI number for micro-markets, but he gets a page with what's known, what isn't, and week 50 circled as the date to hold Tamzin to, instead of proposing to pull the capability.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this question again for a different product. Which one, and what's the equivalent "not yet"?
Tap to flip
ANSWER
Rainmark, Netherwood Underwriters' satellite drought-index model. The equivalent "not yet" is whether a new region's payout trigger cuts basis risk, unprovable until a real severe-drought season has been observed.

Check yourself Score: 0 / 0

True or false
1. True or false: the rise in micro-market hold percentage over eleven weeks, from a 4.1 percent baseline to about 4.4, is enough evidence that Vigwatch's micro-market pricing is working.
  • True
  • False
Show hint
Compare that gap to the size of the week-to-week swings already seen, including week 6.
Show answer
False. The legacy baseline swings about ±1.8 points week to week on its own, and week 6 alone moved the number more than nine points in a single week. A 0.3-point average lift after eleven weeks sits well inside that noise.
Fill in the blank
2. Micro-markets graded about ___ settled events a week. By week 11 that's about ___ graded events, against a threshold of about ___ needed for a trustworthy read.
Show hint
Check "Let's learn" and the anchor step in the walkthrough.
Show answer
640 a week; 7,040 graded by week 11; about 32,000 needed. That gap, not the model's accuracy, is the real reason the honest answer is "too early to tell."
Multiple choice
3. Which of these belongs in the "what we know now" half of Tamzin's report, safe to report even on eleven weeks of thin data?
  • A. The real hold-percentage lift.
  • B. The trader override rate.
  • C. Whether the model beats the old trader heuristic long-term.
  • D. The total board-approved ROI figure.
Show hint
Think about which metrics need a big settled sample and which don't.
Show answer
B. Override rate is a process metric measured directly on every market, without waiting for anything to settle, so it's trustworthy at low volume. Hold percentage is an outcome metric buried in variance until enough events settle.
Short answer, name the rejected alternative
4. What alternative did Tamzin's team reject instead of building the two-part report, and why did it lose?
Show hint
Look at the "three things worth stating directly" paragraph near the end of Section 4.
Show answer
Model answer: Blending marquee and micro-market performance into one "Vigwatch Impact" ROI figure for the board deck. It lost because a solid marquee win would quietly vouch for a micro-market number nobody had actually earned yet, and if micro-markets turned out net-negative later, the blended figure would already have told the board it was a win.
Short answer, apply it yourself
5. Think of a project at your own job where someone asked for a result before there was enough data to give one. What's one thing you already knew that you could have reported instead of staying quiet?
Show hint
Look for a process signal you had on day one, separate from the outcome nobody could measure yet.
Show answer
Model answer: A support team rolled out an AI ticket-triage tool and got asked for a resolution-time improvement after three weeks. There wasn't enough closed-ticket volume yet, but the team already knew the tool's routing-accuracy rate and how often agents overrode its category tag, both real signals that didn't need to wait.
Short answer, work the number
6. If Vigwatch's micro-markets only graded 400 events a week instead of 640, roughly how many more weeks would it take to reach the 32,000-event threshold, counting from week 11?
Show hint
Subtract what's already graded, then divide by the new weekly rate.
Show answer
About 62 more weeks, landing around week 73. At 400 a week, the 24,960 events still needed (32,000 minus 7,040) take about 62 weeks, not 39, pushing the resolving date from week 50 out to about week 73.
Before you close the answer
Why this works
Tests whether you'll admit real uncertainty out loud in front of leadership without it reading as failure, and whether you know which numbers are trustworthy at a low sample size and which ones aren't.
Follow-up traps
"Why not just wait until week 50 and say nothing until then?" Response: silence for 39 more weeks reads exactly like a failing project. The leading indicators are real and defensible right now, and reporting them is what keeps "not yet" from sounding like stalling.

"Isn't $86,400 in one week proof enough that something's wrong?" Response: one week's loss in a market with barely a season of history is the same size of swing the legacy baseline already produced on its own. It's a reason to fix the stale-price gap that caused it, not proof the whole model failed.
If pressed
The 32,000-event threshold isn't a round number picked for comfort. It's what Chibundu's power analysis says is needed at 90 percent confidence, once you account for how fat-tailed the loss distribution gets from a single correlated trade like the cricket bet. A plain binomial assumption would have suggested closer to 9,000 events were enough, and that number would have been wrong.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more