How would you present ROI when the honest answer is that it is too early to tell?
Vigwatch is Sableridge Sportsbook's engine for pricing in-play bets and managing how much risk the book carries, updated live on every market. Tamzin Barrowclough owns its product story. Eleven weeks ago Vigwatch took on a new job: pricing thin, rarely-watched markets like darts and county cricket on its own. Benedikt Pemberton-Hale, the Chief Risk Officer, wants the ROI for the board deck. Tamzin doesn't have one yet, and has to say so without it sounding like the project failed.
- Split the ROI report by which number actually has enough graded volume behind it, marquee real now, micro-markets not yet.Why: mixing them lets a solid marquee win quietly vouch for a micro-market number nobody has earned.
- Hand over the leading indicators already known, in the same breath as the "not yet."Why: override rate, repricing speed, and pricing sharpness are real signals right now; staying silent about them is what makes "too early" sound like an excuse.
- Name the exact threshold and the week it resolves.Why: about 32,000 graded events, around week 50, gives leadership a date to hold you to instead of a mood to worry about.
- Never blend marquee and micro-market numbers into one figure.Why: a real shortcut that got rejected, because it would let this quarter's marquee win mask a micro-market loss nobody's found yet.
- Treat one loud week, good or bad, as noise until the threshold clears.Why: the week 6 stale-line loss swung the number more than nine points in seven days; that's what thin samples do, not proof of anything.
- Fix the actual gap a near miss reveals, not just the metric that flagged it.Why: watching the number harder wouldn't stop the next stale line; a kill switch that pulls a market during a surprise event does.
How to answer this, stage by stage
Nobody is grading whether you can say "we'd be transparent." They're grading whether you can hand someone a real page instead of a shrug, and whether you know which numbers a small sample can and can't support.
Let's learn
Here's what happens when eleven weeks of promising numbers and one very bad week land on the same desk, before anyone can say which one is real.
Vigwatch is Sableridge Sportsbook's engine for pricing in-play bets and managing how much risk the book is carrying, on every sport, updated live as each match moves.
Before Vigwatch's newest job, micro-markets, darts, table tennis, county cricket, the bottom tier of English football, got priced by hand. A trader set a rough multiplier off the pre-match line and checked it once a day, if that, because the volume never justified more attention than that.
With Vigwatch's new job, every micro-market reprices itself within a couple of seconds of any real move, automatically, all day. Traders stopped babysitting markets nobody was watching closely enough. Their override rate on these markets fell from 38 percent to 6 in eleven weeks.
On marquee sports, football, basketball, the tennis majors, Vigwatch has been live for 14 months and graded about 380,000 settled markets. CLV moved from -0.3 percent before Vigwatch to +1.1 percent after. Hold percentage climbed from 6.8 to 7.4. Those numbers are solid enough to defend in front of anyone, and Tamzin reports them every quarter without a second thought.
Micro-markets are eleven weeks old. Here's the turn. The headline number, hold percentage up from a 4.1 percent baseline to about 4.4, isn't really the story yet, and neither is the one bad week that dragged it to negative 3.1 for seven days. The real risk sits three weeks later, in a boardroom, when Benedikt asks Tamzin for the ROI on micro-markets and gets either silence or a shaky guess instead of a plan for finding out.
What it costs at its worst: silence reads as failure. If Tamzin doesn't answer, Benedikt fills the gap with his own worst guess, and proposes handing micro-market pricing back to manual trader control at the very meeting the number was supposed to inform, before the real answer has had a single extra week to arrive. That would also end the only thing capable of ever producing the honest number: the data still being collected.
What I would leave alone: marquee sports' pricing report doesn't need any of this. Football, basketball, and the tennis majors already clear 380,000 graded markets over 14 months, more than enough for a real number every quarter. Building a leading-indicator dashboard for markets that already have a trustworthy answer would just add paperwork.
The lesson: an honest "not yet" is not a weaker answer than a number. It's a different kind of answer, and it only holds up if it always ships with what's already known and a real date attached. Without those two things, nobody can tell it apart from stalling.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a stale price, not a bad model, is what almost cost Tamzin the board's trust.
For fourteen months, the best part of Tamzin Barrowclough's job was a number that always arrived early.
Tamzin has led product for Vigwatch since before it went live, back when the whole idea was still a slide deck: reprice every in-play market in real time, on every sport Sableridge carries, and manage how much risk the book takes on while it does. Marquee sports went first, sports with years of graded history behind them and enough weekly volume that a real answer showed up within a month of any change. Every quarter, Tamzin walked into Benedikt's office with a number, and the number held up.
Eleven weeks ago, Vigwatch took on something harder: pricing the markets nobody had bothered to watch closely. A trader used to set these by hand, a rough multiplier off the pre-match line, glanced at once a day if that. It made sense. Nobody was going to give a real trader's full attention to a market that settled sixty times a week.
The first weeks looked almost boring, the good kind of boring. Overrides fell from 38 percent to under 10 within a month. The model's opening price crept closer and closer to where each market actually closed. Tamzin watched the hold rate too, out of habit more than anything: 4.0 percent, 5.3, 4.7, a touch above the old 4.1 baseline, nothing dramatic yet.
Then, on a Thursday in week six, rain stopped a county cricket match for forty minutes.
Nothing about that should have mattered. Rain delays happen. But Vigwatch's live price didn't know the delay had changed the shape of the match, not for twenty-six seconds, well past its usual two, on a sport where Vigwatch had barely a season of history to lean on. A professional bettor, watching for exactly that kind of gap, hit the stale line across three connected markets before it moved. Sableridge lost $86,400 that afternoon. Chibundu Kowalik, who runs the numbers behind the micro-market rollout, pulled the trade log before lunch and found it in minutes. There was nothing subtle about it.
Here's what almost got lost in that one number. Week six's hold rate cratered to negative 3.1 percent, for that week alone. On its own, dropped into a slide with no context, it reads as proof the whole rollout had failed. It wasn't proof of that. It was one professional bettor, one delay, one gap in the model's reaction time, on a sport where eleven weeks of data barely covers a season.
Three weeks later, prepping the quarterly board deck, Benedikt asked Tamzin the question he'd asked every quarter for fourteen months: "What's the ROI?" This time she didn't have one. Eleven weeks of a genuinely thin, genuinely volatile market isn't fourteen months of football. She could say the hold rate looked a little better, 4.4 percent against the 4.1 baseline. She could not honestly say that gap was real and not just the same kind of noise that had swung the number nine points in a single week already.
What Tamzin didn't do was stay quiet and hope the next few weeks looked better. She built the report Benedikt actually needed: what Vigwatch already knows for certain, override rate down to 6 percent, opening prices landing within 1.9 percent of the close, and what it still can't know, the real hold-percentage lift, alongside the number that would make it knowable. Chibundu ran the arithmetic: with the variance micro-markets carry, you need about 32,000 graded events before a real signal separates from noise this size. At 640 graded a week, that's roughly week 50. Eleven weeks in, thirty-nine to go.
The decision Tamzin would take back traced to how Sableridge built its own reporting habit in the first place: every AI capability gets a quarterly ROI number starting the first quarter after launch, no exceptions, no waiting for enough data to say something true. That rule made sense when every capability launching was another marquee sport, backed by years of trader-priced history that made a month of live data plenty. It stopped making sense the day a capability launched into a market with no deep history behind it and sixty settled events a day, not sixty thousand.
Run the same board meeting again, with that rule fixed. Benedikt still asks for the ROI. Tamzin still doesn't have one for micro-markets. But now she has a page: what's known, what isn't, and a date, week 50, that Benedikt can hold her to. He doesn't propose pulling micro-market pricing back to manual trading. He circles the date on his own calendar and asks what happens if week 50 comes and the number still isn't good.
One version of that meeting reads a bad week as a verdict. The other reads it as one point on a range that was labeled ahead of time.
What Tamzin would tell herself, back when Sableridge first wrote the quarterly-ROI-no-exceptions rule into how it reports on Vigwatch: a rule built for a market with fourteen months of trader history behind it was never going to survive being handed to a market with eleven weeks of anyone's history at all, and nobody had written down what the rule should do when that happened.
SPARK: the report that turns "not yet" into a real answer
Not a way to dress up "we don't know" so it sounds better. SPARK forces the actual page you'd hand over, and makes you prove it still holds up the week a real number goes wrong.
And if you want to be sure it really works, try it somewhere else
Same five letters, a drought insurer instead of a sportsbook, and this time the rare event isn't a rain delay. It's the severe drought year the model hasn't seen yet.
Netherwood Underwriters sells parametric drought insurance to smallholder farmers through Rainmark, a satellite index that reads vegetation and rainfall signals and decides when a payout triggers, without anyone filing a claim or an adjuster visiting a field. Otoniel Ochieng, who owns Rainmark's model performance, onboarded a new farming region eight months ago. The region has been through exactly one growing season since. A reinsurance partner wants to know whether Rainmark's new drought trigger, tuned for that region's soil and rainfall pattern, actually cuts basis risk, the gap between what farmers really lost and what the index paid, better than the old, blunter regional index it replaced.
Same rank, different lever, mapped straight onto SPARK: the situation is a reinsurance partner asking for a basis-risk number the same way Benedikt asked for a hold-percentage number, before enough seasons exist to answer honestly. The payoff is the same trust: the partner accepts "not yet" because it comes with real numbers already in hand, how closely Rainmark's rainfall readings track ground-truth gauges this season, how fast payouts reached farmers. The anchor is the same shape: what's known now, and the number of full growing seasons, probably three, spanning at least one genuinely severe drought year, before the basis-risk claim is trustworthy. The risk is the same shape too: one unusually mild season could get mistaken for proof the trigger works. And what stays out is the same discipline: no borrowing the old, decade-proven regional index's track record to vouch for the new one early.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: report what you can measure now, name the number of full cycles still needed, and never let one loud result, good or bad, stand in for the real answer.
Cost: no budget this quarter to wait three full seasons before saying anything. Ship the honest interim version: leading indicators only, with the same "not proof yet" label attached, updated every season instead of staying silent until the end.
The model got better, for real: say Rainmark's satellite resolution doubles overnight. The two-part shape doesn't change. A sharper sensor narrows how many seasons are needed, it doesn't excuse skipping straight to a claim before a real severe-drought year has actually been observed.
Where people run it wrong.
They wait for total certainty before saying anything at all, and leadership reads the silence as a worse verdict than the honest one would have been.
They blend a new, thin result into an old, proven number to get one clean line for a deck.
They let one loud period, good or bad, decide the story before enough of them have actually happened.
How to use it live. Ask the coverage question before naming a number: "How much of the real variance has this actually been tested against yet, and what's the smallest bad case it hasn't seen?" That question alone usually separates a real answer from a guess dressed up as one.
Three things worth stating directly, since this is where the real judgment sits. The alternative Tamzin's team considered, and rejected, was blending marquee and micro-market performance into one "Vigwatch Impact" figure for the board deck, the same shortcut Netherwood's team was tempted by too. It lost, because a statistically solid marquee win would quietly vouch for a micro-market number nobody had actually earned yet, and if micro-markets turn out net-negative once real data arrives, that blended number will have already told the board it was a win. The AI-specific failure worth naming is distribution shift under a rare event: Vigwatch's live price kept pricing off the pre-delay distribution for twenty-six seconds because nothing had told it the match had genuinely changed shape, and a model that doesn't know to say "I'm less sure right now" keeps serving a confident, stale number. The guardrail is a stale-price kill switch: any flagged surprise, a rain delay, a red card, an injury, pulls that specific market off the board for a few seconds until Vigwatch recalculates, instead of continuing to quote the old price. And the trade-off is real: pulling a market costs Sableridge a few seconds of betting handle on it, every time the switch fires, even on the vast majority of surprises that turn out harmless. That's accepted on purpose, because the alternative is finding out about the next stale line from an $86,400 trade log instead of a switch that fired quietly and cost nothing but a few seconds of handle.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't $86,400 in one week proof enough that something's wrong?" Response: one week's loss in a market with barely a season of history is the same size of swing the legacy baseline already produced on its own. It's a reason to fix the stale-price gap that caused it, not proof the whole model failed.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.