CaseAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #22
Explain how you would handle an ROI model whose assumptions your CFO rejects.
PICK · autonomous bug hunting and QA testing for game studios
Snagline is Nettlebrook Systems' AI agent that plays a studio's game build the way a tireless tester would, hunting for crashes, softlocks, and broken quest states before a human ever touches it. Odilia Winterscale owns its ROI story. Three days before a quarterly investment review, Ignatius Thackerston, Nettlebrook's CFO, pulls Cavernlight Games' own numbers on their open-world RPG Emberreach, and finds a catch rate nothing like the 85 percent sitting in Odilia's deck.
The direct answer
Concede the blended 85 percent catch rate does not hold for Emberreach's genre, on the spot, and rebuild the ROI live with Snagline's own genre-segmented number, 61 percent, not "85 on average." Never say "let's take this offline": pull the data that already exists and show the real, smaller, still-positive return in the room.
Do this, in order
Concede the blended catch-rate assumption live, and rebuild the number with genre-segmented data, in the room.Why: an assumption the eval data itself contradicts is the cheap mistake to admit fast, not the one worth defending.
Know, before the meeting, which assumptions are genuinely defensible and which are an average hiding a worse number.Why: that's what tells you whether to concede or hold the line, instead of guessing under pressure.
Name who pays for each kind of mistake before picking a side.Why: caving on a number the data supports teaches finance to discount every future ask; defending a number it doesn't costs one meeting, not the whole relationship.
Keep genre-segmented eval numbers on hand for every deal, not just the disputed one.Why: the concession only works live because the real number was already tracked, not invented under pressure.
State the kill criteria before the meeting, not after.Why: a number that can't say what would change it isn't a position, it's a hope wearing a spreadsheet.
Leave the blended number alone for genres where it's actually true.Why: forcing every deal into cautious segmented reporting underprices the strong numbers the same way over-claiming underprices the honest ones.
How to answer this, stage by stage
Nobody is grading whether you can stay calm when a number gets challenged. They're grading whether you check it against real data before you decide whether to fold or hold.
1
Ground it in one deal, one number, one name
Say it like this
"Let's ground this in one deal. Snagline is Nettlebrook Systems' AI agent that plays a studio's build like a tireless tester, hunting bugs before a human touches it. Odilia Winterscale owns its ROI story. Ignatius Thackerston, Nettlebrook's CFO, is the one who rejected her number."
Why this works
Keeps every claim checkable against a real deal, instead of a general policy on "handling pushback."
2
Say the method out loud
Say it like this
"I'll run this as PICK. Take a position on what actually happens in the room, say who pays for each kind of mistake, name which one costs more, then say what would change my mind."
Why this works
Two seconds of structure signals a plan is running, not a scramble for an answer.
3
Give the position, committed, before any reasoning
Say it like this
"Here's my position. The moment a CFO rejects an assumption with real evidence behind the objection, I concede it on the spot and rebuild the number live, using whichever data is actually true. I don't say 'let's take this offline.'"
Why this works
This is the direct answer, stated before anyone has to dig for it.
4
Name who pays for each kind of mistake
Say it like this
"Two ways this goes wrong, and two different people feel it. If I cave on a number the data actually backs, Ignatius learns to discount every figure I bring him from now on. If I dig in on a number the data doesn't back, Cavernlight's own QA lead has already told their account manager Snagline missed most of the questline bugs, and I look like I never checked my own work."
Why this works
Turns an abstract "handle pushback" question into two real costs, not one generic worry about looking bad.
5
Name which mistake is the expensive one
Say it like this
"Digging in on a weak number is cheap. Ignatius checks it against Cavernlight's real QA log and corrects me in the same meeting, five embarrassing minutes and it's over. Caving on a strong number is expensive. It doesn't cost me that meeting, it costs every number I bring him for the next two years, because now he assumes they're all soft until proven otherwise."
Why this works
The hardest step in PICK. Naming which direction actually compounds is what makes the position defensible, not just confident.
6
Prove it, rebuilt live with the real number
Say it like this
"Here's what actually happened. Snagline's blended catch rate, 85 percent, comes from every genre we've shipped against, mostly linear and mobile titles. Open-world titles like Emberreach test differently, and our own genre-tagged data says 61 percent there, not 85. So I said that, out loud, and rebuilt the deck on the spot: at 61 percent, Emberreach's contract still nets Cavernlight about 373,000 dollars a year against the license fee. Smaller number. Still real."
Why this works
Shows the concession isn't a guess, it's a number that was already sitting in the eval logs, waiting to be pulled.
7
State the kill criteria and close on the decision
Say it like this
"One more thing before I sit down. If genre-segmented catch rate for open-world titles hasn't cleared 70 percent within two release cycles of shipping the traversal module, I retire the module as its own pitch and every open-world deal gets the segmented number by default, no exceptions. So: concede fast when the data says I'm wrong, defend fast when it says I'm right, and never leave the room without saying what would change my mind."
Why this works
Closes on the decision and shows the position can move, which is what makes it a real position instead of stubbornness.
Let's learn
For two years, every ROI deck Odilia Winterscale built for Nettlebrook Systems said the same number: eighty-five percent.
Snagline is the AI agent Nettlebrook sells to game studios. It plays a build of a game for hours at a time, the way a tireless QA tester would, and files a bug report the moment something breaks.
Five steps, and the middle one is the model checking its own confidence, calibrated against 2,300 human-confirmed bugs across 60 shipped titles, mostly linear and mobile games.
Before Snagline, a studio the size of Cavernlight Games ran a contract QA team through a title's last six months, at roughly 1.3 million dollars for a project the size of Emberreach. Snagline runs the same build every night, unpaid overtime a human team can't match, and never gets bored on hour six of clicking the same door.
Averaged across every genre Nettlebrook has ever shipped against, mostly linear levels, puzzle titles, mobile games, Snagline catches 85 percent of the bugs a matched human QA team would find, in the same window. That's the number in Odilia's deck. That's the number in every deck.
Catch rate by genre, against the blended average
Genres near the blended averageEmberreach's genre, open-world RPGBlended average, 85%
Four genres sit close enough to 85 percent that the average is basically honest for them. Open-world titles are 24 points below it, and nobody built a line for that until Ignatius asked.
Same nightly test budget, about 40 agent-hours either way. A linear level's whole map fits inside that budget. An open world's quest branches don't.
Knowledge spark: what's exploration coverage?
How much of everything a player could possibly do, Snagline actually tries in the time it gets. A linear level has a few hundred real branches. An open world can have millions: quests, item combos, dialogue paths. Same testing hours, a much smaller slice actually gets checked.
Here's the turn. The 85 percent isn't wrong. It's an average, and an average is quietly promising that every game tests the same way. Emberreach doesn't. It's an open-world RPG with something like forty hours of branching quest content, and Snagline's real catch rate there, the genre-segmented number nobody put in the deck, is 61 percent.
The 85 percent wasn't wrong. It just wasn't Emberreach's number.
At its worst, that gap doesn't just make one slide wrong. Cavernlight's own QA lead had already told their account manager Snagline "missed a ton of stuff in the questline branches." If Odilia's next answer to Ignatius is a shrug, or a promise to check and get back to him, every number she's ever brought him gets quietly discounted from that afternoon on, whether or not it deserved to be.
Not a bug in the calculator. A default that made sense the year it was built, and nobody's job to come back and check.
The choice that mattered
Snagline's ROI calculator was built, in its first year, to report exactly one number: the blended catch rate, averaged across the whole customer base. That was fine then, almost every early customer shipped a linear or mobile title, and the average and the truth were close enough not to matter. Nobody revisited that default once open-world deals started showing up in the pipeline.
What I would leave alone: the blended number itself is fine, left alone, for the deals it was always honest about. A puzzle-game studio or a mobile-game studio gets a catch rate within a point or two of 85 percent either way. Building genre-segmented reporting for them would be paperwork nobody asked for.
The lesson: an ROI model is only as honest as its most averaged number. The day your product's accuracy stops being uniform across the customers you sell it to is the day defending that average, in front of someone whose whole job is checking averages, becomes a losing move. Segment it before someone makes you.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why nine words in a Slack channel did more damage than the actual gap in the number.
Odilia Winterscale can read a genre-segmented eval table the way some people read a box score: straight to the row that's actually moving, past the average sitting on top of it.
She built Snagline's confidence-check step herself, three years ago, back when Nettlebrook had four customers and all four of them shipped tight, linear levels. The system was simple then. Play the build, flag anything odd, check how sure you are, and hand anything under the line to a person. It worked, because at four customers, 85 percent meant 85 percent for everyone.
For two years, that number carried every deck Odilia built. New studios signed. Old studios renewed. The board slide never needed a second line.
Nettlebrook's pipeline changed slowly enough that nobody marked the day it happened. First it was one open-world deal, small, a proof of concept nobody expected to renew. Then two more, bigger. By the time Cavernlight Games signed for Emberreach, an open-world RPG with something like forty hours of branching quest content, open-world titles were a third of the pipeline, and the ROI calculator still had exactly one output field.
The trigger was nine words in a Slack channel nobody thought mattered. A solutions engineer, closing out the Cavernlight deal retro, wrote: "their QA lead says we missed a ton of stuff in the questline branches." Nobody replied. It sat there for four days.
Neither of them was being unreasonable. One of them just checked first.
Ignatius Thackerston reads deal retros the way other people read the news, out of habit, three days before he has to sign off on anything. He read that line on a Tuesday. Wednesday, he pulled Cavernlight's actual bug counts himself, matched them against the human QA team's own findings for the same build, and got 61 percent. Not 85.
He didn't wait for the quarterly review to ask about it. He forwarded Odilia the comparison at 6:40 that evening with one line: "Walk me through this before Thursday."
The whole answer to this question, in one picture. One dial can't describe a world that keeps changing shape underneath it.
Thursday came fast. Odilia had two days, and the easy version of Thursday was the one where she asked for more time. "Let me dig into that and get back to you" would have bought her a week, and cost her something she couldn't get back: the next time she said a number in that room, Ignatius would check it before he believed it.
So she didn't ask for the week. She opened Snagline's own eval dashboard, live, in the meeting, and did the thing the blended number had let everyone skip for two years: filtered by genre. Open-world titles, 61 percent, not 85. She said so before he could. "You're right. The 85 is real, it's just an average, and it was never Emberreach's number. Here's Emberreach's number."
The decision she'd take back happened three years earlier, in a fifteen-minute stretch of a roadmap meeting nobody remembers clearly. Someone had asked whether the ROI calculator should report catch rate by genre from day one. The room said no, there wasn't enough variance yet to bother, every customer at the time shipped roughly the same shape of game. That was true, then. Nobody wrote down when to come back and check whether it was still true.
Run Thursday's meeting again, with that one thing fixed. Ignatius still finds Cavernlight's real number before the meeting, the same way, the same nine words in the same Slack channel. But now the deck he's holding already has a line for it: open-world, 61 percent, Emberreach's actual figure, not the blended one. He doesn't have to catch anything. He just checks a number that was already honest. The meeting runs eleven minutes instead of forty, and it ends with him asking about the traversal module budget, not about whether he can trust the deck.
One version of that Thursday spends forty minutes rebuilding trust. The other spends eleven minutes talking about what to build next.
What Odilia would tell her past self, back in that fifteen-minute roadmap meeting: an average is a promise that the thing behind it doesn't change shape. The day Nettlebrook started selling into open worlds, that promise stopped being true, and nobody came back to check whether the calculator still knew it.
PICK, or how an average survives being checked
Not a script for staying calm when a CFO pushes back. PICK forces a real commitment about which way you move when the pushback lands, then makes you say, out loud, which mistake actually costs more.
PPosition. Your pick, in one sentence, before any reasoning.
When a rejected assumption is backed by real data, concede it on the spot and rebuild the number live, using the data that already exists, not a promise to check later. Ignatius rejected the blended 85 percent behind Emberreach's ROI, and it wasn't a bluff, Cavernlight's own numbers backed the objection, so the position was to concede it inside the meeting and show the real number, 61 percent, immediately.
Say the position before the number, or the room spends five minutes guessing whether this is a fight or a correction.
IImpact. Who feels each kind of error, and in what units.
Book it wrong toward stubbornness, and Cavernlight's own QA lead has already told their account manager Snagline missed the questline bugs, so digging in reads as either dishonest or unaware. Book it wrong toward caving, and Ignatius starts discounting every figure Odilia brings him, for years, not just this deal.
Naming both people, the one deciding whether to trust the deck and the one who already caught the gap, keeps this from turning into a one-sided caution story.
CCost asymmetry. The heart of it.
Digging in on a weak number is cheap. Ignatius checks it against Cavernlight's real QA log and corrects it in the same meeting, five embarrassing minutes and it's over. Caving on a strong number is expensive. It doesn't cost the meeting, it costs every number brought to that room for years afterward, because now they're all assumed soft until proven otherwise.
This is the step that earns the position. Anyone can say "be honest with finance." Naming which direction actually compounds is what survives a follow-up question.
One of these mistakes gets caught in the room. The other one gets caught two years later, spread across every number after it.
KKill criteria. What evidence would flip the position.
Open-world catch rate sits at 61 percent right now and has climbed about three points a quarter for the last four quarters on its own, before any new investment. The traversal module Odilia is asking for, about 720,000 dollars in engineering time aimed at the eleven more open-world deals already in the pipeline, worth roughly 3.1 million dollars combined, only makes sense if that curve keeps moving. If genre-segmented catch rate doesn't clear 70 percent within two release cycles of the module shipping, the module gets retired as its own pitch, and every open-world deal gets the segmented number by default, no blended figure, ever again. If it clears 80 percent and holds for two quarters, the number moves from modeled to measured, a stronger claim, not a weaker one.
A position with no kill criteria is a hope you're defending forever. This makes it a live number that can move in either direction.
Open-world catch rate, against the kill line
Open-world catch rate, no module yetKill line, 70% within two release cycles of shipping it
The line is already climbing on its own, about three points a quarter. The module has to prove it moves that line faster than doing nothing would.
And if you want to be sure it really works, try it somewhere else
Same four letters, a garment mill instead of a game studio, and this time the average was hiding a print, not a genre.
Flawtrace is Hallowick & Combe's vision system for the last stop on a fabric line: a camera reads every roll for broken threads, dye spots, and weave slubs before it ships to a cutting floor. Perdita Casterbrook owns its ROI story. Blended across the whole product line, mostly solid and lightly textured fabrics, Flawtrace catches 91 percent of the defects a trained human inspector would find. On patterned fabric, busy prints that visually hide a flaw inside the pattern's own noise, the real number is 68 percent.
Same shape of gap as Snagline's. The camera never gets less accurate. The pattern just gives it more places to hide a flaw.
Gwyneira Rookwood, Hallowick & Combe's finance lead, rejected the blended 91 percent the same way Ignatius rejected Snagline's 85: a new upholstery contract was entirely patterned fabric, and the 91 percent in Perdita's case had nothing to do with what that specific line would actually catch.
The decision Perdita would take back
Flawtrace's dashboard shipped with one blended accuracy number from day one, the same shape of default as Snagline's calculator. Nobody revisited it once patterned-fabric contracts entered the pipeline.
Same rank, different lever: Perdita conceded the blended number live, in the room, and rebuilt with the segmented 68 percent. At 650,000 dollars a year in inspection labor for that line against a 180,000 dollar Flawtrace license, 68 percent still nets Hallowick & Combe about 262,000 dollars a year, a smaller number than the 411,000 the blended figure implied, and a real one.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: concede fast when the data backs the objection, defend fast when it doesn't, and never leave without saying what would change your mind.
Cost: no time this quarter to build genre-segmented reporting company-wide. Ship the cheap version: flag just the disputed segment with its own honest number, expand the rest later.
The model got better, for real: say Flawtrace's pattern-matching zone improves and patterned-fabric catch rate jumps to 80 percent overnight. The position doesn't change. A bigger number still gets defended with the same segmented honesty, not folded back into one blended figure.
Where people run it wrong.
They defend the average instead of conceding fast, and lose credibility for every number that comes after.
They quietly lower every deal's assumption to the worst-case number to avoid ever being caught again, which underprices the honest deals right along with the risky one.
They wait to bring the real number until after the meeting, which reads exactly like getting caught, not like preparation.
How to use it live. Ask which number the objection is actually about before answering: "Is this pushback on the average, or on this specific segment?" That question alone tells you whether to defend or concede, and it buys a few real seconds to check.
Three things worth stating directly, since this is where the real judgment sits. The alternative Nettlebrook considered, and rejected, was locking the ROI calculator to always quote the conservative 61 percent floor, for every deal, regardless of genre. It lost, because it would underprice the real, defensible 90-plus percent number on most of the customer base, puzzle titles, mobile games, tightly bounded levels, punishing the honest deals to protect against the one that wasn't. The AI-specific failure worth naming is a coverage gap that tracks genre, not accuracy: Snagline explores each nightly build under a fixed budget, about 40 agent-hours regardless of genre, and a linear level's reachable-state graph gets fully walked inside that budget in a handful of passes, while an open-world quest graph with branching flags does not, so the same 40 hours buys a much smaller slice of coverage without the model ever reporting itself as less sure. The guardrail is genre-tagged eval reporting as the calculator's default output, plus a coverage-confidence flag on any build where exploration didn't reach a set fraction of known reachable states. And the trade-off is real: Nettlebrook could raise the open-world exploration budget to close the gap faster, but more agent-hours per nightly build means a slower nightly turnaround and a higher compute bill per run, and the team accepts a lower catch rate on purpose rather than make studios wait past sunrise for a build that used to be ready by 7am.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits this question, and what's its one job?
Tap to flip
ANSWER
PICK: commit to a position when an assumption gets challenged, then show which of two mistakes, caving or defending, actually costs more.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Odilia Winterscale, who owns Snagline's ROI story at Nettlebrook Systems, and Ignatius Thackerston, Nettlebrook's CFO, who checked her number himself before believing it.
3 · THE POSITION
What's the P step here, in one line?
Tap to flip
ANSWER
Concede a rejected assumption on the spot when the data backs the objection, and rebuild the number live in the room, never "let's take this offline."
4 · THE COST ASYMMETRY
Which mistake is cheap and visible, and which is hidden and expensive?
Tap to flip
ANSWER
Defending a weak number is cheap: corrected in the same meeting once checked against the real data. Caving on a strong one is expensive: it quietly discounts every number you bring for years after.
5 · THE OLD DECISION
What decision would Odilia take back?
Tap to flip
ANSWER
Building Snagline's ROI calculator to report one blended catch rate by default, because early customers all shipped similar, linear-shaped games. Nobody revisited it once open-world deals entered the pipeline.
6 · THE NUMBER
Fill in the blank: Snagline's blended catch rate is ___ percent. Its genre-segmented catch rate for open-world titles like Emberreach is ___ percent.
Tap to flip
ANSWER
85 percent blended. 61 percent for open-world titles, the number that actually applies to Emberreach.
7 · THE REPLAY
Same Thursday meeting, new default, what changes?
Tap to flip
ANSWER
Ignatius still finds the real number first, the same way, but the deck already shows it. The meeting runs about eleven minutes instead of forty, and ends on the next investment, not on rebuilding trust.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this question again for a different product. Which one, and what's the disputed number there?
Tap to flip
ANSWER
Flawtrace, Hallowick & Combe's fabric-defect inspection system. The disputed number is its blended 91 percent defect-catch rate, versus 68 percent on patterned fabric.
Check yourself Score: 0 / 0
Multiple choice
1. The moment Ignatius names a specific, checkable objection to her ROI assumption, what should Odilia do?
A. Ask for a week to review the numbers and follow up by email.
B. Defend the blended number, since it's the one already in the deck.
C. Check whether the objection is backed by real data, then concede or defend live, in the room.
D. Escalate to her manager before answering.
Show hint
Look at the P step in the framework recap.
Show answer
C. Asking for a week stalls and costs trust the same way caving does. The position is to check and respond live, not defer to "let's take this offline."
Fill in the blank
2. Snagline's blended catch rate is ___ percent. Emberreach's genre-segmented number, the one that actually applies to Cavernlight's deal, is ___ percent.
Show hint
Check the bar chart in Let's Learn.
Show answer
85 percent; 61 percent. The gap between the two is the entire reason Ignatius rejected the number in the first place.
True or false
3. True or false: because the blended 85 percent turned out to be wrong for Emberreach, Nettlebrook should stop quoting it to every prospect.
True
False
Show hint
Check "what I would leave alone" in Let's Learn.
Show answer
False. The blended number is honest for genres like puzzle and mobile games where it sits within a point or two of the truth. Forcing segmented reporting everywhere would be paperwork nobody asked for.
Short answer, name the rejected alternative
4. What alternative did Nettlebrook consider instead of genre-segmented reporting, and why was it rejected?
Show hint
Look at the "three things worth stating directly" paragraph near the end of Section 4.
Show answer
Model answer: Locking the ROI calculator to always quote the conservative 61 percent floor, for every deal, regardless of genre. Rejected because it would underprice the real, defensible 90-plus percent number on most of the customer base, punishing the honest deals to protect against the one that wasn't.
Short answer, apply it yourself
5. Think of a number you've defended at work that turned out to be an average hiding a worse case underneath. What would you have needed on hand to catch it before someone else did?
Show hint
Think about a segment or category breakdown you could have pulled ahead of time, not a bigger sample of the same average.
Show answer
Model answer: A support team quoted one blended satisfaction score across every ticket type, when billing disputes scored far lower than everything else. They needed the category breakdown pulled and ready before anyone thought to ask for it.
Short answer, work the number
6. If Emberreach's genre-segmented catch rate had come in at 50 percent instead of 61, would the deal still clear its 420,000 dollar license cost?
Show hint
Multiply the catch rate against the 1.3 million dollar QA labor pool, then subtract the license fee.
Show answer
Yes, barely. 0.50 times 1,300,000 is 650,000 in avoided cost, minus the 420,000 license, leaves 230,000 net. Still positive, but a much thinner margin than the 373,000 the real 61 percent number produces.
Before you close the answer
Why this works
Tests whether you'll fold the moment a number gets challenged, or dig in out of pride. The strongest answer does neither: it checks the objection against real data, live, and moves fast in whichever direction that data actually points.
Follow-up traps
"Doesn't conceding in the room make you look unprepared?" Response: only if the concession comes with nothing else. Rebuilding the number live, off data that was already being tracked, reads as prepared, not caught out.
"What if the segmented data isn't on hand when the objection lands?" Response: then the real failure happened weeks earlier, the calculator should have reported genre from the point the customer base stopped being one shape of game, not the day someone got challenged on it.
If pressed
The 70 percent kill threshold isn't arbitrary. It's set at the point where the traversal module's projected coverage gain stops paying for its own compute cost, about 900,000 dollars a year in added exploration compute across the current open-world pipeline, so anything short of 70 percent means the module costs more to run than the extra bugs it catches are worth.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.