CaseIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #13

How do you communicate an inconclusive spike result to leadership?

PICKa spike that would not say yes or no

Basepath Analytics builds a pitch-recommendation tool for minor league clubs. Teodor Marchetti is the AI PM who ran a two-week spike for the Tallow Creek Anchors and has to walk into General Manager Renata Solis's office with an answer that is not a clean yes or a clean no.

The direct answer
Report it as a calibrated range, never a verdict: what the spike ruled in, what it ruled out, the exact evidence gap left, and the cheapest next test that would close it. Name a bounded next step, not a shrug. An inconclusive spike rounded into a clean yes or no is not honesty, it is a guess wearing a lab coat, and it costs you the one thing an interviewer is actually checking: whether you can be trusted with an unclear number.
Do this, in order
  1. Report the result as a range with a named gap, not a rounded yes or no.Why: rounding an unclear result either kills a real idea or greenlights an untested one.
  2. Say exactly what evidence would close the gap, and what it costs to get it.Why: leadership can only act on ambiguity if you hand them the next decision, not just the feeling.
  3. Separate "the sample was too small" from "the model can't do this."Why: these call for opposite next moves, and blending them is where trust breaks.
  4. Give a real number for what a wrong read costs each way.Why: a shelved good idea and an underprepared bad launch are not the same size of mistake.
  5. Name the one thing that would flip you to a clean go or no-go.Why: without a kill line, "inconclusive" can drag on forever and never becomes a decision.

How to answer this, stage by stage

Nobody is scoring whether you know the word "inconclusive." They are scoring whether you can hand someone in charge a decision they can actually use when the data will not commit to an answer.

Stage 1
Scope it to one real result
Say it like this
"Let's ground this in a real spike: a pitch-recommendation tool for a minor league club, tested for two weeks, where the result was genuinely a mixed bag."
Why this works
Keeps the answer from turning into a lecture about "managing up," which is what this question sounds like if you let it float.
Stage 2
Name the position before the reasoning
Say it like this
"My pick is: report it as calibrated uncertainty with a named next step, never as a false verdict. I'll walk through why that beats rounding it either direction."
Why this works
PICK opens on the position, not the tour of the evidence. An interviewer wants to see you commit before you explain.
Stage 3
Name who feels each kind of wrong read
Say it like this
"If Renata hears 'inconclusive' as failure, a genuinely promising tool gets shelved and nobody ever finds out it was close. If she hears it as 'basically fine,' it launches without the guardrails a real gap deserves."
Why this works
This is where a strong candidate stops talking about "communication style" and starts talking about who actually eats each mistake.
Stage 4
Say which error you are optimizing against, and why
Say it like this
"The visible, cheap error is Renata asking one more clarifying question and losing a week. The hidden, expensive one is her killing a tool that would have worked with one more sharpened test. I'd rather eat the week."
Why this works
This is the actual judgment PICK is testing. Naming the asymmetry out loud is what separates a decision from a shrug.
Stage 5
Prove it with the compressed failure
Say it like this
"We ran the spike on 40 real at-bats. Fastball counts, the tool's calls agreed with the coaching staff 91 percent of the time. On breaking-ball counts, only 58 percent, on a sample too small to call that real or noise. I told Renata 'looks promising, not proven,' and named the 200-pitch follow-up spike that would settle it."
Why this works
A real number, split by the exact place it got shaky, is worth more than any adjective about "communicating uncertainty well."
Stage 6
Name the kill line, then close on the pick
Say it like this
"If the follow-up spike still splits like this past 200 pitches, I'd call it a real capability gap on breaking balls, not bad luck, and I'd say so plainly. Until then: report the range, name the gap, name the next test. Never round it to a verdict you don't actually have."
Why this works
Closing on the kill criteria shows you are not just hedging forever. It shows judgment has an end point.

Let's learn

Picture a tool that watches a pitcher's last twenty pitches and suggests the next call: fastball away, curveball in the dirt, whatever the data favors.

Before Basepath's tool, the Tallow Creek Anchors' pitching coach made every call by memory and gut, drawing on maybe 200 pitches he could actually recall clearly from a hitter's last at-bat against the club. The new tool draws on every pitch that hitter has ever seen against anyone, thousands of them, and returns a call in under a second.

Hand sketched icon list titled Signs this result is genuinely unclear, not just a bad spike. Four rows: sample too small to trust either way. Signal is there but noisy pitch to pitch. Test conditions did not match a real game. Strong on fastballs, unclear on breaking pitches.
Four honest reasons a spike can come back unclear, none of them "we ran a bad test."

Here's the turn: the two-week spike came back split. Strong agreement with the coaching staff on fastball counts. Weak, noisy agreement on breaking-ball counts, on a sample too small to say whether that gap was real or just forty at-bats of bad luck. Teodor's habit up to now had been simple: spikes were either clearly good or clearly bad, and he'd walk in and say so. This one would not sort itself that way.

Model agreement with the coaching staff, by pitch type
100% 50% 0 91% Fastball (24 at-bats) 58% Breaking ball (16 at-bats)
Same tool, same two weeks. One pitch type looked strong. The other looked shaky on a sample too small to call it either way.

At its worst, Teodor rounds this into "it works" and the Anchors run it into a real September call that costs a game, or he rounds it into "it doesn't work" and the club walks away from a tool that would have been solid with three more weeks of breaking-ball data.

An inconclusive spike is not a smaller yes. It is a different kind of answer, and rounding it into a yes or no throws away the one honest thing you actually learned.
The choice I would take back Months earlier, when Basepath first pitched the tool to the Anchors, someone promised, "every spike will tell you clearly whether it's ready." That made the pitch easy. It stopped making sense the day a real spike came back genuinely split, because now an honest report reads like breaking a promise instead of just being a report.

What I would leave alone: I wouldn't build a special escalation process for every close spike result. Most spikes in this shop land clearly one way or the other; forcing a heavy "confidence band" ritual onto all of them would slow down the easy calls to protect against the rare hard one.

The lesson: the promise that every result would be clear was never true, and the day it wasn't true, the fix wasn't a better spike, it was a better sentence for reporting the honest one.

Now here is the same thing as a story

The short version above is what you'd say defending your report in a five-minute hallway conversation. Read this one for what it felt like the week "inconclusive" almost got treated as a synonym for "no."

Teodor is good at reading a room before he's good at reading a stat sheet. Two years running spikes for minor league front offices taught him that a GM's first question after any pitch is never about the model. It's "so do we use it or not."

For most of that two years, the answer was easy either way. A tool would clear 90 percent agreement across the board and Teodor would say so, plainly, and the club would adopt it. A tool would flop under 60 percent everywhere and he'd say that too, and the club would pass. Clean, both times.

Hand sketched quadrant titled Where this spike actually sits. Axes: sample size from a handful to a full season, consistency across pitches from scattered to steady. Past clear go sits high on both axes. Past clear no-go sits high on sample size but low on consistency. This spike sits in the middle on both, neither corner.
Past spikes landed in a corner. This one landed in the middle, and the middle has no ready-made sentence.

The pitch-recommendation spike for the Anchors did not land in a corner. It landed in the middle. Strong on fastballs. Shaky on breaking balls, on a sample too small to call real.

Then, three days before the readout meeting, a colleague on the eng side leaned over Teodor's desk and said, half joking, "you're not really going to tell Renata it's inconclusive, are you? Just call it a win on fastballs and leave it there."

Knowledge spark: why can't you just average the two numbers? Averaging 91 percent and 58 percent gives you a tidy 75 percent that describes nothing real. No pitcher throws an average pitch. The tool is genuinely strong in one place and genuinely unproven in another, and a blended number hides exactly the gap a leader needs to see.

It was tempting. Rounding up would have made for an easier fifteen minutes in Renata's office. Teodor didn't do it. He printed the split, the sample sizes, and one line about what would close the gap, and walked in with the honest, harder-to-say version.

Hand sketched comparison titled The asymmetry, drawn. Left panel, question mark box icon, labeled Reads inconclusive as ask more, caption one more week, one sharper spike, cheap. Right panel, box icon, labeled Reads inconclusive as no, caption a real idea shelved for good, hidden, expensive.
Two ways to hear the same word. One costs a week. The other costs the whole idea, quietly.

Renata read the printed sheet twice. Her first question was the dangerous one: "so is it ready or not." Teodor didn't flinch into either answer. He said: ready for fastball counts, not yet proven on breaking balls, and here's the exact test that would tell us which in three weeks, at the cost of running it through one more September series.

Hand sketched decision tree titled What leadership does with the word inconclusive. Root: Renata hears inconclusive. Three branches: hears it as failure leads to shelves the idea for good. Hears it as maybe leads to funds it anyway underprepared. Hears a calibrated range leads to funds the exact next test.
Same word, three different meetings. Only one of them ends with a decision anyone can actually stand behind.

The real cost was never going to be the fifteen minutes in that office. It was going to be whichever club, months from now, hears "inconclusive" from some other AI PM, reads it as failure, and never finds out the tool was three weeks of data away from working.

Hand sketched labeled parts diagram titled What a good spike readout holds. A document icon at the center labeled Spike Readout, with four labeled callouts around it: Ruled in, Ruled out, Evidence gap, Next test and its cost.
Four parts, every time. Skip any one of them and "inconclusive" turns back into a coin flip for whoever hears it.

Renata funded the three-week follow-up. It came back clean: 88 percent agreement on breaking balls too, once the sample was big enough to trust. The tool shipped for the whole rotation, not just fastball counts, and nobody had to pretend the first spike said something it didn't.

What I'd tell my past self, the one who almost took the eng colleague's advice: the easy sentence was never the honest one, and Renata could handle the harder sentence just fine. She couldn't have handled finding out later that I'd rounded it for her.

PICK, the pick that survives the follow-up questionNot a script for hedging every result. PICK is what tells you exactly which kind of wrong read costs more, and to build your report around avoiding that one.

P
Position. Your pick, before any reasoning.
Report the range and the gap, never a rounded verdict. Say it before the evidence, not after.
Commit first. The reasoning that follows is defense, not discovery.
I
Impact. Who feels each kind of wrong read?
Renata feels a lost week if she asks a clarifying question over nothing. The whole club feels a shelved tool that would have worked, or an unready one launched live, if the word "inconclusive" gets rounded either direction.
This is the step that turns a vague worry about "communication" into a real, nameable cost on each side.
C
Cost asymmetry. Which error is hidden and expensive?
A rounded-up "it works" that ships too early is visible fast, someone notices a bad call in a real game. A rounded-down "it doesn't work" that gets a good tool shelved is invisible forever. Optimize against the second one.
The whole PICK answer turns on naming this asymmetry out loud, not just noticing it privately.
K
Kill criteria. What would flip the pick?
A second, larger spike on breaking-ball counts. If the split still holds past 200 pitches, that becomes a real capability gap worth saying plainly, not more hedging.
Without a kill line, "inconclusive" can drag on forever. This is what turns honesty into a real decision with an end date.

The recap, one line per letter: position is report the range, never the rounded verdict. Impact is Renata's lost week versus a shelved good idea or an underprepared launch. Cost asymmetry is optimizing against the shelved idea, because it never shows up on any dashboard. Kill criteria is the 200-pitch follow-up spike that would turn "inconclusive" into a real yes or no.

And if you want to be sure it really works, try it somewhere elseSame four letters, a veterinary imaging tool instead of a pitch call. Different flip family entirely, same discipline about the word "inconclusive."

Dr. Priya Balan runs diagnostics at Windle Hollow Animal Diagnostics, piloting an imaging tool that flags likely tumors on X-rays for a second look. A ten-case spike came back split: strong agreement with radiologists on soft-tissue masses, weak on bone lesions, on only three cases either way. Mapped onto PICK: position is report the split, not a blended accuracy number. Impact is a vet who trusts a rounded-up "it works" missing a bone lesion, against a clinic that abandons a genuinely useful tool over three unlucky cases. Cost asymmetry favors naming the bone-lesion gap loudly, since a missed lesion costs an animal's health, not just a clinic's convenience. Kill criteria is a twenty-case follow-up read specifically on bone lesions before the tool touches anything but soft tissue.

The same hand sketched icon list reused: Signs this result is genuinely unclear, not just a bad spike, reapplied to a veterinary imaging spike with the same four honest reasons a result can come back split.
The same four honest reasons apply here too: small sample, real but noisy signal, mismatched test conditions, uneven strength across categories.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "report the range and the gap, name the next test, never round it to a verdict," and stop.
Cost: there's no budget for a proper follow-up spike. Say so honestly, and scope the cheapest version, ten more at-bats logged during games already being played, rather than skipping the follow-up entirely.
The result is a clear improvement, for real: a spike that comes back strongly positive everywhere still deserves the same discipline, since "clearly great" on forty samples can still hide a rare failure mode nobody tested for.

Where people run it wrong.
They round an unclear result up to look decisive, or down to look cautious, instead of reporting what was actually found.
They give leadership a feeling instead of a number and a next step.
They let "inconclusive" become a permanent state with no test that would ever resolve it.

How to use it live. The moment an interviewer asks how you'd communicate an unclear result, buy two seconds by asking yourself: what specifically would leadership do differently if I rounded this up versus down? Naming that gap out loud is the whole answer.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Abandonment flip: leadership can't tell "promising but unresolved" from "failed," so an honestly unclear result risks getting the whole idea quietly shelved.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Teodor Marchetti, the AI PM at Basepath Analytics, reporting a split spike result to Tallow Creek Anchors GM Renata Solis.
3 · THE HABIT
What did Teodor stop being able to do, that used to work every time?
Tap to flip
ANSWER
Give a clean, one-word yes or no. Every past spike had landed clearly one way or the other, until this one landed in the middle.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Leadership hearing "inconclusive" as a calibrated range worth funding one more test on, versus hearing it as a euphemism for failure and shelving the idea. No middle ground once the word lands the wrong way.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Promising early on that "every spike will tell you clearly whether it's ready," which made an honest, split result feel like breaking a promise instead of just being a report.
6 · THE NUMBER
Fill in the blank: the tool agreed with the coaching staff ___ percent of the time on fastball counts, and ___ percent on breaking-ball counts.
Tap to flip
ANSWER
91 percent on fastballs, 58 percent on breaking balls, the second one on a sample too small to call it real.
7 · THE REPLAY
Same split result, reported honestly instead of rounded. What changes?
Tap to flip
ANSWER
Renata funds a three-week follow-up spike instead of killing the idea or launching it half-tested. The follow-up comes back at 88 percent on breaking balls, and the tool ships for the whole rotation.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Windle Hollow Animal Diagnostics' imaging spike. The underlying flip there is a verification flip: a vet either trusts a rounded "it works" outright or checks every read personally, with no report format in between.

Check yourself Score: 0 / 0

True or false
1. True or false: Teodor's spike found the tool performed equally poorly across every type of pitch.
  • True
  • False
Show hint
Look at the chart splitting fastball versus breaking-ball agreement.
Show answer
False. It performed strongly on fastball counts (91 percent) and only weakly, on a small sample, on breaking-ball counts (58 percent). The result was split, not uniformly bad.
Multiple choice
2. Why did Teodor refuse to average the 91 percent and 58 percent into one number?
  • A. Averaging percentages is mathematically impossible.
  • B. A blended number would hide exactly the gap Renata needed to see, since no pitch is actually "average."
  • C. Renata specifically asked him not to use averages.
  • D. The coaching staff disagreed with the sample sizes.
Show hint
Look at the knowledge spark about averaging the two numbers.
Show answer
B. A tidy blended number describes a pitch that doesn't exist and buries the real gap between strong and unproven performance.
Fill in the blank
3. Fill in the blank: the visible, cheap error is Renata asking one more question and losing a week. The hidden, expensive error is a real idea getting ___ for good.
Show hint
Look at the cost asymmetry step, C, in the framework recap.
Show answer
Shelved. A rounded-down "inconclusive means no" can end a genuinely promising tool with nobody ever finding out how close it actually was.
Short answer, where it wouldn't matter
4. Name a situation where you wouldn't need this careful a readout process for a spike result.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A spike that comes back clearly strong or clearly weak across the board doesn't need a special escalation ritual. Save the careful, split-result process for results that actually land in the middle.
Short answer, apply it yourself
5. Think of a time you had to tell someone in charge that a test or a trial run came back genuinely unclear. What did you say, and did you name a next step or just describe the ambiguity?
Show hint
Think about whether your report ended in a decision someone could act on, or just a feeling.
Show answer
Model answer: A strong version names what's proven, what's not, and the cheapest way to close the gap, the same shape as Teodor's printed sheet for Renata.
Short answer, work the number
6. If the breaking-ball sample had been 160 at-bats instead of 16, and still showed 58 percent agreement, would Teodor's "inconclusive" framing still be the right call?
Show hint
Think about what makes a result genuinely unclear versus just a small sample.
Show answer
Model answer: No. At 160 at-bats, 58 percent agreement is a large enough sample to call it a real capability gap, not noise. The honest report at that point is a clean "not ready on breaking balls," not "inconclusive."
Before you close the answer
Why this works
Tests whether you'll protect an honest but unclear result from getting rounded into a false verdict, under real pressure to sound decisive in the room.
Follow-up traps
"Isn't 'it's complicated' just a way of avoiding a real answer?" Response: no, because the report still ends in a concrete decision, a named next test and its cost, not a shrug. That's what makes it different from a dodge.

"What if leadership just wants a yes or no, and won't accept nuance?" Response: give them the range anyway, but frame it as a decision: "here's what I'd bet on, here's what I'm not sure of yet, here's the cheapest way to find out." That's still an answer, just an honest one.
If pressed
The follow-up spike design mattered as much as the report: it specifically added 200 more breaking-ball counts across three additional games rather than just re-running the same 16 at-bats, since repeating a too-small sample would have just produced a differently unclear result.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more