Artifact critiqueIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #6
What does a good spike report contain?
GUARDnobody in the room could answer the new hire's question
Verdant Sort Systems builds SortEye, a camera system that identifies recyclable material on a sorting line so contaminated items get flagged before a bale ships. Xiomara Peralta is the AI PM whose two-week spike report said one clean number, 94 percent accuracy, and nothing else, until a new hire asked what that number was actually hiding.
The direct answer
A good spike report contains five things, not one headline number: the overall result, a breakdown of that result by whatever segment is most likely to hide a real weak spot, the worst-performing segment named explicitly with its own number, at least one real failure example, and a stated path for whoever is downstream of a wrong call to get it reviewed. SortEye's first report had only the first item. It took a new hire asking "what about black plastic" to reveal the other four were missing, and that the workers whose pay depended on the sorting line's numbers had no way to contest a wrong flag at all.
Do this, in order
Break the headline number down by the segment most likely to hide a weak spot, before writing anything else.Why: an overall accuracy number can look great while hiding one category that's badly wrong.
Name the worst-performing segment explicitly, with its own number, even if it's uncomfortable.Why: a report that only shows good news isn't a report, it's marketing for the launch decision already made.
Say who is downstream of a wrong call, and whether they have any way to contest it.Why: the person reading the report and the person affected by its mistakes are often not the same person.
Include at least one real failure example, not just a summary statistic.Why: a specific wrong case tells a reader what "wrong" actually looks like in practice.
Recommend a concrete guardrail for the worst segment, not just a rollout decision.Why: naming the weak spot without proposing what to do about it leaves the risk sitting there anyway.
How to answer this, stage by stage
Nobody is scoring whether you can name five report sections from memory. They're scoring whether you'd notice the report in front of you is missing the four that matter.
Stage 1
Scope it to one real report
Say it like this
"Let's ground this in SortEye's actual first spike report, one number, 94 percent, and a recommendation to roll out. That's the report a new hire's one question was enough to break."
Why this works
Keeps the answer from becoming a generic list of "good report" traits with nothing real behind it.
Stage 2
State the structure out loud
Say it like this
"I'll run this as GUARD. Groups, who's affected. Unequal, where the harm actually lands. Ability to contest, who has no lever. Reduce, the concrete fix. Detect, how you'd catch it in production."
Why this works
Signals a repeatable way to audit any spike report, not a one-off complaint about this one.
Stage 3
Reframe: it isn't "is the number good," it's "who can't check the number for themselves"
Say it like this
"The question isn't whether 94 percent sounds like a good result. It's whether the person whose pay depends on SortEye's flags has any way to know when it got one wrong, or any way to say so."
Why this works
This is where a strong answer separates from someone who just says "reports should be thorough."
Stage 4
Give the reduce decision, what the report must contain
Say it like this
"Here's the fix: every spike report gets a breakdown by material category and by shift, the worst segment named with its own number, one real failure photo, and a stated review path for a worker who disagrees with a flag. Not optional sections, required ones."
Why this works
This is the direct answer, concrete enough to check any future report against.
Stage 5
Prove it with the question nobody could answer
Say it like this
"Once the team actually ran the breakdown, black plastic, about 8 percent of line volume, was right only 61 percent of the time, against 97 percent on the common, well-lit materials the headline number was quietly averaging over."
Why this works
Compresses the whole failure into the one gap a single clean number was built to hide.
Stage 6
Say the detect mechanism
Say it like this
"Past the spike, we'd track false-flag rate by material category and by shift monthly, not just watch one overall number, since that's exactly the split that was missing the first time."
Why this works
Shows the fix isn't a one-time report edit, it's an ongoing habit the first version never had.
Stage 7
Close on the one line
Say it like this
"A good spike report isn't proof the idea works. It's proof someone went looking for where it doesn't, and said what to do about the person who'll feel it first."
Why this works
Restates the direct answer in one breath, tying the whole critique back to a single sentence.
Let's learn
Say we build a camera system that watches a recycling sorting line and flags contaminated or misplaced items before a bale of material gets shipped out to a buyer.
Before SortEye, a line supervisor manually spot-checked bales for contamination, catching maybe a fraction of real issues in the time available, and buyers occasionally rejected whole bales over contamination found only after shipping.
Five things a report needs. SortEye's first draft had exactly one of them.
Here's the turn: the wrong flags on black plastic were never really the first problem. The first problem was a spike report that said "94 percent accuracy, ready for rollout" and nothing else, so nobody in the room, including Xiomara, actually knew there was a weak spot until someone happened to ask about one directly.
SortEye accuracy by material category
A 36-point gap, sitting quietly underneath the one number the first report ever printed.
At its worst, SortEye ships on the strength of one clean number, black plastic keeps getting misflagged for months, and workers whose incentive pay is tied to their line's contamination rate quietly lose money to a category nobody ever told them was the model's weak spot, with no way to ask why.
The report wasn't wrong. It just never went looking for the one number that would have changed the recommendation.
The choice I would take back
Xiomara wrote the first report around a single headline number because that's what leadership had asked to see fast, and a fast answer felt like the priority. That made sense under deadline pressure. It stopped making sense the moment the number quietly decided a rollout for a group of workers who'd never get a say in whether it was fair to them.
What I would leave alone: for a low-stakes internal dashboard showing SortEye's numbers to engineers only, one clean summary number is genuinely fine. The full breakdown matters for the report that decides whether real workers' pay gets affected.
The lesson: a spike report's job isn't to prove the number is good. It's to prove someone actually went looking for where it isn't, before a person downstream finds out the hard way.
Now here is the same thing as a story
The short version above is what you'd say defending the rewritten report template to leadership. Read this one for what the original review meeting actually felt like.
Xiomara Peralta had shipped the first version of the SortEye spike report the way she'd always seen it done: one clean number, one clear recommendation, ready to move fast.
The review meeting went smoothly for the first ten minutes. Ninety-four percent sounded strong, and the room was ready to approve a pilot rollout across two facilities.
One question, on day three, is what turned a one-page report into a real one.
Callum Reyes, three weeks into the job and still asking questions a veteran might have learned to skip, raised his hand. "What does it do on black plastic? I read that's usually hard for these cameras." Nobody in the room had an answer, because nobody had ever run that specific breakdown.
Knowledge spark: why is black plastic hard for a sorting camera?
Black or very dark plastic absorbs most visible light instead of reflecting it, so a camera has far less contrast and color information to work with than it does on a bright PET bottle or a light cardboard box, making the material genuinely harder to classify correctly.
Xiomara ran the breakdown that afternoon. Black plastic, about 8 percent of the line's total volume, was right only 61 percent of the time. The rest of the line looked fine at 97 percent. The clean 94 percent headline had simply been averaging over both.
One number, and four things quietly missing behind it. Callum's question only found the first one.
Digging further, the same gap showed up by shift: 95 percent accuracy on the well-lit day shift, 79 percent on the dimmer night shift. And nobody had ever asked what happened to a worker whose bale got wrongly downgraded by a black-plastic misflag on a night shift, at a facility where sorting accuracy affected incentive pay.
The person who reads the report and the person who lives with its mistakes were never the same person.
The real question was never whether SortEye's overall number was good enough to launch. It was whether anyone had checked which group of workers would carry the cost of the 36-point gap the headline number never showed.
Three steps happen automatically. The fourth box, where a worker could push back, was never built.
When the first report was drafted, someone said, "let's keep it simple, one number, one recommendation, leadership doesn't have time for a wall of statistics," and it sounded reasonable, since a clean summary really is easier to act on fast.
A rare, badly-wrong segment sits exactly where an average headline number is built to hide it.
Rerun the same review meeting with the required breakdown already built in: the 61 percent black-plastic number and the day-night gap surface on day one, at the cost of one extra afternoon of analysis, instead of surfacing three weeks in because a new hire happened to ask the right question.
What I'd tell myself, hearing how close that report came to shipping with nobody checking: a clean number that nobody had to work for to get is exactly the kind that hides the thing that would have changed the decision.
GUARD, the report that names who can't push backNot a script for distrusting every good result. GUARD is what tells you exactly who's downstream of the number you didn't check.
G
Groups. Who is affected?
The operator: a line supervisor who reads the report and decides on rollout. The subject: a sorting-line worker whose pay depends on flags they never see coming.
Naming both, not just the reader of the report, is what makes the question real instead of abstract.
U
Unequal. Where does the harm land unevenly?
On black-plastic-heavy loads and on the night shift specifically, where accuracy drops to 61 and 79 percent while the rest of the line sits at 95 to 97.
An average headline number is exactly what hides an unequal harm like this one.
A
Ability to contest. Who never gets to push back?
A sorting-line worker whose bale gets downgraded has no way to know a flag was wrong, or to ask for it to be rechecked, before the belt has already moved on.
This is the hardest step, and the one that turns a technical gap into a fairness problem the report has a duty to name.
R
Reduce. The specific design change.
Require a breakdown by material and shift in every spike report, name the worst segment explicitly, and add a stated review path for a worker who disagrees with a flag.
A concrete report requirement, not a policy statement about "being thorough."
D
Detect. How would you know in production?
Track false-flag rate by material category and shift monthly, the exact split the first report never ran, instead of watching one overall number.
The same gap that a new hire's question caught once needs a standing check, not a lucky question twice.
The recap, one line per letter: groups is the supervisor versus the worker, unequal is black plastic and night shift carrying the real gap, ability to contest is a worker with no lever once a bale is downgraded, reduce is a required breakdown and a review path, and detect is a monthly false-flag check by segment.
And if you want to be sure it really works, try it somewhere elseSame five letters, a public library instead of a sorting line. The missing appeal path looks almost identical.
Béatrice Toure runs cataloguing at the Penderyn Public Library Consortium, testing ShelfSense, a tool meant to flag books for restricted-access review when their content might not suit the section they're shelved in. Mapped onto GUARD: groups are librarians, who read the spike report and decide whether to trust ShelfSense's flags, and patrons, especially children, who lose access to a book with no idea it was ever flagged. Unequal is health-education and identity-related topics getting flagged roughly three times more often than other topics, a pattern the first spike report never broke out by subject. Ability to contest is a flagged book quietly missing from a search result, with most patrons never even learning a flag happened, let alone how to ask for a review. Reduce is requiring every spike report to show flag rate by topic category and a visible "why was this flagged, request a review" link on any restricted result. Detect is tracking how often a librarian's appeal review reverses a flag, which turned out to be 4 times in 10 once anyone checked.
A different shelf entirely, and the same missing piece: nobody had measured how often the flag itself was wrong.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "break the number down by segment, name the worst one, say who can't contest a wrong call," and stop.
Cost: no time for a full breakdown before a deadline. Say so honestly, and report the single most likely weak segment specifically, rather than ship a headline number with no caveat at all.
The number really does hold up everywhere, for real: if the breakdown shows every segment performing evenly, that's a genuinely stronger report for having checked, and saying so plainly is what makes the discipline worth keeping even when nothing was hiding.
Where people run it wrong.
They ship a single headline number because it's faster to write and easier to approve.
They never name who's downstream of a wrong call, only who's reading the report.
They treat "we didn't get any complaints" as proof nothing's wrong, instead of proof nobody affected had a way to complain.
How to use it live. The moment you're asked what a good spike report contains, ask yourself: who is this number going to make a real decision about, and would they be able to tell if it got them wrong? Build the report around answering that, not around looking finished fast.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a risk or fairness question, and its one-line job?
Tap to flip
ANSWER
GUARD: name who can't push back. Groups, Unequal, Ability to contest, Reduce, Detect.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Xiomara Peralta, the AI PM at Verdant Sort Systems, who rewrote SortEye's spike report after a new hire's question exposed what it was missing.
3 · THE GROUPS
Who reads a spike report, and who actually lives with its mistakes?
Tap to flip
ANSWER
A line supervisor reads it and decides on rollout. A sorting-line worker, whose pay may depend on the line's numbers, lives with any mistake it makes.
4 · WHAT WAS MISSING
Name two of the four things the first spike report never included.
Tap to flip
ANSWER
A breakdown by material category and a stated appeal path for a worker who disagrees with a flag. It also lacked a lighting/shift breakdown and any named worst-case example.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing the first report around one headline number because a fast, clean summary felt like the priority under deadline pressure.
6 · THE NUMBER
Fill in the blank: the headline accuracy was ___ percent overall, but only ___ percent on black plastic specifically.
Tap to flip
ANSWER
94 percent overall; 61 percent on black plastic.
7 · THE DETECT STEP
What ongoing check replaced the one-time report breakdown?
Tap to flip
ANSWER
A monthly tracked false-flag rate by material category and by shift, instead of watching only one overall accuracy number.
8 · CROSS PRODUCT TRANSFER
Section 4 runs GUARD again for a different product. Which one, and what's the missing appeal path there?
Tap to flip
ANSWER
Penderyn Public Library Consortium's ShelfSense. A flagged book quietly disappears from search results with no visible way for a patron to know it was flagged or ask for a review.
Check yourself Score: 0 / 0
Short answer, apply it yourself
1. Think of a report or dashboard you've seen that showed one clean overall number. What breakdown, if it existed, might have changed the decision made from it?
Show hint
Think about which segment of the underlying data was most likely to be different from the average.
Show answer
Model answer: A customer satisfaction dashboard showing one average score hid that a specific product line scored far worse; breaking it down by product would have surfaced the real issue much sooner.
Multiple choice
2. Why did SortEye's overall 94 percent accuracy number fail to reveal the black-plastic problem?
A. Black plastic wasn't actually part of the line's volume.
B. An overall average blends a large share of easy, well-lit materials with a smaller, much harder segment, so the strong majority hides the weak minority.
C. The camera hardware was broken during the spike.
D. Workers manually corrected all the black-plastic errors before they were counted.
Show hint
Look at the bar chart comparing common materials to black plastic.
Show answer
B. 97 percent on the common majority and 61 percent on a smaller, harder segment average out to a headline number that looks fine either way.
True or false
3. True or false: before Callum's question, someone at Verdant had already checked SortEye's accuracy by shift and by material type.
True
False
Show hint
Look at "now here is the same thing as a story."
Show answer
False. Nobody had run that breakdown before the review meeting. It only happened after Callum's question, the same afternoon.
Fill in the blank
4. Fill in the blank: SortEye's accuracy was ___ percent on the day shift, but only ___ percent on the dimmer night shift.
Show hint
Look at "digging further, the same gap showed up by shift."
Show answer
95 percent day shift; 79 percent night shift. A second gap the first report also never surfaced, hiding behind the same single headline number.
Short answer, where it wouldn't matter
5. Name a situation where a single clean headline number, with no further breakdown, would genuinely be enough in a spike report.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A low-stakes internal dashboard shown only to engineers for their own awareness, with no rollout decision or worker impact riding on it.
Short answer, work the number
6. If black plastic were only 2 percent of line volume instead of 8 percent, would its 61 percent accuracy still be worth naming explicitly in the report?
Show hint
Think about who is affected even when a segment's overall volume is small.
Show answer
Model answer: Yes. Even a small share of volume still means real workers handling real misflagged bales regularly; volume changes how big the problem is, not whether it deserves to be named at all.
Before you close the answer
Why this works
Tests whether you treat a spike report as proof an idea works, or as the place you go looking for exactly where it doesn't, and for whom.
Follow-up traps
"Isn't breaking every report down by segment just slower and more work?" Response: one extra afternoon of analysis, against three weeks of a wrong number quietly deciding a rollout that affects real workers' pay.
"What if the worst segment turns out fine after all this scrutiny?" Response: then the report says so plainly, with the number to prove it, which is a stronger, more trustworthy report than one that never checked at all.
If pressed
The rewritten guardrail routes every black-plastic and night-shift flag to a quick human glance before a bale gets downgraded, specifically because those two segments were the ones the breakdown showed carrying almost the entire error gap.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.