ConceptAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #12

What are the signs a spike is answering the wrong question?

TRACEthe score was fine, the harm was not

Reclaimed Goods Exchange is a marketplace for buying and selling used furniture. TrustFrame is the spike meant to flag fraudulent listings before a buyer gets scammed. Casimir Duquette is the trust-and-safety AI PM who had to figure out why a 95 percent score never once caught the fraud actually costing people money.

The direct answer
The clearest sign is a spike's own score staying flat and healthy while the real problem it's supposed to prevent keeps getting worse. That gap means the eval set and the actual failure are two different things wearing the same name. TrustFrame held 95 percent for four months while dollar losses from item swaps and misrepresented listings kept climbing, because its eval set had quietly been filtered down to easy, obvious fraud only, stolen photos and duplicates, never the harder fraud actually costing buyers money. When the metric and the real-world harm stop moving together, the spike is answering a question nobody actually asked.
Do this, in order
  1. Watch whether the spike's score and the real-world harm move together, not just whether the score is high.Why: a flat, healthy score next to a rising real problem is the single clearest sign of a mismatch.
  2. Recut the score by the type of case that actually matters, not just the overall rate.Why: an average can hold steady while one specific, costly case type gets caught almost never.
  3. Check who built the eval set, and what got filtered out before scoring even started.Why: a clean, easy eval set can hide the exact cases the spike was actually supposed to catch.
  4. Rule out a tracking or reporting change before trusting that a rising problem is real.Why: a definitions change can look exactly like a real trend on a dashboard.
  5. Run one evidence test that separates your top two suspects, not a vague "look into it."Why: a real audit of unfiltered cases, scored honestly, is what actually proves the mismatch.
  6. Say plainly when a spike's score really does track the real problem, and trust it there.Why: shows judgment about where the gap exists, instead of doubting every metric on principle.

How to answer this, stage by stage

Nobody is scoring whether you can list generic reasons a metric might be wrong. They're scoring whether you can find the one gap between what a spike measured and what it was actually supposed to catch.

Stage 1
Scope it to one real spike
Say it like this
"Let's ground this in TrustFrame at Reclaimed Goods Exchange. That's the spike where a 95 percent score sat completely still while real losses kept climbing underneath it."
Why this works
Keeps the answer from becoming a generic list of "things that can go wrong with metrics."
Stage 2
Say the structure out loud
Say it like this
"I'll run this as TRACE. Timeline, when it actually started. Recut, slice it by segment. Assume nothing, rule out tracking first. Cause candidates, three named suspects. Evidence test, the one check that decides it."
Why this works
Signals a disciplined way to diagnose a metric mismatch, instead of a hunch dressed up as an answer.
Stage 3
Reframe: it isn't "is the score high," it's "does the score track the actual harm"
Say it like this
"This isn't really a question about whether 95 percent is a good number. It's a question of whether that number has ever moved together with the real losses it's supposed to be preventing."
Why this works
This is where a strong answer separates from someone who just accepts a high score at face value.
Stage 4
Rule out the boring explanation first
Say it like this
"Before blaming the spike, I'd check whether dispute definitions or reporting categories changed at all in this window. They hadn't. The rise in disputed listings was real, not a side effect of counting things differently."
Why this works
Shows discipline: a real investigation clears the easy, boring explanation before reaching for a bigger one.
Stage 5
Name the three suspects
Say it like this
"Three possibilities: buyers just got pickier about condition, sellers changed how they photograph listings, or TrustFrame's eval set never actually tested the fraud that was rising. All three are plausible. Only one explains why the score never moved."
Why this works
Narrows the field to named, testable hypotheses instead of a vague sense that "something's off."
Stage 6
Run the evidence test
Say it like this
"I pulled a fresh, unfiltered sample of every flagged listing, not the cleaned-up eval set, and scored TrustFrame against it directly. On stolen photos and duplicates, 95 percent, matching the reported number. On item swaps and misrepresented condition, 11 percent."
Why this works
This is the single strongest move in the whole method: one test that separates the real cause from the other two suspects cleanly.
Stage 7
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't a generic reporting problem is that a model's score only ever reflects the eval set it was measured against, never the full space of real cases. We accepted a lower headline number, and a slower path to a rebuilt eval set, in exchange for a score that finally means what everyone assumed it already meant."
Why this works
This is the load-bearing, AI-specific judgment: a model's accuracy is only ever accurate about the data it was actually tested on.
Stage 8
Close on the one line
Say it like this
"A spike answering the wrong question looks fine on paper the entire time. The sign to watch for isn't a bad score, it's a good score that stops explaining anything happening in the real world."
Why this works
Restates the direct answer and leaves the interviewer with the exact test to apply to any other metric they hear.

Let's learn

Here is what happens when a tool gets very good at answering a question nobody actually asked.

Before TrustFrame, trust-and-safety analysts manually reviewed every flagged furniture listing, catching most obvious fakes, stolen stock photos, duplicate posts, by eye. With TrustFrame, a flagged listing gets scored in under a second, and analysts were told to trust a clean score and focus their time elsewhere.

Hand sketched icon list titled TRACE the five letters. Five rows: Timeline when it actually started. Recut slice it by segment. Assume nothing rule out tracking first. Cause candidates three named suspects, shown in a different color. Evidence test the one check that decides it.
The five letters, held up as one page. Cause candidates is the step that stops a hunch from becoming an answer too early.

Here's the turn: TrustFrame's eval set was built by pulling a sample of flagged listings and quietly filtering out the messy, disputed ones to make labeling faster. What was left was mostly obvious fraud, stolen photos and duplicate posts, the exact cases already easiest to catch. The 95 percent score was real. It was also answering a much smaller question than anyone realized.

Hand sketched comparison titled What we tested, versus what we needed. Left panel, a document icon labeled What we tested, caption can it spot already obvious fakes. Right panel, a question mark box icon labeled What we needed, caption can it catch fraud that's actually hard to see.
Two very different questions, and only one of them ever got measured.

At its worst, a genuinely fraudulent, high-value listing sails through untouched, month after month, while the dashboard everyone trusts keeps saying 95 percent, right up until a buyer catches it by luck.

TrustFrame's recall, by fraud type
100% 50% 0 97% Stolen photo 94% Duplicate listing 34% Misrepresented 11% Item swapped
The overall blended number stayed near 95 percent because the easy categories carried it. The two categories driving real disputes barely got caught at all.
The score was never lying. It was just telling the truth about a question much smaller than the one everyone thought it answered.
The choice I would take back Once the eval set's filtering rule, drop messy or disputed listings before scoring, was baked into how listings got pulled for review, the only way to fix its blind spot was to rebuild the whole pipeline from scratch. There was no way to patch just the sanitization step. That rule made sense when the team needed a fast, clean way to measure early progress. It stopped making sense the moment it quietly became the permanent definition of what TrustFrame was tested against.

What I would leave alone: for stolen stock photos and duplicate listings, the exact cases the eval set was built around, TrustFrame's 95 percent was, and still is, an honest, trustworthy number.

The lesson: a spike doesn't have to be wrong to be dangerous. It just has to be very good at answering a question that quietly stopped being the one that mattered.

Now here is the same thing as a story

The short version above is what you'd say defending an audit finding in a leadership review. Read this one for what it felt like the sixteen weeks before a near miss forced the question.

The headset sits on Casimir's desk between shifts, still warm from the last review call.

Hand sketched flow diagram titled How a listing reached the eval set, third step emphasized. Five steps left to right: Listing flagged for review. Analyst pulls a sample. Messy listings filtered out, shown in a different color. Clean sample scored. 95 percent reported up.
The third step is where the eval set quietly stopped representing the fraud that actually mattered.

TrustFrame launched at 95 percent, and the number held rock steady for months. Nobody complained. The dashboard looked, by every visible measure, like a genuine success.

Hand sketched timeline titled What shipped, and when the real losses moved, second milestone emphasized. Four milestones: Spike hits 95 percent on eval set, week 1. Filter rule added quietly, week 3, shown in a different color. Disputed listings keep rising, week 10. Near miss on a 4,200 dollar sale, week 16.
Week three is where a quiet labeling shortcut became the permanent shape of the eval set.

Underneath that steady number, disputed listings involving item swaps and misrepresented condition kept climbing, quarter over quarter, while TrustFrame's own reported accuracy never budged.

A flat score, next to a rising real problem
high low reported accuracy, flat real weekly losses, rising
One line never moved. The other one kept climbing underneath it, unwatched.

When disputes were first noticed rising, someone said, "the model's accuracy hasn't changed, it's probably just buyers getting fussier," and it sounded reasonable, since the dashboard genuinely showed nothing wrong.

Knowledge spark: why would a model's accuracy stay flat while real fraud losses rise? A model's reported accuracy only ever describes how it performs against the specific eval set it was scored on. If that eval set quietly stopped representing the real mix of cases, the score can stay perfectly steady while the model's performance on the cases that matter most silently stays terrible the whole time.

A near miss finally forced the question: a $4,200 dining set swap at pickup almost got approved as a legitimate return, caught only because a buyer happened to photograph a serial number sticker that didn't match the listing.

Hand sketched labeled parts diagram titled The three suspects. A question mark box icon at the center labeled Why Disputes Rose, with three labeled callouts around it: Buyers just got pickier, Sellers changed how they photograph, The eval never tested real fraud.
Three named suspects, and only one of them explains why the score itself never moved.

Casimir pulled a fresh, unfiltered sample of every flagged listing from the last month, not the cleaned-up eval set, and scored TrustFrame against it directly.

Hand sketched quadrant titled Where TrustFrame actually catches fraud. X axis how obvious the fraud is, obvious to subtle. Y axis revenue at risk, small to large. Items placed: Stolen stock photo obvious, small risk. Duplicate listing obvious, small risk. Swapped item at pickup subtle, large risk. Misrepresented condition subtle, large risk.
TrustFrame's whole strength sat in the bottom-left corner, exactly where the real financial risk did not.

The real question was never whether TrustFrame was 95 percent accurate. It was whether that 95 percent had ever been measured against the fraud that was actually costing anyone money.

Rerun the same sixteen weeks with an eval set built from a genuinely unfiltered sample, disputed and messy cases included: the item-swap blind spot shows up in week one, not week sixteen, and the near miss on that $4,200 dining set never gets close to shipping at all.

What I'd tell myself, hearing about that near miss: 95 percent was never the number that mattered. The number that mattered was 11, and it had been sitting there, unmeasured, since week three.

TRACE, the audit that found the real questionNot a script for distrusting every clean-looking score. TRACE is what tells you exactly which real-world number it was never actually tracking.

T
Timeline. When exactly did it start?
The filtering rule that shaped TrustFrame's eval set was quietly added in week three, three weeks before disputed listings began their steady climb.
Starting the clock at the filter, not at the first complaint, is what makes the rest of the trace make sense.
R
Recut. Slice it by segment.
Overall recall looked fine at 95 percent. Recut by fraud type, stolen photos and duplicates carried that average while item swaps and misrepresented condition sat at 11 and 34 percent.
The average was hiding exactly the two categories that mattered most.
A
Assume nothing. Rule out instrumentation first.
Dispute definitions and reporting categories had not changed in this window, so the rise in disputed listings was a real trend, not a side effect of counting things differently.
Clearing the boring explanation first is what makes the real finding credible instead of a guess.
C
Cause candidates. Three named suspects.
Buyers got pickier about condition, sellers changed how they photograph listings, or TrustFrame's eval set never tested the fraud types actually rising.
Naming three specific, testable hypotheses instead of one vague hunch is what makes the next step possible.
E
Evidence test. The one check that separates the top hypotheses.
Scoring TrustFrame against a fresh, unfiltered sample of every flagged listing showed 95 percent on the easy categories and 11 percent on item swaps, matching the eval-set explanation exactly and ruling out the other two.
This is the hardest step, and the strongest move in the whole method: a single test that decides the question, not a debate about which suspect feels most likely.

The recap, one line per letter: timeline is the filter rule added in week three, recut is the fraud-type breakdown hiding under one clean average, assume nothing is ruling out a reporting change, cause candidates is buyers, sellers, or the eval set itself, and evidence test is the unfiltered audit that pinned it on the eval set.

And if you want to be sure it really works, try it somewhere elseSame five letters, a ticketing platform instead of a furniture marketplace. The suspect list changes entirely, the audit step doesn't.

Yara Bergström-Adeyemi runs product at GateFlow Ticketing, where ScalperScan flags likely bot-driven resale listings for a review queue. As ticket volume grew, analysts quietly started reviewing only small sample batches from that queue instead of every flagged listing, since the queue had outgrown what anyone could check by hand. Mapped onto TRACE: timeline is the week ticket volume crossed a threshold analysts could no longer keep up with. Recut by ticket category shows sports and concert resales still get reviewed at the old rate, while a newer theater-ticket category gets almost no review at all. Assume nothing rules out a change in how bot listings get flagged in the first place. Cause candidates are more bots overall, a change in bot behavior, or reduced review coverage. Evidence test is comparing ScalperScan's own flag rate against a fully staffed, one-week manual audit of the theater-ticket category specifically, which is what actually confirms coverage, not bot behavior, had quietly shrunk.

Hand sketched comparison titled How many tickets get a human look. Left panel, a person icon labeled Before, caption every flagged reseller checked. Right panel, a person icon labeled After, caption only small batches, as volume grew.
A different flip than TrustFrame's: not a filtered eval set, but a review process that quietly shrank as volume outgrew it.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "check whether the score and the real harm actually move together, not just whether the score is high," and stop.
Cost: there's no time to rebuild a full eval set before a launch review. Say so honestly, and run a smaller, honest audit sample instead of trusting a filtered eval set's number as if it were complete.
The real problem actually improves on its own, for real: if disputed listings had fallen instead of risen, that's still worth checking against an unfiltered sample, since a lucky improvement can hide the same blind spot just as easily as a rising problem reveals it.

Where people run it wrong.
They trust a clean, high score without ever checking who built the eval set behind it.
They treat a flat metric as proof nothing is wrong, instead of asking what it was actually measured against.
They jump straight to a cause without first ruling out a boring change in how something got counted.

How to use it live. The moment a spike's score looks suspiciously clean, ask yourself: has this number ever moved together with the real-world problem it claims to track. If it hasn't, go find out what it was actually built to measure instead.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Pre-editing flip: the eval set quietly filtered out messy, disputed listings before scoring, leaving only easy, obvious fraud for TrustFrame to be judged against.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Casimir Duquette, the trust-and-safety AI PM at Reclaimed Goods Exchange, who audited TrustFrame after a near miss on a high-value listing.
3 · THE HABIT
What did analysts stop doing once TrustFrame's score looked consistently healthy?
Tap to flip
ANSWER
They stopped manually spot-checking listings outside the categories TrustFrame's eval set had actually been built around, trusting the flat 95 percent as if it covered everything.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Labeling from a full, messy sample of real flagged listings, versus quietly filtering out the disputed ones to make labeling faster. No in-between once the filter rule got baked into the pipeline.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Filtering messy, disputed listings out of the eval set to make labeling faster, with no way afterward to patch just that one rule without rebuilding the whole pipeline.
6 · THE NUMBER
Fill in the blank: TrustFrame's reported accuracy held at 95 percent, but its recall on item-swap fraud specifically was only ___ percent.
Tap to flip
ANSWER
11 percent.
7 · THE REPLAY
Same sixteen weeks, with an eval set built from a genuinely unfiltered sample from day one. What changes?
Tap to flip
ANSWER
The item-swap blind spot shows up in week one instead of week sixteen, and the near miss on the $4,200 dining set never gets close to shipping.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
GateFlow Ticketing's ScalperScan. The flip is scope: analysts started reviewing only small sample batches from the flagged queue once volume outgrew what they could check by hand.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: the clearest sign a spike is answering the wrong question is a score that stays flat and healthy while the real ___ it's supposed to prevent keeps getting worse.
Show hint
Look at the direct answer.
Show answer
Harm. When a metric and the actual problem stop moving together, the metric has quietly stopped measuring the thing that matters.
Multiple choice
2. Why did TrustFrame's 95 percent score fail to predict the near miss on the $4,200 dining set?
  • A. The model's accuracy had actually dropped, but nobody had checked the dashboard recently.
  • B. The eval set behind that 95 percent had been filtered to easy, obvious fraud, and never tested item swaps or misrepresented condition at all.
  • C. The buyer who caught the swap was an unusually experienced furniture collector.
  • D. TrustFrame was never actually deployed to real listings.
Show hint
Look at the grouped bar chart, "TrustFrame's recall, by fraud type."
Show answer
B. A score only ever describes performance against its own eval set, and that eval set had quietly stopped representing the fraud types that mattered most.
True or false
3. True or false: the rise in disputed listings turned out to be caused by a change in how disputes were counted, not a real increase in fraud going uncaught.
  • True
  • False
Show hint
Look at the A step, "assume nothing," in the TRACE recap.
Show answer
False. Dispute definitions hadn't changed. The rise was real, which is exactly why it needed a real cause, not a reporting explanation.
Short answer, where it wouldn't matter
4. Name a fraud category in this same tool where the 95 percent score really was trustworthy, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Stolen stock photos and duplicate listings, the exact cases the eval set was built around. TrustFrame's 95 percent there was, and still is, an honest number.
Short answer, apply it yourself
5. Think of a metric you've seen reported as healthy while a real problem kept getting worse underneath it. What might that metric have quietly stopped measuring?
Show hint
Think about who built the metric and what got left out of it.
Show answer
Model answer: A customer support team's "average resolution time" staying flat while complaints about unresolved issues grew, because the metric only counted tickets that got formally closed, not the ones quietly abandoned.
Short answer, work the number
6. If item-swap fraud made up 25 percent of all flagged listings instead of a small slice, would the overall 95 percent blended score still look believable?
Show hint
Think about how a blended average shifts as the mix of case types changes.
Show answer
Model answer: No. With item swaps at only 11 percent recall making up a quarter of all cases instead of a small share, the blended score would drop well below 95, and the mismatch would likely have surfaced far sooner.
Before you close the answer
Why this works
Tests whether you'll accept a healthy-looking score at face value, or check whether it has ever actually moved together with the real problem it claims to track.
Follow-up traps
"Couldn't the rise in disputes just mean buyers got pickier, like the first suspect?" Response: that's exactly why the evidence test mattered, an unfiltered audit showed TrustFrame's own recall on those specific fraud types was 11 percent, which a change in buyer pickiness alone can't explain.

"Isn't rebuilding the whole eval set an overreaction to one near miss?" Response: no, the near miss was just the moment anyone finally looked. The 11 percent recall had been true since week three, whether or not that dining set got caught.
If pressed
The rebuilt eval set now stratifies by fraud type deliberately, holding a fixed minimum share of item-swap and misrepresented-condition cases, so no future filtering step can quietly shrink coverage of the hardest, most expensive fraud again without it showing up immediately in the reported number.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more