Artifact critiqueAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #19
Describe the decision criteria you would set before the spike begins.
LEADbefore Petra started sorting the hard ones herself
Cobalt Ledger sells expense software to mid-size companies. SnapAudit is its feature that reads a photographed receipt and files it into the right expense category. Marisol Deschamps is the AI PM who ran the spike that decided whether SnapAudit shipped. Renske Bakhuizen runs the finance desk that has to live with what it got wrong.
The direct answer
Before the spike runs, write down three sentences: which real, unfiltered slice of production data it will be graded on, the exact number it has to clear on that data, and what a miss costs in the worst realistic case, not the average one. Put a name next to who can say no even if the demo looks great. Do this before anyone sees a result, or the team will quietly grade the spike against whatever number it happens to hit.
Do this, in order
Write the pass bar and the data source down before the spike runs.Why: a number chosen after seeing the result was never really a bar, it was a description.
Build the eval set from a full, real slice of production data, not a hand-picked sample.Why: a curated sample only tests the cases that were already easy to get right.
Name the cost of a miss in the worst realistic case, not the typical one.Why: a wrongly flagged coffee receipt and a wrongly flagged HR-sensitive one are not the same mistake.
Name who signs off on pass or fail, and get their agreement on the bar first.Why: a bar with no owner gets renegotiated by whoever is in the room when the number lands.
Watch a leading signal, like how often a person quietly overrides it, not just the headline accuracy.Why: the headline number can hold steady while a person is silently doing the hard part by hand.
Say plainly when a lighter bar is fine, like a purely internal draft nobody acts on without a human reading it first.Why: shows judgment about where the extra rigor earns its cost, not fear applied everywhere.
How to answer this, stage by stage
Nobody is scoring whether you can name a good accuracy number. They're scoring whether you'd write the bar down before you're standing in front of a result you already like.
Stage 1
Scope it to one real spike
Say it like this
"Let's ground this in SnapAudit, Cobalt Ledger's receipt categorizer, and the exact spike that decided whether it shipped."
Why this works
Keeps the answer from turning into a generic list of best practices with nothing real behind it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as LEAD. Link it to the real outcome we're chasing. Find the early signal that moves first. Say how it gets gamed. Say the decision at each threshold."
Why this works
Signals a repeatable way to set criteria, not a one-off gut call about one number.
Stage 3
Reframe: it isn't "pick a good accuracy number," it's "decide what that number has to be true on"
Say it like this
"This isn't really a question of picking one impressive accuracy figure. It's a question of what data that figure has to be measured against, before anyone's excited about it."
Why this works
This is where a strong answer separates from someone who just says "set a high bar."
Stage 4
Give the criteria, concretely
Say it like this
"Here's what I'd write down before the spike starts: test on a full month of real receipts, not sixty hand-picked ones. Clear 90 percent on that real set, not the curated one. And say out loud that a miss on an HR-sensitive vendor line costs more than a miss on a coffee receipt, so the bar can't just be one flat number."
Why this works
This is the direct answer to what actually goes wrong, stated as something you could write on a whiteboard.
Stage 5
Prove it with the compressed failure
Say it like this
"We skipped that step here. The spike ran on sixty clean receipts, hit 97 percent, and shipped on the strength of that one number. Eight weeks later, the real accuracy across a full month of actual receipts turned out to be 71 percent, and Petra had already quietly started sorting a quarter of them by hand herself."
Why this works
Compresses the whole case into the one gap that a written bar would have caught before launch.
Stage 6
Name the AI-specific reasoning, then close
Say it like this
"The honest reason this isn't a generic launch-checklist question is that a model's number is only ever a number about the data you ran it on, and a curated sample and real production traffic are not the same data. I wouldn't demand this rigor for a purely internal draft nobody acts on unsupervised. For SnapAudit, the bar has to exist in writing before the number does, or the number will quietly write its own bar."
Why this works
Closes with real judgment about where the concern doesn't apply, and restates the direct answer in one breath.
Let's learn
Every month, about 600 employee expense receipts landed on Cobalt Ledger's own finance desk. Someone in accounts payable read each one and typed in a category by hand: travel, meals, supplies, or one of a dozen others. Three minutes a receipt, thirty hours a month, done the same way for years.
With SnapAudit, a photographed receipt gets a category back in under two seconds. Marisol's team ran a first spike on sixty receipts, hand-picked from the tidiest folder in the archive, and it hit 97 percent. The team was thrilled, and nobody had written down what number, on what data, would have counted as a real pass.
The four letters, held up as one page. Early signal is the step this question is really testing.
Here's the turn: the spike's 97 percent was never wrong, exactly. It was just a fact about sixty receipts nobody had disagreed about the category of. The real 600-a-month distribution held handwritten tips, foreign currency, split bills, and vendor names that could mean two different things depending on the trip. None of that showed up in the curated sixty.
Petra's manual override rate, week by week
Nobody was watching this line. The dashboard everyone watched was the accuracy score, which barely moved.
At its worst, an employee's legitimate travel dinner gets mis-coded as a policy violation, a manager gets an automated flag about it before anyone checks, and the tool that was supposed to save thirty hours a month has quietly cost someone their afternoon and their patience.
The 97 percent was never a lie. It was just a true fact about the wrong sixty receipts.
The choice I would take back
The team agreed, informally, in a Slack thread, that whatever number the first spike produced would be good enough, since sixty curated receipts felt like a generous sample. That made sense when nobody wanted to slow the team down over a demo. It stopped making sense the moment that number became the entire reason SnapAudit shipped unsupervised.
What I would leave alone: I wouldn't demand a full production-data eval for an internal-only draft feature that a person already reads and edits before anything reaches a customer, since a curated spike is plenty to greenlight more building there.
The lesson: a spike's number is a fact about the data it ran on, not a fact about the feature. Decide the data and the bar before you see the number, or the number will happily decide them for you.
Now here is the same thing as a story
The short version above is what you'd say defending a spike plan in a review meeting. Read this one for what it felt like the eight weeks nobody wrote the bar down at all.
Renske Bakhuizen can smell a mis-coded travel expense before she's finished reading the vendor line. Nine years on Cobalt Ledger's own finance desk will do that.
The third step is where a curated spike's promise either holds up on a real receipt or quietly doesn't.
When SnapAudit launched, Petra loved it. A photo came in, a category came back, and for the first few weeks she checked a handful at random just to see it behave. It always did. She let it run.
Same word, test set, and only one of the two versions ever touched a hard case.
Then came a run of vendor names that could be either a client dinner or a personal one, and a foreign hotel bill split three ways. SnapAudit filed them fast and filed some of them wrong. Petra didn't complain. She just quietly started re-checking that category by hand before it posted.
Knowledge spark: why would a curated sample and real data give two different numbers for the same model?
A curated sample tends to hold the cases someone already agreed on, the clean ones. Real production data holds every ambiguous, foreign, split, and handwritten case that never made it into anyone's tidy folder. The model didn't change between the two. The data it was actually tested on did.
Four things a written bar needs, and the spike that shipped SnapAudit had none of them on paper.
By week eight, Petra was quietly re-doing about a quarter of SnapAudit's work by hand, and nobody upstream knew it, because the dashboard everyone watched still said the model was 97 percent accurate. Then a legitimate client dinner got coded as a personal expense on an employee already under a performance review, and it nearly reached that employee's manager as a flag before Petra caught it, purely because she happened to be looking that morning.
The number everyone watched sat in the obvious corner. The number that actually mattered sat in the hidden one.
The real question was never whether SnapAudit's model was good. It was whether anyone had ever written down, before the spike ran, what data that goodness had to survive.
When the spike was first proposed, someone said, "let's just run it on whatever receipts we've got handy and see how it does," and it sounded reasonable, since nobody expected a small sample to lie.
Accuracy, curated sample versus full real distribution
Same model, same week it launched. Only the data underneath the number changed.
Rerun the same spike with the bar written down first: test on a full real month, require 90 percent there, and name the cost of an HR-sensitive miss out loud. The 71 percent shows up before launch, not eight weeks after it, and the fix, a stronger rule for ambiguous vendor names, ships before Petra ever needs to quietly cover for it.
What I'd tell myself, hearing how close that flag came to reaching a manager: sixty clean receipts were never a test. They were a compliment the model paid itself.
LEAD, the bar that has to exist before the number doesNot a script for distrusting every good spike result. LEAD is what tells you exactly which data that result is actually describing.
L
Link. The real outcome, not the model's score.
Thirty hours a month of finance-desk time actually saved, without creating new manual work somewhere else.
Not "97 percent accurate," which is a fact about a sample, not about anyone's month.
E
Early signal. What moves weeks before the outcome does.
Petra's manual override rate, climbing from 4 percent to 26 percent over eight weeks while the headline accuracy score barely moved at all.
This is the hardest step, and the one the written criteria exist to force someone to watch.
A
Abuse. How this metric gets gamed.
A curated, hand-picked eval sample quietly inflates the score, without anyone intending to cheat, just by leaving out every case that was ever hard to agree on.
Nobody lied. The sample did the lying for them.
D
Decision. What you'd do at each threshold.
Below 90 percent on a real, full sample: don't ship unsupervised. Above it: ship, but keep watching the override rate as the real pass/fail signal, not the launch-day score.
A metric with no attached decision is a number on a dashboard nobody acts on.
The recap, one line per letter: link is the thirty hours of finance-desk time actually being saved, early signal is Petra's override rate climbing while the accuracy score stayed flat, abuse is a curated sample quietly inflating the launch number, and decision is a written 90 percent bar on real data, watched by the override rate afterward, not the launch-day score.
And if you want to be sure it really works, try it somewhere elseSame four letters, an event-ticketing marketplace instead of an expense desk. Different flip family entirely, the same missing bar.
Conrad Aldous runs trust and safety product at Foxglove Ticketing, where a model flags likely fraudulent resale listings for a review team to confirm. The spike ran on a curated batch of obvious fakes and hit a strong score. Mapped onto LEAD: link is real chargebacks avoided, not the flag rate itself. Early signal here is a concealment flip, not an override rate: once a review analyst's flagged case got auto-closed the moment a form was filed, analysts stopped bothering to double check a flag that looked routine, and the form-filed count looked healthy while a slower, subtler class of fake listing kept slipping through unflagged. Abuse is the same shape as Cobalt Ledger's: the curated batch of fakes was all obvious ones, so the score never touched the subtle cases. Decision is requiring the eval set to include listings a human analyst originally disagreed about, not just ones everyone already agreed were fake.
A different flip entirely: not a person doing more manual work quietly, but a person doing less checking, quietly, because a form made it look done.
The gap between the launch number and the real number was never a surprise. It was just never measured until the near miss forced it.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "write the pass bar and the real data source down before the spike runs, not after," and stop.
Cost: no time to build a full production eval set before a deadline. Say so honestly, and pull even 150 real, unfiltered receipts instead of relying on sixty curated ones.
The spike result is a genuine improvement, for real: if a later model scores even higher on the real set, that's worth trusting more, since the bar was always about the data, not about being suspicious of good news.
Where people run it wrong.
They agree on a pass bar informally, after the number is already sitting in front of them.
They build the eval set from whatever data was easiest to gather, not the data the feature will actually face.
They watch the launch-day score forever and never build a second signal for what happens after.
How to use it live. The moment an interviewer asks what criteria you'd set before a spike, ask yourself: what real data will this be graded on, what number does it have to clear there, and what does a miss cost in the worst case. Say those three out loud, in that order, and the rest of the answer follows on its own.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Pre-editing flip: once trust cracked, Petra began quietly cleaning the hard cases out before SnapAudit ever saw them, instead of letting the tool's real failure rate show.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Deschamps, the AI PM at Cobalt Ledger, who ran the spike that shipped SnapAudit on a curated sample's score.
3 · THE HABIT
What did Petra stop doing, in the first weeks after launch?
Tap to flip
ANSWER
She stopped fully re-checking SnapAudit's categorizations, spot-checking a handful at random instead, since it kept behaving well on the cases she happened to check.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Feeding SnapAudit the real, messy receipt versus quietly re-sorting the hard ones by hand before they ever reached it. No middle setting once Petra stopped trusting the ambiguous cases.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Agreeing informally that whatever number the first spike produced would count as a pass, instead of writing the real bar down before running it.
6 · THE NUMBER
Fill in the blank: the curated spike scored ___ percent, and the real, full distribution later scored ___ percent.
Tap to flip
ANSWER
97 percent on the curated sixty, 71 percent on the real six hundred.
7 · THE REPLAY
Same spike, written criteria in place first. What changes?
Tap to flip
ANSWER
The 71 percent real-data gap shows up before launch, not eight weeks after. A fix for ambiguous vendor names ships before Petra ever has to quietly cover for it by hand.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Foxglove Ticketing's fraud-listing flag. The flip is concealment: analysts stopped double-checking flags once a filed form made a case look closed.
Check yourself Score: 0 / 0
True or false
1. True or false: the model itself got worse between the curated spike and the real-data measurement eight weeks later.
True
False
Show hint
Look at "the choice I would take back" and what actually changed.
Show answer
False. The model never changed. Only the data it was actually being measured against did, from a curated sixty to the real, messy six hundred.
Fill in the blank
2. Fill in the blank: before a spike runs, write down the real data it will be tested on, the ___ it must clear on that data, and the ___ of a miss in the worst case.
Show hint
Look at the direct answer.
Show answer
The number (pass bar); the cost. All three have to exist in writing before anyone sees a result, or the bar gets set by whatever the result happens to be.
Multiple choice
3. Why did SnapAudit's real accuracy come in so far below its spike score?
A. The model was retrained between the spike and launch.
B. The spike's sixty receipts were hand-picked and clean, while real production receipts included ambiguous vendors, foreign currency, and split bills the curated set never had.
C. Employees started submitting receipts less carefully after launch.
D. Petra stopped checking anything at all after week one.
Show hint
Look at the knowledge spark about curated samples versus real data.
Show answer
B. The curated sample only ever held the easy, already-agreed cases. The real data held every messy one that never made it into that folder.
Short answer, where it wouldn't matter
4. Name a situation where a quick, curated-sample spike would be reasonable, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: An internal-only draft feature that a person already reads and edits before anything reaches a customer. A curated spike is plenty when a human is checking every output anyway.
Short answer, apply it yourself
5. Think of a time a tool worked great in a demo and then struggled once real, everyday use began. What would a written bar, agreed before the demo, have needed to include to catch that gap?
Show hint
Think about what made the demo data different from your actual daily use.
Show answer
Model answer: A note-taking app that summarized tidy typed notes beautifully in the demo, then struggled with real handwritten, messy notes. A written bar would have required testing on a sample of your own actual notes, not the vendor's clean example set.
Short answer, work the number
6. If the real-data accuracy had come in at 88 percent instead of 71 percent, would the same written 90-percent bar still have caught the problem before launch?
Show hint
Look at the decision step, and what the 90 percent bar was actually set to require.
Show answer
Model answer: Yes. Even 88 percent falls short of a 90 percent bar measured on real data, so the spike would still have failed and stayed unsupervised until fixed, just by a smaller, less dramatic margin than the actual 71 percent gap.
Before you close the answer
Why this works
Tests whether you'll set a real, checkable bar before you're standing in front of a number you already like, or let a good-looking result quietly write its own definition of success.
Follow-up traps
"Isn't this just generic project planning, writing requirements down first?" Response: no, because the specific risk is that a model's number is only ever true of the data it ran on, and a curated sample can look perfect while quietly never touching the cases that will actually break it in production.
"What if you don't have enough real data yet to build a proper eval set?" Response: then say so honestly and use the largest real, unfiltered sample you can get, even 150 receipts, rather than defaulting to a small curated one and calling it enough.
If pressed
The real fix that followed this incident split the pass bar by category risk: 95 percent required on routine categories like office supplies, but 99 percent on categories tied to HR or compliance flags, since a miss there costs far more than a miscoded coffee receipt.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.