ConceptIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #3

What question should a feasibility spike answer, and what should it explicitly not try to answer?

PICKa field Tomasz only walked out of habit

Solmark Agritech is building CropSentry, which reads a phone photo of a tomato leaf and flags likely disease. Solveig Amundsen is the AI PM deciding what a two-week feasibility spike is actually for, after a six-week attempt to prove it works on everything at once produced no answer at all.

The direct answer
A feasibility spike should answer exactly one question: can the model tell apart the two things that look alike and cost real money when confused, under the real conditions it will actually see, not a curated best case. For CropSentry, that question was early blight versus a harmless leaf spot, tested in real field lighting. It should explicitly not try to answer whether the model works on every crop Solmark eventually wants to support, whether the interface is easy to use, or whether it plugs into the farm's other systems. Those are real questions. They are just not the one that decides whether to keep building at all.
Do this, in order
  1. Name the one confusable pair that costs real money if missed, and test only that first.Why: an average across easy and hard cases hides exactly the one that matters.
  2. Test it under the real, messy conditions it'll actually be used in, not a clean sample.Why: a demo photo taken in good light proves nothing about a photo taken in harsh midday sun.
  3. Write down, in plain words, what the spike will not try to answer.Why: a spike with no stated boundary quietly grows until it's the whole product.
  4. Set the kill bar around the costly error, not the cheap one.Why: a missed disease and an unnecessary spray are not the same mistake, and shouldn't be measured as if they were.
  5. Say what evidence would change your answer, before you see any result.Why: deciding the bar after the number exists means the number always looks like enough.

How to answer this, stage by stage

Nobody is scoring whether you can list every possible thing a spike could test. They're scoring whether you can say, in one sentence, which single thing it actually has to.

Stage 1
Scope it to one real decision
Say it like this
"Let's ground this in CropSentry at Solmark, where a spike that tried to answer everything ended up answering nothing for six weeks."
Why this works
Keeps the answer from turning into an abstract list of things a spike could theoretically cover.
Stage 2
State the structure out loud
Say it like this
"I'll run this as PICK. Position, the one thing the spike will answer, stated before any reasoning. Impact, who feels each kind of wrong answer. Cost asymmetry, which error is the expensive one. Kill criteria, what would change my mind."
Why this works
Signals a repeatable way to scope any spike's question, not a one-off judgment call about tomatoes.
Stage 3
Reframe: it isn't "does it work," it's "which wrong answer costs the most"
Say it like this
"The question was never whether CropSentry works in general. It was whether it can tell a disease that wipes out a field apart from a leaf spot that doesn't matter, under the exact lighting a field scout actually shoots in."
Why this works
This is where a strong answer separates from a spike plan that just lists features to test.
Stage 4
Give the position, what the spike will and won't answer
Say it like this
"Here's the position: two weeks, one confusable pair, blight versus leaf spot, tested across four real lighting conditions. It will not touch the other eleven crops, the app's screens, or the farm records integration. Those come later, if this passes."
Why this works
This is the direct answer, narrow enough that someone could argue with it specifically.
Stage 5
Prove it with the near miss
Say it like this
"An agronomist double-checking a low-risk flag out of habit found early blight labeled 'harmless leaf spot' with high confidence. Under harsh midday sun specifically, the same confusable pair only separated correctly 54 percent of the time, against 91 percent overall."
Why this works
Compresses the whole argument into the one number the broad, unscoped spike never found in six weeks.
Stage 6
Say the kill criteria
Say it like this
"The bar was catching at least 85 percent of blight cases correctly, in every lighting condition tested, not just on average. Midday sun came in at 54, so the spike killed the plan to ship as-is and sent the model back for retraining on harsh-light photos specifically."
Why this works
Shows the bar was set before the result existed, which is what makes a kill decision credible instead of convenient.
Stage 7
Close on the one line
Say it like this
"A feasibility spike answers one question about the one confusion that actually costs money. Everything else, on purpose, waits."
Why this works
Restates the direct answer in one breath, tying the whole plan back to a single sentence.

Let's learn

Every day, Tomasz Wysocki walks about three acres of tomato fields on foot, spending roughly 25 minutes an acre checking leaves that look off, mostly telling early blight apart from a common, harmless leaf spot that looks almost identical in its first days.

Say we build a tool that reads a phone photo of a leaf and flags likely disease in seconds, instead of a scout walking every row by eye.

Hand sketched flow diagram titled The near miss, step by step, fourth step emphasized. Five steps left to right: Photo taken, Flagged low risk, Normally trusted, Tomasz checks anyway, Blight caught by luck.
Step four only happened because Tomasz still walks fields out of old habit. Nothing in the system made it happen.

Here's the turn: the extra false flags on hard cases were never the real problem. The real problem was that Solmark's first spike tried to prove CropSentry worked across all 12 crops it might eventually support, and six weeks in, still had no answer, while the one confusion that could actually cost a field, blight versus a harmless look-alike, had never been tested on its own.

Cost of each kind of wrong answer, per acre
$9,000 $4,500 $0 $140 Unnecessary spray (false positive) $8,400 Missed blight (false negative)
Sixty times the cost, and both errors look identical on a dashboard that only tracks overall accuracy.

At its worst, CropSentry ships trained mostly on well-lit sample photos, quietly misreads a real blight case as harmless under ordinary midday glare, and by the time symptoms are obvious enough for a human to catch by eye, a whole section of the field is already lost.

The spike was never short on time. It was long on questions it was never going to be the right tool to answer.
The choice I would take back Solveig let the first spike's scope grow to "prove CropSentry works" in general, since that felt like the honest, thorough version of the question. It made sense when the goal was reassuring stakeholders broadly. It stopped making sense the moment six weeks passed with no answer to the one question that actually decided anything.

What I would leave alone: whether CropSentry eventually supports all 12 crops is a real question, just not this spike's question. Answering it here would have meant never finishing at all.

The lesson: a spike that tries to prove everything proves nothing on any deadline. A spike scoped to one costly confusion proves something in two weeks.

Now here is the same thing as a story

The short version above is what you'd say defending the narrower scope in a review. Read this one for what almost missing the near miss actually felt like.

Solveig Amundsen had run three model launches before CropSentry, each time waiting for a clean, unambiguous "it works" before greenlighting a build.

The first spike attempt was ambitious on purpose: test all 12 crops Solmark eventually wanted to support, every growth stage, every lighting condition, before saying anything definitive. Six weeks in, the team had covered about 40 percent of that plan, with a growing spreadsheet of partial results and no clear answer for leadership.

Hand sketched quadrant titled Which confusion is worth a spike. Axes, how visually confusable versus cost if confused. Items placed: blight versus leaf spot in the far corner, look alike and a lost field. Healthy versus sunburned leaf near obvious and cheap. Wilt versus drought stress and pest holes versus hail damage in the middle.
Only one confusion sat in the corner that actually mattered. The six-week plan was testing all four equally.

Tomasz, an agronomist who still walks fields out of old habit even with CropSentry running, pulled up a batch of 50 leaves the tool had flagged "low risk" and checked them by eye anyway, the way he'd done long before any app existed.

Knowledge spark: why would blight and leaf spot confuse a model early on? In their first days, both diseases show as small, dark spots with a similar shape and color. The differences that matter, a slightly different edge pattern and how the spot spreads over the next few days, are subtle enough that even an experienced eye sometimes needs a second look.

One leaf stood out: CropSentry had confidently called an early blight case "harmless leaf spot, low priority." Tomasz caught it purely because that particular row happened to be on his usual walking route that week, not because any part of the system flagged it for review.

Hand sketched comparison titled The asymmetry, drawn. Left panel, a small gauge icon labeled Unnecessary spray, caption cheap, visible, an annoyance. Right panel, a larger question mark box icon labeled Missed blight, caption hidden, expensive, a lost field.
Drawn to scale on purpose. One of these mistakes is loud and cheap. The other is silent and costs a field.

Solveig scrapped the 12-crop plan and scoped a new, two-week spike around exactly one question: could CropSentry reliably tell blight from leaf spot, specifically under the lighting conditions field photos actually get taken in.

Hand sketched labeled parts diagram titled What the spike must decide. A question mark box icon at the center labeled Feasibility Spike, with four labeled callouts around it: One question, A kill number, Real field photos, What we skip.
Four things, decided before day one. Everything the six-week plan lacked, in one page.

The real question was never whether CropSentry could tell diseases apart in general. It was whether it could tell apart the one pair that actually costs a farmer a field, in the actual light a phone camera sees at two in the afternoon.

Hand sketched timeline titled Two weeks, the narrow scope, first milestone emphasized. Four milestones: Freeze the confusable pair set, day 2. Test all four lighting conditions, day 7. Find the weak condition, day 10. Kill or go, on that pair only, day 14.
Fourteen days, one confusion, four lighting conditions. Nothing else was allowed onto the calendar.

When the first plan was proposed, someone said, "let's be thorough and test the whole crop lineup before we commit to anything," and it sounded responsible, since nobody wanted to greenlight something narrow and get blindsided later.

Hand sketched icon list titled What this spike will NOT try to answer. Four rows: Whether it works on all 12 crops eventually planned. Whether accuracy holds in every lighting condition on earth. Whether the scouting app's screens are easy to use. Whether it integrates with the farm's existing records system.
Saying what a spike won't answer turned out to matter as much as saying what it will.

Rerun the six weeks with the narrow scope already chosen: the 54 percent midday-sun gap surfaces on day ten, against a tightly built test set, instead of surfacing whenever the next scout happens to walk that particular row by habit.

What I'd tell myself, hearing how close that field came to being lost by luck alone: thoroughness that never finishes isn't caution. It's just a different way of never actually finding out.

PICK, the tradeoff that scopes the questionNot a script for testing less. PICK is what tells you exactly which one thing has to be tested first.

P
Position. The one thing the spike will answer, stated first.
Can CropSentry tell blight from leaf spot, specifically under real field lighting.
This is the hardest step, and the one the six-week plan skipped by trying to answer everything at once.
I
Impact. Who feels each kind of wrong answer?
A false positive costs a farmer an unnecessary spray. A false negative costs a farmer the whole affected section of the field.
Naming both sides in real terms is what keeps the tradeoff from staying abstract.
C
Cost asymmetry. Which error is the expensive, hidden one?
A missed blight case costs about $8,400 an acre in lost yield. An unnecessary spray costs about $140. The spike optimizes against the first.
This is what turns "test the model" into a specific, defensible bar to test against.
K
Kill criteria. What would change the position?
Catching at least 85 percent of blight cases in every lighting condition, not just on average. Midday sun's 54 percent killed the plan to ship as-is.
Stating the bar before the result exists is what makes a kill decision credible.

The recap, one line per letter: position is the one confusable pair under real light, impact is a wasted spray versus a lost field, cost asymmetry is $140 against $8,400 an acre, and kill criteria is 85 percent in every lighting condition, which midday sun failed at 54.

And if you want to be sure it really works, try it somewhere elseSame four letters, a pharmacy counter instead of a tomato field. The costly confusion is still the one thing the spike has to nail.

Delphine Okwuosa runs product at Larkspur Pharmacy Group, testing ScriptScan, a tool meant to flag a photographed prescription as likely altered before a pharmacist fills it. Mapped onto PICK: position is whether ScriptScan can tell a genuinely altered prescription apart from a legitimate one written in unusual handwriting, not whether it can catch every conceivable fraud scheme that might exist. Impact is a delayed, embarrassed patient with a legitimate script versus a controlled substance fraud that goes through unflagged. Cost asymmetry is a delay costing a patient an inconvenient extra call to their doctor, against a missed controlled-substance fraud costing a real health and legal risk that can't be undone once the medication leaves the counter. Kill criteria is catching at least 95 percent of confirmed past fraud cases without flagging more than a small, stated share of legitimate ones.

Hand sketched decision tree titled What ScriptScan's spike should decide. Root, altered versus unusual handwriting. Three branches: catches 95 percent or more of known fraud leads to go pharmacist review tool, confuses odd handwriting with fraud too often leads to kill false alarms erode trust, misses known fraud cases leads to kill not safe for controlled scripts.
A different counter, the same shaped tradeoff: the costly error is the one that can't be quietly corrected after the fact.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "name the one costly confusion, test it under real conditions, set the bar around the expensive error," and stop.
Cost: no time to build four lighting conditions into the test. Say so honestly, and test the single worst condition you already suspect, rather than an average across easier ones.
The model really does handle it, for real: if every lighting condition clears the 85 percent bar, that's a genuine green light, and saying so plainly is what makes the method trustworthy instead of reflexive doubt.

Where people run it wrong.
They scope a spike to "prove it works," which never has a stated finish line.
They test on the easiest, cleanest version of the data and report an average that hides the case that matters.
They treat every wrong answer as equally bad instead of naming which one actually costs something.

How to use it live. The moment you're asked what a spike should answer, ask yourself: which two things does this model most need to tell apart, and which mistake between them actually costs money? Scope the spike to exactly that, and say out loud what it deliberately leaves for later.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a tradeoff or scoping-boundary question, and its one-line job?
Tap to flip
ANSWER
PICK: commit, then show the asymmetry. Position, Impact, Cost asymmetry, Kill criteria.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Solveig Amundsen, the AI PM at Solmark Agritech, who rescoped CropSentry's spike after a near miss in the field.
3 · THE POSITION
What was the one question the rescoped spike committed to answering?
Tap to flip
ANSWER
Can CropSentry tell early blight apart from harmless leaf spot, specifically under real field lighting conditions.
4 · THE ASYMMETRY
Which of the two possible errors is the expensive, hidden one?
Tap to flip
ANSWER
A missed blight case, a false negative. It costs about $8,400 an acre in lost yield, versus about $140 for an unnecessary spray.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting the first spike's scope grow to "prove CropSentry works" broadly, instead of naming the one costly confusion to test first.
6 · THE NUMBER
Fill in the blank: overall accuracy on the confusable pair was ___ percent, but under harsh midday sun it dropped to ___ percent.
Tap to flip
ANSWER
91 percent overall, 54 percent under harsh midday sun.
7 · WHAT IT WON'T ANSWER
Name two things this spike explicitly did NOT try to answer.
Tap to flip
ANSWER
Whether CropSentry works on all 12 planned crops, and whether the scouting app's screens are easy to use. Both are real questions for later.
8 · CROSS PRODUCT TRANSFER
Section 4 runs PICK again for a different product. Which one, and what's the position there?
Tap to flip
ANSWER
Larkspur Pharmacy Group's ScriptScan. The position is telling a genuinely altered prescription apart from unusual legitimate handwriting, not catching every conceivable fraud scheme.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: a feasibility spike should answer one question about the ___ that costs real money if confused, under ___ conditions, not a curated best case.
Show hint
Look at the direct answer.
Show answer
Confusion (or confusable pair); real (or field/real-world). Testing the easy, clean case proves nothing about the case that will actually happen in a real field.
True or false
2. True or false: the six-week, all-12-crops spike attempt produced a clear kill or go answer.
  • True
  • False
Show hint
Look at "now here is the same thing as a story."
Show answer
False. Six weeks in, the team had only covered about 40 percent of the plan, with no clear answer for leadership either way.
Multiple choice
3. Why does the spike set its kill bar around missed blight cases rather than unnecessary sprays?
  • A. Unnecessary sprays are illegal under agricultural regulations.
  • B. A missed blight case costs about 60 times more per acre than an unnecessary spray, and the loss can't be undone once the field is lost.
  • C. Farmers never notice an unnecessary spray happening.
  • D. The model is naturally biased toward false positives.
Show hint
Look at the bar chart comparing the cost of each kind of wrong answer.
Show answer
B. $8,400 an acre for a missed case versus about $140 for an unnecessary spray, and only one of those losses can be reversed after the fact.
Short answer, where it wouldn't matter
4. Name a question about CropSentry that genuinely doesn't need to be answered by this two-week spike, and say why.
Show hint
Look at the icon list, "what this spike will NOT try to answer."
Show answer
Model answer: Whether CropSentry eventually works on all 12 planned crops. That's a real question, just not the one gating whether to keep building this feature at all.
Short answer, apply it yourself
5. Think of an AI feature you've seen pitched. What two outcomes does it need to tell apart, and which mistake between them would actually cost someone something real?
Show hint
Think about which error is loud and cheap versus which one is quiet and expensive.
Show answer
Model answer: A spam filter needs to tell junk apart from an urgent client email marked as junk by mistake; the missed client email is the quiet, costly error worth testing for specifically, not the overall spam-catch rate.
Short answer, work the number
6. If the model had scored 85 percent under midday sun instead of 54, would the spike still have flagged a problem worth fixing before shipping?
Show hint
Compare that number to the stated kill bar.
Show answer
Model answer: No, 85 percent would have cleared the stated bar in every lighting condition, which means the spike would have supported shipping rather than sending the model back for retraining.
Before you close the answer
Why this works
Tests whether you can narrow a feasibility question to the one costly confusion that matters, or let "let's be thorough" quietly expand a spike until it never finishes.
Follow-up traps
"Isn't testing only one confusable pair leaving too much untested?" Response: yes, on purpose. Everything else waits for a later spike; this one exists only to answer whether the costliest confusion is safe to ship on.

"What if leadership wants proof it works on all 12 crops before funding more?" Response: then that's a real, separate ask, and it deserves its own scoped spike with its own kill bar, not folding into this one until this one never finishes.
If pressed
The retrained model's harsh-midday-sun accuracy on the confusable pair was rechecked against the same 85 percent bar four weeks later, using a fresh set of field photos the original spike had never seen, specifically to rule out the model having simply memorized the first test set.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more