CaseAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #10
How do you run a spike when you do not yet have the data you would need in production?
SPARKa spike that borrows data on purpose
AgroFleet Exchange is a marketplace where farmers buy, sell, and rent used tractors, sprayers, and other equipment. FieldScout is the add-on that lets a farmer photograph a leaf and get a quick read on whether it needs spraying. Kwame Delacroix-Nkomo is the AI PM who had to spike it before a single real farm photo existed.
The direct answer
Run the spike on the closest honest proxy you can find, but design the spike to know it is guessing. Give every result a confidence band, flag anything unlike the proxy data as unseen instead of fine, route the unseen cases to a person, and quietly bank every real example the spike sees, right or wrong. FieldScout's first version trained on public, studio-lit crop photos because no real farm photo existed yet. Because it said "no local data yet" instead of a confident wrong answer, it started building the exact dataset it lacked from day one. That is the real output of a spike run this way, not a finished model, a running start on the data it never had.
Do this, in order
Run the spike on the closest honest proxy, and make it know when it is guessing.Why: a confident wrong answer on data the model never saw does far more damage than an honest "not sure."
Route every low-confidence case to a person, never a silent pass.Why: a spike without real data will always meet a case it has never seen, and that case needs somewhere safe to land.
Design the spike to save every real example it sees, right or wrong, from day one.Why: this turns the spike into the very data-collection process it is missing, instead of waiting on a second project to build one.
Watch how people feed it, not only what it outputs.Why: people will quietly reshape their own behavior to get a confident answer, which can hide the real gap until it is your only signal.
Name upfront which cases are day one and which are deliberately held for later.Why: shows judgment about exactly what the proxy data can and cannot honestly stand in for.
Say plainly when the proxy really is close enough, like a common, well-photographed condition.Why: not every case needs the confidence-band treatment, and treating all of them the same wastes the spike's real signal.
How to answer this, stage by stage
Nobody is scoring whether you found a clever workaround for missing data. They're scoring whether the spike itself can tell the difference between a case it has seen and one it hasn't.
Stage 1
Scope it to one real spike
Say it like this
"Let's ground this in FieldScout, at AgroFleet Exchange. That's the spike that had to run before a single real farm photo existed."
Why this works
Keeps the answer from becoming an abstract debate about data availability with nothing real underneath it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as SPARK. Situation, how the job gets done today. Payoff, the habit I want to build. Anchor, the one design decision. Risk, what breaks the first time it's wrong. Keep out, what I won't build on day one."
Why this works
Signals a repeatable way to design around missing data, instead of a one-off trick.
Stage 3
Reframe: it isn't "do we have the data," it's "can the spike say when it doesn't"
Say it like this
"This isn't really a question about whether real farm photos exist yet. It's a question of whether the spike can tell you, honestly, the moment it's looking at something it's never seen."
Why this works
This is where a strong answer separates from someone who just says "use public data" and stops.
Stage 4
Give the anchor
Say it like this
"The one decision everything else hangs on: every reading gets a confidence band and a 'no local data yet' flag instead of a bare answer, low-confidence cases go to a person, and every real photo, right or wrong, gets saved to build our own first dataset."
Why this works
Names one concrete, inspectable decision instead of a vague promise to "be careful with the data."
Stage 5
Prove it with the compressed evidence
Say it like this
"FieldScout hit 91 percent on the public proxy set. On the first real batch of farm photos, it dropped to 63. That gap is exactly why the confidence band existed, and exactly what it was built to catch."
Why this works
Compresses the whole argument into the one number the anchor was designed to survive.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't a generic bootstrapping problem is that a model trained on clean proxy photos will still answer confidently on real, messy ones, it just won't be right. We accepted slower, human-reviewed answers on unfamiliar cases in exchange for never handing a farmer a confident wrong read on data we'd never actually tested."
Why this works
This is the load-bearing, AI-specific judgment: a model doesn't know what it doesn't know unless you build that knowledge in on purpose.
Stage 7
Say what's kept out for later, then close
Say it like this
"We're not auto-diagnosing rare regional blights or recommending sprays automatically on day one, that waits for real local data. I'm not defending a finished product. I'm defending a spike that builds the exact dataset it started without."
Why this works
Shows judgment about scope and restates the direct answer in one breath.
Let's learn
The tablet is bolted to the tractor's dash with two zip ties, cracked screen protector, dust in every seam.
Before FieldScout, a farmer walked the rows looking for wilting or spotting by eye, and when something looked wrong, drove into town to ask the local agronomist, losing most of a day. With FieldScout, a photo gets a read in under three seconds, and the farmer decides on the spot whether to spray or wait.
The second step is the one FieldScout tried to shortcut, before it had ever seen a real leaf.
Here's the turn: no real farm photo existed when the spike had to run, so FieldScout trained on public, professionally lit crop-disease photos instead, and hit 91 percent against that proxy set. The real question was never whether that number was good. It was whether the spike would know the difference once a real, messy farm photo showed up.
The one design decision everything else hangs on: it never claims a diagnosis, only a confidence band and an honest flag.
At its worst, a farmer trusts a confident-sounding answer the model was never actually tested on, sprays for the wrong thing, and loses a week and a chemical bill on a blight the tool had never once seen a real photo of.
A model trained on clean stock photos does not know it has never seen a real field. Only the spike around it can know that, if you build it in.
The choice I would take back
Nothing, really, in this one. The anchor held. What I would flag instead: the proxy dataset assumed every real photo would arrive centered and well-lit, like the stock photos it trained on, since no real farm photo existed yet to prove otherwise. That assumption made sense on day one. It stopped making sense the moment real farmers started photographing leaves in the field.
What I would leave alone: for the one or two most common conditions, well represented in the public proxy data and easy to spot even under bad lighting, the spike's confident reads held up fine against real farm photos from the very first week.
The lesson: a spike without real data is not a smaller version of the real product. It is the actual process of getting the real data, disguised as a demo, if you design it that way on purpose.
Now here is the same thing as a story
The short version above is what you'd say defending the design in a planning review. Read this one for what it felt like the week a new hire's question exposed what the tablet had never actually been tested against.
Every Thursday, Kwame pulled the week's FieldScout photos and skimmed them for anything odd before the Monday product review.
Same words, photo of a leaf, describing two very different kinds of image.
The spike launched well. Farmers who tried it liked the speed, and for the first few weeks the confidence bands mostly read high, green, trustworthy.
Knowledge spark: why would farmers' own photos start looking like the demo set?
When a photo gets a confident, high-band reading, people notice what worked and repeat it. If centered, single-leaf shots against a blank hand or shirt happened to score higher, farmers learned to frame their photos that way, without anyone telling them to, since it's what got them a clear answer.
A new agronomy intern, reviewing a batch of low-confidence photos, asked Kwame plainly: "Why does it do great on the demo photos but barely says anything useful on real farmers' own shots?"
Week four is where the gap between proxy data and real data first became visible.
Looking closer, Kwame found real farmers had quietly learned to photograph a single leaf, centered, held against a shirt for a plain background, since that framing reliably got a confident read. Whole-plant shots and natural, cluttered backgrounds, what a farmer would show an agronomist in person, scored low far more often.
Accuracy, proxy data versus the first real batch
A 28-point gap between clean stock photos and messy real ones, exactly the case the confidence band was built to flag.
The real question was never whether farmers were framing photos "wrong." It was whether the spike's own design had quietly taught them what the model could and couldn't handle, without anyone meaning to.
The anchor's whole job: even on a photo the model has never seen, the farmer never gets a confident wrong answer.
When the confidence band was first designed, someone said, "let's keep the threshold generous so it feels useful," and it sounded reasonable, since at the time nobody knew real photos would look so different from the proxy set.
Naming what waits for real data is what kept the spike honest about what it actually was.
Rerun the same ten weeks with the local photo set already growing from day one, exactly as the anchor intended: by week ten, four hundred real farm photos have been banked through the low-confidence flow, and real-batch accuracy climbs from 63 to 84 percent, without anyone waiting on a second data-collection project that was never going to get funded on its own.
What I'd tell myself, hearing that intern's question: the spike was never really about hitting 91 percent on photos we already had. It was about building the pile of photos we didn't.
SPARK, the plan for a spike with no data yetNot a script for pretending a proxy dataset is the real thing. SPARK is what makes the spike say so, out loud, every time it isn't sure.
S
Situation. How does the job get done today, without you?
A farmer walks the rows, spots wilting by eye, and drives into town to ask the local agronomist when something looks wrong, losing most of a day.
Grounds the spike in one real task before any data question comes up.
P
Payoff. What habit do you want this to build?
Stop driving into town for every uncertain leaf. Start trusting a quick photo read for the obvious cases, and knowing exactly when it's not one of those.
Names the actual behavior change the spike exists to earn, not just the time saved.
A
Anchor. The one design decision everything hangs on.
Every reading carries a confidence band and a "no local data yet" flag, low-confidence cases route to a person, and every real photo gets saved to build the spike's own first local dataset.
This is the hardest step, and the one that turns missing data from a blocker into the spike's actual job.
R
Risk. What breaks the first time you're wrong?
A real farm photo the proxy data never resembled gets a low-confidence flag and goes to a person, instead of a silently confident wrong answer reaching the farmer.
Proves the anchor survives its own failure case, not just the cases it was built and tested on.
K
Keep out. What you won't build on day one.
Auto-diagnosing rare regional blights, fully automatic spray recommendations, and cross-farm yield comparisons, all deferred until real local data actually exists.
Shows judgment about scope, instead of a wish list dressed up as a roadmap.
The recap, one line per letter: situation is a farmer's day without FieldScout, payoff is trusting a quick read for the obvious cases, anchor is the confidence band plus the "no local data yet" flag plus saving every real photo, risk is an unfamiliar case routing to a person instead of a confident guess, and keep out is everything that waits on real local data.
And if you want to be sure it really works, try it somewhere elseSame five letters, a library consortium instead of a farm. The proxy data changes shape entirely, the honest flag doesn't.
Wilhelmina Osei-Bonsu runs product at Fenwick County Library Consortium, where CatalogMind drafts subject tags for new books before a human cataloger confirms them. No real patron search data existed yet when the spike ran, so the team trained CatalogMind on public library metadata records instead. Mapped onto SPARK: situation is a cataloger manually assigning tags from a printed style guide, one book at a time. Payoff is catalogers reviewing drafted tags instead of writing every one from scratch. Anchor is showing which tags CatalogMind is confident about versus guessing, with every guess reviewed before it ever reaches a public search. Risk is a tag drafted for a book type the public records never covered, like a local self-published history, getting flagged low-confidence instead of published silently wrong. Keep out is fully automatic tagging for anything outside the consortium's most common categories.
A different flip than FieldScout's: once a few drafted tags looked embarrassing, catalogers quietly stopped marking which ones were AI-drafted at all.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "run it on the closest honest proxy, and make it say when it's guessing," and stop.
Cost: there's no budget yet to collect real data separately. Say so honestly, and design the spike itself to collect that data as a byproduct, rather than treating data collection as a future, unfunded project.
The proxy data turns out to be closer than expected, for real: if real farm photos match the stock set better than feared, that's still worth confirming against the confidence bands, not assumed safe just because the headline accuracy looked fine.
Where people run it wrong.
They treat a spike's proxy-data accuracy as if it already describes real production conditions.
They let a model answer confidently on data it has never actually seen, with nothing catching the gap.
They wait for a second, separate project to collect real data instead of designing the spike to collect it from day one.
How to use it live. The moment someone asks how to spike without production data, ask yourself: what's the closest honest proxy, and how will the spike admit the moment a real case looks nothing like it. Name both, and the rest of the plan follows.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Input flip: farmers didn't complain about low readings, they quietly changed how they photographed leaves, framing shots to match the demo set instead of showing what they'd naturally show an agronomist.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Kwame Delacroix-Nkomo, the AI PM at AgroFleet Exchange, who spiked FieldScout on public proxy data before any real farm photo existed.
3 · THE HABIT
What did farmers start doing, without being told, to get a confident reading?
Tap to flip
ANSWER
They started photographing a single leaf, centered, against a plain background, mimicking the demo set's framing, instead of the whole-plant, cluttered shots they'd naturally take.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Photographing a leaf however felt natural, versus framing it the way that reliably got a confident reading. No middle setting once farmers learned which framing worked.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Assuming real farm photos would arrive centered and well-lit like the proxy stock photos, since no real farm photo existed yet to prove otherwise.
6 · THE NUMBER
Fill in the blank: FieldScout hit ___ percent on the proxy dataset, and only ___ percent on the first real batch of farm photos.
Tap to flip
ANSWER
91 percent, and 63 percent.
7 · THE REPLAY
Same ten weeks, with the local photo set growing from day one as the anchor intended. What changes?
Tap to flip
ANSWER
By week ten, 400 real farm photos have been banked through the low-confidence flow, and real-batch accuracy climbs from 63 to 84 percent, without waiting on a separate data-collection project.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this again for a different product. Which one, and what stays the same?
Tap to flip
ANSWER
Fenwick County Library Consortium's CatalogMind. The same anchor idea stays: flag what the spike is guessing on, using public records as a proxy before real patron data exists.
Check yourself Score: 0 / 0
Short answer, the anchor
1. What is the one design decision this answer says everything else hangs on?
Show hint
Look at stage 4 of the walkthrough, or the A step in the SPARK recap.
Show answer
Model answer: Every reading carries a confidence band and a "no local data yet" flag, low-confidence cases go to a person, and every real photo gets saved to build the spike's own dataset.
Multiple choice
2. Why did FieldScout's accuracy drop from 91 percent on the proxy set to 63 percent on real farm photos?
A. The tablet's camera hardware was lower quality than expected.
B. Real farm photos looked very different from the clean, centered, studio-lit proxy photos the model had trained on.
C. Farmers stopped using the app entirely after the first week.
D. The public proxy dataset was mislabeled from the start.
Show hint
Look at the metaphor scene, "two things called a photo of a leaf."
Show answer
B. The proxy dataset and real farm conditions were simply different distributions, which is exactly the gap the confidence band was built to catch.
True or false
3. True or false: farmers were explicitly instructed to photograph leaves centered against a plain background.
True
False
Show hint
Look at "locate the habit" and the knowledge spark about why real photos started to resemble the demo set.
Show answer
False. Nobody instructed them. They learned it themselves, since that framing reliably produced a confident reading.
Fill in the blank
4. Fill in the blank: after ten weeks of banking real photos through the low-confidence flow, real-batch accuracy climbed from 63 percent to ___ percent.
Show hint
Look at the replay, near the end of the FLIPS-style story section.
Show answer
84 percent. Built from 400 real farm photos collected as a byproduct of the anchor's design, not a separate data project.
Short answer, apply it yourself
5. Think of a tool you've used that gave a confident-sounding answer on something unusual. What would an honest "I haven't seen this before" flag have looked like instead?
Show hint
Think about the difference between a model refusing to guess and a model guessing anyway.
Show answer
Model answer: A translation app that confidently mistranslates a rare regional phrase. An honest flag might read "uncommon phrasing, translation may be unreliable" instead of presenting the guess as settled.
Short answer, where it wouldn't matter
6. Name a case in FieldScout where the proxy data was already close enough, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The one or two most common, easy-to-spot conditions well represented in the public proxy data. Those held up fine against real farm photos from the very first week.
Before you close the answer
Why this works
Tests whether you'll treat missing production data as a blocker to work around quietly, or design the spike itself to know exactly when it's out of its depth.
Follow-up traps
"Isn't a confidence band just a technical detail, not a real design decision?" Response: no, it's the entire anchor, it's the one thing standing between a farmer trusting a guess the model was never tested on and a farmer getting routed to a person instead.
"What if the proxy data is actually nothing like the real world at all?" Response: then most cases will read low-confidence at first, which is the honest, if unglamorous, correct behavior, and it's exactly the signal that tells you to slow down before real farmers ever see a confident wrong answer.
If pressed
The confidence threshold for "no local data yet" was set using how far a real photo's pixel statistics sit from the proxy set's own distribution, not a fixed guess, so it tightens automatically as the local dataset grows and the model's true blind spots shrink.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.