CaseAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #15

What should you spike first when both data quality and model capability are in doubt?

ORDERtwo hundred photos of the same leaf, and a real disagreement underneath

Cornbelt Growers Cooperative is piloting RowSense AI, a computer-vision tool that reads drone and phone photos of corn leaves for early signs of disease. Tunde Okonkwo is the AI PM who has to decide what to spike first when nobody can say yet whether the problem is the photos or the model.

The direct answer
Spike the data first. Pull a raw, unfiltered sample of real field images and check whether a human expert can reliably diagnose disease from them at all, using two independent reviewers and measuring where they agree. If the data itself doesn't carry a clear signal, no amount of model testing afterward tells you anything real, and every hour spent tuning the model before that check is an hour that might get thrown away.
Do this, in order
  1. Audit a raw sample of real field images before touching the model at all.Why: a model can only be as good as the signal in what it's shown, and nobody yet knows if that signal exists.
  2. Have two independent experts label the same sample and measure their agreement rate.Why: if two trained agronomists can't agree on a photo, no model score on that photo means anything either.
  3. Only test model capability once the data has passed that bar.Why: this is the decision that's hardest to undo. A bad model can be swapped; bad data compounds into every future spike.
  4. Check field-collection conditions, not just image quality in isolation.Why: staged, clean sample photos can hide exactly the lighting and angle problems real field collection will produce.
  5. Say plainly when the data audit alone is enough to stop the project early.Why: shows judgment about when the answer is already clear, instead of running every stage regardless.

How to answer this, stage by stage

Nobody is scoring whether you know that both data and model can be bad. They are scoring whether you can say, concretely, which one to check first and why getting that order wrong is the expensive mistake.

Stage 1
Scope it to one real tool
Say it like this
"Let's ground this in RowSense AI, a disease-detection tool for corn leaves at Cornbelt Growers Cooperative, where both the photos and the model were genuinely in doubt at the same time."
Why this works
Keeps the answer from turning into a generic "data quality matters" lecture with nothing real behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as ORDER: outcome, reversibility, dependency, evidence, rank. Reversibility is the one that actually decides this question."
Why this works
Naming the structure up front tells the interviewer you have a repeatable way to prioritize, not just an instinct.
Stage 3
Name what everything is actually competing to move
Say it like this
"Both candidates are competing for the same thing: confidence that we can greenlight this project cheaply, before committing real engineering weeks to fine-tuning a model on data nobody has checked yet."
Why this works
Without naming the shared outcome first, ranking two very different risks is just opinion.
Stage 4
Give the one decision, and why it's ranked first
Say it like this
"Spike the data first. A bad model decision is reversible, you swap the model and the data is still useful. A bad data decision compounds, every hour spent tuning a model on unreliable photos gets thrown away once the real problem surfaces."
Why this works
This is ORDER's actual test: ranking by which mistake is hardest to undo, not by which task feels more urgent.
Stage 5
Prove it with the compressed failure
Say it like this
"We had two agronomists independently label the same 200 field photos. On clean, well-lit shots they agreed 94 percent of the time. On the blurrier, real-world half of the sample, agreement dropped to 61 percent. No model score on that second half was ever going to mean anything."
Why this works
A real disagreement rate is worth more than any argument about "data quality" as an abstract concern.
Stage 6
Say what depends on what
Say it like this
"You can't meaningfully test model capability on a dataset human experts can't agree on themselves. The data audit isn't just first by preference, it's a real dependency: the model test has no ground truth to be scored against until the labels are trustworthy."
Why this works
Naming the dependency, not just the priority, is what separates a reasoned order from a gut-feel ranking.
Stage 7
Name the AI-specific reasoning and the trade-off
Say it like this
"The reason this isn't a generic prioritization question is that a model's apparent capability is only ever as trustworthy as the labels it's measured against. We accepted two extra weeks up front for a proper data audit, in exchange for never fine-tuning a model against ground truth that wasn't actually true."
Why this works
This names the load-bearing, AI-specific judgment: model capability can't be evaluated in a vacuum from the data quality underneath it.
Stage 8
Say what wouldn't need this order, then close
Say it like this
"If we'd already been running a mature version of this tool for a season with a known, clean image pipeline, I'd spike model capability first, since the data question would already be settled. Here, with both in doubt: data first, model second, always in that order."
Why this works
Closing with when the order would flip shows real judgment, not a rule applied blindly, and restates the ranked decision in one breath.

Let's learn

The phone in Colette Vranken's hand held two hundred photos of the same rows of corn, taken across three mornings, meant to teach a tool the difference between a healthy leaf and one showing the first signs of blight.

Before RowSense AI, Colette walked the rows herself every Tuesday, checking maybe 40 plants by eye across two hours, catching disease early enough to treat about 85 percent of the time. The tool promised to scan every plant a drone could photograph, thousands per hour, and flag the ones that needed a second look.

Hand sketched flow diagram titled What has to happen before what, first step emphasized. Four steps left to right: Audit raw images. Human labelable question mark. Test model capability. Test at field scale.
Four steps, in order. Skip the first one and the third step has nothing real to be scored against.

Here's the turn: when the first spike came back with mixed results, the instinct in the room was to ask "is the model good enough." That is the wrong question to ask first. Nobody yet knew whether the two hundred photos themselves were even reliable enough for any model, good or bad, to learn something true from.

Two agronomists' disease-label agreement rate, by image condition
100% 50% 0 94% Clean, well-lit photos 61% Typical field photos
Same two experts, same disease, same 200 photos split by condition. Almost 40 percent of the real-world sample had no reliable ground truth at all.

At its worst, engineers spend three months fine-tuning a model against labels that two human experts couldn't even agree on themselves, produce a model that scores well on paper against those same shaky labels, and ship a tool that's confidently wrong on exactly the blurry, real-world photos a drone actually produces.

A model can only be as honest as the labels it was scored against. If the labels are a coin flip, the model's accuracy number is measuring the coin, not the disease.
The choice I would take back When the pilot was scoped, the team defaulted to using drone images captured at a fixed altitude and time of day, since that was the easiest setting to standardize across test flights. That made sense when the goal was a clean first demo. It stopped making sense once real growers started flying drones at whatever height and hour actually fit their schedule, producing photos the fixed-setting default never anticipated.

What I would leave alone: I wouldn't run this same two-step data-then-model order for every future model update. Once the data pipeline and label agreement are established and stable, a routine model refresh can be spiked directly against that known-good baseline.

The lesson: we asked whether the model was ready, when the real first question was always whether the data even had a stable, agreed-upon answer to be ready against.

Now here is the same thing as a story

The short version above is what you'd say defending the spike order in a five-minute stand-up. Read this one for what it looked like the week two bad readings in a row nearly sent the team straight to blaming the model.

Every Tuesday, Colette printed the week's photo batch and walked out to the yard with it, checking the physical plants against what the drone had caught, because nine years of doing this by eye had taught her that some things a photo just doesn't show.

RowSense's first pilot batch looked promising on the surface: the model flagged 34 plants as showing early blight symptoms out of 2,000 scanned. Tunde and the engineering team were pleased. Then Colette walked two of the flagged plants herself and found nothing there at all, just normal leaf shadow that the model had mistaken for a lesion.

Hand sketched icon list titled Signs the data is the real problem. Four rows: blurry images, inconsistent lighting. Two agronomists label the same photo differently. Missing metadata on crop growth stage. Photos staged clean, not natural field shots.
Four honest signs it was never really about the model's accuracy score.

Two in a row, both false alarms, both traced back to the same shadow pattern under low morning light. The easy read in that room was "the model needs more training." Tunde didn't take that read.

He pulled two of Cornbelt's own agronomists, Colette and a colleague from the co-op's second site, and had them independently label the same 200 photos the model had been trained on, without telling either of them what the model had said. On the clean, well-lit half of the sample, they agreed 94 percent of the time. On the blurrier, real-world half, shot at odd angles and low light, they agreed only 61 percent of the time.

Hand sketched labeled parts diagram titled What's in a data quality audit. A document icon at the center labeled Data Audit, with four labeled callouts around it: Label agreement rate, Image sharpness, Lighting variance, Natural versus staged shots.
The audit that answered the real question before the model ever got blamed for anything.
Knowledge spark: what does "ground truth" actually mean here? Ground truth is just the answer key a model gets scored against, in this case, whether a leaf really has blight. If two trained experts looking at the same photo can't agree on the real answer, there is no reliable answer key for that photo, and a model's accuracy against it doesn't mean what it looks like it means.

The real question was never whether the model could tell a shadow from a lesion. It was whether the photos themselves carried a clear enough signal for anyone, human or model, to tell the difference reliably in the first place.

Hand sketched quadrant titled Which risk to spike first. Axes: cost to test from cheap to expensive, how much it blocks everything else from small to blocks all downstream work. Data quality sits low cost, high block. Model capability sits higher cost, lower block. Label agreement and field lighting variance sit near data quality.
The cheapest test to run also turned out to be the one blocking every other test downstream of it.

Once the audit named the real gap, the fix wasn't a bigger model. It was a stricter field-photo protocol: a minimum lighting threshold and a required angle range before any photo entered the training set at all. Re-run against 200 new photos collected under that protocol, agreement between the two agronomists rose to 89 percent, and the model's flagged false-alarm rate on real field conditions dropped from roughly 1 in 15 to about 1 in 60.

Hand sketched comparison titled Reversible or not. Left panel, box icon labeled A bad data decision, caption keeps compounding, expensive to unwind later. Right panel, gauge icon labeled A bad model decision, caption swap the model, the data is still good.
The asymmetry that decided the order: one mistake compounds quietly for months, the other one is a Tuesday afternoon fix.

What I'd tell myself, watching those two false alarms almost get pinned on the model: the model was never lying. It was answering exactly the question the data gave it, and nobody had checked yet whether that was a fair question to ask.

ORDER, the rank that survives the follow-up questionNot a script for auditing data before every single model change. ORDER is what tells you exactly which risk is hardest to undo, and why that one goes first.

O
Outcome. What are all the candidates competing to move?
Cheap, early confidence about whether RowSense AI is worth a full engineering investment, before real weeks get spent fine-tuning against untrustworthy labels.
Without naming this first, ranking two different kinds of risk is just gut instinct dressed up as a decision.
R
Reversibility. Which mistake is hardest to undo?
A bad model decision is reversible, swap the model and the data is still useful. A bad data decision compounds, every hour of model tuning against unreliable labels gets thrown away once the real problem surfaces.
This is the hard step, and the one that decides the whole ranking: reversibility, not urgency, sets the order.
D
Dependency. What unblocks what?
Model capability can't be honestly tested until the labels it's scored against are themselves trustworthy. The data audit isn't a preference, it's a real prerequisite.
Some of the order here is forced by reality, not chosen. That's worth saying out loud.
E
Evidence. What's cheap to learn before committing the quarter?
Two independent agronomists labeling the same 200 photos, a two-day exercise, before any engineering time gets spent on the model itself.
Cheap evidence that changes the whole plan is worth more than an expensive test that only confirms what you already believed.
R
Rank. State the order, defend the top pick.
Data quality first, model capability second, because the first one is both cheaper to test and the one whose mistakes compound the longest if skipped.
A countable result: label agreement rose from 61 to 89 percent after fixing the field-photo protocol, and false alarms dropped roughly fourfold.

The recap, one line per letter: outcome is cheap confidence before committing real engineering time, reversibility is a bad model swap versus a compounding bad-data problem, dependency is that model testing has no honest ground truth until labels are trustworthy, evidence is a two-day agreement audit before touching the model, and rank is data first, model second, defended by the fourfold drop in false alarms once the field-photo protocol shipped.

And if you want to be sure it really works, try it somewhere elseSame five letters, a fish-quality grading tool instead of a corn field. Different flip family entirely, the same audit-first order.

Niamh Colby runs quality control at Saltbourne Fisheries, piloting a tool that grades fish freshness from photos taken dockside by crew as catches come in. Mapped onto ORDER: outcome is cheap confidence before investing in a full grading-line rollout. Reversibility favors the data question again: a bad model swap is cheap, but a bad labeling protocol, disagreement between graders on what "grade A" even looks like, compounds across every catch already logged. Dependency is that model accuracy against inconsistent human grades measures nothing real. The flip here is different from Cornbelt's: it's a concealment flip, crew who worry the tool will replace them start quietly using their own private grading system and only reporting the tool's number when it happens to match, which hides the real disagreement rate from management entirely. Evidence is a blind, independent regrading of 150 recent catches by two senior graders. Rank still lands data first, though the reason getting there runs through a different, more human problem.

The same hand sketched comparison reused: reversible or not, a bad data decision keeps compounding while a bad model decision is a simple swap, applied here to a fish grading tool instead of a corn disease tool.
The same asymmetry holds in a completely different industry: the data mistake is the one that keeps costing you.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "spike the data first, because a bad model swap is reversible and a bad data decision compounds," and stop.
Cost: there's no budget for a full two-agronomist relabeling exercise. Say so honestly, and run the smallest real version, twenty photos, two reviewers, rather than skipping the check entirely.
The model tests better than expected, for real: even a strong early model score is worth auditing the data behind, since a model can score well against labels that are consistently wrong in the same direction.

Where people run it wrong.
They treat "the model isn't good enough yet" as the default explanation before checking whether the data even has a stable answer.
They standardize test conditions so cleanly that the pilot never meets the messy real-world inputs it will actually face.
They let one clean spike result convince them the whole pipeline is solid, instead of testing the next batch of real conditions too.

How to use it live. The moment an interviewer asks what to spike first, ask yourself: if I test the wrong one first and I'm wrong, which mistake can I still undo cheaply next week, and which one have I already baked in? Whichever one you can't undo goes first.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Input flip: once field conditions weren't standardized, growers and drone operators risked staging cleaner shots to satisfy the model rather than capturing the messy real conditions it actually needed to learn from.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tunde Okonkwo, the AI PM at Cornbelt Growers Cooperative, deciding what to spike first for RowSense AI, alongside agronomist Colette Vranken.
3 · THE HABIT
What did the team almost stop doing, that would have hidden the real problem?
Tap to flip
ANSWER
Checking the raw data itself. After two false alarms, the instinct was to blame the model directly and skip straight to retraining it.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Photographing real, messy field conditions naturally versus quietly staging cleaner, standardized shots to keep the model happy. Once staging starts, the model stops learning from real conditions at all.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Defaulting to drone photos captured at a fixed altitude and time of day, which was easy to standardize for a demo but never matched how real growers would actually fly their drones.
6 · THE NUMBER
Fill in the blank: two agronomists agreed on disease labels ___ percent of the time on clean photos, and only ___ percent on typical field photos.
Tap to flip
ANSWER
94 percent on clean, well-lit photos; 61 percent on typical, blurrier field photos.
7 · THE REPLAY
Same audit, a stricter field-photo protocol already in place. What changes?
Tap to flip
ANSWER
Agronomist agreement rises to 89 percent under the new lighting and angle protocol, and the model's real-world false-alarm rate drops roughly fourfold, from about 1 in 15 flagged plants to about 1 in 60.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Saltbourne Fisheries' dockside fish-grading tool. The flip there is concealment: crew worried about being replaced quietly run their own private grading system, hiding the real disagreement rate from management.

Check yourself Score: 0 / 0

Multiple choice
1. Why does Tunde spike the data before the model, even though both are in doubt?
  • A. Data audits are always cheaper than model tests, in every situation.
  • B. A bad model decision is reversible, a swap; a bad data decision compounds and can't be honestly tested against until fixed.
  • C. Colette insisted on it personally.
  • D. The model had already failed a benchmark test.
Show hint
Look at the Reversibility step, R, in the ORDER recap.
Show answer
B. Reversibility, not urgency or personal preference, is what actually sets the order in ORDER.
True or false
2. True or false: the two false alarms were caused by the model being poorly trained on plenty of good data.
  • True
  • False
Show hint
Look at what the label-agreement audit actually found.
Show answer
False. The false alarms traced back to unreliable labels on blurry, real-world photos, not a poorly trained model working from good data.
Fill in the blank
3. Fill in the blank: a model's accuracy score can only be trusted as far as the ___ it was scored against.
Show hint
Look at the knowledge spark about ground truth.
Show answer
Ground truth (the labels/answer key). If two experts can't agree on the real answer for a photo, a model's accuracy against that photo doesn't mean what it looks like it means.
Short answer, where it wouldn't matter
4. Name a situation where this same data-first order wouldn't be necessary.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine model refresh once the data pipeline and label agreement are already established and stable can be spiked directly against that known-good baseline.
Short answer, apply it yourself
5. Think of a project you've worked on where two different things were both uncertain at once. How did you decide which one to check first, and would reversibility have changed your order?
Show hint
Think about which mistake, if made, would have been cheap to undo versus expensive to undo.
Show answer
Model answer: A strong version names the harder-to-undo risk explicitly, the same way the corn-leaf case names bad data as the one that compounds.
Short answer, work the number
6. If agreement on the messy field photos had come back at 85 percent instead of 61 percent, would the data audit still have been worth running first?
Show hint
Think about what counts as "good enough" ground truth for a disease-detection tool.
Show answer
Model answer: Yes, though the urgency would be lower. Even 85 percent agreement means roughly 1 in 7 labels is disputed, still worth knowing before committing engineering weeks, just less likely to change the plan entirely.
Before you close the answer
Why this works
Tests whether you'll default to blaming the model first, the easy instinct, or check the harder-to-undo risk underneath it before spending real engineering time.
Follow-up traps
"Isn't checking the data first just generic due diligence, nothing AI-specific?" Response: no, because a model's accuracy is only meaningful relative to its labels, a normal software bug doesn't have this problem, there's no "ground truth" a login button is scored against.

"What if the data audit itself takes too long and the deadline is fixed?" Response: scope it down, not away, even twenty photos and two reviewers gives a real signal about label reliability in under a day.
If pressed
The final field-photo protocol specifically required a minimum resolution and a shooting angle within 30 degrees of directly overhead, rebuilt from the exact conditions the two false-alarm photos had violated, not a generic image-quality checklist.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more