ConceptIntermediateAI Opportunity & Model Strategy / Data strategy as product strategy / #7
Describe the difference between data volume, data quality and data relevance for AI products.
LEAD/what a running coach learned about clean data pointed at the wrong terrain
What happens when a model has plenty of data, all of it accurate, and it still gets the call wrong? Cinder Trail Coaching runs an app that builds each runner's weekly training plan from their GPS watch history. Marisol Vance coaches thirty of those runners and is the one who finds out, one season in, that having enough clean data is not the same question as having the right data.
The direct answer
Volume is how much data you have. Quality is how clean and correct that data is. Relevance is whether it actually matches the decision the model has to make. You can have huge volume and near-perfect quality and still ship a bad AI feature, because the data was clean and correct for a different question than the one your product is answering. Fix relevance first, quality second, and treat volume as the cheapest of the three to buy.
Do this, in order
Name the exact decision the model has to make, then check whether your biggest data pile actually informs that decision.Why: relevance is the gap that volume and quality both quietly hide.
Weight or slice training data by relevance to that decision, not by how much of it exists.Why: an abundant, accurate, irrelevant pile still trains a model that answers confidently and wrong.
Set a quality bar, clean values and correct labels, as a pass or fail gate before anything gets weighted.Why: quality keeps garbage out, but it never guarantees the data is about the right thing.
Track relevance coverage as its own number, separate from total volume.Why: a relevance gap hides behind a healthy-looking blended average for weeks.
Only once relevance and quality are covered, spend more effort growing volume.Why: volume is the cheapest lever and the least likely to be the actual problem.
Say plainly where raw volume really is the right lever, like giving a broad, general feature wider coverage.Why: shows judgment instead of treating relevance as the fix for every data question.
How to answer this, stage by stage
Nobody is scoring whether you can define three words. They're scoring whether you can catch the moment someone mistakes a big, clean pile of data for the right one.
Stage 1
Scope it to one product and one number
Say it like this
"Let's ground this in Cinder Trail Coaching's training-plan generator, and the actual season where volume and relevance quietly pulled apart."
Why this works
Keeps a definitional question from turning into a dictionary recitation.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as LEAD. Link, the outcome that matters. Early signal, what moves first. Abuse, how it gets gamed. Decision, what you'd do at each level. That's how I'll actually separate these three words instead of just defining them."
Why this works
Signals a repeatable way to tell volume, quality, and relevance apart, not three memorized definitions.
Stage 3
Reframe: it isn't "enough data," it's "enough of the right data for this call"
Say it like this
"The mistake people actually make isn't confusing quality with volume. It's assuming a mountain of clean, accurate data is automatically the right data for the specific decision the model has to make."
Why this works
This is where a live answer stops sounding like a vocabulary quiz.
Stage 4
Give the direct answer, applied to two real athletes
Say it like this
"For the road racers, we had huge volume, clean quality, and genuine relevance, since road pacing is exactly what the plan needs to predict. For the trail runners, the GPS quality was just as clean, but almost none of it touched technical descents, the one thing that actually predicts overtraining risk for them."
Why this works
Makes the three-way distinction concrete instead of abstract.
Stage 5
Prove it with the compressed failure
Say it like this
"We had forty two thousand miles of clean road data and three thousand of clean trail data. The season dashboard looked fine for fourteen weeks. Underneath, six of eight trail runners had quietly gone back to hand-built plans, because the one time it mattered, the data covering their terrain just wasn't there."
Why this works
Compresses the whole story into the one number that shows quality and volume were never the problem.
Stage 6
Name the AI-specific reasoning, then close
Say it like this
"The honest reason this isn't just a vocabulary question is that a model trained on clean-but-wrong data doesn't fail loudly, it answers confidently and gets it wrong on exactly the cases it never really saw. I'd leave the road-racer pipeline alone, since volume and relevance already line up there. For the trail runners, I'd track relevance coverage as its own number instead of folding it into total volume."
Why this works
Closes with real judgment and restates the direct answer in one breath.
Let's learn
The four letters, held up as one page. Early signal is the step this question is really testing.
Cinder Trail Coaching's app reads a runner's GPS watch history and builds their next week of training automatically. Before it existed, Marisol Vance built all thirty of her athletes' plans by hand every Sunday, reading each watch export and adjusting by feel, about three hours total. With the tool, all thirty plans come back in under a minute.
Three different questions about the same pile of data. A pile can pass two of them and still fail the third.
Where the season's logged miles actually came from
Both piles were clean GPS exports. One was fourteen times bigger, and that gap alone tells you nothing about which one the plans actually needed.
Here's the turn: the trail runners' extra mistakes were not about how much data existed or how clean it was. Both pools passed a quality check the same way. What broke was that almost none of the accurate data on hand covered technical descents, the exact terrain that predicts overtraining for a trail runner. Marisol's next move wasn't to demand more data. It was to quietly stop trusting the tool for that slice of her roster.
The data was not thin. It was thin about the one thing that mattered for these eight runners.
The choice I would take back
Nobody ever decided to treat terrain as its own data dimension. Every athlete's watch history got pooled into one training set weighted straight by volume, because at launch twenty seven of the first thirty athletes were road racers and pooling everything was the simplest thing to ship. That made sense then. It stopped making sense the day trail runners joined and their data got graded on the same volume-only scale.
What I would leave alone: I wouldn't touch the road-racer half of the pipeline. Volume-weighted road data really is the relevant data for predicting road pacing, so there's no terrain gap to fix there.
The lesson: relevance is not a fancier word for quality. A pile of data can be completely clean and completely wrong for your product, and the only way to catch that is to ask what specific decision the data is supposed to answer, not how much of it you collected.
Now here is the same thing as a story
The short version above is what you'd say defining these three words in a phone screen. Read this one for how the gap between volume and relevance actually opened up, one Sunday at a time.
Marisol Vance has coached runners for nine years, and can read a GPS trace the way other people read weather, one glance and she knows if a week was too hard.
The third step is where every athlete's data got treated the same, no matter what terrain it actually came from.
When the plan generator launched, Marisol reviewed every one of the thirty plans against each athlete's watch history, carefully, for the first two weeks. They matched what she'd have written by hand, every time, for the road racers especially. By week two she was mostly spot-checking. By week three, after one close call, she had quietly pulled six of her eight trail and ultra runners off the automated plans entirely and gone back to building theirs by hand.
Knowledge spark: why would clean, accurate data still be the wrong data?
A GPS watch records pace, distance, and elevation with real precision, so the numbers themselves are correct. But if almost none of those correct numbers ever came from a steep technical descent, the model has nothing true to learn about what a safe week looks like on that terrain. Accurate and relevant are two separate checks, and a dataset can pass the first while failing the second.
Nobody decided, on any single day, to trust the tool less for trail runners. It just stopped earning that trust one close call at a time.
The night before a technical trail race, Marisol double-checked a plan out of old habit and caught a forty percent week-over-week mileage jump scheduled two days before the race, on the exact descent that had put a different athlete in a walking boot the year before. She rewrote it by hand that night. Nothing about the model changed. What changed was which eight runners she was willing to let it decide for.
Volume and relevance move on separate axes. App logins were abundant and useless for this decision. Trail descents were rare and exactly what mattered.
By week fourteen, the two trail runners still on automated plans both needed an emergency manual override in the same week, and that was the first time the company-wide dashboard, blending all thirty athletes together, showed any sign at all that something had been wrong since week three.
The early signal moved eleven weeks before the dashboard did
Blue is trail-relevant coverage share, the leading signal. Amber is the season-wide override rate everyone was actually watching. Eleven weeks separate the two.
When the plan generator was first built, someone said, "let's just pool everything, we'll have plenty of data soon enough," and it sounded reasonable, since almost every athlete at the time was a road racer anyway.
Rerun the same season with terrain tagged as its own dimension, weighted separately from raw volume: the trail runners' thin coverage shows up as an honest, visible flag in week one instead of hiding inside a healthy blended average. The team hand-labels four hundred more technical-descent examples by week four. The season-wide override rate never clears three percent, because the model was never confidently wrong about terrain it had barely seen. It said so instead.
What I'd tell myself, watching Marisol quietly rebuild eight plans by hand every Sunday again: the forty two thousand clean miles were never in question. It was always going to be the three thousand that decided whether this worked.
LEAD, the four letters that actually separate these three wordsNot a script for deciding you need more data. LEAD is what tells you whether the data you already have is even pointed at the right question.
L
Link. The outcome that actually matters.
Athletes staying injury-free and remaining subscribers through the season, not a model accuracy score on its own.
Without naming the real outcome, volume, quality, and relevance all sound interchangeable.
E
Early signal. What moves first.
The share of an athlete's own training data that is actually terrain-relevant, tracked apart from total volume, which dropped to 25 percent by week three.
This is the hardest step, and the one that actually separates relevance from the other two words.
A
Abuse. How it gets gamed.
Tagging any elevation gain, even a road hill repeat, as trail-relevant would inflate the coverage number without fixing anything real.
A relevance metric needs an honest terrain check, not a self-reported tag.
D
Decision. What you'd actually do at each level.
If an athlete's relevant-coverage share falls under fifteen percent, route their plan to manual review before it ever ships.
A number nobody acts on is a dashboard decoration, not a metric.
The recap, one line per letter: link is athletes staying injury-free and subscribed, early signal is relevant-coverage share dropping to 25 percent three weeks in, abuse is a self-reported terrain tag inflating that number for free, and decision is routing any athlete under fifteen percent to manual review before their plan ships.
And if you want to be sure it really works, try it somewhere elseSame four letters, a public library's reading recommender instead of a running app. Different flip family entirely, the same volume that never was the answer.
Halima Suleiman runs teen services at Fennimore Public Library Consortium, where a reader-advisory tool recommends the next book to check out. Mapped onto LEAD: link is teens actually finishing what they check out, not raw checkout counts. Early signal is the share of the tool's training checkouts that come from the teen section itself, a small slice next to the library's much larger adult-fiction volume. Abuse would be counting any checkout by a card registered to a teen, even one made for a school report nobody chose to read, as a real preference signal. Decision is holding back any recommendation built mostly from adult-fiction volume until enough teen-specific signal exists. The flip here is delegation, not scope: Halima had handed teen picks over to the tool entirely, and reclaimed the job by hand once its huge, accurate, mostly-adult checkout history kept recommending titles no teen actually finished.
A different flip entirely: not a coach narrowing which slice she trusts, but a librarian taking a whole job back by hand.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "volume is how much, quality is how clean, relevance is whether it's about the actual decision, and only the third one was ever missing here," and stop.
Cost: there's no time to hand-label anything before the next release. Say so honestly, and ship a visible warning on any recommendation built from low-relevance data instead of pretending the gap isn't there.
The data turns out to be relevant after all: if an audit shows the "irrelevant" pile actually does cover the decision, that's real news, not a shortcut, and it changes what you'd spend effort fixing next.
Where people run it wrong.
They treat a big number of records as proof the model is ready, without asking what decision those records actually describe.
They chase more volume the moment something goes wrong, when the honest problem was always relevance.
They let a self-reported or loosely-defined tag stand in for real relevance, which quietly recreates the same gap under a new name.
How to use it live. The moment an interviewer asks you to define these three words, skip straight to naming the actual decision the data has to support, then ask whether the biggest pile you have is even about that decision. The rest of the definition falls out of that one question.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Scope flip: the tool ran on the whole roster at first, then the coach's trust in it narrowed down to just the slice of athletes the data actually covered well.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Vance, a running coach at Cinder Trail Coaching with nine years of experience building plans by hand before the tool existed.
3 · THE HABIT
What did Marisol stop doing once the road-racer plans kept matching her own?
Tap to flip
ANSWER
She stopped checking every one of the thirty plans line by line against each athlete's watch history and moved to spot-checking instead.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Running the plan generator across all thirty athletes at once, versus running it only on the twenty two road racers and building the eight trail runners' plans by hand.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pooling every athlete's watch data into one training set weighted straight by volume, with no terrain tag, because twenty seven of the first thirty athletes were road racers.
6 · THE NUMBER
Fill in the blank: trail and ultra runners logged about ___ miles across the season, next to 42,000 from road racers.
Tap to flip
ANSWER
3,100 miles. Fourteen times less volume, and the volume gap alone was never the real problem.
7 · THE REPLAY
Same season, terrain tagged and weighted separately from volume from day one. What changes?
Tap to flip
ANSWER
The trail runners' thin coverage shows up as a visible flag in week one, the team hand-labels 400 more technical-descent examples by week four, and the season-wide override rate never clears 3 percent.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Fennimore Public Library Consortium's teen reading recommender. The flip is delegation: Halima handed teen picks to the tool, then took the job back by hand once mostly-adult checkout volume produced bad teen recommendations.
Check yourself Score: 0 / 0
Multiple choice
1. According to this answer, what actually explains why the trail runners' plans went wrong, even though the underlying GPS data was accurate?
A. The model itself was poorly built and needed retraining.
B. There simply was not enough data volume overall, across every athlete.
C. The accurate data that existed barely covered the terrain that mattered for trail runners.
D. Marisol stopped trusting the tool for no particular reason.
Show hint
Look at the direct answer and the quadrant diagram.
Show answer
C. Quality was fine on both pools. The trail data just was not about the decision, technical descents, that mattered most for those athletes.
True or false
2. True or false: this answer argues that Cinder Trail Coaching should focus on collecting a lot more total training data across all thirty athletes.
True
False
Show hint
Look at the priority list and "what I would leave alone."
Show answer
False. The fix is tagging and weighting by relevance, not growing raw volume. The road-racer pipeline is left alone precisely because its volume was already relevant.
Fill in the blank
3. Fill in the blank: the season-wide override rate stayed near 2 to 3 percent until week ___, eleven weeks after the leading indicator had already dropped.
Show hint
Look at the line chart, "the early signal moved eleven weeks before the dashboard did."
Show answer
Week 14. The relevant-coverage share, the leading indicator, had already dropped to 25 percent by week 3.
Short answer, where it wouldn't matter
4. Name a part of Cinder Trail's pipeline where pooling all the data by volume, with no relevance tagging, genuinely isn't a problem, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The road-racer pipeline. Volume-weighted road data is already the relevant data for predicting road pacing, so there's no terrain gap to fix there.
Short answer, apply it yourself
5. Think of a tool you use that has plenty of accurate data behind it. Name one case where that data might be accurate but not actually about the decision you need it to make.
Show hint
Look for a place where a large, clean history exists but was never tagged for the one distinction that actually matters.
Show answer
Model answer: A spending app's transaction history is accurate and large, but if it's never tagged as planned versus impulse, it can't answer the one question you actually want, no matter how many years of history it holds.
Short answer, work the number
6. If Cinder Trail had caught the relevant-coverage drop the same week it happened, week 3, instead of week 14, roughly how many weeks earlier would the fix have started?
Show hint
Compare the week the leading indicator moved to the week the lagging dashboard moved.
Show answer
Model answer: About 11 weeks earlier, since the leading indicator dropped by week 3 while the season-wide dashboard didn't move until week 14.
Before you close the answer
Why this works
Tests whether you'll treat "we have lots of data" and "our data is clean" as proof you have the right data, or actually check what the data is supposed to be about.
Follow-up traps
"Isn't relevance just another word for quality?" Response: no. Quality asks whether the data is clean and correct. Relevance asks whether that clean, correct data is actually about the decision at hand. A pile can pass one and fail the other.
"Couldn't you just collect a lot more trail data and skip the tagging step?" Response: more raw trail volume helps eventually, but without tagging you still can't tell which examples are actually terrain-relevant, so you'd be guessing which weeks of data to trust in the meantime.
If pressed
The tag that actually shipped wasn't a simple trail-versus-road flag. It graded each segment by descent grade and technicality, since a flat gravel trail behaves like a road for training purposes and a steep technical descent doesn't.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.