Explain why the data you collect today determines the products you can build in two years.
Fretwell is a marketplace for buying and selling used instruments. Marisol Quintrell has run trust and safety there for five years, resolving disputes when a buyer says an instrument did not arrive as described. When leadership asked for two years of clean defect data to train a new feature, she found out exactly how much of that history was usable.
- Decide what future product needs proving, then instrument for that exact outcome now.Why: unlabeled history from the past cannot be rebuilt later at any price.
- Track a leading data-health number, like tag or label completion rate, next to volume.Why: volume can look fine for months while the part that actually teaches a model quietly disappears.
- Make structured capture the default step, never one a busy person can skip.Why: whatever is optional under time pressure is the first thing people stop doing.
- Treat "we have years of history" as a claim to test, not a fact to plan around.Why: a roadmap promise built on assumed data can turn into a five-month rebuild the moment someone checks it.
- Budget backfill and relabeling as its own line item, apart from the new feature's timeline.Why: old records rarely turn into clean labels for free, and pretending they will blows up ship dates.
- Revisit what fields are required at least once a year as the roadmap changes.Why: what counted as enough data for last year's feature will not match what next year's feature needs.
How to answer this, stage by stage
Nobody is scoring whether you know that "data matters." They are scoring whether you can name, out loud, the exact number that would have told you two years early.
Let's learn
Here is what a two-setting switch, buried inside a support tool, can do to a roadmap two years later.
Fretwell lets people buy and sell used instruments. When a buyer says an item didn't arrive as described, a support agent resolves it: full refund, partial refund, or the buyer keeps it. When the feature launched, every resolved ticket also got a "defect reason" tag, a dropdown next to the close button, so the team could see patterns: cracked bodies, wrong finish, missing hardware, broken electronics. In year one, ticket volume was light, about 40 a week, and 92 percent of them got tagged.
Two years later, volume had grown fifteen times over, to about 600 tickets a week. Somewhere in between, an SLA metric came in, timing how fast an agent closed a ticket. The tag dropdown added a few seconds to every close. So a redesign made tagging optional, defaulting to "skip" unless an agent went out of their way to fill it in.
Here's the turn: the untagged tickets themselves were never the problem. Nobody was harmed by a missing dropdown value. The real problem showed up two years later, the day a new data hire was asked to pull two years of labeled defect examples for a new feature, and found that only 11 percent of tickets carried a usable tag.
At its worst, the project to build a condition-risk model, sold to leadership as a six-week build, becomes a five-month scramble to hand-relabel old tickets, since only one in nine of the 230,000 resolved tickets in that window carried anything a model could learn from.
What I would leave alone: I wouldn't force the same structure onto the free-text notes field on the same ticket. Nobody was ever going to bulk-train a model on raw prose, so there was never a leading signal worth protecting there in the first place.
The lesson: data collected today is only worth what it can prove two years from now. A field that's optional under pressure will get skipped under pressure, and by the time anyone checks, the two years you needed are already gone.
Now here is the same thing as a story
The short version above is what you'd say defending a data-instrumentation budget in a planning review. Read this one for how a two-second dropdown quietly disappeared over two years with nobody deciding it on purpose.
Marisol's laptop has had the same crack running along its hinge for three years. She keeps a sticky note over the webcam and never once bothered fixing the crack, because the machine still opens the ticket queue fine, and that's all it needs to do.
For the first year, tagging a ticket before closing it was just part of the job. Forty tickets a week, each one closed with a refund decision and a defect reason. Marisol liked the pattern reports it produced. She could tell leadership, with a straight face, exactly which instrument categories had the worst packaging problems that month.
Then the marketplace grew. Ticket volume climbed toward 600 a week, and an SLA metric arrived, timing how fast a ticket got closed. The tag dropdown, once a two-second habit, started to feel like the one extra click standing between an agent and their number. A redesign made it skippable. Nobody announced it as a data decision. It was announced as a productivity fix.
By the second year, tagging had thinned from a habit into an exception. A handful of senior agents still filled it in out of old habit. Most didn't, because nothing downstream ever seemed to care whether they did.
Then a new data hire, building a model to flag risky listings before a human ever saw them, asked Marisol for two years of labeled defect examples by category. She pulled the number. Eleven percent.
Nobody at Fretwell had lied about the data. Nobody had hidden anything. The tickets were all there, sitting in the database, resolved and closed exactly as they should have been. What wasn't there was the one field that turned a closed case into a lesson.
When the SLA redesign was proposed, someone in the room said, "let's just make the tag optional so agents can hit their close-time targets," and it sounded completely reasonable, since the close-time number was the one everyone was being measured on that quarter.
Rerun the same two years with tagging required, not skippable: the SLA redesign still ships, agents still hit their close-time target, but the tag takes two seconds inside the same flow instead of being optional. Two years later, 92 percent of 230,000 tickets carry a usable label, about 211,600 of them, instead of roughly 25,300. The condition-risk model starts training in week one instead of week seventeen.
What I'd tell myself, watching a new hire discover an eleven-percent number nobody had been watching: the close-time metric was never wrong to add. It was wrong to let it quietly override the one field the next two years of the roadmap would depend on, without anyone ever being asked to make that trade on purpose.
LEAD, the four letters that would have caught this in month threeNot a lecture on why data matters. LEAD is the specific test for whether today's number would have warned you before the roadmap needed it.
The recap, one line per letter: link is shipping a condition-risk model without a multi-month relabeling delay, early signal is the tag completion rate that started falling four quarters before anyone noticed, abuse is mistaking raw ticket count for usable training data, and decision is making tagging required again the moment completion drops below 60 percent.
And if you want to be sure it really works, try it somewhere elseSame four letters, a translation agency instead of a marketplace. A different flip family entirely, the same missing signal.
Henrik Aasen leads review at Meridian Language Group, a translation agency. Their pipeline runs a machine translation draft, a human post-edit pass, then a final review before a document ships to a client. Two years from now, Henrik's team wants to build a quality-estimation model that flags which drafts need a full post-edit and which are close enough to ship lightly reviewed, trained on how much editors actually changed each draft. Mapped onto LEAD: link is shipping that quality-estimation feature without guessing at what "close enough" means. Early signal is the size of the edit, measured consistently, between the raw draft and the final shipped text, for every single document. Abuse is counting "number of documents reviewed" as proof of quality, when the edits themselves, the actual substance, might never get captured at all. Decision is requiring every edit to happen inside the tracked tool, not a side document, so the signal exists at all.
The flip here is delegation, not abandonment. When the machine translation engine improved, Henrik started handing the first post-edit pass down to junior linguists instead of doing it himself, since the drafts needed less fixing. Some juniors, more comfortable in their own word processor, started pulling drafts out of the tracked tool, editing there, and pasting the final text back in. The edit itself, the actual signal the future model needed, vanished the moment it left the tool. Henrik didn't notice until a client complained about inconsistent terminology, pulled the work back onto his own desk, and discovered two junior linguists' worth of edits had never been captured anywhere at all.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "name the future product, then track the one number that would tell you months early whether today's data supports it," and stop.
Cost: no budget to build a full data-health dashboard before the next feature ships. Say so honestly, and track the single leading number by hand in a weekly export until there's budget for more.
The model got better, for real: Fretwell's dispute rate actually dropped as instrument photos improved. That's still a perturbation. Fewer disputes meant fewer chances to tag anything at all, so the leading signal needed watching even harder, not less.
Where people run it wrong.
They watch volume, like ticket count or documents processed, and mistake it for a proxy for usable data.
They let a structured field go optional under time pressure without ever asking what future capability depended on it staying required.
They discover the gap only when someone downstream asks for the data, instead of watching a leading number the whole time.
How to use it live. The moment an interviewer asks why today's data shapes tomorrow's product, ask yourself: name the thing you'd want to build in two years, then ask what number, tracked today, would have told you months early whether you're actually collecting what that thing needs. Answer those two, and the rest of the response writes itself.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't this just 'log everything,' which creates its own kind of noise?" Response: no, LEAD says instrument for the named future outcome, not everything, which is exactly why the free-text notes field was deliberately left alone. Only the fields tied to a real future capability need the discipline.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Data strategy as product strategy
- #2 What is a data flywheel and what are its preconditions?
- #3 Describe how you would instrument a product to generate training or eval data as a byproduct.
- #4 Your company has ten years of unstructured documents. Is that an asset? Interrogate the claim.
- #5 How do you evaluate whether proprietary data is actually a moat?
- #6 What are the product implications of not owning your own data?
- #7 Describe the difference between data volume, data quality and data relevance for AI products.