ConceptFoundationalAI Opportunity & Model Strategy / Data strategy as product strategy / #1

Explain why the data you collect today determines the products you can build in two years.

LEAD · the afternoon a new hire asked for two years of clean defect data at Fretwell

Fretwell is a marketplace for buying and selling used instruments. Marisol Quintrell has run trust and safety there for five years, resolving disputes when a buyer says an instrument did not arrive as described. When leadership asked for two years of clean defect data to train a new feature, she found out exactly how much of that history was usable.

The direct answer
Decide now which product you might want to build two years out, and instrument specifically for that outcome today, because unlabeled history cannot be recovered later. A closed ticket, a finished order, or a resolved case is not training data on its own. It only becomes training data if someone captured the outcome in a form a model can learn from, at the moment it happened. Years of volume with no structure behind it is not an asset. It is years you cannot get back.
Do this, in order
  1. Decide what future product needs proving, then instrument for that exact outcome now.Why: unlabeled history from the past cannot be rebuilt later at any price.
  2. Track a leading data-health number, like tag or label completion rate, next to volume.Why: volume can look fine for months while the part that actually teaches a model quietly disappears.
  3. Make structured capture the default step, never one a busy person can skip.Why: whatever is optional under time pressure is the first thing people stop doing.
  4. Treat "we have years of history" as a claim to test, not a fact to plan around.Why: a roadmap promise built on assumed data can turn into a five-month rebuild the moment someone checks it.
  5. Budget backfill and relabeling as its own line item, apart from the new feature's timeline.Why: old records rarely turn into clean labels for free, and pretending they will blows up ship dates.
  6. Revisit what fields are required at least once a year as the roadmap changes.Why: what counted as enough data for last year's feature will not match what next year's feature needs.

How to answer this, stage by stage

Nobody is scoring whether you know that "data matters." They are scoring whether you can name, out loud, the exact number that would have told you two years early.

Stage 1
Ground it in one product and one afternoon
Say it like this
"Let's put this at Fretwell, an instrument marketplace, on the exact afternoon a new data hire asked for two years of defect data and got a very quiet answer back."
Why this works
Stops the answer from turning into a general lecture about how "data matters."
Stage 2
Name the future product out loud, before judging today's data
Say it like this
"Before I say anything about today's data, I need to name what we're building in two years. Say it's a model that flags a listing's condition risk before a human ever looks at it."
Why this works
You can't judge whether today's data is enough until you know exactly what it has to be enough for.
Stage 3
Reframe: this is a leading-signal problem, not a volume problem
Say it like this
"This isn't really a 'do we have enough data' question. It's a 'what would have told us we were losing the usable part, months before we needed it' question. That's a leading indicator, not a row count."
Why this works
This is the line that shows you're running LEAD, not just restating the prompt back at the interviewer.
Stage 4
Give the direct answer
Say it like this
"I'd decide the target capability now, and instrument specifically for it, because two years of unlabeled tickets can't be turned into labeled ones after the fact."
Why this works
This is the one sentence an interviewer should be able to write down and walk away with.
Stage 5
Prove it with the compressed failure
Say it like this
"At Fretwell, tag completion on dispute tickets fell from 92 percent to 11 percent over two years, because the field defaulted to 'skip' once an SLA metric came in. Nobody noticed, because ticket volume itself kept climbing the whole time."
Why this works
This is the number the whole argument would fall apart without.
Stage 6
Say what you'd measure going forward
Say it like this
"I'd put tag completion on the same dashboard as ticket volume, with a floor, so a drop below 60 percent triggers a fix before two more years pass."
Why this works
Shows you think past the one incident, toward a standing signal someone actually watches.
Stage 7
Say what you'd leave alone
Say it like this
"I wouldn't force structure onto the free-text notes field. Nobody was ever going to bulk-train a model on raw prose, so there's no leading signal to protect there."
Why this works
Shows judgment instead of blanket paranoia about every unstructured field in the product.
Stage 8
Close on one line
Say it like this
"The real asset was never the ticket count. It was whether each ticket carried a label somebody could train on, and that number had been disappearing in plain sight for two years."
Why this works
Restates the direct answer in one breath, with the number doing the closing work.

Let's learn

Here is what a two-setting switch, buried inside a support tool, can do to a roadmap two years later.

Fretwell lets people buy and sell used instruments. When a buyer says an item didn't arrive as described, a support agent resolves it: full refund, partial refund, or the buyer keeps it. When the feature launched, every resolved ticket also got a "defect reason" tag, a dropdown next to the close button, so the team could see patterns: cracked bodies, wrong finish, missing hardware, broken electronics. In year one, ticket volume was light, about 40 a week, and 92 percent of them got tagged.

Hand sketched icon list titled LEAD in one screen. Four rows: Link, the outcome that actually matters, gauge icon. Early signal, what moves weeks before it, document icon. Abuse, how the number gets gamed, question mark box icon. Decision, what you would do at each threshold, scale icon.
The four letters, held up as one page. Early signal is the step this question is really testing.

Two years later, volume had grown fifteen times over, to about 600 tickets a week. Somewhere in between, an SLA metric came in, timing how fast an agent closed a ticket. The tag dropdown added a few seconds to every close. So a redesign made tagging optional, defaulting to "skip" unless an agent went out of their way to fill it in.

Here's the turn: the untagged tickets themselves were never the problem. Nobody was harmed by a missing dropdown value. The real problem showed up two years later, the day a new data hire was asked to pull two years of labeled defect examples for a new feature, and found that only 11 percent of tickets carried a usable tag.

Tag completion on dispute tickets, by quarter
100% 50% 0 Q1 Q4 Q8 SLA redesign, tag defaults to skip 11%
Tag completion started dropping the same quarter the SLA redesign shipped, four quarters before anyone thought to check it.

At its worst, the project to build a condition-risk model, sold to leadership as a six-week build, becomes a five-month scramble to hand-relabel old tickets, since only one in nine of the 230,000 resolved tickets in that window carried anything a model could learn from.

The real asset was never the ticket count. It was whether each ticket carried a label somebody could train on. And that number had been vanishing for two years before anyone thought to look.
The choice I would take back When the SLA redesign shipped, tagging defaulted to "skip" instead of "required," to save a few seconds per ticket and hit the new close-time target. That made complete sense when the team was drowning in a volume spike. It stopped making sense the moment "skip for now" quietly became "skipped for two years, for almost every ticket."

What I would leave alone: I wouldn't force the same structure onto the free-text notes field on the same ticket. Nobody was ever going to bulk-train a model on raw prose, so there was never a leading signal worth protecting there in the first place.

The lesson: data collected today is only worth what it can prove two years from now. A field that's optional under pressure will get skipped under pressure, and by the time anyone checks, the two years you needed are already gone.

Now here is the same thing as a story

The short version above is what you'd say defending a data-instrumentation budget in a planning review. Read this one for how a two-second dropdown quietly disappeared over two years with nobody deciding it on purpose.

Marisol's laptop has had the same crack running along its hinge for three years. She keeps a sticky note over the webcam and never once bothered fixing the crack, because the machine still opens the ticket queue fine, and that's all it needs to do.

Hand sketched timeline titled The two years the tag quietly disappeared, second milestone emphasized. Four milestones: Field launched, 92 percent tagged. SLA redesign, tag defaults to skip, shown in a different color. Year 2, 11 percent tagged. New hire asks, no answer exists.
Nobody decided, on any single day, to stop tagging tickets. The default just quietly won, one skipped dropdown at a time.

For the first year, tagging a ticket before closing it was just part of the job. Forty tickets a week, each one closed with a refund decision and a defect reason. Marisol liked the pattern reports it produced. She could tell leadership, with a straight face, exactly which instrument categories had the worst packaging problems that month.

Then the marketplace grew. Ticket volume climbed toward 600 a week, and an SLA metric arrived, timing how fast a ticket got closed. The tag dropdown, once a two-second habit, started to feel like the one extra click standing between an agent and their number. A redesign made it skippable. Nobody announced it as a data decision. It was announced as a productivity fix.

Hand sketched comparison titled Two clocks ticking at different speeds. Left panel, a gauge icon labeled Ticket volume, caption looks healthy the whole time. Right panel, a gauge icon labeled Tag completion, caption already collapsing months earlier, shown in a different color.
Ticket volume told a fine story the whole time. Tag completion was already telling a different one, months before anybody checked it.

By the second year, tagging had thinned from a habit into an exception. A handful of senior agents still filled it in out of old habit. Most didn't, because nothing downstream ever seemed to care whether they did.

Then a new data hire, building a model to flag risky listings before a human ever saw them, asked Marisol for two years of labeled defect examples by category. She pulled the number. Eleven percent.

Knowledge spark: why doesn't "we have 230,000 tickets" mean "we have training data"? A resolved ticket only teaches a model something if the outcome was captured in a form the model can learn from, at the time it happened. A closed ticket with no reason code is just a record that something ended, not a record of what the pattern was. Volume and usable signal are two different numbers, and they can move in opposite directions at once.
Hand sketched metaphor scene titled 230,000 tickets is not 230,000 lessons. Left, a document icon labeled RESOLVED, caption 230,000 tickets closed. Right, a person icon walking away labeled USABLE, caption 1 in 9 has a real label, shown in a different color.
The count on the dashboard kept climbing. The part that actually taught anything was walking away the whole time.

Nobody at Fretwell had lied about the data. Nobody had hidden anything. The tickets were all there, sitting in the database, resolved and closed exactly as they should have been. What wasn't there was the one field that turned a closed case into a lesson.

Hand sketched labeled parts diagram titled What a usable record actually needs. A document icon at the center labeled Usable Record, with four labeled callouts around it: Seller photo. Defect tag. Verified outcome. Linked ticket ID.
Four small parts. Fretwell had three of them on every ticket. The one it defaulted to skipping was the one the whole model depended on.

When the SLA redesign was proposed, someone in the room said, "let's just make the tag optional so agents can hit their close-time targets," and it sounded completely reasonable, since the close-time number was the one everyone was being measured on that quarter.

Planned build time vs. actual, once the data was checked
24 wks 12 wks 0 Planned, 6 wks Actual, 22 wks
Sixteen extra weeks, almost all of it spent hand-relabeling two years of tickets that a required field would have labeled for free.

Rerun the same two years with tagging required, not skippable: the SLA redesign still ships, agents still hit their close-time target, but the tag takes two seconds inside the same flow instead of being optional. Two years later, 92 percent of 230,000 tickets carry a usable label, about 211,600 of them, instead of roughly 25,300. The condition-risk model starts training in week one instead of week seventeen.

What I'd tell myself, watching a new hire discover an eleven-percent number nobody had been watching: the close-time metric was never wrong to add. It was wrong to let it quietly override the one field the next two years of the roadmap would depend on, without anyone ever being asked to make that trade on purpose.

LEAD, the four letters that would have caught this in month threeNot a lecture on why data matters. LEAD is the specific test for whether today's number would have warned you before the roadmap needed it.

L
Link. The outcome that actually matters.
Being able to ship a condition-risk model in two years that flags listings before a human review, without a multi-month relabeling delay.
Without naming this first, "is the data enough" has no target to be measured against.
E
Early signal. What moves weeks before the outcome does.
Tag completion rate on resolved tickets, which started falling the same quarter the SLA redesign shipped, four quarters before anyone checked it.
This is the hardest step, and the one nobody at Fretwell was watching.
A
Abuse. How this metric gets gamed.
Reporting "230,000 resolved tickets" as if raw ticket count proves the data is ready, when only 25,300 of them carry a usable tag.
Every metric has a way to look healthy without doing the actual work; volume is the easiest one to fake yourself out with.
D
Decision. What you'd actually do at each threshold.
Below 60 percent tag completion, tagging becomes required again immediately and a backfill budget gets opened; above it, the field stays as is.
A metric nobody acts on is just a number on a dashboard nobody reads.

The recap, one line per letter: link is shipping a condition-risk model without a multi-month relabeling delay, early signal is the tag completion rate that started falling four quarters before anyone noticed, abuse is mistaking raw ticket count for usable training data, and decision is making tagging required again the moment completion drops below 60 percent.

And if you want to be sure it really works, try it somewhere elseSame four letters, a translation agency instead of a marketplace. A different flip family entirely, the same missing signal.

Henrik Aasen leads review at Meridian Language Group, a translation agency. Their pipeline runs a machine translation draft, a human post-edit pass, then a final review before a document ships to a client. Two years from now, Henrik's team wants to build a quality-estimation model that flags which drafts need a full post-edit and which are close enough to ship lightly reviewed, trained on how much editors actually changed each draft. Mapped onto LEAD: link is shipping that quality-estimation feature without guessing at what "close enough" means. Early signal is the size of the edit, measured consistently, between the raw draft and the final shipped text, for every single document. Abuse is counting "number of documents reviewed" as proof of quality, when the edits themselves, the actual substance, might never get captured at all. Decision is requiring every edit to happen inside the tracked tool, not a side document, so the signal exists at all.

The flip here is delegation, not abandonment. When the machine translation engine improved, Henrik started handing the first post-edit pass down to junior linguists instead of doing it himself, since the drafts needed less fixing. Some juniors, more comfortable in their own word processor, started pulling drafts out of the tracked tool, editing there, and pasting the final text back in. The edit itself, the actual signal the future model needed, vanished the moment it left the tool. Henrik didn't notice until a client complained about inconsistent terminology, pulled the work back onto his own desk, and discovered two junior linguists' worth of edits had never been captured anywhere at all.

Hand sketched flow diagram titled Where the edit signal quietly got lost, third step emphasized. Five steps left to right: Draft, machine translation. Junior post-edit. Off tool rewrite. Senior review. Ship to client.
The third step is where two years of potential training signal quietly left the building, one pasted-in paragraph at a time.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "name the future product, then track the one number that would tell you months early whether today's data supports it," and stop.
Cost: no budget to build a full data-health dashboard before the next feature ships. Say so honestly, and track the single leading number by hand in a weekly export until there's budget for more.
The model got better, for real: Fretwell's dispute rate actually dropped as instrument photos improved. That's still a perturbation. Fewer disputes meant fewer chances to tag anything at all, so the leading signal needed watching even harder, not less.

Where people run it wrong.
They watch volume, like ticket count or documents processed, and mistake it for a proxy for usable data.
They let a structured field go optional under time pressure without ever asking what future capability depended on it staying required.
They discover the gap only when someone downstream asks for the data, instead of watching a leading number the whole time.

How to use it live. The moment an interviewer asks why today's data shapes tomorrow's product, ask yourself: name the thing you'd want to build in two years, then ask what number, tracked today, would have told you months early whether you're actually collecting what that thing needs. Answer those two, and the rest of the response writes itself.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Abandonment flip: agents quietly stopped filling in the defect tag once it became optional, with no complaint and no ticket ever filed about it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Quintrell, who has run trust and safety at Fretwell for five years and resolves buyer disputes.
3 · THE HABIT
What did agents stop doing because nothing downstream ever seemed to need it?
Tap to flip
ANSWER
Filling in the defect reason tag before closing a dispute ticket, once the SLA redesign made it skippable.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Tagging every resolved ticket versus tagging almost none of them. No middle setting once the field defaulted to skip and nothing pulled agents back to it.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Defaulting the defect tag to "skip" during the SLA redesign, to save a few seconds per ticket against the new close-time target.
6 · THE NUMBER
Fill in the blank: tag completion on Fretwell's dispute tickets fell from 92% to ___ over two years.
Tap to flip
ANSWER
11 percent, meaning only about 25,300 of 230,000 resolved tickets carried a usable label.
7 · THE REPLAY
Same two years, tagging required instead of optional. What changes?
Tap to flip
ANSWER
About 211,600 of the 230,000 tickets carry a usable label instead of 25,300. The condition-risk model starts training in week one instead of week seventeen.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Meridian Language Group's translation pipeline. The flip is delegation: Henrik reclaimed post-edit review after junior linguists' edits stopped being captured at all.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: tag completion on Fretwell's dispute tickets fell from 92% to ___ over two years.
Show hint
Look at the line chart, "Tag completion on dispute tickets, by quarter."
Show answer
11%. The drop started the same quarter the SLA redesign shipped, four quarters before anyone checked it.
Multiple choice
2. What was the actual two-setting switch in Marisol's team, according to this answer?
  • A. The model's accuracy at predicting instrument defects.
  • B. Whether agents tagged the defect reason on a resolved ticket, or skipped it.
  • C. How many hours Marisol personally worked each week.
  • D. Whether buyers filed disputes at all.
Show hint
Look at "the flip, in this story" flashcard.
Show answer
B. Tagging every ticket versus tagging almost none of them, with no middle setting once the field defaulted to skip.
True or false
3. True or false: this answer argues every field in Fretwell's support tool, including free-text notes, should have been made mandatory.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The free-text notes field was left alone on purpose, since nobody was ever going to bulk-train a model on raw prose.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense at the time it was made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Defaulting the defect tag to "skip" once the SLA redesign shipped. It made sense because the team was drowning in a volume spike and needed to hit a new close-time target, but it quietly cost two years of usable labels.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one signal, tracked today, that would tell you months early whether you're collecting what you'd need to build something new with it two years from now?
Show hint
Think about a field or step that's optional today but that a future feature would quietly depend on.
Show answer
Model answer: A recipe app's "did you cook this" follow-up prompt completion rate. If it's falling, a future "predict what you'll actually cook" feature has no ground truth to learn from, long before anyone asks for that feature.
Short answer, work the number
6. If tag completion had stayed at 92% instead of falling to 11%, roughly how many more of the 230,000 tickets would carry a usable label?
Show hint
Compare 92% of 230,000 against 11% of 230,000.
Show answer
Model answer: About 211,600 tickets would carry a usable label at 92%, versus about 25,300 at 11%. That's roughly 186,000 more labeled tickets, the entire training set the new feature actually needed.
Before you close the answer
Why this works
Tests whether you can name today's actual leading indicator instead of gesturing at "more data is better," and whether you understand that unstructured history can't be recovered after the fact.
Follow-up traps
"Couldn't you just have someone relabel two years of tickets after the fact?" Response: partly, at real cost and real error, since a relabeler two years later can only guess at what actually happened from a closed ticket's text. That's a slow, expensive stand-in for capturing the label at the time, not a real substitute.

"Isn't this just 'log everything,' which creates its own kind of noise?" Response: no, LEAD says instrument for the named future outcome, not everything, which is exactly why the free-text notes field was deliberately left alone. Only the fields tied to a real future capability need the discipline.
If pressed
The rule Fretwell put in afterward requires a defect tag before any ticket can close, with a weekly-reviewed exception queue for the rare ticket that genuinely can't be tagged, so completion never silently drops again without someone seeing it inside a week.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more