CaseAdvancedAI Opportunity & Model Strategy / Data strategy as product strategy / #3

Describe how you would instrument a product to generate training or eval data as a byproduct.

SPARK · the return that went viral because Switchback couldn't explain why a boot recommendation failed twice

Switchback Outfitters sells hiking boots and packs online. Its "Fit Assistant" recommends a size from a few photos and measurements. Callista Doyle owns that feature, and for a while she personally read every return email tagged "wrong size" to look for patterns. Then volume grew twenty six times over, and she couldn't anymore.

The direct answer
Attach the byproduct capture to a step almost everyone already completes, not to an optional survey nobody fills in. At Switchback, that means making the actual fit outcome, too big, too small, wrong width, a required part of the return flow itself, tied to the exact recommendation, measurements, and model version that produced it. If your training and eval data depends on someone volunteering extra effort, you will only ever learn from the small slice of people who had time to complain.
Do this, in order
  1. Attach data capture to a step nearly everyone already completes, not an optional extra one.Why: an optional field gets skipped by exactly the people whose data you need most.
  2. Link every captured outcome back to the specific recommendation, input, and model version that produced it.Why: an outcome with no link back to its cause can't be used to fix the cause.
  3. Make the required field short and structured, not a free-text box nobody wants to fill in.Why: structure is what a model can actually learn from; a paragraph of prose usually isn't.
  4. Design so a silent, no-explanation return still produces some signal.Why: the customers who say nothing are often the ones the model failed hardest.
  5. Keep the capture scoped to what the next model version actually needs.Why: collecting more than you'll use just adds friction without adding signal.
  6. Revisit the schema whenever the product's failure modes change.Why: the fields that caught today's failures won't automatically catch tomorrow's.

How to answer this, stage by stage

Nobody is scoring whether you know the word "instrumentation." They're scoring whether you can name the exact step in the product where the data would actually get captured.

Stage 1
Scope it to one concrete flow
Say it like this
"Let's ground this in Switchback Outfitters' Fit Assistant, and the exact return flow a customer goes through when a boot recommendation was wrong."
Why this works
Keeps the answer from turning into a generic essay about "logging more data."
Stage 2
Say your structure out loud
Say it like this
"I'll run this as SPARK. Situation, how it works today with no instrumentation. Payoff, the habit we want to build. Anchor, the one design decision. Risk, what breaks the day it's wrong. Keep out, what we deliberately don't build yet."
Why this works
Signals a repeatable design method, not a list of logging ideas invented on the spot.
Stage 3
Reframe: the question is "what step does everyone pass through," not "what should we log"
Say it like this
"This isn't really a question of what fields to add. It's a question of which existing step in the product nearly every affected customer already walks through, whether they feel like helping you or not."
Why this works
This is where a strong answer separates from someone who just says "add a feedback survey."
Stage 4
Give the anchor
Say it like this
"Make the actual fit outcome a required part of the return form itself, tied automatically to the recommendation ID, not an optional survey that shows up in someone's inbox two days later."
Why this works
This is the concrete design decision an interviewer should be able to picture on an actual screen.
Stage 5
Prove it with the compressed failure
Say it like this
"A customer's wrong-size return went viral, twice, and Switchback couldn't explain why, because only 22 percent of past wrong-size returns had ever included any real explanation. The optional note field had quietly become useless the moment return volume outgrew what one person could read by hand."
Why this works
Compresses the whole argument into the one number the anchor is specifically designed to fix.
Stage 6
Say what you'd deliberately leave out
Say it like this
"I wouldn't build a full customer preference profile on day one. I'd capture the recommendation, the measurements, and the outcome, and stop there, because that's the only thing the next model version actually needs."
Why this works
Shows judgment about scope, not a wish list of everything that could theoretically be logged.
Stage 7
Close on one line
Say it like this
"Byproduct data isn't a survey people fill out because they feel generous. It's a required field sitting inside a step they were already going to complete anyway."
Why this works
Restates the direct answer, using the anchor itself as the closing line.

Let's learn

Here is what happens when the only path to training data runs through an optional box nobody has a reason to fill in.

Switchback's Fit Assistant recommends a boot or pack size from a few photos and measurements. About 14,000 orders a month use it, and roughly 9.5 percent come back tagged "wrong size." The return form has always had a free-text note for "tell us what happened," entirely optional. Only 22 percent of wrong-size returns ever used it.

Hand sketched flow diagram titled Today, without instrumentation, fourth step emphasized. Four steps left to right: Boots returned. Pick reason code. Optional note, often blank. Never linked to the recommendation.
Four steps, and the fourth one is where two years of potential lessons quietly went nowhere.

For a while, that 22 percent was enough. Callista read every one of those notes herself, about fifty a week, and could tell you which boot model ran narrow and which pack's size chart was off by a size for tall torsos. Then a marketing push tripled order volume in two months, and wrong-size returns climbed from fifty a week toward thirteen hundred.

Here's the turn: the missing notes were never really the problem while volume was small. The real problem showed up the day a customer's video, ordering the "right" size twice on the tool's own recommendation and getting it wrong both times, reached a few hundred thousand views, and Switchback's team went looking for an explanation that thousands of past returns had simply never captured.

Share of wrong-size returns Callista personally reviewed, by week
100% 50% 0 Wk 1 Wk 10 4%
Volume grew twenty six times over in ten weeks. One person's reading capacity did not grow at all.

At its worst, the model kept shipping updated recommendations with nobody actually able to say which specific failure mode it was still getting wrong, because the only record of that failure lived in an optional box that almost nobody, understandably, bothered with.

Byproduct data doesn't fail because customers refuse to help. It fails because it was built to depend on someone's spare effort, and spare effort is the first thing that runs out when volume grows.
The choice I would take back Making the "what went wrong" field optional on the original return form, to keep the form short and quick. That made sense when volume was fifty a week and Callista read every single one anyway. It stopped making sense the moment volume outgrew what any one person could read, and the optional field became the only place that information could have lived.

What I would leave alone: I wouldn't require a detailed explanation for returns tagged "changed my mind" or "ordered a gift, wrong recipient," since those aren't fit-model failures at all, and forcing structure there would just add friction for no signal.

The lesson: if a byproduct signal only exists when someone volunteers it, you don't have an instrumentation plan. You have a hope.

Now here is the same thing as a story

The short version above is what you'd say defending the return-form redesign in a product review. Read this one for how a fifty-person habit quietly stopped scaling with nobody deciding to let it fail.

Callista's Monday mornings used to start the same way: coffee, then the "wrong size" queue, forty or fifty return notes to read before the rest of the team was even online.

She was good at it. She could tell, from two sentences, whether a boot ran narrow in the toe box or whether a customer had simply measured their foot wrong. She kept a running list in her head of which product pages needed a sizing warning, updated almost weekly.

Hand sketched labeled parts diagram titled The anchor, close up, one required field. A document icon at the center labeled Return Form, with four labeled callouts around it: Recommendation ID. Actual fit outcome. Measurements used. Linked automatically.
This is the one screen the whole answer turns on. Not a survey. A required field inside a form the customer was already filling out.

Then the marketing team ran a push that tripled order volume almost overnight. Return volume followed a few weeks later, the way it always does. Fifty notes a week became two hundred, then six hundred, then thirteen hundred. Callista kept reading, then kept reading faster, then started skimming, then started only opening a random handful to "stay in touch with the pattern."

Nobody told her to stop reading everything. She just physically couldn't anymore, and nothing in the product forced anyone else to pick up the slack, because the note field had never been anyone's job to fill in. It was optional for the customer, and reading it had only ever been Callista's personal habit, not a system.

Hand sketched comparison titled The day the recommendation is wrong. Left panel, a question mark box icon labeled BEFORE, caption silent return, nothing captured, ever. Right panel, a gauge icon labeled AFTER, caption outcome logged even if the customer skips the survey, shown in a different color.
Before, a wrong recommendation and a customer who says nothing produce exactly the same record: none at all.

Then a customer named the boot model in a post that went past a few hundred thousand views: recommended a size using measurements and photos, sent back for being too small, sent a second pair in the "corrected" size, also too small. No note field filled in either time, because why would a frustrated customer do Switchback's product research for free.

Knowledge spark: why is a silent return often the most important one? A customer who takes the time to explain a mistake usually still trusts the brand enough to bother. A customer who says nothing and just returns the item may already be the one who's given up. If your only signal comes from people who explain themselves, you are systematically missing the failures that made someone stop caring enough to complain.

The team pulled two years of "wrong size" returns looking for a pattern on that specific boot model. Only 22 percent of them had any usable explanation at all. The rest were closed tickets with a reason code and nothing else, no width, no length direction, no way to tell a narrow-toe complaint from a short-in-the-heel one.

Hand sketched timeline titled The return nobody could explain, second milestone emphasized. Four milestones: Return goes viral, wrong size twice. Team searches records, only 22 percent explained, shown in a different color. Required field ships, tied to recommendation ID. Next quarter, 96 percent captured.
The gap between the first milestone and the second is two years of returns that could have already answered this question.

When the original return form was designed, someone said, "let's keep the note optional so we don't slow people down at checkout," and it sounded completely reasonable, since a fast return flow really was worth protecting at the time.

Hand sketched decision tree titled What ships day one. Root: New Fit Assistant capability idea. Four branches: needs the outcome ground truth leads to Build now. Needs a full preference profile leads to Later. Needs cross category matching leads to Later. Needs predictive reorder timing leads to Later.
Only one branch actually needed building right away. The other three were exactly the kind of scope creep that would have delayed it.

The fix Switchback shipped made one thing required: for any "wrong size" return, pick too big, too small, or right length but wrong width, tied automatically to the recommendation ID and the measurements that produced it. No essay. No survey email two days later. Just one required tap, inside a form the customer was filling out anyway.

Wrong-size returns with a usable fit-outcome record
100% 50% 0 Before, 22% After, 96%
Not because customers suddenly cared more. Because the field stopped being something you had to remember to fill in.

Rerun the same two years with that field required from the start: the boot model's narrow toe box shows up in the data within the first few months, well before volume triples, well before a customer has to make a public video to get anyone's attention. Callista's Monday-morning reading habit becomes a dashboard instead of a bottleneck, and it survives volume growth without anyone needing to personally keep pace with it.

What I'd tell myself, watching a team search two years of records for an answer that was never captured: an optional field isn't a lighter version of a required one. It's a coin flip on whether the exact customers you need to hear from will bother, and the customers who've had the worst experience are the ones least likely to flip in your favor.

SPARK, the anchor that survives the day it's wrongNot a brainstorm of things to log. SPARK is what forces the design to actually work the day the recommendation fails.

S
Situation. How it works today, without instrumentation.
A customer returns a boot, picks a reason code, and maybe types a note nobody is required to read structure into.
Grounds the whole design in a real, existing workflow instead of an imagined one.
P
Payoff. The habit we want to build.
Customers get the right size on the first try, and stop hedging by ordering two sizes "just in case," which is itself a workaround for not trusting the recommendation.
Names what the product should make people stop doing, not just what it should log.
A
Anchor. The one design decision.
A required fit-outcome field inside the return flow itself, automatically tied to the recommendation ID, the input measurements, and the model version.
This is the hardest step, and the one that turns "we'd like more data" into an actual, inspectable screen.
R
Risk. What breaks the first time the model is wrong.
A customer who's frustrated enough to say nothing still produces a structured record, because the field is required, not a survey they can ignore.
Proves the anchor survives exactly the case it exists to catch.
K
Keep out. What we deliberately don't build yet.
No full preference profile, no cross-category style matching, no predictive reorder timing, just the recommendation, the measurements, and the outcome.
Shows judgment about scope instead of a data-collection wish list.

The recap, one line per letter: situation is a return flow where the real reason for a wrong size usually goes uncaptured, payoff is customers trusting one recommendation instead of hedging with two sizes, anchor is a required, linked fit-outcome field inside the return form, risk is that even a silent, frustrated customer still produces a usable record, and keep out is stopping at exactly what the next model version needs.

And if you want to be sure it really works, try it somewhere elseSame five letters, a county permits office instead of an outdoor retailer. Different flip family entirely, the same missing signal.

Ingrid Solvang runs the review team at Cascade County Permits, where reviewers approve, reject, or request changes on building permit applications, writing free-text notes to explain each decision. Two years from now, the county wants a pre-screening assistant that flags likely rejections before a reviewer even opens the file, trained on why past applications actually got rejected. Mapped onto SPARK: situation is a reviewer writing whatever notes they feel like, with no required structure. Payoff is applicants fixing common mistakes before submitting at all, instead of waiting weeks for a rejection to find out. Anchor is a required, structured rejection-reason code, tied to the specific application fields the reviewer flagged, built into the same screen where the reviewer already records their decision. Risk is that the anchor survives an unusual, one-off rejection just as well as a routine one, since the structured code still gets attached either way. Keep out is not trying to auto-approve anything on day one, only flagging likely issues for a human to check.

The flip here is input, not scope. A rumor spread that supervisors were quietly judging reviewers by how detailed their notes were, using note length as a stand-in for effort. Reviewers started writing shorter, blander, more defensible notes, performing for an audience that was never actually reading them that closely, instead of documenting the real judgment call they'd made. The honest, specific reasoning that would have made great training signal, "the survey shows the wrong setback because it was measured from the wrong property line," quietly turned into "does not meet setback requirements," true, but useless for training anything.

Hand sketched flow diagram titled Where the reviewer's real judgment got lost, second step emphasized. Four steps left to right: Reviewer's honest note. Rumor, notes are graded, shown in a different color. Notes shrink to boilerplate. Training signal vanishes.
The second step is where Cascade County's real training signal quietly started dying, months before anyone noticed the notes had gotten shorter.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "attach the capture to a step everyone already completes, tied back to the exact input that caused it, and stop there," and stop.
Cost: no engineering time to build automatic linking before the next release. Say so honestly, and require a manually-entered reference ID as a stopgap until the automatic link ships.
The model got better, for real: if Fit Assistant's overall return rate genuinely drops, that's still worth instrumenting, since the remaining wrong-size returns are now the hardest, rarest cases, exactly the ones you can least afford to lose the signal on.

Where people run it wrong.
They build an optional survey and call it instrumentation, then wonder why response rates are low.
They capture an outcome with no link back to the specific input or model version that produced it, making the data unusable for debugging anything.
They try to capture everything a future feature might someday want, instead of scoping to what the next model version actually needs.

How to use it live. The moment an interviewer asks how you'd instrument a product for training data, ask yourself: what step does nearly every affected user already walk through, no matter how they feel about the product that day? Design the capture into that exact step, and the rest of the answer follows.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Scope flip: Callista went from reading every wrong-size return in full to only skimming a shrinking random sample, once volume outgrew what one person could hold.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Callista Doyle, who owns the Fit Assistant feature at Switchback Outfitters and used to personally read every wrong-size return.
3 · THE HABIT
What did Callista stop doing as volume grew, because it worked for a while?
Tap to flip
ANSWER
Reading every wrong-size return note in full. She went from fifty a week to skimming a random handful out of thirteen hundred.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Reading the whole queue versus reading a small slice of it. No middle setting once volume outran what any one person could physically read.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Making the "what went wrong" note optional on the original return form, to keep checkout fast, back when Callista read every one anyway.
6 · THE NUMBER
Fill in the blank: only ___ % of wrong-size returns had a usable explanation before the required field shipped.
Tap to flip
ANSWER
22 percent. After the required field shipped, that rose to 96 percent.
7 · THE REPLAY
Same two years, the fit-outcome field required from day one. What changes?
Tap to flip
ANSWER
The narrow-toe pattern on the boot model surfaces within months, well before volume triples and before a customer needs to go public to get attention.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Cascade County Permits' review tool. The flip is input: reviewers started writing shorter, boilerplate notes once a rumor spread that note length was being graded.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer recommends adding an optional post-delivery survey as the main way to capture fit-outcome data.
  • True
  • False
Show hint
Look at the direct answer and the "anchor" step.
Show answer
False. The anchor is a required field inside the return flow itself, not an optional survey that depends on the customer's spare effort.
Multiple choice
2. Why does this answer say a silent return is often the most important one to capture data on?
  • A. Silent returns are always fraud.
  • B. The customers who say nothing may already be the ones who've given up, and their reason for failure would otherwise never get captured.
  • C. Silent returns cost the company more money than explained ones.
  • D. They are easier to process for the returns team.
Show hint
Look at the knowledge spark about silent returns.
Show answer
B. If your signal only comes from people who explain themselves, you systematically miss the failures that made someone stop caring enough to complain.
Fill in the blank
3. Fill in the blank: Switchback's return volume grew from about 50 a week to about ___ a week over ten weeks.
Show hint
Look at the line chart, "Share of wrong-size returns Callista personally reviewed, by week."
Show answer
1,300 a week. A twenty-six times increase, while one person's reading capacity stayed exactly the same.
Short answer, where it wouldn't matter
4. Name a return reason where this answer says you would NOT force a detailed structured explanation, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Returns tagged "changed my mind" or "wrong recipient." Those aren't fit-model failures, so forcing structure there adds friction without adding any useful signal.
Short answer, apply it yourself
5. Pick a product you use. What's one step you already pass through, no matter your mood, where a required (not optional) field could capture real training or eval data as a byproduct?
Show hint
Think about a step you complete regardless of whether the experience was good or bad, like closing an app, checking out, or ending a call.
Show answer
Model answer: A rideshare app's "rate this trip" screen, if the specific complaint category were required whenever a rating is below three stars, instead of an optional comment box most riders skip.
Short answer, work the number
6. If the required field captures a usable explanation on 96% of the roughly 133 wrong-size returns Switchback now sees in a typical week (9.5% of 14,000 orders is about 1,330 monthly, or roughly 307 weekly; wrong-size specifically is a smaller slice), roughly how many usable records would you expect versus the old 22% rate on that same weekly volume?
Show hint
Multiply the weekly wrong-size return count by 96% and by 22%, then compare.
Show answer
Model answer: On roughly 130 wrong-size returns a week, 96% gives about 125 usable records versus about 29 at 22%, roughly four times as many usable lessons every single week.
Before you close the answer
Why this works
Tests whether you'll design capture into a step people already complete, or default to bolting on a survey and hoping enough people answer it.
Follow-up traps
"Won't a required field just annoy customers during returns?" Response: one tap on a reason they're already selecting from isn't the same as a paragraph they have to write, which is why the anchor stays short and structured, not a free-text essay.

"What if customers just pick the fastest option instead of the true one?" Response: possible, which is why the field is paired with the actual measurements and recommendation on file, so a mismatched answer can be cross-checked rather than trusted blindly.
If pressed
The shipped version also silently logs the customer's original measurements against the final size that actually fit, when a follow-up purchase in a different size succeeds, giving a second, even quieter byproduct signal with zero extra taps required.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more