CaseIntermediateAI Opportunity & Model Strategy / Data strategy as product strategy / #10
What data would you need to collect before you could personalize an AI feature?
BOUNDthe notebook a stylist kept because the system never asked
Every fitting at Thimble & Hale used to end with Selin Aksoy writing three numbers on an index card by hand, then a fourth thing nobody asked her to write down: what the client actually said about the fabric. Selin is an in-store stylist, and she is the one who ends up keeping the data the company's new AI stylist was never built to collect.
The direct answer
You need four kinds of data, not one: who the customer is, their measurements. What they say they want, stated fabric and fit preferences. What they actually do, their real order and correction history. And enough repeat visits, at least two completed orders, before a pattern is more than a fluke. Skip the second and third, and you haven't built personalization, you've built an automatic recommendation that happens to know someone's name.
Do this, in order
Capture explicit fit-correction history at every fitting, not just the original measurements.Why: a static measurement taken once can't tell you what actually changed the second or third time.
Log stated preferences the moment a client says them out loud, in ten seconds or less.Why: a preference mentioned twice and never recorded is indistinguishable, to the system, from a preference that never existed.
Set a minimum repeat-visit threshold before personalization kicks in, and fall back to a clearly labeled generic recommendation below it.Why: one data point looks like a pattern and usually isn't; two completed orders is where a real correction starts repeating.
Give a range for how much data you need company-wide before trusting personalized output over the generic baseline, not a single invented number.Why: a false-precise figure collapses the first time it's wrong; a range you can defend doesn't.
Sanity-check that range against something the business already knows, like how long a new stylist takes to learn a client base.Why: an estimate that doesn't survive a smell test is a guess wearing a decimal point.
Say plainly where the generic recommendation is still the right call, like a customer's first order.Why: there's no history to personalize from yet, and pretending otherwise is worse than being honest about it.
How to answer this, stage by stage
Nobody is scoring whether you can name "more data" as the answer. They're scoring whether you can count exactly which four things, and defend the number that says when you have enough of them.
Stage 1
Scope it to one company and one recommendation
Say it like this
"Let's ground this in Thimble and Hale's AI stylist, and the actual four kinds of data it would need before a recommendation is really personalized instead of just automatic."
Why this works
Turns an open "what data" question into something you can actually count.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as BOUND. Break it down into the kinds of data needed. Own the numbers behind each one. Use a range instead of a fake-precise figure. Nail a sanity check against something real. Say which assumption would move the answer most."
Why this works
Signals a real estimation method instead of listing data types off the top of your head.
Stage 3
Reframe the question
Say it like this
"The trap is answering with whatever's already in the database, purchase history and a measurement taken once. The real question is what you'd need that a generic size chart doesn't already give you for free."
Why this works
Separates a real answer from restating what any e-commerce system already logs.
Stage 4
Give the build-up, out loud
Say it like this
"Four things: who they are, their measurements. What they say they want, stated fabric or fit preferences. What they actually do, their real order and correction history. And enough repeat visits, I'd say two completed orders at minimum, before a pattern is more than a fluke."
Why this works
This is the direct answer, said as an actual list instead of a vague nod to "more data."
Stage 5
Prove it with the compressed failure
Say it like this
"Thimble and Hale personalized on just two of those four, measurements and purchase category. For repeat customers, recommendations got accepted with no fit change only 38 percent of the time, worse than the 40 percent baseline for brand-new customers getting the plain generic recommendation."
Why this works
Shows what happens, concretely, when you skip two of the four inputs instead of just naming them as important.
Stage 6
Give the range and the sanity check, then close
Say it like this
"Back-testing puts the real threshold somewhere between 150 and 400 logged fit corrections company-wide before personalized suggestions beat the generic baseline, roughly two to four and a half weeks at current order volume, about how long a new stylist takes to get a feel for a client base this size. If fit corrections turn out to track fabric more than individual bodies, that number drops by more than half."
Why this works
Closes with an honest range and the one thing that would move it, instead of a confident guess.
Let's learn
The five letters, held up as one page. Use a range is the step this question is really testing.
Thimble and Hale's AI stylist reads a customer's profile and order history and recommends fabric, fit, and a next-order suggestion. At launch it used exactly two fields: measurements taken once at signup, and a customer's last three purchase categories.
Four things a real personalization answer needs. Thimble and Hale launched with exactly two of them.
Recommendations accepted with no fit change, new customers versus repeat customers
A "personalized" recommendation, missing two of the four real inputs, did worse than the plain default it was supposed to beat.
Here's the turn: the extra misses were not about the model being poorly built. Say plainly what happened next: repeat customers kept getting the same confident, unchanged recommendation their measurements had always produced, and Selin, without being asked to, started keeping her own record of what actually needed correcting.
Automatic and personalized are not the same word. The system was automatic for every customer, and personalized for none of them.
The choice I would take back
The launch scope only captured measurements and purchase category, because those already existed in the database. Nobody added a field for fit corrections or stated preferences, since collecting them meant asking stylists to log something extra. That made sense when personalization was a nice-to-have add-on. It stopped making sense the moment the feature started making confident recommendations to repeat clients based on nothing that had actually changed about them.
What I would leave alone: a brand-new customer's first recommendation should stay generic. There's no history to personalize from yet, and the fallback there was never the problem.
The lesson: personalization isn't a modeling problem first. It's a "did we ever collect the one signal that makes this person different from the average" problem, and no amount of model tuning fixes a feature that was never given that signal to begin with.
Now here is the same thing as a story
The short version above is what you'd say estimating this live, on a whiteboard. Read this one for how Selin's private notebook actually filled up, client by client, while the official system stayed at two fields.
Selin Aksoy has fitted clients at Thimble and Hale for four years, and can usually tell from how someone stands in a jacket whether the shoulder needs work before they say a word.
The fourth step is where the real signal exists, every single fitting, and the system never once asked for it.
When the AI stylist launched, Selin used it the way it was meant to be used, pulling up its recommendation before every consultation. For new clients it looked reasonable, a fair guess from a body type and a stated budget. For her repeat clients, the ones she'd fitted three or four times, it kept suggesting the same fabric weight and cut it had suggested the first time, no matter what she'd actually adjusted at pickup since.
Knowledge spark: why would a static measurement stop being enough?
A measurement taken once describes a body on one day. A fit correction, letting out a seam, shortening a sleeve, choosing a lighter fabric weight than the size chart suggests, describes what that measurement actually missed. Without capturing corrections, the system keeps repeating its first guess forever, no matter how many times a person tells it, through their actions, that the guess was slightly off.
By her third month with the tool, Selin had started jotting client fit notes on index cards again, quietly, the same habit she'd had before the AI stylist existed, except now she kept them for eighty five of her roughly one hundred forty assigned repeat clients, since the system had no field for any of it.
The real personalization data existed the entire time. It just lived on index cards instead of in the system that needed it.
The near miss came with a longtime client ordering a large custom suit set for a family event. The AI stylist recommended the same heavier wool weight it always had, for the third order running, even though the client had twice, out loud, told Selin she preferred something lighter for warm-weather events. Selin caught it the night before cutting began, from her own notes, not from the system, and the order was corrected before anything shipped.
The honest answer was never one number. It was a range, and a way to know which end you were closer to.
When the personalization feature was first scoped, someone said, "let's start with what we already have, measurements and order history, and add more later if we need to," and it sounded reasonable, since building new data-entry steps for forty stylists is real work nobody wanted to ask for without proof it mattered first.
Back-tested personalization accuracy as fit-correction records accumulate
Somewhere between 150 and 400 logged corrections, the personalized model actually earns the word. Below that, it's a guess with a name attached.
Rerun the same launch with fit corrections and stated preferences captured from day one, and the two completed orders threshold in place: below that threshold, a client sees a clearly labeled generic recommendation, no worse than what they'd have gotten anyway. Above it, the model has real signal to work from, not just a static number from a first fitting. The heavier wool suggestion never reaches a third order unchallenged, because the system already knows what Selin's notebook knew.
What I'd tell myself, watching a stylist quietly rebuild the exact tracking habit the AI tool was supposed to replace: the model was never the missing piece. The missing piece was ever asking for the data that would have made it personal.
BOUND, the four things and the number that says when you have enoughNot a script for demanding infinite data before shipping anything. BOUND is what tells you exactly which four inputs matter and how much of each is enough.
B
Break it down. The equation, out loud.
Personalization quality equals identity data, plus stated preference, plus fit-correction history, plus enough repeat visits for a pattern to be real.
Naming the four parts is what stops "more data" from being the whole answer.
O
Own numbers. Where each figure comes from.
Two completed orders as the minimum before a fit-correction pattern is more than a fluke, since that's when a repeated correction starts looking intentional.
Stating where an assumption comes from is what makes it defensible instead of invented.
U
Use a range. Low and high, not one guess.
Somewhere between 150 and 400 logged fit-correction records company-wide before personalized suggestions reliably beat the generic baseline for a new fabric type.
This is the hardest step, and the one that keeps the estimate honest instead of falsely precise.
N
Nail the sanity check. Compare to something known.
That range is roughly two to four and a half weeks of current order volume, close to how long a new stylist takes to get a feel for a client base this size.
A number that survives comparison to something real is a number worth trusting.
D
Direction. What assumption moves it most.
If fit corrections track fabric type more than individual bodies, the requirement drops by more than half, since you could personalize by fabric cohort instead of by person.
Naming the one lever that would change the estimate most is what a good estimator says out loud.
The recap, one line per letter: break it down is the four kinds of data personalization actually needs, own numbers is two completed orders as the pattern threshold, use a range is 150 to 400 records company-wide, nail the sanity check is two to four and a half weeks against a stylist's own learning curve, and direction is fabric-driven corrections cutting the requirement in half.
And if you want to be sure it really works, try it somewhere elseSame five letters, a container port instead of a tailoring shop. Different flip family entirely, the same four kinds of data still doing the real work.
Soren Falk schedules berths at Kilbride Port Authority, where a tool personalizes maintenance and scheduling alerts to each vessel captain based on their history. Mapped onto BOUND: break it down is vessel identity, a captain's stated scheduling preference, their real observed behavior, and enough past port calls, Soren settled on three, before trusting a pattern. Own numbers ties that three-call minimum to how long it takes a captain's actual berth-time choices to repeat rather than look random. Use a range says somewhere between 40 and 90 logged calls fleet-wide before the model's personalized timing suggestions beat a fixed schedule. Nail the sanity check compares that to roughly a season of regular traffic on the busier routes. Direction says the estimate would fall sharply if preference data turned out to be more reliable than it actually was. The flip here is input, not workaround: once captains realized a stated "flexible" preference got them priority queueing, several started marking every visit flexible regardless of their real constraints, performing for the form instead of reporting an honest preference, and the system had no way to tell a true answer from a gamed one until Soren cross-checked it against what each vessel actually did on arrival.
A different flip entirely: not a stylist quietly filling a gap, but captains quietly gaming a form the system had no way to check.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "identity, stated preference, real behavior, and enough repeats to trust a pattern, that's the whole build-up," and stop.
Cost: there's no budget to add new data-entry fields before the next release. Say so honestly, and ship the repeat-visit threshold and the honest generic fallback first, since that alone stops the worse-than-baseline problem.
The data turns out to need less: if back-testing shows fabric type predicts corrections better than individual history, that's real news, and it should shrink the collection requirement, not just add to it.
Where people run it wrong.
They personalize on whatever's already in the database and call it done, without checking whether that data actually differs from the average.
They skip the repeat-visit threshold entirely, so a single fluke correction gets treated as a confirmed pattern.
They collect a stated preference once and never check it against what a person actually does, which is exactly what let the gaming go unnoticed.
How to use it live. The moment an interviewer asks what data you'd need, count out loud: identity, stated preference, real behavior, and a repeat threshold. If you can only name one or two of those four, say so, and say which one you'd go collect first.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Workaround flip: with no field for fit corrections in the official system, Selin quietly rebuilt her own private notebook to track what the tool was never asked to capture.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Selin Aksoy, an in-store stylist at Thimble and Hale who has fitted clients there for four years.
3 · THE HABIT
What did the official system never ask Selin to do?
Tap to flip
ANSWER
Log a client's fit corrections or stated preferences anywhere the AI stylist could actually see them.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Trusting the official system's recommendation as it stood, versus quietly keeping a private notebook of fit corrections the system never asked for.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Launching personalization on only the two fields already in the database, measurements and purchase category, because adding new data-entry fields felt like unneeded work at the time.
6 · THE NUMBER
Fill in the blank: repeat customers accepted the AI's recommendation with no fit change only ___ percent of the time, worse than the 40 percent baseline for new customers.
Tap to flip
ANSWER
38 percent.
7 · THE REPLAY
Same launch, all four data types and the two-order threshold built in from day one. What changes?
Tap to flip
ANSWER
Clients below the threshold see a clearly labeled generic recommendation, no worse than before. Above it, the model has real correction data to work from, and a third unchallenged wrong recommendation never happens.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Kilbride Port Authority's berth-scheduling alert. The flip is input: captains started marking every visit "flexible" to game priority queueing, instead of honestly reporting their real constraints.
Check yourself Score: 0 / 0
Short answer, name the four
1. Name the four kinds of data this answer says a real personalization feature needs, without looking back at the direct answer.
Show hint
Think identity, what they say, what they do, and how much history exists.
Show answer
Model answer: Identity and measurements, stated preferences, real order and fit-correction behavior, and enough repeat visits, at least two completed orders, for a pattern to be trustworthy.
Multiple choice
2. Why did the AI recommendation actually do worse than the generic baseline for repeat customers?
A. The underlying machine learning model was outdated and needed retraining.
B. It confidently repeated a static first-fitting profile with no fit-correction or stated-preference data to learn from.
C. Repeat customers were pickier and harder to satisfy than new customers in general.
D. Selin stopped using the tool for her repeat clients entirely.
Show hint
Look at the bar chart and "here's the turn."
Show answer
B. The tool had no fit-correction or preference data to draw on, so it kept confidently repeating the same static profile, which felt worse than an honest generic guess.
True or false
3. True or false: this answer argues that a brand-new customer's first recommendation should also be held to a personalized standard.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. A first-time customer has no history to personalize from, so a clearly labeled generic recommendation is the right call there, not a shortcoming.
Fill in the blank
4. Fill in the blank: back-testing suggests somewhere between 150 and ___ logged fit-correction records, company-wide, before personalized suggestions reliably beat the generic baseline.
Show hint
Look at the line chart, "back-tested personalization accuracy as fit-correction records accumulate."
Show answer
400 records. At current order volume that's roughly two to four and a half weeks of company-wide data once the missing fields start being captured.
Short answer, apply it yourself
5. Think of an app that claims to personalize something for you. Which of the four data types, identity, stated preference, real behavior, or enough repeat history, do you think it's actually missing?
Show hint
Look for a recommendation that keeps repeating something you've since corrected or ignored.
Show answer
Model answer: A music app that keeps recommending a genre you skipped past a dozen times is likely missing real behavior data, treating an early stated preference as permanent instead of updating from what you actually do.
Short answer, work the number
6. If Thimble and Hale's order volume doubled to about 180 orders a week, roughly how long would it take to reach the low end of the 150-record range?
Show hint
150 records divided by the new weekly order volume.
Show answer
Model answer: About one week, roughly half the original 1.7 weeks at 90 orders a week, since doubling the order volume roughly halves the time to reach the same number of records.
Before you close the answer
Why this works
Tests whether you can turn "we need more data" into an actual, countable answer: which kinds of data, how much of each, and what happens below that threshold.
Follow-up traps
"Isn't purchase history already personalization?" Response: purchase history says what category someone bought, not what was wrong with it afterward, so it can't learn from a mistake it was never told about.
"Why not just collect everything and let the model figure out what matters?" Response: collecting fit corrections and stated preferences costs a stylist real time at every fitting, so the honest answer names the four things worth that cost, not an unlimited wish list.
If pressed
The two-order threshold isn't fixed forever. It's set per fabric category, since a heavier fabric with more visible fit issues reaches a trustworthy correction pattern faster than a simple, forgiving one.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.