ConceptAdvancedAI Opportunity & Model Strategy / Opportunity identification for AI / #20
How do you avoid the trap of building an AI feature that demos well and gets no repeat usage?
GUARD · a spike that looked like a win
Larderwell is a grocery delivery and personal-shopping app. Its AI plans a week of meals and fills your basket for you. Sixten Groenewald is the product manager who owns that AI shopper. Anzhelika Farkas, Larderwell's VP of Growth, championed the app's newest feature, Surprise Basket, off the strength of one incredible demo. Dulcie Holmqvist, a nurse who has planned her family's meals by hand for six years, tried it once.
The direct answer
Don't grade a new AI feature on how many people try it, or how well it plays in a demo. Set the real bar before you build it: does someone who tries this come back and use it again inside two weeks. Track that number from launch day, and don't call the feature a win, or fund a bigger version of it, until it clears that bar. A feature that dazzles once and never gets opened again isn't a small win. It costs exactly what never building it would have cost, plus the roadmap slot it just spent.
Do this, in order
Set a real retention bar before you build anything, and hold the launch to it.Why: this is the direct answer. Everything below just protects it.
Track day-14 return for every new feature from the day it ships, not just week-one trial.Why: Surprise Basket's 61 percent trial number looked like a win for six weeks, because nobody was watching the number that told the real story.
Feed the model the household's real constraints as a hard filter, before it ever ranks anything for novelty.Why: this is the actual break. A model built to impress once and a model built to fit someone's week need different inputs.
Don't let a demo score or a launch-week spike stand in for proof a feature works.Why: a one-time delight and a repeat habit are answers to two different questions.
Leave the retention bar off features that were only ever meant to be used once.Why: gating a one-time treat the same way just teaches the team the whole rule is decoration.
When a feature misses the bar, say so out loud and move the roadmap, instead of quietly building a bigger version of the thing that missed it.Why: a quiet miss teaches the next launch nothing at all.
How to answer this, stage by stage
Nobody is grading whether you'd feel bad for Dulcie. They're grading whether you can turn "avoid the trap" into a specific bar you'd actually enforce.
1
Put a real feature and a real person on it
Say it like this
"Let me make this concrete. Say a grocery app builds an AI feature that dreams up a new dinner every week and fills your basket for it. Sixty-one percent of shoppers try it in week one. Almost none of them open it a second time. That's the exact shape I'll answer, because 'avoid this trap' means nothing until there's an actual feature, and an actual person who tried it once and put their phone down."
Why this works
Keeps the answer from turning into general advice about shipping fast.
2
Name the method, out loud
Say it like this
"I'll run this as GUARD. Groups, who actually pays when a feature demos well and dies quietly. Unequal, which number gets celebrated and which one gets ignored. Ability to contest, does anyone even check what happens after the first try. Reduce, the actual fix to how we define a win. Detect, the shape in the data that tells you this already happened."
Why this works
Two seconds of structure tells the interviewer you have a method, not a hunch.
3
Say what the trap actually is
Say it like this
"This isn't really about the feature being bad. It demoed great because it was genuinely good at one thing, being new and unexpected. The trap is that 'exciting once' and 'something someone needs every week' get measured by two completely different numbers, and only one of them ever shows up on a launch dashboard."
Why this works
Separates this from a generic launch-failure story and names the real mechanism.
4
Give the fix, plainly
Say it like this
"Here's what I'd actually do. Before we build it, I write down the bar: does someone who tries this come back and use it again inside two weeks. I track that number starting the day it ships. And nobody, including me, gets to call this a win, or fund a bigger version, until it clears that bar."
Why this works
This matches the direct answer word for word. If it doesn't, the interviewer notices before you do.
5
Run the failure forward, in four sentences
Say it like this
"Without that bar, here's what happens. Sixty-one percent of shoppers try the new feature in week one, and leadership scores the demo a nine out of ten. Six weeks later, a routine report finds that only six percent of those shoppers ever opened it again. By then, nine weeks of engineering time is already spent, and the team is two weeks from starting a bigger version of the same idea."
Why this works
This is the story below, compressed to four sentences, so the interviewer hears the whole shape first.
6
Close on the number that would have caught it
Say it like this
"You'll know it's working when nobody can get a big roadmap slot approved off a launch-week number alone, they need the two-week return rate too. You'll know it's slipping when a feature everyone's still talking about at the next all-hands turns out, the moment you actually pull the data, to have a return rate under ten percent that nobody happened to check."
Why this works
Ends on something the interviewer could actually go verify, not a promise that it's handled.
Let's learn
Larderwell is an app that plans your week of groceries and fills the basket for you, so you're not building the list from scratch every Sunday.
Four things Dulcie checks every single week, without fail, because missing even one of them isn't a small mistake for her family.
For six years, Dulcie Holmqvist did that planning herself. She's a nurse, she works rotating shifts, and every Sunday she'd spend about ten minutes with her own list: nothing with peanut in it, nothing that pushes her past a hundred and forty dollars, and mostly the same dozen meals on rotation, because rotation is what actually gets cooked on a Tuesday night after a twelve-hour shift.
Knowledge spark: what's a day-14 return rate?
Of the shoppers who try a feature once, the share who come back and use it again inside the next two weeks. A feature can have a great first week and a terrible day-14 number at the same time, and the two numbers are telling you two completely different things.
For years, Larderwell judged a new feature by one number: how many shoppers try it in week one. That worked fine, because most of what they shipped was small and useful, things like The Regulars, a feature that quietly reorders whatever a household actually buys most weeks. Nineteen percent of shoppers tried The Regulars in its first week. Of those, sixty-eight percent were still using it two weeks later.
Then Larderwell shipped something genuinely new. Surprise Basket is an AI feature that dreams up an unfamiliar dinner every week and fills your basket with what it needs, ingredients you'd never have bought yourself. Sixty-one percent of shoppers tried it in its first week, more than three times The Regulars' number. Anzhelika Farkas put it up on the all-hands screen and called it the best first week any feature had ever had at Larderwell.
Share of week-one triers still opening the feature, weeks 1 through 6
The RegularsSurprise Basket
Both features start at 100 percent by definition, everyone in the line just tried it. One line settles into a real habit. The other one falls off a cliff before week two is even over.
Here's the turn. A feature a lot of people try once is not the same thing as a feature people need. Of the shoppers who tried Surprise Basket in week one, only six percent ever opened it again inside two weeks.
We didn't lose six percent of Dulcie's baskets. We lost the habit before it ever had a chance to start.
The Regulars checks all four of these, every single week. Surprise Basket's model was never told to.
Here's why. Surprise Basket's model was scored on one thing: how new and exciting a suggestion looked to someone scrolling past it. It never checked Dulcie's actual list, no peanut, under a hundred and forty dollars, meals that repeat on purpose. The Regulars checks all four of those, every week, because that's the only way it gets anything right twice in a row.
At its worst, this is nine weeks of engineering time already spent on Surprise Basket, with seven more weeks already booked for a bigger version of it: more themes, more seasonal rotations, more surprise. Meanwhile The Regulars, the one feature that actually earned a place in people's weeks, never got a single one of those weeks, because nobody had a number that said it deserved them.
The choice I would take back
Larderwell's launch dashboard only ever showed one number: the share of eligible shoppers who tried a feature in week one. That was a fine bar back when almost everything Larderwell shipped was small and modest, a feature only really appealed to shoppers who'd go on to use it. It stopped being a fine bar the day they shipped something built to be exciting on the very first try.
What I would leave alone
Larderwell also ships small seasonal extras, like a countdown of stocking-stuffer snack ideas each December. Those were never meant to be opened every week, so a two-week return bar doesn't belong on them at all. Holding a one-time treat to a repeat-usage bar just teaches the team the whole rule is decoration.
The lesson: a feature that wows someone once and a feature that earns a place in their week are answers to two different questions. A demo can only ever answer the first one.
Now here is the same thing as a story
Stage five above compresses this into four sentences. Here's the six weeks underneath it, the part a stand-up answer skips.
Dulcie Holmqvist can build a week of meals for her family in about ten minutes, standing at the kitchen counter with her phone before her Sunday shift starts. Six years of doing it by hand taught her the shortcuts: which dozen meals repeat well, what her son can't touch, and what a hundred and forty dollars actually buys at Larderwell's prices.
When The Regulars shipped, Dulcie turned it on in her first week and mostly forgot about it. It quietly reordered the same rotation she already trusted it with, milk, the peanut-free bread, the chicken thighs, and it kept getting the small stuff right: the week she ran out of coffee two days early, it caught it. She still built the exciting parts of her list by hand. The boring third of it just happened now.
Then, on a Tuesday, Surprise Basket showed up on her home screen. A dinner she'd never have picked herself: a harissa lamb kefta with a yogurt sauce, ingredients she'd never bought, styled in a photo good enough that she ordered it that same afternoon.
The product team decides what gets built next. Dulcie, on the other end of the app, has whatever they decided and nothing else.
It was genuinely great. She and her son ate well that night, and she told a friend about it the next morning.
The following Sunday, Surprise Basket suggested a peanut noodle bowl. Nothing about the recipe knew her son couldn't touch peanut at all. She spent four extra minutes checking every ingredient by hand before she'd trust it near her own kitchen, more time than her old list ever cost her, and skipped it. The Sunday after that, she didn't open Surprise Basket at all. She built her list the way she always had, and let The Regulars handle the rest.
Nobody at Larderwell noticed any of that happening. There was no complaint, no support ticket, nothing on a dashboard turning red. Six weeks after launch, Torgeir Draganescu, a data analyst on Sixten's team, pulled a routine cohort report, not because anything looked wrong, but because it was the report he ran every six weeks for a completely different feature. Surprise Basket had a column in it almost by accident.
Four of these five weeks looked completely fine, if the only thing anyone was watching was the trial number.
Of the sixty-one percent of shoppers who'd tried Surprise Basket in week one, six percent had opened it again by day fourteen. Torgeir read the number twice before he forwarded it to Sixten.
Nine weeks bought the company a great first Tuesday for a lot of people. It never bought a second one.
It was never really about the recipes being wrong. The recipes were fine, sometimes genuinely great, the harissa lamb kefta was real. The real cost was nine weeks of engineering time already spent, and a second version of Surprise Basket, more themes, more seasonal rotations, already seven weeks into its own roadmap slot, being built on top of a number nobody had checked.
The decision Sixten would take back sits in a meeting eight months before any of this, when Larderwell first wrote down what "a successful launch" meant. Week-one trial rate above forty percent counted as a win. It was a sensible bar at the time. Every feature the team had shipped up to that point was the kind of thing you either needed or you didn't, and trial rate and real use tended to move together. Nobody in that meeting had a reason yet to ask what happens when a feature is built to be exciting rather than useful.
So here's what changed. Run the same six weeks again, with the bar already written down: day-14 return has to clear twenty-five percent before Surprise Basket, or anything like it, gets called a win. Week one still lands the same, sixty-one percent try it, the demo still scores a nine. But by day fourteen, the number sits on Sixten's dashboard automatically now, not waiting on a data analyst's unrelated report. He sees the six percent in week three instead of week six, three weeks earlier, before the second version's roadmap slot ever gets approved.
One design hands the team a number that only ever tells them how good the first Tuesday was. The other hands them a number that tells them if there's a second one.
What I'd tell myself, sitting in that meeting eight months earlier: a bar that only measures the first try isn't a lower bar, it's a different question, and we'd been answering it for so long we forgot to notice when the question changed under us.
GUARD, weighed against one great Tuesday and the quiet cliff right after it
This was never really about whether Surprise Basket's recipes were good. They were. GUARD is for naming who pays when the number that gets celebrated isn't the number that turned out true.
GGroups. Who actually pays for a feature that demos well and dies quietly.
Larderwell's own product and engineering team, who spent nine real weeks building Surprise Basket and already had seven more weeks earmarked for a bigger version of it. And every shopper who tried it once, like Dulcie, who got a genuinely good evening out of it and then, without meaning to, taught the team nothing at all, because nobody was reading the number that would have told them the truth.
Neither group did anything wrong. The team built what the roadmap asked for. Dulcie tried something new and moved on, the way anyone would.
UUnequal. Which number gets a spotlight, and which one doesn't.
Sixty-one percent trial and a 9.1 demo score went straight into a leadership deck and a company-wide announcement. Six percent day-14 return sat quietly inside a dataset nobody had a standing reason to open, because nothing on the launch dashboard ever asked for it.
The harm isn't that one number is fake. It's that only the flattering one ever had a light pointed at it.
AAbility to contest. Could anyone downstream even tell the real return rate was this low.
Could an engineer starting on the sequel, or a shopper wondering if her experience was normal, actually find out that day-14 return sat at six percent? No. The dashboard Larderwell's leadership watched every week only ever showed trial rate, broken out by feature. Nobody had built a place for a two-week number to even show up, let alone trigger a conversation.
GUARD's sharpest question here isn't whether the number was bad. It's whether the process gave anyone a reason to go find it before six weeks had already passed.
Four of these five steps happened exactly as designed. The missing one, a required check on what happens after the first try, was never built into the path.
Week-one trial rate versus day-14 return rate, both features
Surprise BasketThe Regulars
Trial rate is measured against everyone who could have tried the feature. Day-14 return is measured only against the people who did, a different base, on purpose, since a feature can't return a shopper who never opened it once.
RReduce. The actual fix, not a slogan about being data-driven.
Every new feature now gets a written retention bar before a single line of it ships: does a shopper who tries this come back and use it again within two weeks. Sixten set that bar at twenty-five percent, comfortably below what a genuinely sticky feature like The Regulars clears, but nowhere close to where Surprise Basket actually landed. Nothing gets called a win, and no roadmap slot gets approved for "more of this," until the real cohort clears that bar. Underneath the process fix sits the actual model fix: the basket generator now runs every candidate dish through the shopper's stored constraint profile, allergy flags, weekly budget, and a rolling list of what they've actually bought before, as a hard filter, before it ever ranks anything for novelty.
The alternative worth naming and rejecting: require every new "delight" feature to pass an eight-week pre-launch retention pilot before it ships to any real shopper at all. Larderwell rejected that. It runs dozens of small feature experiments a quarter, and most get killed within the first two weeks of real usage data anyway. Requiring two months of testing before anything could ship would have stopped the company from trying things fast at all. The fix gates the decision to call something a win, not the decision to ship it.
Only one axis matters for whether a feature needs the bar: is it meant to earn a place in someone's week, or is it meant to be a nice surprise exactly once.
DDetect. How you'd know this trap has already been fallen into.
Watch for the shape, not a single number: a strong first-week curve that flattens hard right after it, a novelty cliff. Before the fix, the gap between Surprise Basket's trial rate and its day-14 return sat at fifty-five points and nobody had ever plotted the two next to each other. After the fix, any new feature whose gap crosses roughly forty points automatically flags for review before its next roadmap ask.
The failure worth naming plainly: if a bad number only ever gets fixed quietly, with the roadmap re-scoped around a smaller plan and nobody told the original launch was oversold, that's not detection. That's just a mistake nobody's allowed to say out loud.
The trade-off, said out loud: running every candidate dish through Dulcie's real constraint profile before ranking it for novelty adds about 300 milliseconds to building a basket, and it cuts the number of genuinely wild, never-tried-before suggestions Larderwell can offer each week from about five down to two. Larderwell took that trade on purpose, for a feature meant to earn a weekly habit, instead of shipping something fast and exciting that only ever works the first time.
And if you want to be sure it really works, try it somewhere else
Same five letters, a language app instead of a grocery app, and this time the ignored constraint is a learner's real mistake history instead of a peanut allergy.
Tonguecraft, built by Northfold Learning, teaches adults a new language through short daily practice. Wioletta Mirosava owns Tonguecraft's AI roadmap the way Sixten owns Larderwell's.
Same trap, a different subject. The branch that wins the demo is still the branch that loses the second week.
Roleplay Mode drops a learner into a spoken scene, ordering coffee in Lisbon, haggling at a market, with a theatrical AI voice playing the other part. Fifty-four percent of learners tried it in week one, and it was the single best-reviewed feature in Tonguecraft's history at launch. Daily Review is a short, unglamorous quiz that only ever tests the exact words and grammar rules a learner has personally gotten wrong before. Twenty-two percent tried it in week one, less than half of Roleplay Mode's number.
Same rank, mapped onto Tonguecraft: size the win by day-14 return, not the launch-week reaction. Roleplay Mode's day-14 return came back at five percent, because the scenes are dramatic but generic, they never draw on what a specific learner actually struggles with. Daily Review's day-14 return came back at sixty-one percent, because by definition it can only ever quiz you on your own real mistakes, so it keeps being useful in a way a fresh scene never has to be.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: never let a launch-week number alone justify a bigger roadmap slot. Ask for the two-week return number before you approve anything.
Cost: there's no budget this quarter for a full cohort study. Pull a rough number from the last twenty triers by hand instead of zero. A small real sample beats trusting a demo score that was never designed to predict repeat use.
The model got better, for real: say Roleplay Mode's dialogue engine improves enough next quarter to actually reference a learner's real mistake history mid-scene. That's a reason to re-test the day-14 number and raise the bar it's held to, on purpose, with new data behind it, not a reason to have assumed it would happen from day one.
Where people run it wrong.
They treat a strong demo reaction as proof the feature works, instead of proof it's interesting the first time.
They let the launch-week number become the story everyone repeats at the next all-hands, even after a real cohort report says something different.
They wait for someone to notice the silence, a feature nobody talks about anymore, instead of gating any bigger investment behind a dated retention check.
How to use it live. Before answering, ask out loud: "does this feature win because it's new, or because it fits something the person actually needs every week?" Naming that split buys you a few seconds of thinking time, and it's most of the real answer.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
Which framework fits "avoid the trap of a feature that demos well and gets no repeat usage"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real test isn't whether the recipes were good, it's whether anyone built a way to catch a flattering number hiding a quiet failure, before nine weeks of engineering time got spent on a sequel.
2 · THE PEOPLE
Who are the people this answer names?
Tap to flip
ANSWER
Sixten Groenewald, the product manager who owns Larderwell's AI shopper. Anzhelika Farkas, VP of Growth, who championed Surprise Basket off its demo. Dulcie Holmqvist, the nurse who tried it once. Torgeir Draganescu, the data analyst who found the real number in a routine report.
3 · THE OLD HABIT
What did the team stop doing, out of habit, once week-one trial started looking so good?
Tap to flip
ANSWER
They stopped asking whether anyone came back. Week-one trial rate had always predicted real use before, so nobody built a habit of checking day-14 return once a launch looked strong.
4 · THE TRAP, IN ONE LINE
What's the actual mismatch this question is testing?
Tap to flip
ANSWER
A feature scored on how new and exciting it looks once is answering a different question than a feature scored on whether it fits someone's real, recurring week. Only one of those two questions gets asked by a launch-week trial number.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Larderwell's launch dashboard only ever tracked week-one trial rate above forty percent as the bar for a win. That was fine when trial rate and real use moved together. It stopped being fine the day a feature got built to be exciting on the very first try.
6 · THE NUMBER
Fill in the blank: Surprise Basket's week-one trial rate was ___ percent. Its day-14 return rate was ___ percent.
Tap to flip
ANSWER
61 percent trial. 6 percent return. That 55-point gap is the whole story: a feature a lot of people liked once, and almost nobody needed twice.
7 · THE REPLAY
Same six weeks, new bar in place. What changes?
Tap to flip
ANSWER
Day-14 return sits on the dashboard automatically instead of waiting on an unrelated report. Sixten sees the real six percent in week three instead of week six, three weeks before the sequel's roadmap slot gets approved.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which one, and who plays the equivalent roles?
Tap to flip
ANSWER
Tonguecraft, Northfold Learning's language app. Wioletta Mirosava plays Sixten's role. Roleplay Mode, a flashy conversation feature, plays Surprise Basket's role: 54 percent trial, 5 percent day-14 return, against Daily Review's 22 percent trial and 61 percent return.
Check yourself Score: 0 / 0
Fill in the blank
1. Surprise Basket's week-one trial rate was ___ percent. Its day-14 return rate was ___ percent.
Show hint
Check flashcard 6, and the bar chart in the GUARD recap.
Show answer
61 percent, and 6 percent. That 55-point gap sat unwatched for six weeks, because the dashboard only ever showed the first number.
Multiple choice
2. Why didn't Larderwell catch this before nine weeks of engineering time went into Surprise Basket?
A. The model's recipe suggestions were too low quality to demo well.
B. The launch dashboard only ever tracked week-one trial rate, with no place for a day-14 or day-30 number to show up.
C. Dulcie never told anyone she'd stopped opening the feature.
D. Larderwell didn't have enough shoppers using Surprise Basket to measure anything real.
Show hint
Check the Ability to contest step in the GUARD recap.
Show answer
B. The recipes were genuinely good, and Dulcie never needed to complain, she just quietly stopped opening it. The real gap was structural: nothing in the process ever asked for the day-14 number.
True or false, with why
3. True or false: Surprise Basket's problem was that its recipe suggestions were bad.
True
False
Show hint
Check the paragraph right after the highlight block in the story section.
Show answer
False. The harissa lamb kefta was a real, good suggestion. The break was that the model was scored on novelty, not on whether a dish fit Dulcie's actual, recurring constraints, so it kept suggesting things she couldn't safely repeat.
Short answer, name the rejected alternative
4. Larderwell considered one other fix besides the two-week retention bar. What was it, and why did they reject it?
Show hint
Check the Reduce step in the GUARD recap.
Show answer
Model answer: Require every new "delight" feature to pass an eight-week pre-launch retention pilot before it ever ships to a real shopper. Rejected because Larderwell runs dozens of small feature experiments a quarter, and most get killed within two weeks of real usage data. Requiring two months of pre-launch testing before anything shipped would have stopped the company from trying things fast at all.
Short answer, apply it yourself
5. Think of an app you use with an AI feature. Name one thing about it that wowed you the first time. Would you actually still use it two weeks later, and why or why not?
Show hint
Look for a feature that impressed you once but never learned anything specific about your real, repeating situation.
Show answer
Model answer: An AI photo filter that reimagines a selfie as a painting. Genuinely delightful the first time, and there's no real reason to open it again two weeks later, because it doesn't solve anything that keeps coming back, it's a trick, not a habit.
Fill in the blank, work the number
6. Larderwell's new rule says day-14 return has to clear 25 percent before a feature counts as a win. Surprise Basket's real day-14 return was 6 percent. By how many percentage points did it miss the bar?
Show hint
Subtract the real number from the bar.
Show answer
19 percentage points. Twenty-five minus six. Not a near miss, a clean sign the feature hadn't earned a bigger roadmap slot yet.
Before you close the answer
Why this works
Tests whether you'll grade an AI feature on the excitement it can generate once, or on the habit it can actually earn. Most candidates answer with "get more feedback" or "keep iterating," neither of which names a real, enforceable bar.
Follow-up traps
"What if the feature just needs more marketing to build the habit?" Response: more marketing raises the week-one trial number again, which was never the problem here. It does nothing for the day-14 return number, which was.
"Isn't a two-week bar arbitrary, why not one week, or a month?" Response: two weeks is long enough that a first-try novelty has genuinely worn off, and short enough to catch a bad bet before a full roadmap slot gets spent on a sequel. Larderwell set its own bar at 25 percent based on what its one proven sticky feature, The Regulars, actually cleared.
If pressed
The novelty score inside Surprise Basket's ranking step wasn't hand-tuned. It came from a model trained on which suggestions shoppers clicked or saved inside the app, and clicking or saving turned out to strongly predict curiosity, not repeat purchase. Larderwell only found the mismatch by pulling a second, separate dataset, actual reorders thirty days out, and matching it against the original click data. The two barely correlated at all.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.