CaseIntermediateResponsible AI & Advanced Practice / AI product case study teardowns / #10

Analyze how a consumer AI app handles the cold start problem.

SPARK the product is Humlet, a consumer AI app that recommends what to cook tonight

Humlet suggests a recipe based on what's in your kitchen and how much time you have. Dagny Solheim runs a small-town veterinary clinic all day and opens the app most nights around 11, once the clinic's closed and she's finally thinking about dinner.

The direct answer
Humlet's cold-start fix isn't a long quiz. It's one question, asked the moment you open the app for the first time, that immediately reshapes what you see, instead of showing every brand-new user the same generic popular-recipes list regardless of what's actually in their kitchen that night.
Do this, in order
  1. Ask one real question at first launch, before showing any recommendation.Why: this single design decision is what separates a genuinely personal first list from a generic one.
  2. Make that question about tonight, not about long-term taste.Why: "what do you have too much of right now" is answerable in five seconds. A full flavor-preference quiz isn't.
  3. Build a fast, one-tap "not tonight" swap for when the first guess is wrong.Why: a first guess will sometimes miss, and the design has to survive that moment, not just hope it doesn't happen.
  4. Leave deep, long-term personalization for later sessions.Why: day one doesn't need to be perfect. It needs to not be generic.
  5. Never let a dietary restriction go unasked, even in a short cold-start flow.Why: that's the one first-day guess that's genuinely expensive to get wrong.
  6. Re-check the cold-start question periodically as new user needs show up.Why: the single best first question can shift as the app's audience grows or changes.

How to answer this, stage by stage

Six moves. This is a design question: the interviewer wants your one concrete anchor decision, not a tour of every screen.

Stage 1
Scope it to one real app and one real first session
Say it like this
"I'll take Humlet, a consumer app that suggests what to cook, and look at exactly what a brand-new user sees the first time they open it, with zero history behind them."
Why this works
Grounds "the cold start problem" in one specific, literal first screen.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, payoff, anchor, risk, keep out."
Why this works
Signals a design answer built forward, not a story about what already broke.
Stage 3
Ground the situation, no app yet
Say it like this
"Right now, without an app, a tired person at 11pm just falls back to one of the same three meals they always make. That's the actual competition Humlet's first screen is up against."
Why this works
Names the real baseline, which is a habit, not a blank slate.
Stage 4
Give the one anchor decision
Say it like this
"One question at first launch: 'what's one thing you have too much of right now?' That single answer reshapes the very first list of recipes a new user ever sees."
Why this works
This is the concrete, arguable design decision, exactly what the question is asking for.
Stage 5
Show the anchor survives being wrong
Say it like this
"If the first suggestion still misses, a one-tap 'not tonight' swap asks one more quick question instead of dumping the user back into a generic list."
Why this works
Proves the anchor isn't just good when it works. It has a real answer for when it doesn't.
Stage 6
Name what you're leaving out, and close
Say it like this
"A full flavor quiz or pantry scan can wait. Day one only needs to feel like it's talking to this person, not to everyone who downloaded the app that week."
Why this works
Shows judgment, not a wish list of every personalization feature imaginable.

Let's learn

Every night for six years, once the clinic's last patient went home, Dagny made one of the same three dinners, mostly because deciding anything new felt like one more thing to figure out.

Humlet is a consumer app that recommends a recipe based on what you actually have and how much time you've got that night.

Knowledge spark: what's the "cold start problem," really? It's the moment a recommendation system has zero history about a person and has to suggest something anyway. Most apps solve it by falling back to whatever's generically popular, which technically works but treats every new user as identical, at exactly the moment a first impression matters most.

Without an app, this was Dagny's actual routine: too tired to plan, she'd default to whatever required the least thought, usually pasta, usually the same jar of sauce.

Day-1 "actually cooked it" rate, generic list vs one-question start
60% 30% 0 22% Generic top-10 list 51% One-question start
More than double the day-one cook-through rate, from a single question asked before showing anything at all.

At its worst: a new user opens the app once, sees a list that could belong to anyone, closes it without cooking a thing, and never opens it again. That's a cold start failure with no error message, just a quiet non-return.

The anchor decision One question, asked before the very first recommendation: "what's one thing you have too much of right now?" A tomato. Some leftover rice. Whatever's about to go bad. That single answer reshapes day one's entire list, instead of every new user seeing the same generically popular recipes.

What I would leave alone: a full flavor-preference quiz, or a complete pantry scan, can wait entirely until later sessions. Day one doesn't need deep personalization. It needs to not feel like it was built for a stranger.

The cold start problem isn't a data problem. It's the five seconds before a new user decides whether this app is going to feel like it's talking to them.

The lesson: a generic first list isn't a placeholder while the app "learns" someone. It's a real first impression, and it's making one whether anyone designed it on purpose or not.

Now here is the same thing as a story

The short version above is what you'd say defending this design live. Read this one for how Dagny's actual first night with the app went.

Most nights, once the clinic closed, Dagny's routine was the same: feed the cat, sit down, and reach for whatever needed the least amount of deciding.

Hand sketched flow diagram titled Dagny's first three taps. Four boxes: opens Humlet, same top-10 list highlighted, not what she needs, closes app.
Four steps, and the second one, a list that could belong to anyone, is where a generic cold start actually loses a new user.

The first version of Humlet she tried opened straight to a top-10 popular recipes list. Nothing in it reflected that she was cooking for one, had about twenty minutes, and had a container of leftover rice she needed to use before it went bad.

Hand sketched comparison diagram titled The anchor close up. Left panel a box icon labeled Old default, caption same list for everyone. Right panel a gauge icon labeled One-question start, caption list built from her answer.
The entire redesign is the difference between these two panels: one question, asked before anything else.

She closed the app after scrolling twice, cooked pasta again, and didn't open it for another week.

Hand sketched quadrant titled Sorting new user moments by risk. Axes how often it happens from rare to common, and cost if the guess is wrong from low to high. First recipe shown sits top right, common and high cost. Diet restriction ignored sits upper middle, less common but very high cost. Notification timing sits middle. Account setup typo sits lower left.
The very first recipe shown sits in the riskiest, most common corner. It's exactly where a generic guess does the most damage.

A later version asked one question the moment she opened the app for the first time: what did she have too much of, right now. She typed "rice," and the very first recommendation was a fried rice recipe built around exactly that, ready in fifteen minutes.

Hand sketched icon list titled What day one shouldn't assume. Three rows: a question mark box icon labeled same taste as everyone else, a box icon labeled no allergy or diet noted, a gauge icon labeled one quick answer fixes both.
Two wrong assumptions, and one small question that quietly fixes both of them at once.

She actually cooked it. Not because the recipe was extraordinary, but because for the first time, the app felt like it had been paying attention, on day one, with zero history behind it.

30-day retention, generic list vs one-question start
70% 35% 0 One-question: 44% Generic: 19% Day 1 Day 30
Both groups start close together. The gap widens steadily over a month, the same slow-decay shape as a habit that never quite forms.

Someone building the original version decided a generic popular-recipes list was the safe, obvious default for a brand-new user, since there was no data yet to personalize anything. That's a reasonable starting assumption, and it's also exactly where the design stopped, treating "no data yet" as a wall instead of a five-second question away from not being true.

I would take that assumption back and build the one-question start instead, cheap to ask, answerable in seconds, and enough to make day one feel specific rather than generic.

We shipped the generic list first because it was the fastest thing to build with zero user data, and it felt like the responsible, safe choice. It took watching real new users open the app once and quietly never come back, with no complaint, no review, nothing but a silence in the retention numbers, to see that "safe" and "actually working" weren't the same thing.

SPARK, in one screenFive letters, and the anchor, one quick question before any recommendation, is the whole design.

S
Situation.
A tired person at 11pm, with zero history in the app, whose real alternative is falling back to the same three meals they always cook.
Names the real competition: an existing habit, not a blank slate.
P
Payoff.
Stop reflexively defaulting to the same tired meal, and start trusting a suggestion enough to actually try it on a low-energy night.
Names the habit the design is trying to build, not just the feature.
A
Anchor.
One quick question at first launch, "what do you have too much of right now," reshaping the very first recommendation list.
The hardest step, and the one concrete decision the whole answer is built on.
R
Risk.
The first guess can still miss. A one-tap "not tonight" swap asks one more quick question instead of dead-ending back into a generic list.
Proves the anchor survives being wrong, not just when it's right.
K
Keep out.
A full flavor quiz and a complete pantry scan wait for later sessions. Day one needs to feel specific, not fully personalized.
Shows restraint, not a wish list dressed up as a launch plan.

The recap, one line per letter: situation is the existing habit of falling back to the same three meals, payoff is trusting a new suggestion instead, anchor is the one-question start, risk is the one-tap swap for a wrong first guess, and keep out is deferring deep personalization past day one.

And if you want to be sure it really works, try it somewhere elseSame five letters, a volunteer-matching app instead of a dinner app. This time the cold start question is about a skill, not a fridge.

A regional volunteer-matching nonprofit built an app that suggests local volunteer opportunities. A brand-new user, with zero history, would otherwise see the same generic "popular this week" list every other city sees.

Mapped onto SPARK: situation is someone who wants to help but has never volunteered through an app before, and would otherwise just not bother searching at all. Payoff is turning a vague, well-meaning intention into an actual first sign-up within the same week. Anchor is one quick question at first launch: "what's one thing you're actually good at?" instead of just asking for a city and a generic interest category. Risk: if the first match still misses, an "this isn't really my thing" tap asks one more quick question about time availability instead of dumping the user back into the full unsorted list. Keep out: a full skills-and-availability intake form, the kind most volunteer platforms front-load, waits until after that very first match, since asking for too much up front is its own way of losing someone before they ever see a single opportunity.

Hand sketched labeled parts diagram titled A volunteer matching app's same cold start. Center person icon labeled New volunteer, with four callouts: one quick question, skills not just city, matched same week, not a generic list.
A different app, a different question, and the same underlying anchor: ask one small thing before showing anyone a generic list.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "ask one real question before the first recommendation, instead of showing every new user the same generic list" and stop.
Cost: no engineering time to build a smart first-question flow this quarter. Say so, and start with the cheapest version: a single multiple-choice question with three or four options, not free text, still personalizes the first list without much build cost.
The model gets better, for real: even if Humlet's underlying recommendation model improves substantially, a generic top-10 list still ignores the one piece of information that actually mattered, what this person has and needs tonight. A better model doesn't fix a missing question.

Where people run it wrong.
They treat the cold start problem as something only more data or a better model can fix, instead of a design decision about what to ask on day one.
They build a long onboarding quiz, assuming more questions means more personalization, when a single well-chosen question usually beats ten generic ones.
They forget to design for the moment the first guess is wrong, leaving new users with no way forward except closing the app.

How to use it live. When asked how a consumer app handles cold start, ask yourself one question first: what's the single cheapest thing you could ask a brand-new user that would meaningfully change what they see first? That question is almost always the real answer, not a list of personalization features.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "analyze how an app handles the cold start problem"?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. A design question run forward, not a story about something that already broke.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Dagny Solheim, who runs a small-town veterinary clinic and used to default to the same three dinners most nights.
3 · THE SITUATION
What's Dagny's real alternative to using Humlet at all?
Tap to flip
ANSWER
Falling back to one of the same three tired meals she always cooks, out of habit, without deciding anything new.
4 · THE ANCHOR
What's the one concrete design decision this whole answer turns on?
Tap to flip
ANSWER
One quick question at first launch, "what do you have too much of right now," that reshapes the very first recommendation list.
5 · THE OLD ASSUMPTION
What decision would you take back?
Tap to flip
ANSWER
Defaulting new users to a generic top-10 popular list, reasonable when there's no data yet, but treated as a wall instead of a five-second question away from not being true.
6 · THE NUMBER
Fill in the blank: users shown the one-question start actually cook something on day one ___ percent of the time, versus 22 percent for the generic list.
Tap to flip
ANSWER
51 percent. More than double the generic list's day-one cook-through rate.
7 · THE REPLAY
Same tired 11pm moment, one-question design in place. What changes?
Tap to flip
ANSWER
Dagny gets a fried rice recipe built around her actual leftover rice and cooks it that night, instead of closing the app after scrolling a generic list.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's the equivalent question there?
Tap to flip
ANSWER
A volunteer-matching nonprofit app. There, the question is "what's one thing you're actually good at," instead of asking about a fridge.

Check yourself Score: 0 / 0

Multiple choice
1. Why does asking one quick question at first launch work better than showing a generic popular-recipes list?
  • A. Because the popular list contains inaccurate recipes.
  • B. Because it turns "no data yet" from a dead end into a five-second answer that meaningfully reshapes the very first list.
  • C. Because generic lists take longer to load.
  • D. Because popular recipes are always more complicated to cook.
Show hint
Look at "the anchor decision" key point.
Show answer
B. The cold start problem isn't unsolvable without data, it's solvable with one cheap, well-chosen question.
True or false
2. True or false: this answer recommends building a full flavor-preference quiz and pantry scan before a new user sees their first recommendation.
  • True
  • False
Show hint
Look at the "keep out" step.
Show answer
False. Deep personalization is deliberately deferred. Day one only needs one quick, specific question.
Fill in the blank
3. Fill in the blank: users shown a generic top-10 list actually cook something on day one only about ___ percent of the time.
Show hint
Look at the grouped bar chart.
Show answer
22 percent. Less than half the rate of users who got the one-question start.
Short answer, name the reversal
4. What old assumption does this answer take back, and why did it make sense when it was made?
Show hint
Look at the paragraph about the "reasonable starting assumption."
Show answer
Model answer: That a generic popular list was the only safe default with zero user data, reasonable at first but treated as a permanent wall instead of a temporary gap.
Short answer, where it wouldn't matter
5. Name a part of Humlet's onboarding where skipping deep personalization genuinely doesn't cost much.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Full flavor-preference personalization can wait past day one; it barely affects whether someone tries their very first recipe.
Short answer, apply it yourself
6. Think of an app you signed up for recently. Did it ask you anything specific before showing you its first recommendation, or did it show you the same thing it shows everyone?
Show hint
Think about the very first screen you saw after signing up for a streaming, shopping, or recommendation app.
Show answer
Model answer: Many apps default to a generic "most popular" list, the exact cold-start gap this answer is built around.
Before you close the answer
Why this works
Tests whether you can turn "the cold start problem" into one concrete, designable decision, instead of describing it as a data limitation that only time or a bigger model can fix.
Follow-up traps
"Isn't one question still just a guess, no more reliable than a popularity list?" Response: it's a guess grounded in something true about this specific person right now, which is categorically different from a guess based on what's popular with everyone else.

"What if the user skips the question entirely?" Response: then fall back to the generic list exactly as before, the one-question flow only ever adds a better option, it never removes the existing safety net.
If pressed
The real version of this question also rotates its phrasing slightly by time of day, since "what do you have too much of" performs differently at 11am than at 11pm, when the honest answer is closer to "the least effort possible."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more