InterviewAdvancedAI Opportunity & Model Strategy / When NOT to use AI / #23
Tell me about an AI feature you would kill today if you owned it.
TRACE96 percent of people never used the caption the model wrote for them, for six straight months
Picwell is a photo-sharing app. CaptionMuse suggests three AI-written captions the moment someone posts a photo. Fenwick Osata is the AI PM who owned it, and this is the feature Fenwick would kill today, with real usage data behind the decision, not a hunch.
The direct answer
Kill CaptionMuse's default-on suggestions. Six months of real usage data shows only 4 percent of posts used a suggested caption unedited, flat, not climbing, the whole time. The feature was solving "writing a caption is hard." The real friction was never that. It was deciding whether a photo was worth posting at all, and wanting the caption, once written, to feel like it actually came from the poster.
Do this, in order
Measure real usage of the AI output itself, not impressions or session time.Why: the launch dashboard tracked everything except the one number that actually mattered, whether people used what the model wrote.
Check whether the number is flat or still ramping before deciding anything.Why: 4 percent in month one could be early adoption. 4 percent in month six, unchanged, is the real number.
Talk to real users about why they skip the suggestion, not just that they do.Why: the usage number tells you something's wrong. Only a real conversation tells you what.
Kill the default-on version once the evidence is this clear, not just "improve" it again.Why: a feature solving the wrong problem doesn't get fixed by writing better captions for the same wrong problem.
Redirect the freed engineering time and inference budget toward the real friction point.Why: killing a feature isn't the end of the decision, it's what you do with what it frees up that matters.
Keep a small, opt-in version for the minority who genuinely want caption ideas.Why: 4 percent is small, not zero. Removing the option for people who do want it would be its own mistake.
How to answer this, stage by stage
Nobody is scoring whether you can name a feature you dislike. They're scoring whether killing it is backed by real evidence, and whether you'd say it out loud about your own work.
Stage 1
Scope it to one real feature, with its real numbers
Say it like this
"I'd kill CaptionMuse, the AI caption suggestions on Picwell. It shows on every post. Six months of real data says only 4 percent of people ever use a suggestion unedited, and that number never moved."
Why this works
Names a real feature with a real, checkable number instead of a vague "something I'd improve."
Stage 2
Say your structure out loud before any content
Say it like this
"I'll run this as TRACE. Timeline, what shipped and what the dashboard actually tracked. Recut, what the low number really means. Assume nothing about what 'engagement' proves. Cause candidates, the real suspects. Evidence test, what actually confirmed it."
Why this works
Signals this is a diagnosed decision, not a personal preference dressed up as analysis.
Stage 3
Say plainly why killing a shipped feature isn't a failure to hide
Say it like this
"This isn't me admitting a mistake I'm embarrassed by. Shipping it was a reasonable bet at the time. Killing it now, with six months of real evidence saying it's solving the wrong problem, is the actually accountable move, not the uncomfortable one."
Why this works
Reframes a kill decision as a sign of judgment, exactly the interview trap this question sets.
Stage 4
Give the one decision: kill the default-on version
Say it like this
"Here's what I'd actually do. Remove CaptionMuse's default-on suggestions from every post. Replace it with a small, opt-in 'give me an idea' button for the minority who genuinely want it. Redirect the freed inference budget and engineering time at the real friction: deciding whether a photo's worth posting at all."
Why this works
This is the direct answer, stated as an actual decision with a redirect attached, not just a removal.
Stage 5
Prove it with the compressed evidence
Say it like this
"87 percent of people ignored the suggestions completely and wrote their own caption from scratch. Another 9 percent used one as a starting point and changed more than half of it. Only 4 percent used one as written, and that held flat for six straight months, not a slow ramp-up."
Why this works
This is the load bearing evidence, specific enough that a follow-up question can't easily poke a hole in it.
Stage 6
Name the real diagnostic insight, not just the number
Say it like this
"We built CaptionMuse assuming the hard part was thinking of words. User interviews said otherwise: people wanted their caption to feel like theirs, and using a model's exact words felt like it wasn't. The real friction, for most people, was deciding whether the photo was even worth posting, something a caption suggestion was never going to touch."
Why this works
This is the load bearing judgment. It wouldn't make sense to ask this about a feature with no model in it, since the failure is specifically that a model's exact output felt inauthentic to use as-is.
Stage 7
Say what you'd keep, then close on one line
Say it like this
"I wouldn't kill AI writing help at Picwell entirely, a small, opt-in version for the 4 percent who genuinely want ideas is worth keeping cheap and simple. What I'd kill is showing it to everyone by default, on every post, for a problem most people never actually had."
Why this works
Closes with real judgment instead of a blanket "AI features don't work," and restates the direct answer in one breath.
Let's learn
CaptionMuse is a feature on Picwell that writes three suggested captions the instant someone posts a photo, meant to save people from staring at a blank caption box.
Nothing about the first month's dashboard predicted this. It took someone actually checking the right number.
At launch, CaptionMuse looked like a clear win. Suggestion impressions climbed steadily. Time spent on the post-composing screen went up too, which the team read as engagement. Nobody was tracking the one number that actually mattered: how many suggestions got used, unedited, as written.
Two metrics that looked like success, and the one that would have told the real story, missing entirely.
When Fenwick finally pulled the real number, six months in, it was 4 percent. Not a typo, not a rounding error. Ninety-six times out of a hundred, someone saw three AI-written captions and used none of them as written.
What actually happens when a caption suggestion is shown
Ignored entirelyHeavily editedUsed as written
The 87 percent isn't people who never saw the feature. It's people who saw three suggestions and typed their own words anyway.
So Fenwick checked whether 4 percent was still ramping up, the way a new feature sometimes does before people discover it. It wasn't. The number sat at 4 percent in month one, 5 percent in month three, 4 percent again in month six. Flat, not early, just genuinely low.
The feature wasn't failing to catch on. It had already found its real ceiling, and the ceiling was 4 percent.
Here's the turn: CaptionMuse was never badly built. Its suggestions were often genuinely decent, sometimes funny, occasionally exactly what someone might have written themselves. The turn is that "decent suggestion" and "something I want to actually use" turned out to be two different bars, and almost nobody was clearing the second one.
Knowledge spark: why does a good suggestion still go unused?
A caption is a small piece of someone's own voice attached to something personal. Even a genuinely well-written AI suggestion can feel like it isn't really theirs, the same way a beautifully worded card someone else wrote for you doesn't feel like your own words. Quality and authenticity are different things, and a model can nail the first without ever touching the second.
Suggestion-used-as-written rate, month by month
Used as written
A genuinely flat line across half a year is its own kind of evidence. This was never an adoption curve waiting to bend upward.
At its worst, this cost showed up as a real, ongoing inference bill, tens of thousands of dollars a month, spent generating three captions for every single post on the platform, for a feature 96 percent of people never used as intended. That budget had a real alternative use the whole time.
The choice I would take back
Launching CaptionMuse as default-on for every post, without first testing whether people wanted AI-written words attached to something as personal as their own caption. That made sense when the team assumed "writing is hard" was the universal blocker. It stopped making sense the moment real usage data said otherwise, and nobody revisited the assumption for six months.
What I would leave alone: Picwell's AI-powered photo enhancement, sharpening and lighting correction, stays exactly as it is. Nobody's identity is wrapped up in whether their photo got auto-brightened, so there's no authenticity cost the way there is with someone else's words standing in for your own.
The lesson: a feature that produces genuinely good output can still be solving the wrong problem. Track whether people actually use what the model made, not just whether they saw it, and don't let six months pass before checking.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what it actually felt like to watch a feature everyone had called a launch success quietly fail a much simpler test.
Fenwick Osata had shipped CaptionMuse eight months earlier to genuine excitement. The launch review had gone well: impressions up, session time up, a few glowing internal Slack messages about how funny some of the generated captions were.
Marlone Byfield, an engineer on the team, asked an offhand question in a retro: "Do we actually know how many people use the suggestions as-is?" Nobody in the room had that number. It had never been on the dashboard.
Three honest guesses, before anyone had the real data to tell which ones were right.
Fenwick pulled it that afternoon, expecting something respectable, maybe 20 or 30 percent. The real number was 4. Not a bad week. Six months, averaged, flat the whole time.
The first instinct, understandably, was to assume the suggestions just weren't good enough yet. Fenwick's team ran a quick internal review of a sample and found most of the captions were fine, sometimes genuinely funny, occasionally better than what people ended up writing themselves.
Not a quality problem. A problem with what the feature had assumed the friction actually was.
So Fenwick did the thing that should have happened at launch: talked to real users. Not a survey, actual short conversations with a dozen people who'd used Picwell that week. The pattern was consistent and, in hindsight, obvious. "It's not really mine if the app wrote it." "I don't struggle with the caption, I struggle with whether the photo's even good enough to post." Nobody said the suggestions were bad. Almost everyone said using one felt wrong.
We had built a very good answer to a question almost nobody was actually asking.
Fenwick never had a fixed threshold for when a shipped feature's usage number crossed from "still finding its audience" to "solving the wrong problem." It came down to a feeling with two settings: the number is still moving, or it's found its real, flat ceiling. Six unchanged months made that second setting unmistakable.
The fork the evidence actually pointed to, once someone finally looked.
Back at launch, shipping CaptionMuse default-on for everyone wasn't an unreasonable call. The team's working theory, that a blank caption box was the real blocker, was a genuinely plausible guess at the time, with no data yet to contradict it. It stopped being reasonable the moment six months of flat, real usage said the theory was wrong, and nobody had checked.
Killing the feature was never the end of the story. What replaced it is the part that actually mattered.
Here's the replay: CaptionMuse's default-on suggestions come off every post. A small, opt-in "give me an idea" button stays for the 4 percent who genuinely want it. The inference budget that used to write three unused captions per post gets redirected toward a "which of your last three photos is your best shot" comparison tool, aimed squarely at the friction the interviews actually surfaced.
One version of this story keeps a quietly unused feature running for years because its impressions chart never dipped. The other spends one honest afternoon of user interviews and redirects real budget toward the problem people were actually stuck on the whole time.
What I'd tell myself, hearing Marlone's offhand question in that retro: a dashboard full of numbers that look good will never tell you the one thing that matters, whether people actually wanted what you gave them. Somebody has to ask that question on purpose, because the launch metrics never will.
TRACE, run on a feature that looked fine on every chart except the one that matteredNot a script for finding fault with old work. TRACE is what turns "I have a hunch this isn't working" into a real, defensible decision to kill it.
T
Timeline. What shipped, and when did the real number actually get checked?
CaptionMuse launched to strong impression and session-time metrics. The real usage-as-written rate, 4 percent, didn't get checked until six months later, when a retro question finally surfaced it.
The gap between "looked like a success" and "actually got checked" is the whole story.
R
Recut. What does the low number actually mean, sliced apart from the obvious read?
Not that the model wrote bad captions, internal review confirmed most were fine. The real slice: a good suggestion and something a person wants to claim as their own words are two different bars.
This is the mechanism, not just the symptom, of why quality alone never predicted usage.
A
Assume nothing. What did the launch team wrongly take "engagement" to mean?
Rising impressions and session time read as success. Neither one measured whether people actually used what the model wrote, the one number that would have told the real story from day one.
A metric climbing is not proof the thing it's supposed to represent is climbing too.
C
Cause candidates. What are the real, honestly named suspects?
Three candidates: the suggestions weren't good enough, people didn't want the model's exact words for something personal, or captions were never the real friction at all. User interviews pointed clearly at the second and third.
This is the direct answer's real evidence, a genuine shortlist tested against real conversations, not assumed.
E
Evidence test. What actually confirmed the cause?
A dozen real user conversations, plus six months of flat usage data ruling out a slow adoption curve. "It's not really mine if the app wrote it" came up unprompted, consistently.
Real conversations, not a survey, are what turned a suspicion into a confirmed diagnosis.
The recap, one line per letter: timeline is the six-month gap between launch and actually checking the real number, recut is separating suggestion quality from usage, assume nothing is the trap of reading impressions as proof of real value, cause candidates is the honest shortlist of what might explain the flat 4 percent, and evidence test is the real user conversations that confirmed which one was true.
And if you want to be sure it really works, try it somewhere elseSame five letters, a fitness app instead of a photo app. This time the kill decision is about a workout-plan generator, not a caption writer.
Priya Wentz owns product at Pacefield, a fitness app. Pacefield shipped an AI feature that generates a full personalized workout plan the moment someone signs up. Mapped onto TRACE: timeline is that signup completion and initial plan-view rates looked strong for months, while the real number, how many people were still following their AI-generated plan after week three, sat unchecked. Recut is that the plans themselves tested as reasonable by trainers reviewing them, the real issue wasn't quality. Assume nothing means not trusting "plan generated" as proof the plan got used. Cause candidates are plan quality, plan rigidity (no easy way to adjust it as life happened), or the real problem being accountability, not planning. Evidence test is user interviews finding that people who quit weren't confused by their plan, they'd simply missed a day, felt they'd "broken" the AI's schedule, and never came back, the same authenticity-adjacent gap as CaptionMuse, but about commitment instead of voice. The kill decision replaces the rigid full-plan generator with a lighter, adjustable week-by-week suggestion that expects and absorbs missed days.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to "the real usage number was flat and low for six months, here's what people actually wanted instead," and stop.
Cost: no time to run real user interviews before this conversation. Say so honestly, and name the interviews as the concrete next step rather than guessing at the cause.
The model got better, for real: say a future version of CaptionMuse wrote genuinely indistinguishable-from-human captions. The authenticity problem likely persists anyway, since the issue was never really about writing quality.
Where people run it wrong.
They track impressions and engagement time as proof of value, without ever checking whether the AI's actual output got used.
They assume a low usage number means the model needs to improve, instead of checking whether it's solving the right problem at all.
They treat killing a shipped feature as an admission of failure, instead of the accountable move once real evidence says it's not working.
How to use it live. The moment an interviewer asks about a feature you'd kill, don't reach for something you personally dislike. Reach for the one where the real usage number, not the impressions number, told a story nobody wanted to look at closely. That's usually where the honest answer is.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits deciding whether a shipped AI feature should actually be killed?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. It separates what shipped and looked fine from when the real problem was actually found and confirmed.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Fenwick Osata, the AI PM who owns CaptionMuse at Picwell. Marlone Byfield is the engineer whose retro question surfaces the real usage number.
3 · THE MECHANISM
Why did a feature with genuinely decent AI output still go almost entirely unused?
Tap to flip
ANSWER
Quality and authenticity are different bars. A caption is personal, and using a model's exact words for it felt, to most users, like it wasn't really theirs, even when the words themselves were fine.
4 · THE DECISION
What's the one concrete thing this answer says to actually do?
Tap to flip
ANSWER
Kill CaptionMuse's default-on suggestions, keep a small opt-in version for the 4 percent who want it, and redirect the freed budget toward the real friction: deciding whether a photo is worth posting.
5 · THE OLD DECISION
What old decision would this answer take back?
Tap to flip
ANSWER
Launching CaptionMuse as default-on for everyone without first testing whether people wanted AI-written words for something personal. Reasonable with no data yet to contradict it. Wrong once six flat months said the theory was off.
6 · THE NUMBER
Fill in the blank: only ___ percent of shown suggestions were used exactly as written, and that number held flat for ___ months.
Tap to flip
ANSWER
4 percent, flat for 6 months. Not a slow-adopting feature, a feature that had already found its real, low ceiling.
7 · THE EVIDENCE TEST
What actually confirmed the real cause, instead of just guessing?
Tap to flip
ANSWER
A dozen real user conversations, where "it's not really mine if the app wrote it" came up unprompted and consistently, ruling out suggestion quality as the real problem.
8 · CROSS PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what's the equivalent evidence test?
Tap to flip
ANSWER
Pacefield's AI workout-plan generator. The equivalent evidence test is user interviews finding people who quit had missed a day and felt they'd "broken" the plan, an accountability problem, not a planning-quality one.
Check yourself Score: 0 / 0
Multiple choice
1. Why didn't Picwell catch CaptionMuse's real problem sooner?
A. The AI model's suggestions were too low quality to notice a pattern.
B. The launch dashboard tracked impressions and session time, but never how many suggestions were actually used as written.
C. Users never saw the feature because it wasn't shown often enough.
D. Fenwick left the company shortly after launch.
Show hint
Look at the icon list in "Let's learn" showing what the dashboard tracked.
Show answer
B. Rising impressions and session time looked like success, but neither one measured whether people actually used what the model wrote.
True or false
2. True or false: internal review found CaptionMuse's suggested captions were generally low quality.
True
False
Show hint
Look at the story section, where Fenwick's team reviews a sample of suggestions.
Show answer
False. Most suggestions were found to be fine, sometimes genuinely good. The problem was never quality, it was that people didn't want to use someone else's words for something personal.
Fill in the blank
3. Fill in the blank: ___ percent of posts ignored the suggestion entirely and wrote a caption from scratch.
Show hint
Look at the stacked bar chart in "Let's learn."
Show answer
87 percent. The overwhelming majority response, and the clearest single signal that the feature wasn't solving the problem it assumed it was.
Short answer, where it wouldn't matter
4. Name a feature at Picwell where this same authenticity problem would NOT apply, and say why not.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: AI-powered photo enhancement, like auto-brightening. Nobody's personal voice is attached to whether their photo got color-corrected, so there's no authenticity cost the way there is with borrowed words.
Short answer, apply it yourself
5. Think of an AI feature you've used that produced good output you still didn't end up using. What was the real reason, beyond quality?
Show hint
Think about something personal, creative, or identity-linked, like a bio, a message, or a photo caption.
Show answer
Model answer: An AI-suggested dating-app bio was well-written but felt generic and impersonal, not because it was inaccurate, but because a first impression meant to represent "me" felt wrong coming from a template, the same authenticity gap CaptionMuse ran into.
Short answer, work the number
6. If the "used as written" rate had climbed from 4 percent in month 1 to 18 percent by month 6, would the kill decision still make sense?
Show hint
Think about what a genuinely climbing number, versus a flat one, would suggest about whether the feature is still finding its audience.
Show answer
Model answer: Probably not as clearly. A real, sustained climb would suggest the feature was still finding its audience rather than having hit a real ceiling, and the honest move would be to keep watching a few more months before deciding, not kill it on an upward trend.
Before you close the answer
Why this works
Tests whether you can name a real, evidence-backed kill decision about your own work without treating it as an admission of failure, and whether you know to measure real usage of an AI feature's output, not just engagement signals around it.
Follow-up traps
"Couldn't you just make the suggestions feel more personal instead of killing it?" Response: user interviews suggested the issue wasn't tone, it was that any words not typed by the poster felt borrowed, a harder problem than prompt tuning can fix.
"Isn't 4 percent still meaningful at Picwell's scale?" Response: yes, which is exactly why the recommendation keeps a small, opt-in version rather than removing the capability entirely, it's sized to match real demand instead of defaulting to everyone.
If pressed
The 9 percent who heavily edited a suggestion, changing more than half its words, were treated as a distinct segment from the 4 percent, because editing that much suggests the suggestion served as a rough starting point, not as the actual desired output.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.