What signals in user research suggest an AI solution rather than a better interface?
Kernelwell recommends recipes and plans a family's week of meals around their allergies, their budget, and their macros. Ionara Vradenburg owns the part of Kernelwell that decides what to recommend and why. This is the quarter she learned that one of her loudest complaints was never about the screen at all.
- Test every "hard to use" note against one question: would it survive a perfect navigation fix?Why: this is the actual signal test. Everything else here just feeds it good evidence.
- Read the real words behind each complaint, not the tag a queue gave it.Why: one tag, "hard to use," can hide two totally different problems wearing the same label.
- Treat "the task itself is hard" as a judgment signal, not a UI signal.Why: checking three lists by hand is a different problem from not finding a button, and no menu redesign touches it.
- Treat a household's own spreadsheet or workaround as proof, not just a complaint.Why: they've already written the spec for the AI feature themselves, by hand, for free.
- Fix the real navigation cases with a plain interface change, nothing fancier.Why: building a model for a findability problem wastes effort and the button is still hard to find.
- Ground any AI recommendation in a real ingredient database before it ships, never the model's own claim alone.Why: a hallucinated "dairy free" label is worse than the manual check it was supposed to replace.
How to answer this, stage by stage
Nobody is grading whether you know recipe apps. They're grading whether you can tell a hard-to-find button apart from a hard-to-do task, using what people actually said.
Let's learn
What does it actually mean when someone tells you a task is hard?
Kernelwell recommends recipes and plans a family's week of meals, matched to their allergies, their budget, and their macros.
For four straight quarters before this one, every single piece of feedback tagged "hard to use" went straight into the interface team's backlog. No exceptions. Nobody read the actual words behind the tag, only the tag itself.
This quarter, a new researcher read every one of them first. Over thirteen weeks, a hundred and eighteen pieces of feedback got the same tag, "hard to use." This time somebody actually reread the words behind each one before it got routed anywhere.
Fifty one of those notes really were about a control someone couldn't find, a buried filter, a setting three taps deep. Sixty seven were not. They were about a task that was still hard once you'd found the right screen: checking a recipe against two allergy lists and a set of macro targets, by hand, every week.
What it costs at its worst: leadership sees "hard to use" and ships another interface pass, more filters, better labels, and the sixty seven judgment cases don't move an inch, because the screen was never their problem. Or worse, someone builds a quick recipe assistant to look responsive, lets it answer "is this safe" in free text with no real check behind it, and it tells a parent a sauce is peanut free when it isn't. A slow, honest "maybe check this one" is annoying. A fast, wrong "yes, safe" is dangerous.
What I would leave alone: the fifty one real navigation cases. A missing filter chip, a buried setting, a label that doesn't say what it means, all get a plain interface fix. Building a model to guess what someone meant by a button they couldn't find would be solving a problem the button already solves, just badly placed.
The lesson: "hard" is not one word with one meaning. It hides two completely different problems, and a tag on a support ticket cannot tell them apart. Only the actual words can.
Now here is the same thing as a story
Say the short version out loud in an interview. Read this one when you want to feel exactly how a filter chip that fixed nothing sat unnoticed for eight weeks.
Ozma Bakare can tell in about four seconds whether a recipe is safe for her son, just from the ingredient list.
She's managed her son's peanut allergy since he was two, her daughter's dairy intolerance since last year, and she's six weeks from a half marathon, chasing a protein target her coach set for her. Every recipe that lands on her table gets checked against all three, in her head, before anyone eats.
Kernelwell arrived in her house in February, recommended by a friend from her running club. For the first few weeks it was genuinely useful. It found her fifty new dinner ideas she'd never have searched for herself, tagged high protein, dairy optional, nut free wherever the recipe's headline ingredients allowed it.
But "wherever the headline ingredients allowed it" turned out to be doing a lot of quiet work in that sentence. Kernelwell tagged recipes by their main ingredients, not every line of the method. A "nut free" curry might still list a nut oil three steps down, in the marinade nobody scrolls to. So Ozma kept doing exactly what she'd always done: reading every recipe's full ingredient list herself, checking it against her son's list, her daughter's list, and her own macro target, by hand, every single week. The tool changed what she found. It never touched what she still had to check.
By March, in every research interview Kernelwell ran with her, she said some version of the same thing: "this app is hard to use." The research team wrote it down, tagged it "hard to use," and moved on to the next interview.
In week 3, the interface team shipped what looked like the obvious fix: allergy filter chips, so you could filter the whole recipe list down to "nut free" and "dairy free" in one tap instead of hunting through menus.
It didn't help. Ozma kept saying the same thing in April, and in May. The filter chips solved a problem she'd never actually had, finding recipes. They left completely untouched the one she did have, checking them.
Nobody noticed the gap for six more weeks, because "hard to use" kept getting logged as one tag, and one tag looked like one problem.
In week 9, Tobit Sarafyan, a researcher who'd joined Kernelwell a month earlier, sat down to prep a quarterly report and did something nobody had done in four straight quarters: he reread the actual words behind every single "hard to use" ticket, instead of trusting the tag. Twenty minutes in, he found Ozma's line from March, word for word: "I have to check every recipe against my son's list, my daughter's list, and my own macros by hand, every week, and the app doesn't help with any of it."
He brought it to Ionara with one question: "why are we treating 'I can't find the filter' and 'I have to check three lists by hand' as the same ticket?"
Ionara ran the recut herself. She pulled every "hard to use" note from the last thirteen weeks, a hundred and eighteen of them, and this time she sorted by what the words actually said, not the tag.
Twenty three households, unprompted, described the same cross check Ozma had: allergies plus macros, done by hand, every week. Nine of those twenty three had already built their own workaround, a shared spreadsheet, a running notes app list, to do it themselves.
Before Ionara acted on any of it, she wrote out three honest guesses, because a real pattern and a convenient one can look identical from across a support queue. One, the "hard to use" count was just noise, unlucky households, no shared cause. Two, it really was navigation, and the March filter chips just hadn't reached everyone's app version yet. Three, it was a genuine task, checking allergies and macros across a recipe, that no filter or menu could ever remove, because the checking itself was the hard part.
Only the third guess survived contact with the actual numbers. The filter chips had reached every user by week 4. The households still describing the cross check by hand, in week 11, all had the update. They just weren't being helped by it, because it was never their problem.
The test that told the two apart was simple to say and took real work to run: if navigation were fixed perfectly, would the hard part still be there? For the fifty one, no. Find the filter, and the complaint disappears. For the twenty three, yes. Ozma could have every allergy chip Kernelwell ever built, tap them all correctly on the first try, and she would still be sitting at her counter on Sunday night, reading every ingredient line by hand, because Kernelwell never actually checked the recipe against her son's allergy list. It only ever checked whether the recipe's headline said "nut free."
Ionara's team had one faster option on the table before this: a single disclaimer on every recipe, "please check ingredients carefully," and call it done. She rejected it. It would have covered Kernelwell, not Ozma, and it still would have left her doing the exact same manual check every week, just with an extra sentence to read first.
The AI-specific risk sat right underneath the fix Ionara actually wanted to build: a model that reads a recipe and a household's allergy and macro profile, and flags a real conflict, has to be right, not just fluent. A model asked "is this dairy free" will answer in the same calm, certain voice whether it's correct or hallucinating a swap that still contains a dairy derivative. Before it could ship, that check had to clear a set pass rate on a labeled set of recipes with known allergens, not just look right on the handful the team happened to try. The guardrail Ionara built in: every flag gets checked against a real, structured ingredient and allergen database before it ever reaches a screen, never just the model's own free text claim. That check adds real time, about two to three seconds per recipe, against a near instant unchecked answer. She took that trade on purpose, because a wrong "safe" is worse than a slow one.
TRACE, so a hard task doesn't get diagnosed as a hard-to-find button
Not a way to decide whether Ozma was right to be frustrated. TRACE is what stops "just add a filter" and "build us a model" from getting the same confident answer, when only one check actually tells them apart.
Three things worth saying plainly, since interviewers push here. The team's first, faster idea was a blanket ingredient disclaimer on every recipe, rejected because it still left Ozma doing the same manual check every week, just with an extra line to read. The AI-specific risk worth naming by name: a model checking a recipe against an allergy list can sound exactly as certain when it's wrong as when it's right, so every flag it raises gets checked against a real structured ingredient database before it ships, never trusted as free text alone. And the trade-off, accepted on purpose: that check costs two to three seconds per recipe against an instant, unchecked answer, worth it because a wrong "safe" is worse than a slow one.
And if you want to be sure it really works, try it somewhere else
Same five letters, a loading dock at six in the morning instead of a kitchen counter. This time most of the "hard to use" pile really was navigation, and only a genuine minority was ever a candidate for AI.
GridWrench is a scheduling and dispatch app HVAC field technicians use to see their day's jobs, reroutes, and parts. Wynfield Kastor owns dispatch product there. Over an eight week stretch, forty pieces of feedback said some version of "I need smarter routing."
Wynfield ran the same recut. He read the actual words behind every one of the forty notes, not the tag. Thirty one were genuinely about not being able to find where a reassigned job had gone, buried three screens deep after a dispatcher moved it. Nine were about something else entirely: three jobs go red at once, and the technician has to decide which one to hit first, weighing drive time, parts on the van, and which customer's contract has the tightest service window. No amount of finding the reassigned job faster removes that decision.
Wynfield's team floated one AI idea early: a model that predicts the "best" next job automatically and just tells the technician what to do. He rejected it for the thirty one navigation cases, because building a smart recommendation on top of a screen someone simply couldn't locate would treat a navigation problem as if it were a judgment problem. He shipped a persistent "reassigned today" list on the home screen instead, and the thirty one cases dropped out almost immediately.
The nine judgment cases stayed open, a real candidate for a future tool that weighs drive time against parts against contract terms. But nine out of forty wasn't worth a model before the thirty one cheap, honest wins had shipped.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it, don't assume every "hard to use" note is a UI problem or an AI problem, check what the words actually say first.
Cost: no time to reread all hundred and eighteen notes. Sample twenty, and ask the same one question of each: would this survive a perfect navigation fix?
The model got better, for real: say a future version of Kernelwell checks allergies with total accuracy. The evidence test still runs exactly the same way, it just moves where the honest line between UI and AI sits, not whether you still need to check.
Where people run it wrong.
They read the tag and skip the words, because the tag is faster and feels like enough.
They find one real judgment case and assume the whole pile is the same, building AI for complaints that were genuinely just about a missing button.
They find a real UI problem and build a clever model around it anyway, because "AI" sounds better in a roadmap review than "we moved a filter."
How to use it live. When an interviewer throws this at you cold, buy two seconds by asking one thing back: "do we know yet whether people mean they can't find it, or they can't do it?" That question alone is usually exactly what a question shaped like this one is listening for.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if a complaint has both signals, someone can't find the filter and also does the cross check by hand?" Response: split it. Fix the finding half with the interface, and still run the evidence test on the doing half separately, since fixing one doesn't retire the other.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Opportunity identification for AI
- #1 What characteristics make a workflow a good candidate for AI? List five.
- #2 Describe a method for finding AI opportunities inside an existing product without starting from the technology.
- #3 How do you distinguish a problem AI solves from a problem AI merely touches?
- #4 Rank these by AI suitability and justify: expense approval, contract review, invoice matching, hiring decisions.
- #5 Explain why high-volume, low-stakes, tolerant-of-error tasks are the best first targets.
- #6 Your support team handles 8,000 tickets a month. Structure a discovery process to find the AI opportunity.