…
AI Product Case Questions

A launched AI feature had real model limits. How did you design the CX?

A worked answer to a real AI PM interview question: a launched AI feature had real model limits. How did you design the experience?

Transcript

Read the full transcript (1,431 words)

[INTERVIEWER] A launched AI feature had real model limits. How did you design the CX? "A launched AI feature had real model limits. How did you design the customer experience around them?" Here's the truth this question is built on. Every real AI feature has limits. The model is wrong some of the time, and often it's confidently wrong, which is worse.

Great AI PMs design honestly around that instead of hiding it. The winning answer shows you surfaced uncertainty, gave users a way out when the model wasn't sure, and built the fallback into the product rather than pretending the model was perfect. The core of this question is whether you design for the model being wrong. A PM who assumes the model is right ships harm, full stop.

They don't fear the imperfect model, they fear the PM who doesn't plan for imperfection. In the next few minutes I'll give you the scaffold, the three design decisions that carry the answer, a full worked example, and the follow ups. Pick a launched feature where the model was genuinely imperfect. Not a hypothetical, something you shipped. Start with Situation, about forty five seconds.

The feature, and the specific limit, named honestly. Something like the model was right about eighty five percent of the time and confidently wrong the rest, on a task where a wrong answer cost the user real money or real time. The honesty about the number is the whole setup. Then Task, thirty seconds. Your job is to ship something useful despite a model that would sometimes be wrong.

Name the tension out loud. It's usefulness versus trust, and if you push either one too far the product fails. Now for the Action, which takes about three and a half minutes and covers three design decisions. The first decision is to surface uncertainty honestly. How did the UX tell the user when the model wasn't sure? Confidence shown, hedged language, a double check this prompt.

What you're avoiding is a false air of certainty, because fluent confident prose is exactly what makes people trust a wrong answer. The second decision is designing the fallback. What happened when confidence was low? A human handoff, showing sources so the user could verify for themselves, or a not sure here are some options state instead of a single wrong answer.

The fallback is the safety net, and a feature without one is just hoping. The third decision is containing the blast radius. Where did you decide the model must not act alone? Reversible actions only. A human in the loop for anything costly. The model suggests, the user confirms. This is where you prove you understand that some mistakes you can undo and some you can't.

Then Result, about a minute and a half. Trust and adoption metrics. A number on satisfaction, task completion, or complaint rate. And the point you want to land is that being honest about the limits grew trust rather than killing usage. Here's the whole thing joined up. We launched an AI feature that automatically drafted replies to customer emails for our support agents.

The real limit was that on nuanced or angry emails, it was right maybe eighty percent of the time. And worse, it was fluent and confident even when it had misread the situation completely, so an agent could fire off a wrong answer without ever noticing. My job was to make it genuinely useful without letting that confidently wrong twenty percent reach customers.

We made three design calls. First, we never automatically sent. Ever. The model drafted, the agent always reviewed and hit send themselves. The model suggests, the human confirms. That one rule contained the entire blast radius. Second, we surfaced uncertainty. When the confidence of the model was low, or when it detected the email was a complaint or a refund request, the draft came with a visible banner saying please review carefully this one is tricky, and no single click send on those.

So the UX literally got more cautious exactly where the model got less reliable. Third, we showed which knowledge base article the draft was based on. So the agent could check the source in two seconds, rather than trusting the fluent prose on faith. And we measured trust, not just speed. Agents used the draft as is on about sixty percent of routine emails, edited it on most of the rest, and here's the good part.

The low confidence banner correlated with a much higher edit rate. Which is exactly what we wanted, agents leaning in harder precisely where the model was shakier. Reply time dropped about thirty percent. And customer satisfaction on AI assisted replies held level with fully manual ones, because the honest review this design stopped the bad drafts going out the door.

The best part was when we split tested hiding the confidence banner. Edit rates fell and complaint rate ticked up. That told us the honesty wasn't decoration. It was doing the actual work. That last test is the strongest thing in the answer, because it proves the honest design earned its place. If you want a second one to keep ready, here's a different flavour of the same idea.

We shipped an AI feature that suggested a category and price band when a seller listed a used item. The model was right about eighty five percent of the time, but on unusual items it would guess confidently and wrong, and a wrong price band costs the seller real money. So we never automatically applied the suggestion. It arrived prefilled but editable, with a quiet line saying suggested from your photos adjust if this isn't right.

When the confidence of the model was low, we didn't show a single guess at all, we showed two or three plausible categories and let the seller pick. And the price was always a range, never a single number, because a range tells the truth about the uncertainty. Sellers accepted the suggestion as is on about seventy percent of listings, correction rates were highest exactly where confidence was low, and listing time dropped by roughly a quarter without the mispricing complaints we'd feared.

Same three moves. Surface the uncertainty, give a fallback when unsure, and never let the model act alone on something costly. Here's what makes them lean in. First, the model suggests and a human confirms, with automatic action switched off for anything costly. That's you showing you understand irreversibility. Second, uncertainty surfaced in the UX, not buried under confident prose.

That's the AI native instinct they're hunting for. And third, a metric showing that honesty grew trust, plus a test that proved the honest design mattered. Anyone can claim honesty is good. You measured it. Here are the ways candidates usually fall down. The first is designing as if the model is right, with no fallback for the wrong answers at all.

If your answer has no what happens when it's wrong in it, you've missed the entire point of the question. The second trap is hiding the limits to look impressive, which burns trust the first time it fails visibly, and AI features always fail visibly eventually. And the third is letting the model take an irreversible or costly action with no human confirm.

The moment you describe an AI automatically sending, automatically charging, or automatically deleting on an eighty percent model, the interviewer stops listening. And expect the follow up. They'll ask whether showing all those warnings and hedges hurt adoption. This is the key question, so have the answer ready. The answer is no, because trust is what drives sustained adoption, and the split test showed hiding the honesty made things worse, not better.

They may also ask how you picked the confidence threshold for the banner. Say you tuned it against the edit rate data, so the banner fired where agents genuinely needed to slow down, not on everything. To pull the whole picture together, name the limit honestly with a real number. Then three design moves. Surface the uncertainty in the UX.

Build a real fallback for the low confidence cases. And contain the blast radius, human confirms on anything costly or irreversible. Then land a trust metric and, if you can, a test that proves the honesty did the work. The one line to carry in is this. Design for the twenty percent where the model is wrong. Surface uncertainty, show the source, keep a human on anything costly.

Being honest about the limits is what earns the trust, not what costs you the adoption.

Keep learning