…
AI Product Case Questions

Tell me about a time you built an AI product from scratch

A worked answer to a real AI PM interview question: tell me about a time you built an AI product from scratch.

Transcript

Read the full transcript (1,425 words)

[INTERVIEWER] Tell me about a time you built an AI product from scratch. "Tell me about a time you built an AI product from scratch." Here is the thing most candidates miss. Zero to one AI is not really a story about a feature. It is a story about the calls you made under ambiguity. What to build, what data you had, and how you knew it worked before you shipped it.

A strong answer owns those three decisions out loud. A weak one narrates a feature and quietly skips every judgement call that actually mattered. The core of this question is whether you can make defensible decisions when nobody hands you the ground truth. In a normal product you have specs and analytics. In a from scratch AI product you often have no spec, no labelled data, and just a rough hunch that users want the thing.

Scoping and eval design are the tells of someone who has genuinely built this before. In the next few minutes I will give you the three calls to structure around, a full worked example, and the follow ups that separate a real builder from someone who just watched one get built. Pick a product or feature you took from nothing to shipped.

Then structure it around STAR, but weight it heavily toward three specific decisions inside the Action. Start with the Situation, taking about forty five seconds. The problem, and why AI was the right tool rather than a bolt on. What made a plain deterministic solution not enough? If a rules engine would have done the job, they will wonder why you reached for a model.

Then Task, thirty seconds. Your remit, and the ambiguity. Say it plainly that you had no spec, no labelled data, just a hunch users wanted this. Naming the ambiguity is what makes the rest impressive. Now the Action, and this is three and a half minutes of walking through three calls. First call is scope. How did you cut it down to a first version that could actually ship?

What did you deliberately leave out? And I mean deliberately. The narrower your version one, the more senior you sound, because scoping down under pressure is the hard part. Second call is data. Where did the training or grounding data come from? Or how did you get by with none? Cold start is the real AI PM problem, so name how you cracked it.

A small curated set, synthetic examples, a rules baseline first, or human labels you gathered yourself. Third call is eval. How did you decide it was good enough to ship? The golden set, the metric, the bar, and the human in the loop for the long tail. This is the load bearing part of the whole answer. If you skip it, you have told them you ship on vibes.

Then Result, about a minute and a half. Adoption, a quality metric, a business outcome, a number. Then what you learned and what version two changed. Here is the whole thing, joined up. "We had a knowledge base of four thousand help articles, and users could not find anything. Deflection was low, tickets were climbing. So I pitched an AI answer box that read the user question and answered from our own docs.

It was properly zero to one, with no spec, no eval set, and a real risk of confident wrong answers on billing questions, which is exactly where a wrong answer costs you. On scope, I cut it hard. Retrieval augmented generation over our docs only. No open ended chat. And it would only answer when it retrieved a good source.

Otherwise it showed the top three articles and a contact link. The rule that it refuses to answer when it cannot ground itself was the most important scoping decision I made. On data, I had no labelled question and answer pairs. So I pulled a hundred and twenty real questions out of past tickets, sat with a support lead, wrote the correct answers, and that became my golden set.

That is the cold start answer. I built the ground truth by hand from real user language. On eval, I measured two things. Retrieval hit rate, meaning was the right article in the top three. And answer faithfulness, meaning did the answer only use the retrieved text, graded by a stronger model plus a human spot check on forty cases.

I set a ship bar before I built anything. Eighty five percent retrieval hit, and zero hallucinated policy claims in the sample. First build came in at seventy two percent retrieval. Not good enough. So I fixed the chunking and added query rewriting, and got it to eighty eight. The result was that after launch, self serve deflection rose from twenty two percent to thirty four over eight weeks.

Ticket volume on covered topics dropped about twenty percent. And because of that no source fallback, we had zero policy hallucination escalations, which for a billing product is the number that keeps you employed. Version two added feedback thumbs, so real misses grew the golden set over time." Notice what that answer never does. It never just describes the feature.

Every beat is a decision with a reason behind it. If they want a second example, keep a lighter one ready. "We had no product for turning long sales calls into follow up summaries. Reps did it by hand and half skipped it. I scoped version one to a three bullet action summary from the transcript, with no automated CRM writes.

For data I used forty real calls where a manager wrote the ideal summary as my golden set. The eval was simple. A rep rated each summary usable or not, and I needed eighty percent usable to ship. The first build hit sixty five because it invented commitments, so I tightened the prompt to only extract spoken facts and reached eighty four.

Adoption jumped from half the team to nearly everyone, since the draft was faster to edit than write." Same three calls in a minute. Scope, data, eval. Here is what makes them lean in. First, a ruthlessly narrow version one with an explicit rule that it does not answer when unsure. That shows you understand AI products fail by being confidently wrong, and you designed against it from day one.

Second, a concrete answer to cold start. You did not wave at getting some data. You said you built the golden set by hand from real tickets. And third, eval defined before launch, with a named bar and a number you missed and then fixed. That arc of missing the bar, diagnosing it, and fixing it is far more convincing than saying it worked the first time.

Now let us look at the ways this goes wrong. The most common one is narrating the feature and skipping how you knew it was good enough to ship. If your answer has no eval in it, you have failed the question, full stop. The second trap is having no real answer to where the data came from, because that is the actual hard part of zero to one AI, and interviewers push on it precisely because it is where weak candidates go vague.

And the third is scope creep in the telling. A version one that tried to do everything reads as junior, because senior PMs are known for what they cut. And expect the follow up. They will ask why not just use a bigger model, or how did you handle the questions your docs did not cover. Your no source fallback answers the second.

For the first, say you measured on your own golden set and the mid size model plus good retrieval cleared the bar, so paying for more capability the task did not need would have just added cost and latency. So here is the whole picture. Pick something you took from nothing to shipped. Walk the STAR, but spend your time on three calls.

What you scoped in, and what you deliberately cut. Where the data came from, especially the cold start move of building it yourself. And the eval bar you set before launch, the number you missed, and how you closed the gap. Then land a real adoption number and one honest lesson. One line to carry in is to own three calls out loud.

What you scoped in, where the data came from, and the eval bar you set before you shipped. Everything else is just the feature.

Keep learning