…
AI Product Case Questions

Prioritize features for an AI writing assistant in Word

A worked answer to a real AI PM interview question: prioritise features for an AI writing assistant inside Microsoft Word.

Transcript

Read the full transcript (1,478 words)

[INTERVIEWER] Prioritize features for an AI writing assistant in Word. This is a prioritisation question, not a design question, and mixing those up is the first way people lose it. The average candidate lists ten features and stops, as if naming ideas is the answer. The strong candidate picks a scoring method, applies it out loud with real numbers, and then commits to a ship order with a reason for each slot.

RICE, used honestly, is what does that. So today is not about inventing clever features. It is about showing you can rank under a fixed capacity and defend the ranking. What exactly is Microsoft testing? Judgement under constraint. Anyone can generate a wish list. They want to know if you can score honestly, sequence on risk and not just raw numbers, and, importantly, tell them what your own framework is hiding from you.

Because a PM who worships RICE is almost as dangerous as one who has never heard of it. By the end of this you will be able to run RICE live, in your head, with stated assumptions, and know exactly when to override it. First, take sixty seconds to pin down two things: the goal and the constraint. Prioritise towards what, exactly?

Activation of the AI feature, retention of Word users, or upselling to a paid AI tier? Say you will optimise for activation and habitual use first, because an AI writing feature only earns its keep when people reach for it weekly. Then the constraint, and this is the one people skip: one squad, one quarter. Prioritisation is meaningless without a fixed capacity to prioritise against.

If you had infinite engineers you would build everything, so name the ceiling before you rank. And clarify one more thing while you are at it: who is the user? A student writing an essay, a lawyer drafting a contract, and a salesperson firing off emails all want different things from a writing assistant, and the same feature scores completely differently across them.

Say you will assume the broad Word base, mostly office professionals writing documents and reports, so your reach numbers mean something. Pin the user, or your whole table is built on sand. Now list the candidates, then define the method. The features on the table: draft from prompt, writing a first draft from a short brief. Rewrite and tone adjustment, meaning make this shorter or more formal.

Grammar and clarity fixes. Summarise a long document. Cite and ground against the user's own files. And translate. Six real candidates. Now RICE. Reach, the users per quarter who actually hit the feature. Impact, the per user effect on the goal, scored on a set scale of 0.25, 0.5, 1, 2, or 3. Confidence, how sure you are, as a percentage.

And Effort, in person months. The score is Reach times Impact times Confidence, divided by Effort. State that formula out loud so the interviewer sees you know it, and then, this is the part that matters, actually apply it. Do not just name drop RICE and move on. On-screen reference block: Let me walk through this table, because the numbers are assumptions and I want you to hear the reasoning behind each.

Rewrite and tone: enormous reach, nearly every document gets edited, so eight million a quarter. High impact, it fixes the most common real job, so a 2. High confidence, ninety percent, because it transforms existing text rather than inventing facts, so the hallucination risk is low. Low effort, three person months. That gives roughly 4.8 million, and it wins outright.

Draft from prompt: high reach and the highest impact, a 3, because a blank page draft is magic when it works. But confidence drops to fifty percent and effort jumps to eight, because a from scratch draft invents content and carries the most hallucination and brand risk. So it scores around 1.1 million, way down the list despite the big impact number.

Grammar plus: huge reach, nine million, but low incremental impact, a 0.5, because Word already does basic grammar, so the AI version adds little. It scores about 1.8 million on reach alone. Summarise: medium reach, high impact for long document users, lands around 1.2 million. And cite and ground scores low today, only 0.45 million, because its reach is small right now.

Hold that number, because in a minute I am going to tell you why you might ship it anyway. Here is a point that separates a candidate who has actually shipped AI features from one who has only read about RICE. For a normal feature, Confidence is about execution risk: will we build it on time, and will users like it.

For an AI feature, Confidence carries a second thing: hallucination risk. Look at draft from prompt again. Its impact is a 3, the highest on the board, because a good first draft is genuinely magic. But its confidence is only fifty percent, and that is not because the engineering is hard. It is because a blank page generation invents facts, and in a Word document that goes to a client or a regulator, an invented fact is a brand damaging incident, not a cute mistake.

So the AI risk lives inside the confidence score, and it also inflates the effort, because you cannot ship draft from prompt without also building grounding, a regenerate button, and an easy undo, and all of that is real engineering. Compare that to rewrite. Rewrite transforms text the user already wrote, so it cannot invent facts out of nothing, which is exactly why its confidence is ninety percent.

Saying this out loud, that AI features carry a risk axis a normal feature does not, and that it shows up in both confidence and effort, is the insight that tells Microsoft you have actually built these things. Now commit, because a ranking with no ship order is just a spreadsheet. Ship rewrite and tone first: best RICE, lowest risk, and it teaches the user the core habit of asking the assistant to change their text.

Summarise second, a distinct high value job at moderate effort. Draft from prompt third, but only once you have built the guardrails, grounding, an easy regenerate, a one tap undo, to manage its hallucination risk. Cite and ground fourth, as the enterprise trust and upsell lever. Grammar and translate go to the backlog. Say the order and the one line why for each slot.

That is the answer, not the table. This is where a senior candidate stands out. RICE is a forcing function, not an oracle, and you should say so. It undervalues strategic bets. Cite and ground scores low on reach today, but it might be the exact reason enterprises pay for the paid tier tomorrow, and reach in a quarter does not capture that.

And it undervalues sequencing dependencies. Rewrite builds the interaction pattern that draft from prompt later reuses, so shipping rewrite first is not just about its high score. It is a dependency. Call both of those out. It shows you use the tool without being ruled by it, which is exactly the judgement they are scoring. Consider what makes the interviewer lean in.

First, that you actually computed RICE with stated assumptions, out loud, instead of just naming the framework and waving at it. Second, that you sequenced on risk and dependency, not only raw score, shipping the low hallucination rewrite before the invent content draft. And third, that you named what the framework misses: strategic bets and dependencies, and adjusted for it.

That last one is the maturity signal. Now for the traps. The first trap is a feature wish list with no scoring and no ship order, which is what most people deliver. The second is fake precise numbers with no stated assumptions, like claiming reach is 7.3 million. Where did that come from? A score built on invented numbers means nothing, so state your assumptions as assumptions.

And the third, the AI specific one, is ignoring that AI features carry a hallucination risk axis your effort and confidence scores have to capture, which normal features do not. A draft feature and a spell check are not the same kind of risk, and RICE only knows that if you tell it. So let us assemble it all. You clarify the goal of activation, and the constraint of one squad one quarter.

You list six candidates and state the RICE formula. You score them out loud with assumptions, and rewrite wins. You commit to an order: rewrite, summarise, draft, cite and ground, and you name why draft waits for guardrails. And you tell them what RICE hides. Carry this one line into the room. Prioritisation is a method plus a committed order: score RICE out loud with stated assumptions, then ship the high reach, low hallucination feature before the invent content one.

Keep learning