Define the north star metric for an AI writing assistant and defend it.
- Name keep rate as the north star, in one sentence, before defending it.Why: an interviewer is testing whether you can commit to a specific number, not list five candidate metrics and shrug.
- Reject drafts-generated-per-week as the north star, by name.Why: it is the metric everyone reaches for first, because it always goes up, and going up feels like winning.
- Say who feels each metric's error, in real minutes and real trust.Why: a metric is only as good as the person it protects, and the two candidates protect different people.
- Name the asymmetry: one error is loud and cheap to fix, the other is silent and expensive.Why: this is the whole argument. Without it, the pick is just a preference.
- Give kill criteria for when keep rate stops being the right number.Why: a metric with no expiry date becomes a religion instead of a tool.
- Pair keep rate with a sampled fact-check gate before any model swap ships wide.Why: a paragraph can survive unedited and still be wrong, so the metric needs a guardrail, not just a threshold.
How to answer this, stage by stage
Nobody is grading whether you can name a metric. They are grading whether you can commit to one number, out loud, and then defend it against the obvious runner-up without flinching. Seven moves get you there.
What happens when a number goes up and nothing gets easier
Here is what happens when a metric climbs for months while the actual job stays exactly as hard. Wordkiln is Penbridge's writing assistant. A writer types a rough note, a half finished thought, a bullet list of what the email needs to say, and gets back a full draft in under a minute.
Before a tool like this, turning a rough note into a finished six hundred word blog post took a marketing writer about forty five minutes: structure it, write it, read it back, fix the parts that sound stiff. With Wordkiln, the first draft lands in under a minute.
The team's first real success metric was simpler than keep rate, and it felt obviously right at the time: drafts generated per active writer per week. It is the easiest number in the whole product to move. Add a "try three more angles" button after every draft, and the number climbs on its own.
It did climb. Four point two drafts a week, then eleven point six, over two quarters. It went on every slide. Here is the turn. That number climbing is not the story. The story is what a writer actually does between hitting generate and hitting generate again: they skim, they don't love it, they try again, and again, hoping the fourth one lands instead of fixing the first one. The habit that quietly took hold was not "read closely and decide." It was "spray and hope one sticks."
At its worst, this costs more than the forty five minutes Wordkiln was supposed to save. A writer chasing a good draft through six regenerations spends closer to an hour just clicking, before they've written a single word of their own. The tool that was meant to save time instead adds a whole extra step nobody asked for: deciding when to stop asking the machine and start typing yourself.
The choice I would take back is not the regenerate button itself. It is that the team's first success metric was chosen for how easily it moved, not for what it actually protected. Drafts generated was the fastest number to put on a slide when Wordkiln was three weeks old and needed a growth story for the board. That was a fine choice for three weeks. It stopped being fine the moment a feature could inflate it without helping a single writer finish a single piece of work.
What I would leave alone: for something like a one line subject suggestion, the stakes are so low that a lightweight "pick one of five, none tracked for editing" flow is genuinely fine. Nobody is going to agonize over fifteen percent of a six word subject line. Keep rate earns its place on paragraphs and full drafts, where a bad pick costs a person real minutes, not on a throwaway suggestion where trying five costs nothing.
The lesson: the metric that always goes up is not the one to trust by default. It is the one to interrogate hardest, because it is the one most likely to be measuring the button instead of the work.
Now here is the same thing as a story
The short version is above. Read on if you want to feel how ordinary the Thursday looked when Cassius almost let the wrong number win.
Cassius Obeng has run product for Wordkiln at Penbridge for a year and a half, long enough to know the difference between a metric that means something and one that just looks good in a deck. He is good at spotting the second kind before it ships, usually.
The quarter Wordkiln added its "try three more angles" nudge, drafts generated per active writer per week climbed from four point two to eleven point six inside eight weeks. It was the best number the team had ever posted. Cassius put it on the slide for the board meeting himself.
Then, in a hallway on a Thursday, a data scientist named Yuki stopped him with one line: "You're not still counting regenerate clicks as adoption, are you?"
He pulled four hundred writers' real sent emails and published posts and matched them, paragraph by paragraph, against what Wordkiln had drafted for them that same week. Sixty one percent of AI paragraphs had survived into final text, unedited, before the nudge shipped. After it shipped: twenty two percent. The tool wasn't drafting worse. Writers were mass generating and discarding, over and over, and the dashboard had no way to tell the difference between that and real progress.
He went looking for what it was actually costing people. He found a content writer named Sitembile who used to finish two blog posts before lunch. She now averaged closer to one, because she'd fallen into a habit of hitting regenerate five or six times per paragraph, waiting for one that felt right instead of fixing the first one herself.
Two years earlier, when the founding team first set Wordkiln's success metric, the choice had been fast and reasonable. The product was three weeks old. It needed one number for the board deck that could only go up, and drafts generated was the obvious pick: cheap to compute, always trending the right direction, easy to explain to anyone in five seconds.
What Cassius actually did that Thursday: he did not remove the regenerate button, writers genuinely needed a second try sometimes. He built keep rate as the real weekly number, sampled four hundred writers a week against their own sent and published text, and took it to the exec team even though it told a worse story in the short term. Sitembile's team's keep rate that quarter was twenty two percent, next to a drafts-generated number that had nearly tripled. He said so plainly in the meeting.
The replay, six weeks later: with keep rate as the tracked number, the team killed the blanket "try three more angles" nudge and replaced it with a single prompt that asked what specifically felt wrong about a draft before offering a rewrite. Keep rate climbed back to fifty four percent. Sitembile was back to finishing two posts before lunch, most weeks closer to two and a half.
What I would tell myself, standing in that hallway before Yuki said anything: a number that only ever goes up isn't good news by default. Sometimes it's just the easiest number to move, and nobody has checked yet what it's actually measuring.
PICK, in one screen
This is a commit-then-defend question, not a story about a habit with two settings, so PICK fits: name the position first, then show the asymmetry that justifies it.
Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. First, the alternative most candidates reach for is a general satisfaction score, a thumbs up or down after each draft. That got ruled out too: a thumbs up costs a writer nothing to give and everything to game, and it tells you how a draft felt in the moment, not whether it actually survived contact with the writer's own judgment an hour later. Second, the real bar here is calibrated, not a hard rule: a model swap only ships to every writer once its keep rate on a held out weekly sample clears ninety percent of the outgoing model's baseline, checked against real sent and published text, not a synthetic eval set. That is a probability the model earns, not a promise anyone signs.
The trade is real too. Chasing a higher keep rate can push a model toward safer, blander drafts, the kind nobody edits because nobody has a strong opinion about them either way. Penbridge accepted a small quality cost to guard against that: every drafted paragraph now gets one extra self critique pass before it reaches the writer, which adds real inference cost and about half a second of latency, in exchange for drafts specific enough that a high keep rate means something rather than just meaning "inoffensive."
Keep rate has one blind spot worth naming plainly: a paragraph can survive unedited and still be wrong. A writer in a hurry can keep a sentence with a made up statistic or a quote attributed to the wrong person, because the sentence reads smoothly and nothing about it looks off. That is a hallucination risk, not a writing quality problem, and keep rate alone cannot catch it. The guardrail is the fact-check gate from the priority list: any paragraph carrying a number, a quote, or a named claim gets sampled and checked against a source before it counts toward the metric at all, so a fluent, false paragraph cannot quietly inflate the number that's supposed to mean "this was actually good."
And if you want to be sure it really works, try it somewhere else
Same four letters, a completely different product and industry, so the method proves itself instead of repeating a story you happened to prepare.
Hallett & Wynne sells Clauseworks, an AI tool that drafts contract clauses for a law firm's transactional team, pulling from the firm's own precedent library.
P, position. Ondrej Feryn, who runs product for Clauseworks, picked clause survival rate: the share of AI-drafted clauses that make it into the executed, signed contract with no substantive change, not clauses drafted per associate per week.
I, impact. If clauses-drafted is the number, a first year associate feels it, generating clause after clause hoping one needs no rework, while the partner reviewing the contract eats the real cost in billable hours nobody bills the client for. If survival rate is wrong, the AI team feels it, undercrediting a clause a partner rewrote in their own house style while keeping the same legal substance.
C, cost asymmetry. A low survival rate with high associate satisfaction shows up inside a single sprint review, easy to patch. Clauses-drafted climbing for a full quarter while nothing actually gets to signature is invisible until a partner finally asks why closings are taking just as long as before Clauseworks shipped.
K, kill criteria. Swap it the day Clauseworks moves from drafting whole clauses to redlining a counterparty's draft line by line, where survival stops being the right unit. Swap it if survival sits above ninety five percent for two quarters straight with no separation left between model versions.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the position: name keep rate, name the rejected alternative, name the asymmetry, done.
Cost: finance says the fact-check gate on every drafted paragraph doubles inference spend. Don't drop the gate to save money. Narrow it to paragraphs carrying a number, a quote, or a named claim, and say plainly that the trade is real: slightly higher cost per draft, in exchange for a keep rate that means something instead of just meaning "nobody bothered to read it."
The model got better, for real: say Wordkiln's base model jumps ten points on a writing benchmark overnight. That still isn't the same claim as "keep rate no longer needs watching." A better model can still get quietly worse at the one thing keep rate actually measures, whether real writers choose to keep what it wrote.
Where people run it wrong.
They pick the metric that's easiest to move, then defend it after the fact instead of before shipping it.
They treat a high number as proof of success without asking who could hit that number without the product actually working.
They pick a metric and never name the condition that would make them swap it, so it outlives its own usefulness by a year.
How to use it live. Say the position before the reasoning: "My north star is keep rate, not drafts generated. Here's why." That single sentence buys you the room to walk the asymmetry calmly, instead of listing five candidate metrics and hoping the interviewer picks one for you.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the team just games keep rate by making the model write blander, safer drafts nobody wants to change?" Response: that is the real risk, which is why every drafted paragraph gets a self-critique pass before it reaches the writer, a real inference cost we accept specifically so a high keep rate reflects a genuinely useful draft, not just an inoffensive one.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?
- #7 Explain the problem with measuring acceptance rate of AI suggestions.