CaseIntermediateQuality, Cost & Token Economics / Success metrics for AI products / #2

Define the north star metric for an AI writing assistant and defend it.

The direct answer
The north star is keep rate: the share of AI-drafted paragraphs that survive into the final sent or published version with less than 15 percent of their characters changed, checked weekly against a held-out sample. Not drafts generated, not sessions, not words written. A metric that rewards typing more text a person might throw away is optimizing for the wrong side of the desk.
Do this, in order
  1. Name keep rate as the north star, in one sentence, before defending it.Why: an interviewer is testing whether you can commit to a specific number, not list five candidate metrics and shrug.
  2. Reject drafts-generated-per-week as the north star, by name.Why: it is the metric everyone reaches for first, because it always goes up, and going up feels like winning.
  3. Say who feels each metric's error, in real minutes and real trust.Why: a metric is only as good as the person it protects, and the two candidates protect different people.
  4. Name the asymmetry: one error is loud and cheap to fix, the other is silent and expensive.Why: this is the whole argument. Without it, the pick is just a preference.
  5. Give kill criteria for when keep rate stops being the right number.Why: a metric with no expiry date becomes a religion instead of a tool.
  6. Pair keep rate with a sampled fact-check gate before any model swap ships wide.Why: a paragraph can survive unedited and still be wrong, so the metric needs a guardrail, not just a threshold.

How to answer this, stage by stage

Nobody is grading whether you can name a metric. They are grading whether you can commit to one number, out loud, and then defend it against the obvious runner-up without flinching. Seven moves get you there.

1
Scope it to one real product
Say it like this
"Let's ground this in something real. Wordkiln is Penbridge's writing assistant. You type a rough note and it hands you back a full email or blog post in under a minute. Cassius Obeng runs product for it."
Why this works
A metric argument about nothing in particular is a metric argument nobody can check.
2
State the position first, before any reasoning
Say it like this
"My north star is keep rate. What share of the AI's own paragraphs make it into the writer's final sent or published text with less than fifteen percent of the characters changed. I'm picking that over drafts generated per week, and I'll tell you why."
Why this works
This is the P step. Naming the pick before the reasoning is what separates a commitment from a meandering list of options.
3
Name who feels each metric's error
Say it like this
"If drafts-generated is wrong, the writer feels it first. They're clicking regenerate over and over, burning their morning, while the dashboard reads like a growth story. If keep rate is wrong, the model team feels it. A writer who rewrote a good draft entirely in their own words shows up as a loss, even though the draft genuinely helped them."
Why this works
This is the I step. Naming both people, not just picking a winner, is what makes the tradeoff sound like a judgment call instead of an opinion.
4
Walk the cost asymmetry with the real numbers
Say it like this
"When Wordkiln shipped a 'try three more angles' nudge, drafts generated per writer per week went from four point two to eleven point six. Everyone celebrated. But keep rate, the number nobody was tracking yet, had actually dropped from sixty one percent to twenty two. Writers weren't getting more value. They were mass generating and throwing most of it away, about three and a half hours a week per writer, and the dashboard never once flagged it."
Why this works
This is the C step, the hardest one, made concrete instead of abstract. A number the argument would fall apart without.
5
Name the alternative you rejected and why it lost
Say it like this
"We seriously considered drafts generated per week as the board metric, because it's the easiest number to move and it always trends up. We rejected it. A metric that rewards typing more text a person might delete isn't measuring the product. It's measuring how many times the generate button got pressed."
Why this works
A named, rejected alternative proves this was a real choice, not the only metric anyone thought of.
6
Give the kill criteria before anyone asks
Say it like this
"I'd swap keep rate out the day Wordkiln stops being about whole drafts and becomes a chat that edits your text a sentence at a time. Verbatim survival stops meaning anything once most of the work is small conversational edits, not full paragraphs. I'd also swap it if it sits above ninety percent for two straight quarters with nothing left for a better model to prove. At that point it's stopped separating good model versions from bad ones."
Why this works
This is the K step. A confident pick with no exit condition is stubbornness wearing a metric's clothes.
7
Close on the defended position, in one line
Say it like this
"So: keep rate, checked weekly on a held out sample, paired with a fact-check gate on any paragraph carrying a number or a quote. It moves slower than drafts generated and it makes for a worse slide. It is also the only one of the two that cannot climb for six months while every writer using it quietly gets nothing done."
Why this works
Closes on the rule itself, a line that would defend any north star pick, not just this one.
If you remember one thing A metric that always goes up is not proof the product is working. It is proof the metric never had to explain why it went down.

What happens when a number goes up and nothing gets easier

Here is what happens when a metric climbs for months while the actual job stays exactly as hard. Wordkiln is Penbridge's writing assistant. A writer types a rough note, a half finished thought, a bullet list of what the email needs to say, and gets back a full draft in under a minute.

Before a tool like this, turning a rough note into a finished six hundred word blog post took a marketing writer about forty five minutes: structure it, write it, read it back, fix the parts that sound stiff. With Wordkiln, the first draft lands in under a minute.

Knowledge spark: what is keep rate, exactly? Every week, Wordkiln compares its own drafted paragraphs against what the writer actually sent or published. If less than fifteen percent of a paragraph's characters changed, it counts as kept. The number is the share of AI paragraphs that cleared that bar, out of everything the model drafted that week.

The team's first real success metric was simpler than keep rate, and it felt obviously right at the time: drafts generated per active writer per week. It is the easiest number in the whole product to move. Add a "try three more angles" button after every draft, and the number climbs on its own.

It did climb. Four point two drafts a week, then eleven point six, over two quarters. It went on every slide. Here is the turn. That number climbing is not the story. The story is what a writer actually does between hitting generate and hitting generate again: they skim, they don't love it, they try again, and again, hoping the fourth one lands instead of fixing the first one. The habit that quietly took hold was not "read closely and decide." It was "spray and hope one sticks."

Drafts generated versus keep rate, before and after the regenerate nudge shipped
12 6 0 Drafts, before Drafts, after nudge Keep %, before Keep %, after nudge 4.2/wk 11.6/wk 61% 22%
Drafts nearly tripled while keep rate dropped by more than half. The dashboard read as a huge win the whole time.
We did not make writers three times more productive. We taught them to generate three times more and finish exactly the same amount of work.

At its worst, this costs more than the forty five minutes Wordkiln was supposed to save. A writer chasing a good draft through six regenerations spends closer to an hour just clicking, before they've written a single word of their own. The tool that was meant to save time instead adds a whole extra step nobody asked for: deciding when to stop asking the machine and start typing yourself.

Hand sketched comparison diagram titled Which error costs more. Left, a small warm grey gauge icon labelled Keep rate looks low, caption shows up fast, easy to fix. Right, a larger red orange tipped scale icon labelled Hours quietly wasted, caption hidden for months, found too late.
Keep rate under-counting a genuine rewrite is loud and cheap to fix, a low number with high satisfaction is easy to spot. Drafts generated hiding wasted hours is quiet and expensive, and by the time it shows up, a quarter of roadmap has already been spent chasing the wrong story.

The choice I would take back is not the regenerate button itself. It is that the team's first success metric was chosen for how easily it moved, not for what it actually protected. Drafts generated was the fastest number to put on a slide when Wordkiln was three weeks old and needed a growth story for the board. That was a fine choice for three weeks. It stopped being fine the moment a feature could inflate it without helping a single writer finish a single piece of work.

What I would leave alone: for something like a one line subject suggestion, the stakes are so low that a lightweight "pick one of five, none tracked for editing" flow is genuinely fine. Nobody is going to agonize over fifteen percent of a six word subject line. Keep rate earns its place on paragraphs and full drafts, where a bad pick costs a person real minutes, not on a throwaway suggestion where trying five costs nothing.

The lesson: the metric that always goes up is not the one to trust by default. It is the one to interrogate hardest, because it is the one most likely to be measuring the button instead of the work.

Now here is the same thing as a story

The short version is above. Read on if you want to feel how ordinary the Thursday looked when Cassius almost let the wrong number win.

Cassius Obeng has run product for Wordkiln at Penbridge for a year and a half, long enough to know the difference between a metric that means something and one that just looks good in a deck. He is good at spotting the second kind before it ships, usually.

The quarter Wordkiln added its "try three more angles" nudge, drafts generated per active writer per week climbed from four point two to eleven point six inside eight weeks. It was the best number the team had ever posted. Cassius put it on the slide for the board meeting himself.

Then, in a hallway on a Thursday, a data scientist named Yuki stopped him with one line: "You're not still counting regenerate clicks as adoption, are you?"

Cassius almost laughed it off. What stopped him was that he didn't actually know the answer.

He pulled four hundred writers' real sent emails and published posts and matched them, paragraph by paragraph, against what Wordkiln had drafted for them that same week. Sixty one percent of AI paragraphs had survived into final text, unedited, before the nudge shipped. After it shipped: twenty two percent. The tool wasn't drafting worse. Writers were mass generating and discarding, over and over, and the dashboard had no way to tell the difference between that and real progress.

He went looking for what it was actually costing people. He found a content writer named Sitembile who used to finish two blog posts before lunch. She now averaged closer to one, because she'd fallen into a habit of hitting regenerate five or six times per paragraph, waiting for one that felt right instead of fixing the first one herself.

Hand sketched comparison diagram titled Which error costs more. Left, a small warm grey gauge icon labelled Keep rate looks low, caption shows up fast, easy to fix. Right, a larger red orange tipped scale icon labelled Hours quietly wasted, caption hidden for months, found too late.
This is the same picture Cassius eventually drew on a whiteboard for the exec team: one error you catch in a week, one error you catch in a quarter, after it has already spent the quarter's roadmap.

Two years earlier, when the founding team first set Wordkiln's success metric, the choice had been fast and reasonable. The product was three weeks old. It needed one number for the board deck that could only go up, and drafts generated was the obvious pick: cheap to compute, always trending the right direction, easy to explain to anyone in five seconds.

What Cassius actually did that Thursday: he did not remove the regenerate button, writers genuinely needed a second try sometimes. He built keep rate as the real weekly number, sampled four hundred writers a week against their own sent and published text, and took it to the exec team even though it told a worse story in the short term. Sitembile's team's keep rate that quarter was twenty two percent, next to a drafts-generated number that had nearly tripled. He said so plainly in the meeting.

The replay, six weeks later: with keep rate as the tracked number, the team killed the blanket "try three more angles" nudge and replaced it with a single prompt that asked what specifically felt wrong about a draft before offering a rewrite. Keep rate climbed back to fifty four percent. Sitembile was back to finishing two posts before lunch, most weeks closer to two and a half.

What I would tell myself, standing in that hallway before Yuki said anything: a number that only ever goes up isn't good news by default. Sometimes it's just the easiest number to move, and nobody has checked yet what it's actually measuring.

PICK, in one screen

This is a commit-then-defend question, not a story about a habit with two settings, so PICK fits: name the position first, then show the asymmetry that justifies it.

P
Position. State the pick before any reasoning.
Keep rate: the share of AI-drafted paragraphs that survive to the final sent or published version with under fifteen percent of characters changed, sampled weekly. Not drafts generated per week.
I
Impact. Who feels each kind of error.
If drafts-generated is wrong, the writer feels it first, in wasted mornings, while the dashboard looks great. If keep rate is wrong, the model team feels it, undercrediting a rewrite that genuinely helped, and might cut something that was working.
C
Cost asymmetry. Which error is cheap and visible, which is hidden and expensive.
Keep rate's error is loud: a low percentage sitting next to high writer satisfaction is obvious within a week and cheap to patch. Drafts-generated's error is silent: it can climb for two full quarters, get celebrated at every all-hands, and only show its real cost when someone finally checks what got finished.
K
Kill criteria. What evidence flips the pick.
Swap it out if Wordkiln moves from whole-draft generation to sentence-level conversational editing, where verbatim survival stops meaning anything. Swap it if keep rate sits above ninety percent for two straight quarters with no room left to separate a better model from a worse one.

Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. First, the alternative most candidates reach for is a general satisfaction score, a thumbs up or down after each draft. That got ruled out too: a thumbs up costs a writer nothing to give and everything to game, and it tells you how a draft felt in the moment, not whether it actually survived contact with the writer's own judgment an hour later. Second, the real bar here is calibrated, not a hard rule: a model swap only ships to every writer once its keep rate on a held out weekly sample clears ninety percent of the outgoing model's baseline, checked against real sent and published text, not a synthetic eval set. That is a probability the model earns, not a promise anyone signs.

The trade is real too. Chasing a higher keep rate can push a model toward safer, blander drafts, the kind nobody edits because nobody has a strong opinion about them either way. Penbridge accepted a small quality cost to guard against that: every drafted paragraph now gets one extra self critique pass before it reaches the writer, which adds real inference cost and about half a second of latency, in exchange for drafts specific enough that a high keep rate means something rather than just meaning "inoffensive."

Keep rate has one blind spot worth naming plainly: a paragraph can survive unedited and still be wrong. A writer in a hurry can keep a sentence with a made up statistic or a quote attributed to the wrong person, because the sentence reads smoothly and nothing about it looks off. That is a hallucination risk, not a writing quality problem, and keep rate alone cannot catch it. The guardrail is the fact-check gate from the priority list: any paragraph carrying a number, a quote, or a named claim gets sampled and checked against a source before it counts toward the metric at all, so a fluent, false paragraph cannot quietly inflate the number that's supposed to mean "this was actually good."

Knowledge spark: why not just track model accuracy on a benchmark instead? A benchmark score can go up while the product gets worse for real writers, because a benchmark grades grammar and structure, not whether a specific person, on a specific Thursday, chose to keep what the model gave them. Keep rate is graded by the writer's own hands, every single week.
Weekly keep rate versus writer pulse score, ten weeks after the fix shipped
80 40 0 Keep rate Writer pulse score
Keep rate climbed steadily and then flattened around week seven, while the writer pulse score kept rising. That gap is the kill signal: it says keep rate is quietly undercounting real help, and a rewritten-in-your-own-words bucket needs adding before the metric goes stale.

And if you want to be sure it really works, try it somewhere else

Same four letters, a completely different product and industry, so the method proves itself instead of repeating a story you happened to prepare.

Hallett & Wynne sells Clauseworks, an AI tool that drafts contract clauses for a law firm's transactional team, pulling from the firm's own precedent library.

P, position. Ondrej Feryn, who runs product for Clauseworks, picked clause survival rate: the share of AI-drafted clauses that make it into the executed, signed contract with no substantive change, not clauses drafted per associate per week.
I, impact. If clauses-drafted is the number, a first year associate feels it, generating clause after clause hoping one needs no rework, while the partner reviewing the contract eats the real cost in billable hours nobody bills the client for. If survival rate is wrong, the AI team feels it, undercrediting a clause a partner rewrote in their own house style while keeping the same legal substance.
C, cost asymmetry. A low survival rate with high associate satisfaction shows up inside a single sprint review, easy to patch. Clauses-drafted climbing for a full quarter while nothing actually gets to signature is invisible until a partner finally asks why closings are taking just as long as before Clauseworks shipped.
K, kill criteria. Swap it the day Clauseworks moves from drafting whole clauses to redlining a counterparty's draft line by line, where survival stops being the right unit. Swap it if survival sits above ninety five percent for two quarters straight with no separation left between model versions.

Same shape, different stakes At Wordkiln, the hidden cost was a marketing writer's morning. At Hallett & Wynne, it's a partner's billable hour and, eventually, a client's trust in what got signed. The asymmetry holds either way: whatever can climb silently for a quarter outranks a number that would embarrass itself within a week.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the position: name keep rate, name the rejected alternative, name the asymmetry, done.
Cost: finance says the fact-check gate on every drafted paragraph doubles inference spend. Don't drop the gate to save money. Narrow it to paragraphs carrying a number, a quote, or a named claim, and say plainly that the trade is real: slightly higher cost per draft, in exchange for a keep rate that means something instead of just meaning "nobody bothered to read it."
The model got better, for real: say Wordkiln's base model jumps ten points on a writing benchmark overnight. That still isn't the same claim as "keep rate no longer needs watching." A better model can still get quietly worse at the one thing keep rate actually measures, whether real writers choose to keep what it wrote.

Where people run it wrong.
They pick the metric that's easiest to move, then defend it after the fact instead of before shipping it.
They treat a high number as proof of success without asking who could hit that number without the product actually working.
They pick a metric and never name the condition that would make them swap it, so it outlives its own usefulness by a year.

How to use it live. Say the position before the reasoning: "My north star is keep rate, not drafts generated. Here's why." That single sentence buys you the room to walk the asymmetry calmly, instead of listing five candidate metrics and hoping the interviewer picks one for you.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits defending a north star metric choice, and why?
Tap to flip
ANSWER
PICK. It's a tradeoff question between two candidate metrics, and the job is to commit to one, then show the asymmetry that justifies it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Cassius Obeng, product lead for Wordkiln at Penbridge, with content writer Sitembile and data scientist Yuki.
3 · THE POSITION
What is the picked north star metric, defined exactly?
Tap to flip
ANSWER
Keep rate: the share of AI-drafted paragraphs that survive into the final sent or published version with less than 15 percent of characters changed, sampled weekly.
4 · THE ASYMMETRY
Which error is cheap and visible, and which is hidden and expensive?
Tap to flip
ANSWER
Keep rate under-crediting a real rewrite is cheap and visible, it shows up fast. Drafts-generated hiding wasted hours is hidden and expensive, it can climb for a full quarter before anyone checks what actually got finished.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Picking drafts generated per week as the first success metric, chosen for how easily it moved. It made sense when Wordkiln was three weeks old and needed a simple growth number for the board.
6 · THE NUMBER
Fill in the blank: after the regenerate nudge shipped, drafts generated went from 4.2 to ___ a week, while keep rate dropped from 61 percent to ___.
Tap to flip
ANSWER
11.6 drafts a week; keep rate dropped to 22 percent.
7 · THE REPLAY
Same product, new design, what changes?
Tap to flip
ANSWER
With keep rate tracked and the blanket nudge replaced by a targeted rewrite prompt, keep rate climbed back to 54 percent within six weeks, and Sitembile went back to finishing two posts before lunch.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the picked metric?
Tap to flip
ANSWER
Clauseworks, the contract-drafting tool from Hallett & Wynne. Its picked metric is clause survival rate, the share of AI-drafted clauses that reach the signed contract with no substantive change.

Check yourself Score: 0 / 0

True or false
1. True or false: drafts generated per writer per week is the stronger north star, because a rising number is always a sign the product is working.
  • True
  • False
Show hint
Check what the number did while writers were mass generating and discarding drafts.
Show answer
False. Drafts generated per week nearly tripled while keep rate, the share of drafts writers actually used, dropped by more than half. The rising number hid the real cost instead of proving success.
Multiple choice
2. Why does keep rate's error rank as cheaper than drafts-generated's error, even though keep rate can also be wrong?
  • A. Because keep rate is always a bigger number, so mistakes matter less.
  • B. Because a low keep rate next to high writer satisfaction is obvious within a week, while a climbing drafts-generated number can hide wasted hours for a full quarter.
  • C. Because keep rate cannot be measured until a model is fully retrained.
  • D. Because drafts-generated only applies to blog posts, not emails.
Show hint
Think about how fast each kind of error would get noticed, not how big either number is.
Show answer
B. The asymmetry is about visibility and speed of discovery, not the size of either metric. A silent, slow-to-surface error is the more expensive one to build your metric against.
Fill in the blank
3. Keep rate is defined as the share of AI-drafted paragraphs that survive into the final sent or published version with less than ___ percent of their characters changed, checked ___ against a held-out sample.
Show hint
Check the direct answer at the top of the page.
Show answer
15 percent; weekly. Both numbers matter: the threshold sets what counts as "kept," and the weekly cadence is what lets the team catch a drop before it costs a whole quarter.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the decision made when Wordkiln was three weeks old, not a dial anyone could just turn back up.
Show answer
Model answer: Choosing drafts generated per week as the first success metric, because it was the cheapest, fastest-moving number to put on a board slide. It made sense for a three-week-old product that needed a simple growth story, and stopped making sense once a feature could inflate it without helping anyone finish real work.
Short answer, apply it yourself
5. Think of a tool you use that could be measured by either "how much you used it" or "how much of what it gave you, you actually kept." Which would you pick as the north star, and what's the asymmetry that decides it?
Show hint
Name who is hurt first by each kind of wrong metric, not just which number sounds nicer.
Show answer
Model answer: A recipe app that suggests substitutions could track "substitutions shown per session" or "substitutions actually cooked with." The first climbs the moment the app gets pushier about suggesting things. The second is harder to game and tells you whether anyone trusted the suggestion enough to use it, so it should win as the north star.
Short answer, the number question
6. If keep rate had dropped from 61 percent to 50 percent instead of 22 percent after the nudge shipped, would the same pick and reversal still hold? Show the reasoning.
Show hint
Think about what the size of the drop changes and what it does not.
Show answer
Yes, the pick would not change, though the urgency would. A drop to 50 percent is smaller, so the fix could wait a sprint instead of shipping immediately. But keep rate would still be the metric worth tracking, because the asymmetry, a silent number can hide real cost while a visible one cannot, does not depend on how big the drop happened to be.
Before you close the answer
Why this works
Tests whether you can commit to one specific metric under pressure and defend it against the metric every team reaches for first, the one that always trends up because it measures activity instead of value.
Follow-up traps
"Isn't keep rate just going to punish writers who genuinely rewrite good drafts in their own words?" Response: yes, a little, which is exactly why it has a kill criterion. The moment the pulse-score gap shows survival is undercounting real help, we add a rewritten-in-your-own-words bucket rather than pretend the metric is perfect forever.

"What if the team just games keep rate by making the model write blander, safer drafts nobody wants to change?" Response: that is the real risk, which is why every drafted paragraph gets a self-critique pass before it reaches the writer, a real inference cost we accept specifically so a high keep rate reflects a genuinely useful draft, not just an inoffensive one.
If pressed
The keep-rate check runs on a stratified weekly sample, not the full population, four hundred writers drawn evenly across content type so one heavy user of short subject lines can't quietly drag the number around. A model swap only rolls out past a five percent canary once its sampled keep rate clears ninety percent of the outgoing model's baseline on that same stratified sample, a probability the new model has to earn, not a fixed pass or fail rule.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more