Define success metrics for a product I name, live.
When an interviewer names a product on the spot, the only wrong move is reaching for the number that needs three months to speak.
- Name the real business outcome before naming any metric.Why: every later number is worthless if it is chasing the wrong finish line.
- Find the signal that would move weeks before that outcome does, and say why it moves first.Why: this is the actual answer to "how do you measure this," not the outcome itself.
- Say out loud how that signal gets faked, before anyone asks.Why: naming the trap yourself is what turns a real metric into a real answer.
- Attach a real action to at least two thresholds on that signal, not just a target number.Why: a metric nobody acts on is a chart nobody opens.
- Pair the leading signal with a check built for AI failure, not a generic edit count.Why: a model can sound sure of itself and still be wrong, and only a fact check catches that.
- Run the same four moves again, in your head, on a totally different kind of product before you stop talking.Why: that is the actual proof you have a method, not a memorized answer for one product.
How to answer this, stage by stage
Nobody is grading whether you know the phrase north star metric. They are grading whether you can build one, out loud, for a product you have never seen before today.
Let's learn
Every quarter, Osahon Amadasun sat down and drafted about ten grant proposals for Bridgemoor Youth Partners, a nonprofit that runs after school programs and mentoring for teenagers. A single proposal took him close to fourteen hours: pulling last year's numbers from a binder, matching them to a funder's own wording, writing the whole narrative from a blank page. He had done this for seven years. Bridgemoor's win rate, the share of proposals that actually got funded, sat around twenty two percent.
Fundstream is a tool built by a company called Tillwork. It reads a nonprofit's past reports and a funder's own guidelines, then drafts the proposal's narrative sections for a person to check before anything goes out.
In his first weeks with Fundstream, Osahon's draft time dropped to about two hours a proposal, mostly reading and checking. He still read every sentence hard, comparing every number in the draft against the real report behind it. He edited close to sixty percent of what Fundstream handed him.
Over the next thirteen weeks, that habit thinned out. Fundstream kept getting the numbers right, so Osahon checked less. The share of paragraphs he submitted with no edit at all, what Tillwork calls the kept as is rate, climbed from thirty five percent in week one to ninety one percent by week thirteen.
In week seven, Tillwork pushed a small update to Fundstream's underlying model, meant to make the narratives read more persuasive. Nobody at Tillwork or Bridgemoor noticed anything wrong at the time. The drafts still sounded right and still matched Bridgemoor's usual voice. Two proposals went out that week with a specific claim in them: a ninety four percent college enrollment rate for the program's graduates. Bridgemoor had never tracked that number, not that year or any year before it. Nobody made it up on purpose. The model did, in a sentence confident enough that Osahon, who had stopped checking every claim by then, read it and moved on. Three weeks later, a program officer at one of the funders called to ask where the ninety four percent came from. Osahon did not have an answer.
The choice I would take back. We watched kept as is rate every week and treated a rising number as good news, full stop. I would take that back. I would have paired it, from day one, with an automatic check: does every number in the draft trace back to a real line in a real report? Below ninety five traced claims out of a hundred, the draft gets held for a person to check, no matter how confident it reads. That check went live the week after the funder's call. It is the reason the next quarter looked like this.
What I would leave alone. Renewal proposals, where Bridgemoor asks a funder to fund the same program again with numbers that same funder already checked and funded last year, barely need the gate at all. Those claims were already verified once. The gate almost never catches anything there, and that is fine. It is not a reason to skip it everywhere else.
The lesson. A number that moves early is not automatically an honest number. It only stays honest if something else is watching what it is built on. Find the metric that moves first, then build the thing that keeps it honest, or the blind spot just moves earlier. It does not close.
Now here is the same thing as a story
The short version is above. Read on for the Tuesday the funder's call actually landed.
Osahon has run grants for Bridgemoor Youth Partners for seven years. Hand him a funder's guidelines and he can tell you inside a minute which of last year's numbers to lead with and which to leave out. That instinct is most of the job.
When Fundstream arrived, the good weeks looked like this. Monday morning, coffee still warm, he would open a blank proposal and have a full first draft by ten. He spent the rest of the morning checking it line by line against the real reports, the way he always had.
The checking thinned out in three beats. By week three, he was still reading every paragraph, but only spot checking the numbers instead of tracing each one back to its source. By week seven, he read the draft once, fast, mostly for tone. By week ten, he barely opened the source reports at all. Fundstream had been right for two straight months. Why would this week be different.
Nothing about week seven felt different either. Same login, same coffee, same two proposals due by Friday. Fundstream handed him a paragraph about the mentoring program's outcomes. It read clean. It matched Bridgemoor's usual voice. It said ninety four percent of graduates went on to college.
He kept reading. He did not stop on that line.
Three weeks later, on an ordinary Tuesday afternoon, his phone rang. A program officer from one of the funders had a question about the college enrollment figure in the proposal. Bridgemoor had never tracked that number, not that year or any year before it. Osahon did not have an answer, and for a moment neither did the story he had been telling himself about how careful he still was.
It was never really about how many paragraphs he edited. He never had a number in his head for how much he trusted Fundstream. He had a feeling, and the feeling only had two settings: worth checking, or not. Two months of clean drafts had quietly flipped it to not, and nothing about week seven told him it had flipped back.
A year earlier, in the meeting where Tillwork decided how kept as is rate would be reported to customers, someone asked whether a rising number always meant the tool was working. The honest answer at the time was yes, mostly, because the model was still new enough that people checked its harder claims out of habit. Nobody in that meeting pictured a model update, two years later, quietly writing a number that existed nowhere in any report.
Run the Tuesday phone call through the fixed design. Fundstream's substantiation check now runs on every draft before Osahon ever opens it. The ninety four percent claim fails the check the moment it is written, because no source document contains it. The draft holds for review with one line flagged. Osahon fixes it in ninety seconds, the same afternoon it was drafted, three weeks before any funder would have seen it.
One design watched a number and called a climb good news on its own. The other watched what the number was standing on.
What I would tell myself, before any of this: the week a number stops needing a person behind it is exactly the week to go check what is actually holding it up.
LEAD, if you want the whole method on one screen
This is a metric question, so LEAD does the real work here, not a framework bolted on after the fact.
Two things are worth saying directly here, since this is where the real judgment sits. We looked at just watching grant win rate every quarter and calling that the metric, full stop, and ruled it out: by the time a bad quarter shows up in win rate, the proposals behind it went out three months earlier and cannot be pulled back. We also looked at requiring a person to check every single AI drafted paragraph forever, no exceptions, and ruled that out too. It is the safest option and the slowest one, and it puts Bridgemoor right back to fourteen hours a proposal, which erases the whole reason to use Fundstream. The risk worth naming by name is hallucination: a model stating a specific number, like a college enrollment rate, that exists nowhere in any source document, in language confident and specific enough that nobody thinks to check it. The guardrail is the substantiation checker paired with a golden set, forty already funded proposals checked by hand, and a rule that any new model or prompt version has to trace at least ninety five claims out of a hundred back to a real source before it reaches a single customer. That check costs something too. It adds a few seconds to every draft, and someone at Tillwork has to keep the golden set current every quarter. That is a small, deliberate cost next to what one fabricated statistic in a funder's inbox can do to a relationship that took years to build.
And if you want to be sure it really works, try it somewhere else
Marlowick runs Swapcrew, a tool that looks at a store's staffing gap for the day and ranks which hourly worker to offer the open shift to first. Fintan Halsey runs product for it. Same four moves, a completely different kind of product: fast, small stakes each time, used dozens of times a day instead of ten times a quarter.
L, link. The business outcome that actually matters is unfilled shifts each week, because an empty shift means paying someone overtime, or closing the floor early.
E, early signal. The share of times the top ranked worker accepts the very first offer, with no manager going down the list. That number moves the same day, not the same week.
A, abuse. A manager keeps that number high by only ever offering the shift to the two or three workers who never say no. The accept rate looks great while everyone else quietly stops getting offered anything at all.
D, decision. When accept rate is high and offers are spread across at least eight different people a week, leave the model alone. When the same three people are absorbing more than half of all offers, throttle repeat offers to any one worker inside a seven day window, even if that means a lower accept rate that week.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to E and A: name the early signal and how it gets faked. That is the actual thing being graded.
Cost: engineering says the substantiation checker cannot ship for two months. Do not turn on unlimited automatic drafting in the meantime and call it done. Hold the automatic draft feature to claims that already have a citation until the checker exists.
The model got better, for real: say Fundstream's underlying model genuinely gets more accurate next quarter. That is still not the same claim as "a rising kept as is rate always means trust was earned." A better model just makes the gaming mode harder to catch, because a careless writer's numbers look fine for longer too.
Where people run it wrong.
They pick the outcome metric because it sounds serious, then wonder why nobody can act on it for three months.
They find a leading signal and stop there, without asking how someone could hit it without the real thing improving.
They wait for a customer, or a funder, to catch the problem, instead of building the check into the product before it ships.
How to use it live. Say the outcome and the timing problem before naming any metric: "The real number here is grant win rate, but a funder can take ninety days to answer, so I need something that moves faster than that." That line buys three or four seconds to actually think of the early signal, and it already sounds like the start of a real answer instead of dead air.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"If the substantiation check slows Fundstream down, wouldn't Tillwork just ship the faster version anyway to hit a deadline?" Response: that is the trade being made on purpose. A few extra seconds a draft costs far less than a fabricated statistic reaching a funder who has backed Bridgemoor for six years, and the bar is a probability, ninety five traced claims in a hundred, not a demand that it hit one hundred every time.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?