InterviewAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #25

Define success metrics for a product I name, live.

When an interviewer names a product on the spot, the only wrong move is reaching for the number that needs three months to speak.

The direct answer
Do not reach for the model's own score. Name the real business outcome first, find the one number that would move weeks before that outcome does, then say out loud how someone could hit that number without the outcome ever getting better. That third step, naming the fake out before the interviewer does, is what makes this sound like judgment instead of a memorized definition, and it works the same way no matter which product gets named.
Do this, in order
  1. Name the real business outcome before naming any metric.Why: every later number is worthless if it is chasing the wrong finish line.
  2. Find the signal that would move weeks before that outcome does, and say why it moves first.Why: this is the actual answer to "how do you measure this," not the outcome itself.
  3. Say out loud how that signal gets faked, before anyone asks.Why: naming the trap yourself is what turns a real metric into a real answer.
  4. Attach a real action to at least two thresholds on that signal, not just a target number.Why: a metric nobody acts on is a chart nobody opens.
  5. Pair the leading signal with a check built for AI failure, not a generic edit count.Why: a model can sound sure of itself and still be wrong, and only a fact check catches that.
  6. Run the same four moves again, in your head, on a totally different kind of product before you stop talking.Why: that is the actual proof you have a method, not a memorized answer for one product.

How to answer this, stage by stage

Nobody is grading whether you know the phrase north star metric. They are grading whether you can build one, out loud, for a product you have never seen before today.

1
Buy yourself two seconds before naming a metric
Say it like this
"Give me one second, I want to pick a real number here, not just say 'engagement.'"
Why this works
Stalls without looking stalled, and signals you are about to build something instead of recite a definition.
2
Ask what business result the product owner is actually paid for
Say it like this
"Say we're talking about Fundstream, the tool that drafts grant proposal narratives for nonprofits. The business Tillwork is actually in is dollars won for its customers, because that's what gets a nonprofit to renew next year."
Why this works
Without this step, every number that follows measures the model instead of the money.
3
Say why that outcome is too slow to run the product by
Say it like this
"Grant win rate is the real answer, but a funder can take ninety days to reply. So by the time win rate moves, whatever broke has already been shipping for three months."
Why this works
Names the actual reason a leading metric is needed, instead of just reciting that one exists.
4
Name the early signal and defend why it moves first
Say it like this
"So I'd watch the share of Fundstream's drafted paragraphs that go out with no edit at all. That number climbs the week trust builds, not ninety days later."
Why this works
This is the literal answer to the question. Everything before it was setup.
5
Immediately say how that number lies
Say it like this
"Here's the catch. That number climbs for two totally different reasons: the tool got good, or the writer stopped checking. Looking at it alone can't tell you which one happened."
Why this works
Interviewers push here first. Naming the trap yourself takes the follow up question off the table.
6
Attach a real action to it, at a real threshold
Say it like this
"So I'd pair it with an automatic check: does every number in the draft trace back to a real source document? Below ninety five traced claims out of a hundred, the draft gets held for review, no matter how confident it reads."
Why this works
A metric with no attached action is decoration. This is where a real, probabilistic bar belongs.
7
Prove it with what actually happened
Say it like this
"This isn't hypothetical. A model update in week seven started stating a college enrollment number nobody had ever tracked. Nobody caught it internally. A funder's program officer called and asked where it came from."
Why this works
A four sentence failure beats a paragraph of theory, every time.
8
Close by running the same checklist on a different product
Say it like this
"If you'd named a shift scheduling app instead, same four moves. The outcome is unfilled shifts, not queries a day. The early signal is whether the AI's top suggestion gets accepted with no manager going down the list. It gets gamed by only ever offering the shift to the two people who always say yes. So the fix throttles repeat offers to the same person inside a rolling window."
Why this works
This is the actual thing being tested: a method, not a memorized story about one product.
If you remember one thing A metric question graded live is never really about the metric. It is about whether you can find the number that moves first, on a product you just met, and then say how it lies.

Let's learn

Every quarter, Osahon Amadasun sat down and drafted about ten grant proposals for Bridgemoor Youth Partners, a nonprofit that runs after school programs and mentoring for teenagers. A single proposal took him close to fourteen hours: pulling last year's numbers from a binder, matching them to a funder's own wording, writing the whole narrative from a blank page. He had done this for seven years. Bridgemoor's win rate, the share of proposals that actually got funded, sat around twenty two percent.

Fundstream is a tool built by a company called Tillwork. It reads a nonprofit's past reports and a funder's own guidelines, then drafts the proposal's narrative sections for a person to check before anything goes out.

Knowledge spark: what is a leading indicator? A number that moves before the thing you actually care about does. Grant win rate is the real goal, but a funder can take ninety days to answer. A leading indicator moves the same week, so a problem gets caught while it is still small.

In his first weeks with Fundstream, Osahon's draft time dropped to about two hours a proposal, mostly reading and checking. He still read every sentence hard, comparing every number in the draft against the real report behind it. He edited close to sixty percent of what Fundstream handed him.

Hand sketched comparison titled Two clocks, same story. Left a grey dial labeled WIN RATE, caption barely moved, funders take ninety days to answer. Right an amber dial labeled DRAFT KEPT AS IS, caption already climbing, this one rings first.
Grant win rate and kept as is rate are telling the same story. One of them just takes three months longer to say it.

Over the next thirteen weeks, that habit thinned out. Fundstream kept getting the numbers right, so Osahon checked less. The share of paragraphs he submitted with no edit at all, what Tillwork calls the kept as is rate, climbed from thirty five percent in week one to ninety one percent by week thirteen.

Thirteen weeks, and the leading number kept climbing
91% 35% funder calls week 1 week 13
Kept as is rate, Osahon's account. The red mark is week seven, when a made up statistic reached a funder undetected.
The number that looked healthiest all quarter was the same number that let a made up statistic walk out the door twice.

In week seven, Tillwork pushed a small update to Fundstream's underlying model, meant to make the narratives read more persuasive. Nobody at Tillwork or Bridgemoor noticed anything wrong at the time. The drafts still sounded right and still matched Bridgemoor's usual voice. Two proposals went out that week with a specific claim in them: a ninety four percent college enrollment rate for the program's graduates. Bridgemoor had never tracked that number, not that year or any year before it. Nobody made it up on purpose. The model did, in a sentence confident enough that Osahon, who had stopped checking every claim by then, read it and moved on. Three weeks later, a program officer at one of the funders called to ask where the ninety four percent came from. Osahon did not have an answer.

The choice I would take back. We watched kept as is rate every week and treated a rising number as good news, full stop. I would take that back. I would have paired it, from day one, with an automatic check: does every number in the draft trace back to a real line in a real report? Below ninety five traced claims out of a hundred, the draft gets held for a person to check, no matter how confident it reads. That check went live the week after the funder's call. It is the reason the next quarter looked like this.

Win rate, before the tool, without the check, and with it
22% 15% 31% Before Fundstream No check yet Check live
Bridgemoor's grant win rate across three quarters. The dip in the middle quarter would not have shown up for ninety days. Watching win rate alone would have caught it three months late.

What I would leave alone. Renewal proposals, where Bridgemoor asks a funder to fund the same program again with numbers that same funder already checked and funded last year, barely need the gate at all. Those claims were already verified once. The gate almost never catches anything there, and that is fine. It is not a reason to skip it everywhere else.

The lesson. A number that moves early is not automatically an honest number. It only stays honest if something else is watching what it is built on. Find the metric that moves first, then build the thing that keeps it honest, or the blind spot just moves earlier. It does not close.

Now here is the same thing as a story

The short version is above. Read on for the Tuesday the funder's call actually landed.

Osahon has run grants for Bridgemoor Youth Partners for seven years. Hand him a funder's guidelines and he can tell you inside a minute which of last year's numbers to lead with and which to leave out. That instinct is most of the job.

When Fundstream arrived, the good weeks looked like this. Monday morning, coffee still warm, he would open a blank proposal and have a full first draft by ten. He spent the rest of the morning checking it line by line against the real reports, the way he always had.

The checking thinned out in three beats. By week three, he was still reading every paragraph, but only spot checking the numbers instead of tracing each one back to its source. By week seven, he read the draft once, fast, mostly for tone. By week ten, he barely opened the source reports at all. Fundstream had been right for two straight months. Why would this week be different.

Nothing about week seven felt different either. Same login, same coffee, same two proposals due by Friday. Fundstream handed him a paragraph about the mentoring program's outcomes. It read clean. It matched Bridgemoor's usual voice. It said ninety four percent of graduates went on to college.

He kept reading. He did not stop on that line.

Three weeks later, on an ordinary Tuesday afternoon, his phone rang. A program officer from one of the funders had a question about the college enrollment figure in the proposal. Bridgemoor had never tracked that number, not that year or any year before it. Osahon did not have an answer, and for a moment neither did the story he had been telling himself about how careful he still was.

We did not lose fourteen hours that week. We nearly lost a funder who had backed Bridgemoor for six years running.

It was never really about how many paragraphs he edited. He never had a number in his head for how much he trusted Fundstream. He had a feeling, and the feeling only had two settings: worth checking, or not. Two months of clean drafts had quietly flipped it to not, and nothing about week seven told him it had flipped back.

A year earlier, in the meeting where Tillwork decided how kept as is rate would be reported to customers, someone asked whether a rising number always meant the tool was working. The honest answer at the time was yes, mostly, because the model was still new enough that people checked its harder claims out of habit. Nobody in that meeting pictured a model update, two years later, quietly writing a number that existed nowhere in any report.

Run the Tuesday phone call through the fixed design. Fundstream's substantiation check now runs on every draft before Osahon ever opens it. The ninety four percent claim fails the check the moment it is written, because no source document contains it. The draft holds for review with one line flagged. Osahon fixes it in ninety seconds, the same afternoon it was drafted, three weeks before any funder would have seen it.

One design watched a number and called a climb good news on its own. The other watched what the number was standing on.

What I would tell myself, before any of this: the week a number stops needing a person behind it is exactly the week to go check what is actually holding it up.

LEAD, if you want the whole method on one screen

This is a metric question, so LEAD does the real work here, not a framework bolted on after the fact.

L
Link. What business result is actually being paid for.
Not the model's own score.
Grant win rate, because that is what gets Bridgemoor's board, and Tillwork's own contract, renewed.
E
Early signal. What moves before that outcome does, and why.
This is the real answer to the question.
Share of Fundstream's drafted paragraphs Osahon submits untouched, climbing weeks before any funder replies.
A
Abuse. How the signal gets hit without the outcome improving.
Every metric has a way to be faked.
It climbs the same way whether the tool got better or the writer just stopped checking, so a made up statistic can ride along underneath a healthy looking number.
D
Decision. What actually happens at each threshold.
A metric nobody acts on is a chart nobody opens.
Below ninety five source traced claims in a hundred, the draft is held for review, no matter how confident it reads.
Hand sketched comparison titled How the metric gets gamed. Left a green document labeled KEPT AS IS, caption looks accepted, nobody checked it fit. Right a red orange figure labeled WIN RATE, caption quietly walks away.
A high kept as is rate can mean the tool earned trust, or that nobody is checking it anymore. The number alone cannot say which.

Two things are worth saying directly here, since this is where the real judgment sits. We looked at just watching grant win rate every quarter and calling that the metric, full stop, and ruled it out: by the time a bad quarter shows up in win rate, the proposals behind it went out three months earlier and cannot be pulled back. We also looked at requiring a person to check every single AI drafted paragraph forever, no exceptions, and ruled that out too. It is the safest option and the slowest one, and it puts Bridgemoor right back to fourteen hours a proposal, which erases the whole reason to use Fundstream. The risk worth naming by name is hallucination: a model stating a specific number, like a college enrollment rate, that exists nowhere in any source document, in language confident and specific enough that nobody thinks to check it. The guardrail is the substantiation checker paired with a golden set, forty already funded proposals checked by hand, and a rule that any new model or prompt version has to trace at least ninety five claims out of a hundred back to a real source before it reaches a single customer. That check costs something too. It adds a few seconds to every draft, and someone at Tillwork has to keep the golden set current every quarter. That is a small, deliberate cost next to what one fabricated statistic in a funder's inbox can do to a relationship that took years to build.

And if you want to be sure it really works, try it somewhere else

Marlowick runs Swapcrew, a tool that looks at a store's staffing gap for the day and ranks which hourly worker to offer the open shift to first. Fintan Halsey runs product for it. Same four moves, a completely different kind of product: fast, small stakes each time, used dozens of times a day instead of ten times a quarter.

L, link. The business outcome that actually matters is unfilled shifts each week, because an empty shift means paying someone overtime, or closing the floor early.
E, early signal. The share of times the top ranked worker accepts the very first offer, with no manager going down the list. That number moves the same day, not the same week.
A, abuse. A manager keeps that number high by only ever offering the shift to the two or three workers who never say no. The accept rate looks great while everyone else quietly stops getting offered anything at all.
D, decision. When accept rate is high and offers are spread across at least eight different people a week, leave the model alone. When the same three people are absorbing more than half of all offers, throttle repeat offers to any one worker inside a seven day window, even if that means a lower accept rate that week.

Same four moves, a different kind of trap Fundstream's abuse case is a model telling a confident lie. Swapcrew's abuse case is a manager quietly leaning on the same few people. Same framework, a completely different failure mode, because the failure mode has to come from the actual product, not the acronym.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to E and A: name the early signal and how it gets faked. That is the actual thing being graded.
Cost: engineering says the substantiation checker cannot ship for two months. Do not turn on unlimited automatic drafting in the meantime and call it done. Hold the automatic draft feature to claims that already have a citation until the checker exists.
The model got better, for real: say Fundstream's underlying model genuinely gets more accurate next quarter. That is still not the same claim as "a rising kept as is rate always means trust was earned." A better model just makes the gaming mode harder to catch, because a careless writer's numbers look fine for longer too.

Where people run it wrong.
They pick the outcome metric because it sounds serious, then wonder why nobody can act on it for three months.
They find a leading signal and stop there, without asking how someone could hit it without the real thing improving.
They wait for a customer, or a funder, to catch the problem, instead of building the check into the product before it ships.

How to use it live. Say the outcome and the timing problem before naming any metric: "The real number here is grant win rate, but a funder can take ninety days to answer, so I need something that moves faster than that." That line buys three or four seconds to actually think of the early signal, and it already sounds like the start of a real answer instead of dead air.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework is this, and what does it exist to find?
Tap to flip
ANSWER
LEAD, for metric questions. It exists to find the signal that moves before the real outcome does.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Osahon Amadasun, grants manager at Bridgemoor Youth Partners, seven years into writing proposals by hand before Fundstream.
3 · THE HABIT
What did he stop doing because it worked?
Tap to flip
ANSWER
Tracing every number in an AI drafted paragraph back to the real source report before trusting it enough to submit.
4 · LEADING AND LAGGING
What are the two numbers in play here, and which one moves first?
Tap to flip
ANSWER
Kept as is rate (leading, moves weekly) and grant win rate (lagging, takes about ninety days). The leading one moves first, but it can climb for a bad reason as easily as a good one.
5 · THE OLD DECISION
What old practice does this answer say to replace?
Tap to flip
ANSWER
Treating a rising kept as is rate as good news on its own, with nothing checking what it was built on.
6 · THE NUMBER
Fill in the blank: kept as is rate climbed from thirty five percent in week one to ___ percent by week thirteen.
Tap to flip
ANSWER
Ninety one percent. It kept climbing right through the week a fabricated statistic reached a funder.
7 · THE REPLAY
Same bad week, new design, what changes?
Tap to flip
ANSWER
The substantiation checker catches the fake college enrollment number the moment it is drafted. Osahon fixes one flagged line in ninety seconds. No funder ever sees it.
8 · CROSS PRODUCT TRANSFER
Section 4 runs the same four moves on a different kind of product. Which product, and what changes?
Tap to flip
ANSWER
Swapcrew, Marlowick's shift offer tool. Same LEAD steps, but the abuse case is a manager leaning on the same few reliable workers, not a model inventing a fact.

Check yourself Score: 0 / 0

True or false
1. True or false: once the substantiation gate was live, checking a renewal proposal against the golden set mattered exactly as much as checking a brand new funder's proposal.
  • True
  • False
Show hint
Check the "what I would leave alone" paragraph.
Show answer
False. A renewal proposal reuses numbers a funder already checked and funded last year, so the gate rarely finds anything there. That does not make the gate pointless everywhere else.
Multiple choice
2. What is the actual early signal this answer says to track for Fundstream?
  • A. The model's own confidence score on each sentence.
  • B. The share of AI drafted paragraphs submitted with no edit at all.
  • C. The number of proposals submitted per week.
  • D. The average time it takes Fundstream to generate a draft.
Show hint
Look for the number that moves the same week trust builds, not the model's own math.
Show answer
B. Kept as is rate is the leading signal, because it moves as soon as a writer's checking behavior changes, weeks before a funder ever replies.
Fill in the blank
3. Bridgemoor's kept as is rate climbed from thirty five percent in week one to ___ percent by week thirteen.
Show hint
Check the first chart in "Let's learn."
Show answer
91 percent. It kept rising the entire quarter, including the week a made up statistic reached a funder.
Multiple choice
4. What old practice does this answer say to replace?
  • A. Watching grant win rate alone as the only success signal.
  • B. Letting grant writers edit any part of a Fundstream draft.
  • C. Showing funders the guidelines Fundstream was trained on.
  • D. Charging nonprofits by the proposal instead of a flat fee.
Show hint
Think about which number is too slow to run the product by.
Show answer
A. Win rate alone takes about ninety days to move, so a problem would already be shipping for three months before anyone caught it that way.
Short answer, apply it yourself
5. Pick an AI product you use yourself. What's one number that would move weeks before you'd ever notice whether it was actually helping you?
Show hint
Look for a habit that changed quietly, not a result you can only see much later.
Show answer
Model answer: A budgeting app that auto categorizes spending. The real outcome, whether you actually save more over a year, takes months to show. But the share of transactions you leave in its automatic category without correcting them would move within weeks, and it can climb because the app got better at reading receipts, or because you got tired of correcting it and started ignoring bad categories.
True or false
6. True or false: a manager could have fixed the risk in this story just by asking Osahon to edit a few more paragraphs each week.
  • True
  • False
Show hint
Ask whether "edit a bit more" is a real threshold or just a vague dial.
Show answer
False. "Check a bit more" is a mood, not a rule, and it fades the same way the original habit did. The fix needed a real, automatic check with a real threshold attached, not a request to try harder.
Before you close the answer
Why this works
Tests whether a candidate can build a real metric on the spot, for a product they just heard named, instead of reciting "north star metric" and stopping. Most candidates name the outcome and never get to the part that actually gets used.
Follow-up traps
"Isn't kept as is rate the same as engagement, doesn't a rising number always look good on a slide?" Response: it looks good either way, which is exactly the problem. It cannot tell you on its own whether trust was earned or checking just stopped, that is why it needs the substantiation check attached, not read alone.

"If the substantiation check slows Fundstream down, wouldn't Tillwork just ship the faster version anyway to hit a deadline?" Response: that is the trade being made on purpose. A few extra seconds a draft costs far less than a fabricated statistic reaching a funder who has backed Bridgemoor for six years, and the bar is a probability, ninety five traced claims in a hundred, not a demand that it hit one hundred every time.
If pressed
The golden set is not the whole archive of past proposals, only the forty that were both funded and hand checked line by line against their source reports. Tillwork retires the oldest ones each quarter, so the check does not quietly start testing against a funder style nobody writes in anymore.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more