ConceptIntermediateQuality, Cost & Token Economics / Success metrics for AI products / #18

What metrics matter for an AI feature in its first two weeks versus its first two quarters?

The direct answer
In week one and two, check whether the tool is telling the truth: run every new role's generated walkthrough against a golden set of real product screens, and watch how many people quietly stall inside it. Only trust the two-quarter numbers, ninety-day retention and account growth, once that correctness check has cleared, because a wrong step teaches a bad habit the day it ships, and a slow quarter is far easier to fix than a habit nobody caught in time.
Do this, in order
  1. Check the generated walkthrough against a golden set of real screens before it reaches a new role.Why: hardest to undo if skipped. A wrong step teaches a bad habit the moment someone signs up, not months later.
  2. Watch the walkthrough's own stall rate, by role, across the first two weeks.Why: a rising stall rate is the leading sign a step stopped matching the real screen, long before anyone files a ticket about it.
  3. Track time to a real action taken outside the tour, still inside week one.Why: shows the walkthrough built a working habit, not just that someone clicked through it.
  4. Do not read the quarter's retention or expansion number as proof of anything until the week-one check has cleared.Why: a quarter number built on top of a broken walkthrough is just a slower way to learn what week one could have shown for free.
  5. Track ninety-day retention and account growth tied to the onboarding path, once correctness is confirmed.Why: this is the real business outcome, but it only means something once the number underneath it can be trusted.
  6. Track new-user support tickets across the quarter, by role, as a hedge.Why: catches whatever the golden set didn't think to test for.

How to answer this, stage by stage

Nobody is grading whether you can name six metrics. They are grading whether you can say, out loud, why one of them has to be checked before the others are allowed to mean anything. Six moves get you there.

1Scope it to one real feature, one real rollout
Say it like this
"Let's ground this. Halvern makes Stepwise, a tool that reads the role a new user picked at signup, ops manager, dispatcher, field tech, and builds them a walkthrough on the spot. Ludgate is a scheduling tool for field service crews, one of Halvern's clients, and Kavi Ilangovan runs product for Stepwise."
Why this works
Grounds the answer in a real product before naming a single metric.
2Name the outcome every metric is competing to serve
Say it like this
"Here's the outcome I'm ranking against: a feature that's both correct and still in real use two quarters from now, not one that just launched clean. Every metric on my list has to earn its spot by how well it protects that."
Why this works
This is the O step. Without a named outcome, a ranked list of metrics is just an opinion wearing numbers.
3Rank by what's hardest to undo, not by when the number arrives
Say it like this
"A silently broken step in week one is the worst thing on this list, because it's teaching someone the wrong habit right now, every single signup. A slow adoption curve in quarter two just needs more runway and a better nudge. Those aren't the same size of problem, so they don't get the same place in line."
Why this works
This is the R step, reversibility. It's the actual reason correctness beats any usage number, not habit or convention.
4Show what depends on what
Say it like this
"You can't trust a ninety-day retention number if nobody checked whether week one's walkthrough was telling the truth. If the dispatcher tour pointed at a button that had already been removed, that quarter number isn't measuring the feature. It's measuring how well people worked around a broken one."
Why this works
This is the D step. It stops a good-looking lagging number from getting credit it hasn't earned.
5Name the cheap two-week test that de-risks the whole quarter
Say it like this
"Before I'd wait ninety days to find out, I'd run every generated walkthrough against a golden set, forty real screens across the client's roles, and check how many steps still point at something real. That test runs in an afternoon, not a quarter, and it catches the exact failure a retention dip would only show me three months late."
Why this works
This is the E step. A cheap eval-based check beats a slow, lagging number every time the two would answer the same question.
6State the ranked call and defend the top pick
Say it like this
"So, in order: correctness against the golden set first, walkthrough stall rate second, both inside week one. Retention and account growth come in the quarters after, and only once I trust what's underneath them. I'd defend correctness as the top pick because it's the only thing on this list that, if wrong, actively teaches the wrong habit instead of just growing slowly."
Why this works
This is the closing R, rank. Ending on a defended top pick turns a list into a decision.
If you remember one thing A quarter number can only be as honest as the week it was built on top of. Check the week before you trust the quarter.

Let's learn

Before Stepwise existed, a product team wrote each new-user walkthrough by hand. Stepwise is the tool that replaced that: it reads the job a new user picked at signup and builds them a short tour on the spot, pointing right at the buttons that job needs first. Under the old, hand-written way, one tour took about six hours to build, and almost nobody went back to fix it once the product changed underneath it. Only four roles out of ten at a typical client ever got a real one. A new user in one of the other six got a generic tour built for nobody in particular, and twenty two out of every hundred of them opened a support ticket in their first week just to ask where something was.

With Stepwise generating a fresh tour the second someone signs up instead, Ludgate's newest signups told a different story fast: the support ticket rate fell from twenty two percent down to nine percent inside the first two weeks. The tour's own completion number looked good too, most people clicked all the way through it.

Two weeks, one number that fell, one that never got built
22% 9% 22% 9% Tickets, week 0 Tickets, week 2
This is the chart Ludgate's leadership actually saw. There was no matching chart for how many generated steps still pointed at a real button, because nobody had built that check yet.

Here is the part that matters. Three weeks before Stepwise shipped, Ludgate had quietly redesigned its dispatch screen and removed a button called Bulk Assign. Nobody told the model that built the tour. For the dispatcher role, Stepwise kept pointing new users at a button that no longer existed. A new dispatcher would tap where the tour said to tap, find nothing there, shrug, and move on to something else on the screen.

That is not a support ticket. It does not show up as an unfinished tour. It shows up nowhere in week one at all.
Knowledge spark: why does a generated tour go stale? Stepwise builds its walkthroughs from a snapshot of what each screen used to look like, not by checking the live product every single time. If the product changes and the snapshot doesn't, the tour keeps sounding just as confident about a button that is already gone.

At its worst, that dispatcher never learns the shortcut that would have saved them forty minutes a day, works the slow way for months without knowing a faster way exists, and rates the product as fine because they never found out otherwise. Multiply that by every dispatcher who signs up before anyone checks the tour against the real screen, and the walkthrough ends up teaching a whole role the wrong habit before quarter one's retention number ever moves enough for anyone to notice.

Hand sketched flow diagram titled What has to be true before the next number counts. Five boxes connected by arrows, reading Golden set, Stall rate, Week one, Retention, Expansion. Golden set is circled in orange as the step everything after it depends on.
Retention and expansion sit at the end of this chain, not the front of it. Neither one is honest until the box before it has actually been checked.

The choice I would take back is not building a nightly check that compares every walkthrough's target screen against the live product. We built Stepwise's tour generator, shipped it, and left the checking for whenever someone happened to notice. That worked fine while Ludgate's screens held still. It stopped working the day dispatch changed.

What I would leave alone: the field tech role's tour never touched a screen that changed, and it never needed the same level of checking. Not every role carries the same risk of going stale, and checking all of them on the same schedule would have just slowed the whole rollout down for no reason.

The lesson: a metric that only tells you people finished the tour will never tell you the tour was true. You have to go check that on purpose, before you let yourself trust anything that comes after it.

Now here is the same thing as a story

The short version sits above. Read on for how ordinary the week this nearly went unnoticed looked from inside Halvern.

Every Stepwise tour looks the same to a new user: five small dots along the top of the screen, one per step, waiting to be tapped through before anyone has met a real person at the company. Kavi Ilangovan runs product for it, and before every release, she walks each new role's tour herself, start to finish, the way a brand new hire would, clicking exactly where the tour says to click instead of trusting a summary of it. It is slow, and for the first year it is the reason nothing shipped broken.

Stepwise's rollout to Ludgate went well by every number Halvern had. Ticket rate down. Tour completion at eighty seven percent. Three weeks in, Kavi got asked to present the numbers at the quarterly review.

By then Halvern had a dozen clients live, and Kavi's personal walkthrough had quietly become a spot check instead of a habit. She still opened a handful of tours a month, the biggest roles by volume, and skipped the smaller ones, because the dashboard for every client already looked healthy and there were only so many hours in a sprint.

The dashboard never once said dispatcher. It just kept saying healthy, for everyone, on average.

A junior customer success manager mentioned it almost as an aside on a call: a dispatcher had asked her over chat why the tour told him to click a button that wasn't there anymore. Kavi almost let it pass. It wasn't a ticket, it wasn't in the dashboard, it was one line in a call recap nobody else would have flagged.

What stopped her was the same habit from before it thinned: watch a real session instead of trusting a summary of it. She pulled the last three weeks of dispatcher sign-ups, about six hours of recordings, skimmed at double speed. Four out of nine dispatchers hit the exact same dead click, in the exact same spot, at the exact same step.

Hand sketched comparison titled Reversible or not, side by side. Left, an orange box labelled Broken tour step, captioned teaches the wrong habit the moment it ships, no way to call back that first impression. Right, a green gauge labelled Slow adoption curve, captioned just needs more runway and a better nudge, comes back on its own.
One of these two problems fixes itself with time. The other one has already happened to four real dispatchers by the time anyone checks.

Two years earlier, when Stepwise's checking pipeline was first designed, the meeting had been short. Building a nightly comparison between every tour and its live screen felt like slowing a launch down for a problem nobody had seen yet. No client had ever redesigned around a tour before. Why would that meeting picture one doing it three weeks after go-live.

What Kavi actually did: she built the golden set, forty real screens pulled fresh from Ludgate's live product, and a nightly job that compares every tour's target element against it. A tour clears the gate when at least ninety percent of its steps check out clean, on a rolling basis, checked again every night rather than trusted once at launch.

Two weeks later, Ludgate touched the field tech onboarding screen in a smaller redesign. The nightly check flagged the mismatch at two in the morning and froze that role's tour on its own. A person on Kavi's team fixed the step by nine, before a single new field tech had signed up against the broken version. Ludgate's dispatcher stall rate, tracked from that point on, fell from eighteen percent back to six within four weeks, once the tour stopped sending anyone toward a button that wasn't there.

What I would tell myself, back before any of this: a number that only measures whether people finished was never going to warn you about what they finished into. You have to go build the number that checks the thing itself.

Ranking it: ORDER, letter by letter

This is a prioritization question, what goes first inside two weeks versus two quarters, so ORDER fits. Not a single north-star metric, and not a risk question about who gets hurt.

O
Outcome. What every candidate metric is competing to serve.
A feature that is both correct and still in real use two quarters out, not one that simply launched without incident.
In this story: correct dispatcher tours that stay trusted well past week two.
R
Reversibility. Which signal is hardest to undo if it's ignored.
A silently wrong step in week one, versus a slow adoption curve in quarter two. The first actively teaches a bad habit right now. The second just needs more runway.
Four dispatchers had already built the wrong habit before anyone checked.
D
Dependency. What has to be true before the next number means anything.
Ninety-day retention can't be trusted until the golden-set check has cleared. A good quarter built on a broken tour is measuring the workaround, not the feature.
Retention and expansion sit at the end of the chain, not the front of it.
E
Evidence. What you could learn cheaply before committing to the long bet.
Run the tour against a golden set of forty real screens, an afternoon's work, instead of waiting three months for a retention dip to say the same thing.
The cheap test would have caught the dead dispatcher click before a single real signup hit it.
R
Rank. State the order, and defend the top pick.
Correctness first, stall rate second, both inside week one. Retention and expansion after, once week one is trusted.
Correctness wins the top spot because it is the only one that actively teaches something wrong if it fails.

Two things worth naming directly, since this is where the AI-specific judgment actually lives. The alternative worth rejecting on purpose is running Stepwise's rollout the way you'd run any ordinary feature launch, and simply watching signup-to-week-four retention like a normal funnel metric. That gets ruled out because retention blends a hundred other things happening in the product at once, and a wrong tooltip doesn't show up in it as a wrong tooltip. It shows up as a slightly worse dispatcher cohort you'd need months of digging to trace back to one broken step. The failure mode worth naming by name is a model generating a confident, specific instruction against a screen that has already moved out from under it, a form of distribution drift most teams only watch for in a model's answers, never in its instructions. The guardrail is the nightly live-screen diff plus the golden set, not a person eyeballing every new tour by hand. The bar itself is calibrated, not absolute: a tour clears the gate at ninety percent of steps verified against the golden set, on a rolling weekly check, not a promise that every step will be perfect forever. And the trade is real. Checking every tour against a live screen on every single generation would catch drift instantly, but it adds real latency and a live lookup cost to every signup, at every client, all day. Halvern's trade is a snapshot refreshed nightly instead of checked live, accepting up to a day of possible staleness in exchange for a walkthrough that still builds itself in under a second.

And if you want to be sure it really works, try it somewhere else

Same five letters, a vet clinic instead of a scheduling tool, so the method proves itself instead of repeating a story I happened to prepare.

Cordway makes Scribevet, a tool that listens to a vet's conversation with a pet owner and drafts the clinical note automatically, including drug names and dosages. Selby Torrance runs product for it.

O, outcome. Notes accurate enough to sit in a permanent medical record, drafted by a tool vets actually keep using instead of quietly going back to their own shorthand.
R, reversibility. A wrong dosage transcribed into a record in week one is nearly impossible to undo once it's fed into a later treatment decision. A slow adoption curve among vets just needs more demos and a better first draft.
D, dependency. Quarter two's hours-saved-per-vet number means nothing if week one's drug-name and dosage accuracy was never checked against real, known-correct notes.
E, evidence. Run Scribevet against two hundred already-transcribed real visits with a known correct answer, a week's work, instead of waiting a quarter for a complaint to surface the same gap.
R, rank. Drug-name and dosage accuracy against the golden set first, vet edit rate second, both inside week one. Hours saved and clinic renewal rate come after, once the first two hold.

Golden-set accuracy by note field, first two weeks
99% 95% 0% 99% 91% Visit summary Drug and dosage
The visit summary field looked strong enough to ship on its own. The one field that actually carries risk, drug and dosage, was the one sitting nine points under the bar, and it would have stayed hidden inside a single blended accuracy score.
Same shape, different stakes At Halvern, the unwatched cost was a dispatcher tapping a dead button. At Cordway, it is a dosage nobody double-checked before it became part of a record. The rank stays the same either way: check the thing that's hardest to undo before you let yourself read the quarter number.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pairing, whatever the two-quarter goal is, name the cheap week-one check that would catch it lying, and rank that check first.
Cost: engineering says the nightly diff pipeline can't ship for two months, not two weeks. Don't roll the feature out to new roles or clinics faster than the manual review can keep up with in the meantime, slow the rollout, not the check.
The model got better, for real: say Stepwise's generator gets meaningfully better at reading a live screen directly instead of a cached one. That still is not proof the two-quarter retention number can be trusted alone. A better model just finds new confident-but-wrong shapes faster than a team updates its golden set to catch them.

Where people run it wrong.
They treat a healthy quarter-two number as proof week one was fine, instead of asking whether week one was ever actually checked.
They rank every metric by how good it looks on a dashboard, instead of by what a wrong number would silently break.
They wait for a quarter of lagging data to tell them what a cheap two-week eval could have told them for the cost of an afternoon.

How to use it live. Say the ranking rule out loud before naming a single metric: "I always ask which number I could get wrong for a whole quarter without anyone noticing, and that's the one I put first." That buys you room to give the real answer, instead of reciting a north-star metric on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits ranking week-one metrics against quarter-two metrics, and why?
Tap to flip
ANSWER
ORDER. It's a prioritization question, rank by what's hardest to undo, not a single metric question about one north star.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Kavi Ilangovan, product manager at Halvern, who built Stepwise's rollout checks for its client Ludgate.
3 · THE HABIT
What did Kavi stop doing because it worked, and what did that cost her?
Tap to flip
ANSWER
Personally walking every new role's tour herself. Once a dozen clients were live she narrowed it to a spot check on the biggest roles, and the dispatcher problem slipped through that gap.
4 · THE TWO SETTINGS
What's the switch in this story, and why is there no safe middle setting?
Tap to flip
ANSWER
Trusting a generated tour because the dashboard looks fine, versus checking it against the live screen before it reaches a role. There's no safe middle once a client's screens can change without warning anyone.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Not building a nightly check comparing every tour's target screen against the live product. It made sense before any client had ever redesigned around a tour.
6 · THE NUMBER
Fill in the blank: Ludgate's new-user ticket rate fell from 22 percent to ___ percent in two weeks, but that number never caught the dispatcher step pointing at a button that no longer existed.
Tap to flip
ANSWER
9 percent. A healthy ticket rate and a broken step lived in the dashboard at the same time.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The nightly diff catches the field tech screen change at 2am and freezes that tour on its own. A person fixes it by 9am, before a single new field tech ever sees the broken step.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its week-one check?
Tap to flip
ANSWER
Scribevet, Cordway's AI clinical-note tool for vet clinics. Its week-one check is drug-name and dosage accuracy against a golden set of two hundred already-transcribed real visits.

Check yourself Score: 0 / 0

Multiple choice
1. In this answer, what gets checked first inside week one, and why does it rank above every quarter-two number?
  • A. The quarter's retention number, because it's the ultimate business goal.
  • B. Whether the generated tour points at real, current buttons, checked against a golden set of screens.
  • C. How many clients have adopted Stepwise so far.
  • D. The average time a new user spends inside the tour.
Show hint
Look for the one thing that, if wrong, actively teaches a bad habit instead of just growing slowly.
Show answer
B. Correctness against the golden set is checked first because it's the hardest thing to undo if it's skipped, not because it's the easiest number to get.
Fill in the blank
2. Ludgate's new-user ticket rate fell from 22 percent to ___ percent in the first two weeks after Stepwise shipped, while the dispatcher tour was still quietly pointing at a button that no longer existed.
Show hint
Check the chart in "Let's learn."
Show answer
9 percent. A number that looked completely healthy the entire time a real step was broken underneath it.
True or false
3. True or false: the field tech tour needed the same nightly correctness check as the dispatcher tour, on the same schedule.
  • True
  • False
Show hint
Check "what I would leave alone" in "Let's learn."
Show answer
False. The field tech tour never touched a screen that changed, so checking it as hard and as often as the dispatcher tour would have slowed the rollout for no real gain.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the meeting memory, not a setting anyone could just turn up.
Show answer
Model answer: Not building a nightly job to compare every generated tour against the live product screen. It made sense at the time because no client had ever redesigned around a tour before, so the risk felt hypothetical rather than real.
Short answer, apply it yourself
5. Think of an AI feature you've used that had a strong-looking launch number. What would the cheap, two-week correctness check have been, the one that could have been run before anyone waited on a slower number to prove the same thing?
Show hint
Look for something you could test against a fixed set of known-correct examples, not something you'd have to wait months to observe.
Show answer
Model answer: A resume-screening tool that reported a strong "time to shortlist" number in week one. The cheap check would have been running it against fifty resumes with a known, human-agreed correct shortlist, before trusting a quarter of hiring-manager satisfaction scores to say the same thing more slowly.
Multiple choice
6. Why can't the team simply wait for the quarter-two retention number to reveal that the dispatcher tour was wrong?
  • A. Because retention numbers take too many engineers to calculate.
  • B. Because by the time retention moves, the wrong tour has already onboarded every dispatcher in that window, and that first impression can't be won back.
  • C. Because retention numbers are always inaccurate no matter what.
  • D. Because quarter numbers only matter to the sales team, not the product team.
Show hint
Think about what reversibility actually means for a metric.
Show answer
B. The whole reason correctness gets checked in week one is that a wrong number in quarter two arrives too late to undo the damage a wrong step already did.
Before you close the answer
Why this works
Tests whether you'll rank metrics by what a wrong one would silently break, or just list every number a dashboard could show. Most candidates name the same six metrics without ever saying which one has to be trusted first.
Follow-up traps
"Isn't checking every tour against a golden set going to slow down every new client rollout?" Response: no, because the check runs in an afternoon against forty screens, not a manual review of every session. It only holds up a role when it actually fails the bar.

"What if the golden set itself goes stale, the same way the tour did?" Response: that's exactly why the check runs nightly against the live product instead of once at launch. A golden set that's never refreshed would fail for the same reason the original tour did.
If pressed
The golden set isn't static either. Halvern rotates ten of its forty screens every month, pulled fresh from whichever client had the most support tickets that period, so the check keeps testing against real drift instead of the same forty screens forever.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more