What metrics matter for an AI feature in its first two weeks versus its first two quarters?
- Check the generated walkthrough against a golden set of real screens before it reaches a new role.Why: hardest to undo if skipped. A wrong step teaches a bad habit the moment someone signs up, not months later.
- Watch the walkthrough's own stall rate, by role, across the first two weeks.Why: a rising stall rate is the leading sign a step stopped matching the real screen, long before anyone files a ticket about it.
- Track time to a real action taken outside the tour, still inside week one.Why: shows the walkthrough built a working habit, not just that someone clicked through it.
- Do not read the quarter's retention or expansion number as proof of anything until the week-one check has cleared.Why: a quarter number built on top of a broken walkthrough is just a slower way to learn what week one could have shown for free.
- Track ninety-day retention and account growth tied to the onboarding path, once correctness is confirmed.Why: this is the real business outcome, but it only means something once the number underneath it can be trusted.
- Track new-user support tickets across the quarter, by role, as a hedge.Why: catches whatever the golden set didn't think to test for.
How to answer this, stage by stage
Nobody is grading whether you can name six metrics. They are grading whether you can say, out loud, why one of them has to be checked before the others are allowed to mean anything. Six moves get you there.
Let's learn
Before Stepwise existed, a product team wrote each new-user walkthrough by hand. Stepwise is the tool that replaced that: it reads the job a new user picked at signup and builds them a short tour on the spot, pointing right at the buttons that job needs first. Under the old, hand-written way, one tour took about six hours to build, and almost nobody went back to fix it once the product changed underneath it. Only four roles out of ten at a typical client ever got a real one. A new user in one of the other six got a generic tour built for nobody in particular, and twenty two out of every hundred of them opened a support ticket in their first week just to ask where something was.
With Stepwise generating a fresh tour the second someone signs up instead, Ludgate's newest signups told a different story fast: the support ticket rate fell from twenty two percent down to nine percent inside the first two weeks. The tour's own completion number looked good too, most people clicked all the way through it.
Here is the part that matters. Three weeks before Stepwise shipped, Ludgate had quietly redesigned its dispatch screen and removed a button called Bulk Assign. Nobody told the model that built the tour. For the dispatcher role, Stepwise kept pointing new users at a button that no longer existed. A new dispatcher would tap where the tour said to tap, find nothing there, shrug, and move on to something else on the screen.
At its worst, that dispatcher never learns the shortcut that would have saved them forty minutes a day, works the slow way for months without knowing a faster way exists, and rates the product as fine because they never found out otherwise. Multiply that by every dispatcher who signs up before anyone checks the tour against the real screen, and the walkthrough ends up teaching a whole role the wrong habit before quarter one's retention number ever moves enough for anyone to notice.
The choice I would take back is not building a nightly check that compares every walkthrough's target screen against the live product. We built Stepwise's tour generator, shipped it, and left the checking for whenever someone happened to notice. That worked fine while Ludgate's screens held still. It stopped working the day dispatch changed.
What I would leave alone: the field tech role's tour never touched a screen that changed, and it never needed the same level of checking. Not every role carries the same risk of going stale, and checking all of them on the same schedule would have just slowed the whole rollout down for no reason.
The lesson: a metric that only tells you people finished the tour will never tell you the tour was true. You have to go check that on purpose, before you let yourself trust anything that comes after it.
Now here is the same thing as a story
The short version sits above. Read on for how ordinary the week this nearly went unnoticed looked from inside Halvern.
Every Stepwise tour looks the same to a new user: five small dots along the top of the screen, one per step, waiting to be tapped through before anyone has met a real person at the company. Kavi Ilangovan runs product for it, and before every release, she walks each new role's tour herself, start to finish, the way a brand new hire would, clicking exactly where the tour says to click instead of trusting a summary of it. It is slow, and for the first year it is the reason nothing shipped broken.
Stepwise's rollout to Ludgate went well by every number Halvern had. Ticket rate down. Tour completion at eighty seven percent. Three weeks in, Kavi got asked to present the numbers at the quarterly review.
By then Halvern had a dozen clients live, and Kavi's personal walkthrough had quietly become a spot check instead of a habit. She still opened a handful of tours a month, the biggest roles by volume, and skipped the smaller ones, because the dashboard for every client already looked healthy and there were only so many hours in a sprint.
A junior customer success manager mentioned it almost as an aside on a call: a dispatcher had asked her over chat why the tour told him to click a button that wasn't there anymore. Kavi almost let it pass. It wasn't a ticket, it wasn't in the dashboard, it was one line in a call recap nobody else would have flagged.
What stopped her was the same habit from before it thinned: watch a real session instead of trusting a summary of it. She pulled the last three weeks of dispatcher sign-ups, about six hours of recordings, skimmed at double speed. Four out of nine dispatchers hit the exact same dead click, in the exact same spot, at the exact same step.
Two years earlier, when Stepwise's checking pipeline was first designed, the meeting had been short. Building a nightly comparison between every tour and its live screen felt like slowing a launch down for a problem nobody had seen yet. No client had ever redesigned around a tour before. Why would that meeting picture one doing it three weeks after go-live.
What Kavi actually did: she built the golden set, forty real screens pulled fresh from Ludgate's live product, and a nightly job that compares every tour's target element against it. A tour clears the gate when at least ninety percent of its steps check out clean, on a rolling basis, checked again every night rather than trusted once at launch.
Two weeks later, Ludgate touched the field tech onboarding screen in a smaller redesign. The nightly check flagged the mismatch at two in the morning and froze that role's tour on its own. A person on Kavi's team fixed the step by nine, before a single new field tech had signed up against the broken version. Ludgate's dispatcher stall rate, tracked from that point on, fell from eighteen percent back to six within four weeks, once the tour stopped sending anyone toward a button that wasn't there.
What I would tell myself, back before any of this: a number that only measures whether people finished was never going to warn you about what they finished into. You have to go build the number that checks the thing itself.
Ranking it: ORDER, letter by letter
This is a prioritization question, what goes first inside two weeks versus two quarters, so ORDER fits. Not a single north-star metric, and not a risk question about who gets hurt.
Two things worth naming directly, since this is where the AI-specific judgment actually lives. The alternative worth rejecting on purpose is running Stepwise's rollout the way you'd run any ordinary feature launch, and simply watching signup-to-week-four retention like a normal funnel metric. That gets ruled out because retention blends a hundred other things happening in the product at once, and a wrong tooltip doesn't show up in it as a wrong tooltip. It shows up as a slightly worse dispatcher cohort you'd need months of digging to trace back to one broken step. The failure mode worth naming by name is a model generating a confident, specific instruction against a screen that has already moved out from under it, a form of distribution drift most teams only watch for in a model's answers, never in its instructions. The guardrail is the nightly live-screen diff plus the golden set, not a person eyeballing every new tour by hand. The bar itself is calibrated, not absolute: a tour clears the gate at ninety percent of steps verified against the golden set, on a rolling weekly check, not a promise that every step will be perfect forever. And the trade is real. Checking every tour against a live screen on every single generation would catch drift instantly, but it adds real latency and a live lookup cost to every signup, at every client, all day. Halvern's trade is a snapshot refreshed nightly instead of checked live, accepting up to a day of possible staleness in exchange for a walkthrough that still builds itself in under a second.
And if you want to be sure it really works, try it somewhere else
Same five letters, a vet clinic instead of a scheduling tool, so the method proves itself instead of repeating a story I happened to prepare.
Cordway makes Scribevet, a tool that listens to a vet's conversation with a pet owner and drafts the clinical note automatically, including drug names and dosages. Selby Torrance runs product for it.
O, outcome. Notes accurate enough to sit in a permanent medical record, drafted by a tool vets actually keep using instead of quietly going back to their own shorthand.
R, reversibility. A wrong dosage transcribed into a record in week one is nearly impossible to undo once it's fed into a later treatment decision. A slow adoption curve among vets just needs more demos and a better first draft.
D, dependency. Quarter two's hours-saved-per-vet number means nothing if week one's drug-name and dosage accuracy was never checked against real, known-correct notes.
E, evidence. Run Scribevet against two hundred already-transcribed real visits with a known correct answer, a week's work, instead of waiting a quarter for a complaint to surface the same gap.
R, rank. Drug-name and dosage accuracy against the golden set first, vet edit rate second, both inside week one. Hours saved and clinic renewal rate come after, once the first two hold.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pairing, whatever the two-quarter goal is, name the cheap week-one check that would catch it lying, and rank that check first.
Cost: engineering says the nightly diff pipeline can't ship for two months, not two weeks. Don't roll the feature out to new roles or clinics faster than the manual review can keep up with in the meantime, slow the rollout, not the check.
The model got better, for real: say Stepwise's generator gets meaningfully better at reading a live screen directly instead of a cached one. That still is not proof the two-quarter retention number can be trusted alone. A better model just finds new confident-but-wrong shapes faster than a team updates its golden set to catch them.
Where people run it wrong.
They treat a healthy quarter-two number as proof week one was fine, instead of asking whether week one was ever actually checked.
They rank every metric by how good it looks on a dashboard, instead of by what a wrong number would silently break.
They wait for a quarter of lagging data to tell them what a cheap two-week eval could have told them for the cost of an afternoon.
How to use it live. Say the ranking rule out loud before naming a single metric: "I always ask which number I could get wrong for a whole quarter without anyone noticing, and that's the one I put first." That buys you room to give the real answer, instead of reciting a north-star metric on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the golden set itself goes stale, the same way the tour did?" Response: that's exactly why the check runs nightly against the live product instead of once at launch. A golden set that's never refreshed would fail for the same reason the original tour did.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?