Critique a dashboard that shows only cumulative usage.
Critique a dashboard built around cumulative usage, and say what you would put next to it before you'd trust it again.
- Swap the headline tile from a running total to weekly active conversations, split by signup cohort.Why: a number that can only go up, by its own arithmetic, will never show a decline, however real the decline is.
- Before trusting any drop, check the event log for a definition change or a double-fired event.Why: a broken tracker produces a chart that looks exactly like a real decline, and chasing the wrong cause wastes the weeks you don't have.
- Cut weekly active by whether the account existed before or after the change you suspect.Why: this is the one check that tells a model regression apart from a weaker crop of new sign-ups, and the two need completely different fixes.
- Gate any change to the AI partner's correction style on a cohort's week-two return rate, not just its offline eval score, before it reaches everyone.Why: the model that scored better in the lab was the one that quietly emptied the app.
- Keep the cumulative total somewhere. Just not as the number anyone reacts to each morning.Why: it's a fine headline for a pitch deck. It's a dangerous headline for a team trying to catch a problem this week.
How to answer this, stage by stage
Nobody is grading whether you can spot that a total only goes up. They're grading whether you can turn that observation into the specific cut of data that would have caught this eight weeks sooner. Seven moves get you there.
Let's learn
What happens when the one number your team checks every morning is built so it can never say the product had a bad week?
Parlence is an app that pairs you with an AI conversation partner to practice speaking a language out loud, the way you'd practice with a patient friend instead of a textbook. You open it, pick a topic, and talk. The AI listens, answers back, and corrects you when you slip.
Zephrine Kanku runs growth analytics at Parlence and built the company's main dashboard two years ago, back when the whole product was a few thousand people. The tile everyone glances at first thing each morning is one big number: total conversations completed, all time. By week twenty it read past two million, and it had never once gone down. It couldn't. Every conversation anyone had ever finished just got added to it.
Here is the turn. Say it plainly: that climbing line was never lying, and it was never going to tell the truth either. It was doing exactly what a running total does, adding whatever came in on top of whatever came before. It has no way to subtract. Underneath it, the number of people who actually opened the app and finished a conversation that week rose from about 3,000 in week one to roughly 9,400 by week eight, then fell for eight straight weeks, down to about 6,100 by week sixteen. The big tile on the dashboard never blinked.
At its worst, this looks like nothing at all for a long time. Nobody at Parlence saw a crisis. They saw a number that kept climbing, a product that seemed to be working, and no reason to look closer. The real cost sat one layer down, in a weekly active count that nobody had ever put on a screen.
What I would leave alone: the cumulative total itself isn't a bad number, it's a bad health number. "Parlence has now hosted two million conversations" is a perfectly good line for a pitch deck or an app store page. Leave it there. The mistake was making it the only number the team reacts to at 9am, not putting it on a slide once a quarter.
The lesson: a total answers "how much has ever happened." A team trying to catch a problem needs the answer to a completely different question: "is this week worse than last week." Building one dashboard and expecting it to answer both is how a real decline hides in plain sight for two months.
Now here is the same thing as a story
The short version sits above. Read on for the Tuesday a board deck forced someone to build the chart that should have existed from day one.
Zephrine had been watching that total climb for two years and had never once doubted it. Neither had anyone else. Every all-hands opened with the same slide: total conversations, up and to the right, a number so reliable it had become the company's unofficial mascot.
The trouble started quietly in week eight, when marketing landed a big influencer partnership and a wave of new downloads hit the app. Everyone was thrilled. The total, already unable to do anything but climb, climbed a little faster. Nobody thought to ask whether these new people would still be around in a month.
Then, in week nine, the AI team shipped an upgrade to the conversation partner's model. The old version let small grammar slips pass and corrected the big ones at the end of a turn. The new version was tuned to jump in mid-sentence and fix things as they happened, and it tested better on every offline benchmark the team had: correction accuracy went up, missed errors went down. It shipped as a quality win, and on paper, it was one.
The habit that had quietly held Parlence together started thinning after that, in three beats nobody noticed at the time. First, a slightly higher share of new users who tried one conversation never opened the app for a second one, up from about one in three to closer to one in two, but nobody was watching that number, because nobody had built a tile for it. Second, some of Parlence's oldest, most loyal users, the ones who'd been practicing daily for months, started skipping days here and there, then skipping weeks. Third, the daily streak reminder, cut that same week as part of an unrelated notification cleanup meant to reduce clutter, stopped nudging anyone who'd already started drifting.
The trigger wasn't a crisis. It was a request. In week seventeen, Parlence's leadership needed a retention slide for a board deck ahead of a funding round, and asked Zephrine to pull one together. She went looking for a weekly active chart and realized, for the first time in two years, that one didn't exist. She built it from raw logs that weekend.
What she found: weekly active conversations had fallen from about 9,400 in week eight to about 6,100 by week sixteen, a drop of roughly a third, while the cumulative total on the wall dashboard had added another half a million conversations in that same stretch. Nobody had lied. Nobody had hidden anything. The dashboard simply wasn't built to notice.
Before believing any of it, Zephrine checked the logging first, out of habit. The "conversation completed" event's definition hadn't changed around week nine, and there was no sign it was firing twice. The numbers were real. That ruled out the easiest, laziest explanation, and meant the real work was still ahead of her.
She named three candidates. Maybe the new correction style was scaring off first-time users. Maybe the influencer wave had simply brought in a crowd who were never going to stick around, no matter what the app did. Maybe the missing streak reminder had quietly stopped catching people who were about to lapse anyway. All three were plausible. Only one evidence check would tell them apart.
She split weekly active users into two groups: people who'd signed up before week nine, and people who signed up after. If the funnel theory was the whole story, only the new group should be falling; the old, established users should be humming along like always. Instead, both groups fell, starting the same week. Users who'd signed up back in week two, who had been opening Parlence four times a week for months, were suddenly down to twice. That pattern doesn't come from a weaker crop of new sign-ups. It comes from something that changed for everyone at once, on the same date the new model shipped.
The old decision that set this up went back two years, to Parlence's earliest weeks, when weekly counts were so small they jumped around by fifty percent for no reason at all. A running total, at that size, was the calmer, more honest-feeling number. Nobody wrote that choice down as temporary. Nobody ever came back to ask whether it still fit a product with two million conversations behind it.
Run the same seventeen weeks through a dashboard built the way I'd build it. A weekly active tile, cut by cohort, sits right next to the total from day one. By week eleven, two weeks after the model shipped, the old-user return rate has already dropped from 61 percent toward the mid-40s, and it's dropping in a cohort that has nothing to do with new signups. That alone is enough to open an investigation two weeks after the change instead of eight. The correction style gets dialed back for a smaller, careful rollout, tested against week-two return rate, not just offline grammar accuracy. Six weeks of quiet churn become two.
One dashboard could only ever tell a good story. The other could tell a true one.
What I would tell myself, back when that first dashboard went up: the day you build a metric that can only move one direction, write down that it can never warn you. Someday it won't, and nothing on the screen will say so.
TRACE, and what it caught that the total missed
This is a diagnosis wearing a design question's clothes. Something changed and a real problem is hiding behind a healthy-looking number, so TRACE runs the investigation the dashboard itself should have been built to run.
Three things worth naming directly, since this is where the real judgment sits. The rejected alternative was leaving the total as the headline tile and adding retention as a smaller chart underneath it. That got ruled out on purpose: as long as the biggest number on the screen is one that can't go down, it's still the number a busy team glances at each morning, and the small chart underneath it becomes wallpaper nobody opens. The AI-specific failure worth naming by name is a silent quality regression: a model version that scores better on an offline eval, grammar-correction accuracy in this case, can score worse on the thing a user actually feels, the tone and frequency of being interrupted mid-sentence. Offline evals and real user behavior are not the same test, and treating a better eval score as proof of a better product is exactly how this kind of drop gets shipped by mistake. The guardrail is simple to state and easy to skip under deadline pressure: gate any change to the AI partner's correction behavior on a small cohort's week-two return rate, checked against that cohort's own trailing baseline, before it reaches everyone, not on the eval score alone. And there's a real trade-off underneath the fix, not a free lunch: a lighter, less aggressive correction style is cheaper to run and less likely to interrupt someone, but it also catches fewer real mistakes, so Parlence is trading a small amount of correction quality for a large amount of retention, on purpose, with eyes open. A single bad week on that cohort metric shouldn't trigger a panic either; the rule that held was a return-rate drop of more than eight points against the cohort's own trailing average, sustained for two straight weeks, not one noisy day.
And if you want to be sure it really works, try it somewhere else
Same five letters, a city permit office instead of a language app, and the total-that-only-climbs problem shows up again in a place with no AI conversation partner in sight.
Fondriel County runs PermitPilot, an AI assistant that reads a homeowner's renovation description and pre-fills most of a building permit application before a human reviewer ever opens it. Havelin Iyassu runs digital services for the county's permit office.
T, timeline. PermitPilot's underlying model was upgraded in month four to catch more code violations automatically, flagging them for the applicant to fix before submission, up from a lighter model that let more borderline cases through to a human reviewer. The county's dashboard shows one number: total permits pre-filled, all time. It never stopped climbing.
R, recut. Weekly completed applications, the share of started applications that actually reach submission, instead of the all-time count. That number fell from 71 percent to 44 percent over eight weeks, hidden entirely behind a total that kept adding every started, unfinished draft to its count.
A, assume nothing. Havelin's team checked first whether "application started" had started counting page-refreshes as new starts. It hadn't; the drop was real applicants, not a counting bug.
C, cause candidates. The stricter model flagging so many code violations that first-time applicants gave up rather than fix them. A city-wide contractor licensing change bringing in a wave of unfamiliar filers who were always going to need more help. A change to the confirmation email that quietly dropped the "resume your application" link.
E, evidence test. Completion rate, split by applicants who'd successfully filed a permit with the county before versus first-timers. Repeat filers, who knew exactly what they were doing, also dropped, from 84 percent completion to 57 percent, which pointed straight at the stricter model rather than at unfamiliar new filers.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever the running total shows, put the weekly active or weekly completed number right next to it, cut by cohort, every single day.
Cost: engineering says a real cohort pipeline is six weeks out. Don't read the total alone as a stopgap in the meantime, pull the cohort cut from raw logs by hand once a week until the real pipeline ships.
The model got better, for real: say the correction accuracy or violation-catch rate genuinely improved. That's still not the same claim as "users are staying." A model that gets more technically correct can get worse at keeping people in the room, and the total will keep climbing through the whole thing either way.
Where people run it wrong.
They add the weekly-active chart, then leave it as a small tile nobody's actually on the hook for reading each morning.
They see the total keep climbing and call that proof the change was fine, instead of checking whether it's climbing only because new signups keep padding it.
They fix a bad rollout with an apology blog post instead of a smaller test cohort and a return-rate gate before the next model change ships.
How to use it live. Open with the reframe before naming a single fix: "a total can only tell you how much has ever happened, never whether this week is worse than last week." That buys you the room to give the real diagnosis instead of reciting "add more metrics" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"How do you know it wasn't just the new, casual signups dragging the average down, with old users fine the whole time?" Response: that's exactly what the cohort cut ruled out. Old users, who had nothing to do with the signup wave, also fell, from 61 percent return to 38 percent, in the same week the model shipped.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #5 What early signal predicts churn from an AI feature?
- #6 How do you build an early warning system for silent quality degradation?