Artifact critiqueIntermediateQuality, Cost & Token Economics / Leading vs lagging indicators for AI / #22

Critique a dashboard that shows only cumulative usage.

Critique a dashboard built around cumulative usage, and say what you would put next to it before you'd trust it again.

The direct answer
A cumulative total is a bad health tile because it cannot fall. Replace it as the number the team reacts to with weekly active conversations, cut by signup cohort, so the dashboard can show the week the product actually started losing people, not just the week someone finally went looking for it.
Do this, in order
  1. Swap the headline tile from a running total to weekly active conversations, split by signup cohort.Why: a number that can only go up, by its own arithmetic, will never show a decline, however real the decline is.
  2. Before trusting any drop, check the event log for a definition change or a double-fired event.Why: a broken tracker produces a chart that looks exactly like a real decline, and chasing the wrong cause wastes the weeks you don't have.
  3. Cut weekly active by whether the account existed before or after the change you suspect.Why: this is the one check that tells a model regression apart from a weaker crop of new sign-ups, and the two need completely different fixes.
  4. Gate any change to the AI partner's correction style on a cohort's week-two return rate, not just its offline eval score, before it reaches everyone.Why: the model that scored better in the lab was the one that quietly emptied the app.
  5. Keep the cumulative total somewhere. Just not as the number anyone reacts to each morning.Why: it's a fine headline for a pitch deck. It's a dangerous headline for a team trying to catch a problem this week.

How to answer this, stage by stage

Nobody is grading whether you can spot that a total only goes up. They're grading whether you can turn that observation into the specific cut of data that would have caught this eight weeks sooner. Seven moves get you there.

1
Ground it in one real product
Say it like this
"Let's make this concrete. Parlence is an app where you practice speaking a new language with an AI conversation partner. Zephrine Kanku is the analyst who owns the growth dashboard there."
Why this works
A critique of an abstract dashboard is a lecture. A critique of one real tile is a decision.
2
Say what's actually wrong with the tile, in one line
Say it like this
"The problem isn't that the number is wrong. It's that a cumulative total is arithmetically unable to fall. If you build your whole health check around a number that can only go up, you've built a check that can never tell you no."
Why this works
This is the reframe the whole critique turns on. Skip it and you're just saying "add more metrics," which answers nothing.
3
Name your structure out loud
Say it like this
"I'd run this through TRACE. Find when it actually started, cut the total into something that can move, rule out a broken tracker, name a few real causes, then find the one check that tells them apart."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion about dashboards.
4
Rule out the boring explanation first
Say it like this
"Before I believe a single thing about this drop, I check the logging. Did the 'conversation completed' event change its definition around the same week? Is it firing twice? If the tracker moved, the chart moves too, and it has nothing to do with anyone's behavior."
Why this works
Interviewers plant this on purpose. Skipping it is the fastest way to look like you chase stories instead of evidence.
5
Name real causes, plural, not just one
Say it like this
"Three things could do this. The AI partner started correcting harder, and people try it once and don't come back. A signup wave brought in casual downloaders who were never going to stick around. Or a reminder that used to pull lapsed users back got quietly switched off."
Why this works
A single guess is a hunch. Three named, distinguishable causes is a diagnosis in progress.
6
Give the one check that tells them apart
Say it like this
"Cut the return rate by whether someone signed up before or after the model changed. If only the new sign-ups are falling, it's the funnel. If people who joined months ago and were doing fine also start dropping off the same week, it's the model, not the crowd."
Why this works
This is TRACE's strongest move. It turns "I have some theories" into "here's the one query that decides between them."
7
Close on what you'd trade and what you ruled out
Say it like this
"We talked about just adding retention as a second, smaller tile and leaving the total as the headline. We ruled that out. As long as the biggest number on the screen is the one that can't go down, that's still the one people glance at, and the small tile becomes wallpaper. Swapping which number is the headline costs us a comfortable morning number. We accepted that trade."
Why this works
Naming a rejected option and a real cost is what separates a design decision from a wish that both were free.
If you remember one thing A total can only tell you how much has ever happened. It cannot tell you what is happening right now. Those are different questions, and a dashboard that only answers the first one will let the second one rot for months.

Let's learn

What happens when the one number your team checks every morning is built so it can never say the product had a bad week?

Parlence is an app that pairs you with an AI conversation partner to practice speaking a language out loud, the way you'd practice with a patient friend instead of a textbook. You open it, pick a topic, and talk. The AI listens, answers back, and corrects you when you slip.

Hand sketched timeline titled What shipped and when the real drop started. Three points on a line. First, big signup push, week 8, influencer campaign. Second, in red, new model version, week 9, harsher corrections. Third, board deck request, week 17, first cohort cut ever built.
Two things shipped close together in week eight and nine. Only one of them explains what the numbers later showed.

Zephrine Kanku runs growth analytics at Parlence and built the company's main dashboard two years ago, back when the whole product was a few thousand people. The tile everyone glances at first thing each morning is one big number: total conversations completed, all time. By week twenty it read past two million, and it had never once gone down. It couldn't. Every conversation anyone had ever finished just got added to it.

Total conversations completed, all time, week 1 to week 20
2.2M 0 Wk 1 Wk 10 Wk 15 Wk 20
Forty thousand conversations in week one, past two million by week twenty. Climbing, every single week, the entire time this story happened.

Here is the turn. Say it plainly: that climbing line was never lying, and it was never going to tell the truth either. It was doing exactly what a running total does, adding whatever came in on top of whatever came before. It has no way to subtract. Underneath it, the number of people who actually opened the app and finished a conversation that week rose from about 3,000 in week one to roughly 9,400 by week eight, then fell for eight straight weeks, down to about 6,100 by week sixteen. The big tile on the dashboard never blinked.

The total wasn't broken. It just wasn't built to say no.
Knowledge spark: what is a signup cohort? Everyone who joined in the same stretch of time, grouped together. Cohort week 9 means everyone who signed up during week nine. Cutting a metric by cohort lets you compare an old group of users against a new one, instead of blending them into one average that hides both.

At its worst, this looks like nothing at all for a long time. Nobody at Parlence saw a crisis. They saw a number that kept climbing, a product that seemed to be working, and no reason to look closer. The real cost sat one layer down, in a weekly active count that nobody had ever put on a screen.

The decision that mattered Building the dashboard around one cumulative tile, with no weekly-active view and no history of it, so nobody could ever compare this week to last week. It made sense two years ago, when volume was so low that a weekly count bounced around too much to mean anything. Nobody rebuilt it once that stopped being true.

What I would leave alone: the cumulative total itself isn't a bad number, it's a bad health number. "Parlence has now hosted two million conversations" is a perfectly good line for a pitch deck or an app store page. Leave it there. The mistake was making it the only number the team reacts to at 9am, not putting it on a slide once a quarter.

The lesson: a total answers "how much has ever happened." A team trying to catch a problem needs the answer to a completely different question: "is this week worse than last week." Building one dashboard and expecting it to answer both is how a real decline hides in plain sight for two months.

Now here is the same thing as a story

The short version sits above. Read on for the Tuesday a board deck forced someone to build the chart that should have existed from day one.

Zephrine had been watching that total climb for two years and had never once doubted it. Neither had anyone else. Every all-hands opened with the same slide: total conversations, up and to the right, a number so reliable it had become the company's unofficial mascot.

The trouble started quietly in week eight, when marketing landed a big influencer partnership and a wave of new downloads hit the app. Everyone was thrilled. The total, already unable to do anything but climb, climbed a little faster. Nobody thought to ask whether these new people would still be around in a month.

Then, in week nine, the AI team shipped an upgrade to the conversation partner's model. The old version let small grammar slips pass and corrected the big ones at the end of a turn. The new version was tuned to jump in mid-sentence and fix things as they happened, and it tested better on every offline benchmark the team had: correction accuracy went up, missed errors went down. It shipped as a quality win, and on paper, it was one.

The habit that had quietly held Parlence together started thinning after that, in three beats nobody noticed at the time. First, a slightly higher share of new users who tried one conversation never opened the app for a second one, up from about one in three to closer to one in two, but nobody was watching that number, because nobody had built a tile for it. Second, some of Parlence's oldest, most loyal users, the ones who'd been practicing daily for months, started skipping days here and there, then skipping weeks. Third, the daily streak reminder, cut that same week as part of an unrelated notification cleanup meant to reduce clutter, stopped nudging anyone who'd already started drifting.

Hand sketched list titled Three ways a weekly number falls that a total never shows. One, in red, model corrects harder now, confirmed, old users fell too. Two, new signups were casual, never coming back anyway. Three, streak reminder got cut as clutter in week 9.
Three real candidates. Only one of them explains why users who had been fine for months suddenly weren't.

The trigger wasn't a crisis. It was a request. In week seventeen, Parlence's leadership needed a retention slide for a board deck ahead of a funding round, and asked Zephrine to pull one together. She went looking for a weekly active chart and realized, for the first time in two years, that one didn't exist. She built it from raw logs that weekend.

What she found: weekly active conversations had fallen from about 9,400 in week eight to about 6,100 by week sixteen, a drop of roughly a third, while the cumulative total on the wall dashboard had added another half a million conversations in that same stretch. Nobody had lied. Nobody had hidden anything. The dashboard simply wasn't built to notice.

Nobody was hiding the drop. The dashboard just wasn't built to see it.

Before believing any of it, Zephrine checked the logging first, out of habit. The "conversation completed" event's definition hadn't changed around week nine, and there was no sign it was firing twice. The numbers were real. That ruled out the easiest, laziest explanation, and meant the real work was still ahead of her.

She named three candidates. Maybe the new correction style was scaring off first-time users. Maybe the influencer wave had simply brought in a crowd who were never going to stick around, no matter what the app did. Maybe the missing streak reminder had quietly stopped catching people who were about to lapse anyway. All three were plausible. Only one evidence check would tell them apart.

She split weekly active users into two groups: people who'd signed up before week nine, and people who signed up after. If the funnel theory was the whole story, only the new group should be falling; the old, established users should be humming along like always. Instead, both groups fell, starting the same week. Users who'd signed up back in week two, who had been opening Parlence four times a week for months, were suddenly down to twice. That pattern doesn't come from a weaker crop of new sign-ups. It comes from something that changed for everyone at once, on the same date the new model shipped.

Week-two return rate, by when the account existed
70% 0% 61% 38% 24% Old users, before wk 9 Old users, after wk 9 New users, signed up after
The same old, loyal users returned 61 percent of the time before the model change and 38 percent after. That fall, in people who had nothing to do with the new signup wave, is what pointed at the model.

The old decision that set this up went back two years, to Parlence's earliest weeks, when weekly counts were so small they jumped around by fifty percent for no reason at all. A running total, at that size, was the calmer, more honest-feeling number. Nobody wrote that choice down as temporary. Nobody ever came back to ask whether it still fit a product with two million conversations behind it.

Run the same seventeen weeks through a dashboard built the way I'd build it. A weekly active tile, cut by cohort, sits right next to the total from day one. By week eleven, two weeks after the model shipped, the old-user return rate has already dropped from 61 percent toward the mid-40s, and it's dropping in a cohort that has nothing to do with new signups. That alone is enough to open an investigation two weeks after the change instead of eight. The correction style gets dialed back for a smaller, careful rollout, tested against week-two return rate, not just offline grammar accuracy. Six weeks of quiet churn become two.

One dashboard could only ever tell a good story. The other could tell a true one.

What I would tell myself, back when that first dashboard went up: the day you build a metric that can only move one direction, write down that it can never warn you. Someday it won't, and nothing on the screen will say so.

TRACE, and what it caught that the total missed

This is a diagnosis wearing a design question's clothes. Something changed and a real problem is hiding behind a healthy-looking number, so TRACE runs the investigation the dashboard itself should have been built to run.

T
Timeline. When it actually started.
Two things shipped close together: a big signup push in week eight, and a new, more corrective AI model in week nine. The weekly active count peaked right at week eight and fell for the next eight weeks straight.
In this answer: not "usage dropped in week seventeen." It dropped in week nine. Week seventeen is just when someone finally looked.
R
Recut. Slice it into something that can fall.
Weekly active conversations, split by signup cohort, instead of one all-time total. The total went from 1.6M to 2.1M over the same stretch weekly actives fell from 9,400 to 6,100.
This is the whole reason one cumulative tile fails here. A number that only sums can hide a number that's shrinking underneath it.
A
Assume nothing. Rule out the tracker first.
Before trusting the drop, Zephrine checked whether the "conversation completed" event's definition had changed, or whether it was double-firing, around week nine. It hadn't, and it wasn't. The drop was real, not a logging artifact.
Skip this and you can spend a month fixing a model that was never broken, chasing a bug in a chart that was never real.
C
Cause candidates. Three, named.
The new model's mid-sentence corrections making first-time users quit after one try. A signup wave bringing in casual downloaders who were never going to return. A streak reminder quietly cut the same week, removing the nudge that used to catch lapsing users.
Three plausible stories. Without a test, all three sound equally likely, and the fix for one does nothing for the others.
E
Evidence test. The one cut that decides it.
Week-two return rate, split by whether the account existed before or after week nine. Old users fell from 61 percent to 38 percent, new users returned at only 24 percent. Because old, established users also fell, the drop can't be only about who signed up, it points at what changed for everyone: the model.
This is the strongest move in the whole framework. One query, and two competing stories stop being equally likely.

Three things worth naming directly, since this is where the real judgment sits. The rejected alternative was leaving the total as the headline tile and adding retention as a smaller chart underneath it. That got ruled out on purpose: as long as the biggest number on the screen is one that can't go down, it's still the number a busy team glances at each morning, and the small chart underneath it becomes wallpaper nobody opens. The AI-specific failure worth naming by name is a silent quality regression: a model version that scores better on an offline eval, grammar-correction accuracy in this case, can score worse on the thing a user actually feels, the tone and frequency of being interrupted mid-sentence. Offline evals and real user behavior are not the same test, and treating a better eval score as proof of a better product is exactly how this kind of drop gets shipped by mistake. The guardrail is simple to state and easy to skip under deadline pressure: gate any change to the AI partner's correction behavior on a small cohort's week-two return rate, checked against that cohort's own trailing baseline, before it reaches everyone, not on the eval score alone. And there's a real trade-off underneath the fix, not a free lunch: a lighter, less aggressive correction style is cheaper to run and less likely to interrupt someone, but it also catches fewer real mistakes, so Parlence is trading a small amount of correction quality for a large amount of retention, on purpose, with eyes open. A single bad week on that cohort metric shouldn't trigger a panic either; the rule that held was a return-rate drop of more than eight points against the cohort's own trailing average, sustained for two straight weeks, not one noisy day.

And if you want to be sure it really works, try it somewhere else

Same five letters, a city permit office instead of a language app, and the total-that-only-climbs problem shows up again in a place with no AI conversation partner in sight.

Fondriel County runs PermitPilot, an AI assistant that reads a homeowner's renovation description and pre-fills most of a building permit application before a human reviewer ever opens it. Havelin Iyassu runs digital services for the county's permit office.

T, timeline. PermitPilot's underlying model was upgraded in month four to catch more code violations automatically, flagging them for the applicant to fix before submission, up from a lighter model that let more borderline cases through to a human reviewer. The county's dashboard shows one number: total permits pre-filled, all time. It never stopped climbing.
R, recut. Weekly completed applications, the share of started applications that actually reach submission, instead of the all-time count. That number fell from 71 percent to 44 percent over eight weeks, hidden entirely behind a total that kept adding every started, unfinished draft to its count.
A, assume nothing. Havelin's team checked first whether "application started" had started counting page-refreshes as new starts. It hadn't; the drop was real applicants, not a counting bug.
C, cause candidates. The stricter model flagging so many code violations that first-time applicants gave up rather than fix them. A city-wide contractor licensing change bringing in a wave of unfamiliar filers who were always going to need more help. A change to the confirmation email that quietly dropped the "resume your application" link.
E, evidence test. Completion rate, split by applicants who'd successfully filed a permit with the county before versus first-timers. Repeat filers, who knew exactly what they were doing, also dropped, from 84 percent completion to 57 percent, which pointed straight at the stricter model rather than at unfamiliar new filers.

Application completion rate, repeat filers vs first-timers
90% 0% 84% 57% 44% Repeat filers, before model Repeat filers, after model All applicants, current average
Fondriel's total permits pre-filled looked fine the whole time. Only the repeat-filer completion rate showed the stricter model was the real cause, not a wave of confused new applicants.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever the running total shows, put the weekly active or weekly completed number right next to it, cut by cohort, every single day.
Cost: engineering says a real cohort pipeline is six weeks out. Don't read the total alone as a stopgap in the meantime, pull the cohort cut from raw logs by hand once a week until the real pipeline ships.
The model got better, for real: say the correction accuracy or violation-catch rate genuinely improved. That's still not the same claim as "users are staying." A model that gets more technically correct can get worse at keeping people in the room, and the total will keep climbing through the whole thing either way.

Where people run it wrong.
They add the weekly-active chart, then leave it as a small tile nobody's actually on the hook for reading each morning.
They see the total keep climbing and call that proof the change was fine, instead of checking whether it's climbing only because new signups keep padding it.
They fix a bad rollout with an apology blog post instead of a smaller test cohort and a return-rate gate before the next model change ships.

How to use it live. Open with the reframe before naming a single fix: "a total can only tell you how much has ever happened, never whether this week is worse than last week." That buys you the room to give the real diagnosis instead of reciting "add more metrics" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a "critique this dashboard" diagnosis question like this one?
Tap to flip
ANSWER
TRACE: find the real timeline, recut the number so it can move, assume nothing about the tracker, name cause candidates, run the one evidence test that decides between them.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Zephrine Kanku, who runs growth analytics for Parlence, an AI language-practice conversation app.
3 · WHY THE TOTAL FAILED
Why can't a cumulative total ever act as a health warning?
Tap to flip
ANSWER
A running total can only add. It has no way to subtract, so no matter how badly weekly engagement falls, the total keeps climbing and never signals a problem.
4 · THE RECUT
What did Zephrine replace the total with, and what did it show?
Tap to flip
ANSWER
Weekly active conversations, split by signup cohort. It showed a fall from about 9,400 a week to about 6,100 a week, hidden completely behind a total that kept rising.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building the dashboard around one cumulative tile with no weekly-active history. It made sense two years earlier, when volume was too low for a weekly number to mean anything. Nobody revisited it once volume made the total meaningless as a health check.
6 · THE NUMBER
Fill in the blank: old users' week-two return rate fell from 61 percent to ___ percent after the model changed, while new users returned at only ___ percent.
Tap to flip
ANSWER
38 percent; 24 percent. Old users falling too is what pointed at the model, not just a weaker crop of new sign-ups.
7 · THE REPLAY
Same seventeen weeks, fixed dashboard, what changes?
Tap to flip
ANSWER
The old-user return rate drop is visible by week eleven instead of week seventeen. The correction style gets a smaller, gated rollout, and six weeks of quiet churn become two.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
PermitPilot, an AI permit pre-fill assistant at Fondriel County. Same TRACE steps, a stricter violation-flagging model instead of a stricter correction style, found the same way: repeat filers also dropped, not just newcomers.

Check yourself Score: 0 / 0

Fill in the blank
1. Parlence's weekly active conversations fell from about 9,400 in week eight to about ___ by week sixteen, while the cumulative total kept climbing the whole time.
Show hint
Check the first chart in "Let's learn," and the paragraph right after it.
Show answer
6,100. A drop of roughly a third in real weekly usage, with no way for the cumulative tile to ever show it.
Multiple choice
2. What is the real problem with using a cumulative, all-time total as a dashboard's main health number?
  • A. It's too expensive to compute every day.
  • B. It only works for products with very few users.
  • C. It can only ever go up, so it can never show a real decline happening right now.
  • D. It ignores how long each conversation lasts.
Show hint
Think about what arithmetic operation a running total can and can't do.
Show answer
C. A sum of everything that has ever happened has no way to subtract, so it cannot signal that this week was worse than last week, no matter how bad this week was.
True or false
3. True or false: Zephrine's evidence test showed the drop was limited to new sign-ups who joined after the marketing push, which meant the funnel, not the model, was the cause.
  • True
  • False
Show hint
Check what happened to old users, the ones who'd already been using Parlence for months.
Show answer
False. Old users, unrelated to the signup wave, also dropped, from 61 percent to 38 percent return. That's what pointed at the model change instead of the funnel.
Multiple choice
4. Before trusting that the drop in weekly active conversations was real, what did Zephrine check first?
  • A. Whether competitors had launched a similar app that quarter.
  • B. Whether the "conversation completed" event's definition had changed or was firing twice.
  • C. Whether the app's servers had gone down at any point.
  • D. Whether the marketing team had changed the app's price.
Show hint
This is TRACE's "assume nothing" step, the one that rules out a boring explanation before a real one.
Show answer
B. A changed or double-firing tracking event produces a chart that looks exactly like a real behavior change, so ruling it out has to come before believing anything else.
Short answer, name the reversal
5. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the "decision that mattered" box in "Let's learn."
Show answer
Model answer: Building the dashboard around one cumulative tile, with no weekly-active view kept anywhere. It made sense two years earlier, when Parlence had so few users that a weekly count bounced around meaninglessly. Nobody came back to rebuild it once the product had millions of conversations behind it.
Short answer, apply it yourself
6. Pick an app you use that shows you an all-time total somewhere (total steps, total minutes listened, total orders placed). What real weekly problem could that total be hiding from you right now?
Show hint
Ask what the total adds up over time, and what would have to happen for that weekly slice to quietly fall without the total ever showing it.
Show answer
Model answer: A fitness app showing "total workouts logged: 340." That number can climb forever even if you've only logged one workout in the last six weeks. It hides exactly the thing a person trying to stay consistent would want to know: not how many workouts you've ever done, but whether this week looks like last week.
Before you close the answer
Why this works
Tests whether you understand that a metric's shape, not just its accuracy, decides what it can ever tell you. Most candidates critique a cumulative dashboard by saying "add more metrics." Fewer explain why this specific metric is structurally incapable of the one job a health tile has.
Follow-up traps
"Couldn't the team have just eyeballed the slope of the total instead of building a new tile?" Response: no, because the slope barely changed. The total added roughly the same amount of growth from new signups every week, even as returning users fell, so the slope of a rising line hid a shrinking one underneath it.

"How do you know it wasn't just the new, casual signups dragging the average down, with old users fine the whole time?" Response: that's exactly what the cohort cut ruled out. Old users, who had nothing to do with the signup wave, also fell, from 61 percent return to 38 percent, in the same week the model shipped.
If pressed
The floor that would actually trigger a review isn't a single bad day. It's a cohort's week-two return rate dropping more than eight points against its own trailing average, sustained across two straight weeks, so one noisy Tuesday doesn't set off a false alarm while a real, sustained drop still gets caught inside a fortnight instead of two months.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more