Which leading indicators break when your user mix changes?
Steadyline's early-warning score for a mental health crisis went quiet on one whole group of patients, and nothing on the dashboard said why.
- Rebuild the leading indicator's baseline per cohort, not once for the whole population.Why: a threshold trained on one crowd quietly stops describing any group whose mix has moved away from that crowd.
- Flag a cohort's baseline stale the moment its share of weekly traffic shifts past a set point, and fall back to a wider default cutoff until a fresh one ships.Why: a stale baseline that keeps quietly running is worse than a broken one, because nobody thinks to check it.
- Track the composition of who is signing up as its own daily number, not just the indicator's own output.Why: the mix shift always happens weeks before the indicator visibly breaks, so watching the crowd catches it before watching the outcome does.
- Give partners and referral sources their own cohort's numbers, not just your aggregate dashboard.Why: the people closest to the affected group are the ones most likely to notice a gap first, if they can actually see it.
- Leave any absolute, rule-based safety trigger alone.Why: a fixed rule doesn't drift with the population the way a relative score does, so it needs no cohort baseline at all.
- Validate every fresh cohort baseline against a held-out sample before trusting it, never against the same check-ins it's about to score.Why: a baseline checked on the data it will later grade always looks calibrated, whether or not it actually is.
How to answer this, stage by stage
Nobody is grading whether you can say the words "distribution shift." They are grading whether you can turn it into a tile someone actually rebuilds before it costs a real person a phone call. Seven moves get you there.
Let's learn
What happens when an early-warning score keeps reporting on a crowd that already left?
Steadyline is a check-in app. Every day it asks how you're doing, offers a small coping suggestion, and watches for signs that someone might be heading toward a real crisis. Before the app had any early-warning score at all, the only sign the team had was silence: no check-in for ten straight days sent a wellness email. That caught real trouble, but late, about nine days after things had already started going wrong.
Then the team built the distress-drift score. It reads each day's check-in, compares the words and tone against that person's own last thirty days, and turns the drift into a number. If the number climbs past a cutoff, a counselor calls within twenty four hours. It cut the time to catch real trouble from about nine days down to about two. For two years it worked exactly the way it was supposed to.
Here is the turn. Say it plainly: the aggregate number was not lying. It was doing exactly what an average does, blending everyone into one crowd. Once Steadyline partnered with three community clinics, clinic-referred patients, mostly first-time app users, mostly writing shorter and more guarded check-ins than the original cohort, made up twenty two percent of weekly volume. Their real distress still moved, week over week, the same way anyone's does. It just never crossed a cutoff built for a different crowd's normal.
At its worst, this doesn't look like a scandal. It looks like nothing. Farhana Chowdhury was referred to Steadyline through her local clinic in her third week using the app. Over three weeks her check-ins got shorter. Flatter. The kind of change that, inside her own history, was real. But her cohort's check-ins already read as terse and guarded against the pooled baseline, so nothing about her drift ever looked unusual next to "everyone." No call happened. Nobody logged a miss, because the system has no way to know what it never caught.
What I would leave alone: Steadyline's explicit self-harm keyword trigger, the one that fires an immediate hotline banner and an always-logged escalation the moment specific crisis language shows up in a check-in, doesn't need a cohort baseline at all. It's a fixed rule, not a score built by comparing you to everyone else, so a shift in who signs up doesn't touch it.
The lesson: a leading indicator that's secretly a statement about "everyone" quietly stops being about anyone in particular the day everyone changes shape. And nothing you built tells you the day that happened, unless you built something to watch the shape itself.
Now here is the same thing as a story
The short version sits above. Read on for the phone call where a clinic counselor's question was the only reason anyone found out.
Katya Vronsky has run Steadyline's crisis-response operations for two years. She was in the room when the distress-drift score first shipped, back when Steadyline's whole user base was people who found the app themselves, downloaded it, signed up, and mostly already knew the language of therapy.
For most of two years, the flagged list on Katya's screen each morning ran between two and four percent of active users. A boring number, and that was the point. It meant the score was catching today what it caught yesterday.
The fading happened in three beats, and none of them looked like a mistake at the time. First, after the clinic partnership launched, the daily flagged list stayed roughly the same size and shape, so nobody thought to break it out by referral channel. Second, the weekly cohort report that existed at launch got quietly folded into a general accuracy dashboard during a reorg, because it kept agreeing with the aggregate anyway. Third, engineering shipped a faster version of the drift model to cut the cost of scoring every check-in, and re-checked it only against the same aggregate numbers everyone already trusted.
The trigger was a routine call. Ten weeks after the clinic partnership launched, one of the partner clinic's on-site counselors asked Steadyline's partnerships lead a plain question, almost in passing: "How many of our patients has your callback team actually reached so far?"
Katya pulled the real number that afternoon. Clinic-referred patients were being flagged for a callback at zero point six percent a week. The original cohort sat at three point one percent, five times higher. The aggregate flag rate had barely moved the whole time, sitting near two point six percent, because clinic-referred patients were still a minority of total check-ins.
Here is the part that actually cost something. Nobody at Steadyline could say how many patients like Farhana Chowdhury had real distress that never crossed the pooled cutoff over those ten weeks, because a check-in that doesn't trigger a call leaves no record anywhere that anyone got missed. When Katya's director asked how long this had been happening, the honest answer was ten weeks, and nobody had known for nine of them.
Two years earlier, building that very first version of the score, the decision took one afternoon with three engineers in a room. One baseline, built from the whole population, updated automatically as more check-ins came in. It was the right call. The user base really was one crowd back then, and slicing it by referral channel would have shown identical charts saying the same thing six times over. Nobody ever put a date on that decision. Nobody ever came back to ask whether the population was still one crowd.
Run the same ten weeks through the fixed design. The referral-mix tracker crosses its stale-baseline trigger by week four, not week ten, and it crosses because clinic referrals had already climbed past the shift threshold, not because of one noisy day. The clinic cohort's baseline gets flagged stale automatically, engineering rebuilds it from a held-out sample of clinic check-ins within days, and by week five Farhana's real drift crosses the new, cohort-specific cutoff. The counselor calls her that same week.
One design watched a number that used to describe everyone. The other watches whether it still does.
What I would tell myself, back in that one-afternoon meeting: the day you build a score around the word everyone, write down what everyone means today. Someday it will stop meaning that, and nothing you built will tell you when, unless you go build that too.
GUARD, and the group the average stopped seeing
This reads like a metrics question, but the honest brief is a fairness check that runs every day, so GUARD does the work here, not a metric-only framework.
Three things worth naming directly here, since this is where the real judgment lives. The easy alternative on offer was to have a person manually review every clinic-referred patient's first thirty check-ins by hand, the way a cautious team might handle any new, unfamiliar population. That got ruled out on purpose: it doesn't scale past a few hundred referrals a month without a triage backlog, and it treats an entire cohort as suspect instead of fixing what's actually broken, which is the baseline itself. The failure worth naming by name is distribution shift: the drift score was never taught a rule that says "trust clinic patients less," it simply started scoring a population whose normal check-in style looks different from the population its baseline was built on, and nobody evaluated it on that split before the clinic partnership scaled up. The guardrail is a cohort-sliced held-out check, a batch of past check-ins from each referral channel with known outcomes, that any new or rebuilt baseline has to clear before it's trusted, the same way the original baseline was checked once, years ago, and then never checked again. And the shift trigger itself isn't a hard alarm that fires on one odd day. It fires on a sustained four-week move in a channel's share of traffic, because one strange week is often just one strange week, and a trigger that fires on noise gets ignored within a month. That protection costs something real too: every new referral channel now runs on a wider, less precise cutoff for the first few weeks while its own baseline gets built and checked, and the team carries one more baseline, one more dashboard, one more thing to watch, for every channel it adds. That's the trade. A slower, less precise start for every new group of patients, in exchange for catching a real miss in weeks instead of months.
And if you want to be sure it really works, try it somewhere else
Same five letters, a car insurance claims queue instead of a check-in app, and the mix shift runs the harm the other direction this time, proof the method isn't a fluke of one industry or one kind of miss.
Ridgeway Mutual runs TrueClaim, a model that reads the written narrative on every auto claim and scores how much it reads like the claims Ridgeway's fraud team has caught before. A high score routes the claim to manual investigation and delays the payout. Sindre Bakke runs Ridgeway's claims fraud operations team.
G, groups. Sindre's team, who work the flagged-claims queue every day. Policyholders like Ximena Manrique, who file a claim and wait to hear whether it clears automatically or gets pulled aside.
U, unequal. Twelve weeks after Ridgeway added a new underwriting region through a regional broker partnership, claims from that region were flagged for manual review at fifteen point eight percent, against four point two percent for the original region. TrueClaim's anomaly score compares each claim's language against a baseline built almost entirely on the original region's narrative style, and the new region's claimants simply describe accidents differently.
A, ability to contest. A flagged policyholder gets a letter saying "additional review needed," with no mention of a language score. The payout waits weeks. Ridgeway's formal dispute process takes months and never reaches the team that tuned the model.
R, reduce. The same fix. A claims-anomaly baseline built per underwriting region, with a wider default cutoff for any region whose baseline hasn't been rebuilt yet.
D, detect. Track each region's share of weekly new claims. A sustained shift past the same eight-point trigger marks that region's baseline stale automatically, instead of waiting for a broker's operations manager to ask why so many of her clients are stuck in review.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever leading indicator your product already trusts, put that same indicator's per-cohort version, and a stale flag, right next to it.
Cost: engineering says a real per-cohort baseline pipeline is two months out. Don't read the aggregate alone as a stopgap in the meantime, pull the cohort split by hand from raw logs once a week until the real pipeline ships.
The model got better, for real: say Steadyline's overall drift-model accuracy genuinely improves next quarter. That still isn't the same claim as every cohort doing fine. A better aggregate model can make a cohort gap easier to miss, because the aggregate number gets healthier at the same time one cohort's real risk keeps climbing underneath it.
Where people run it wrong.
They build the cohort split, then leave it as a chart nobody's on the hook for reading, instead of a number with a trigger and an owner.
They wait for the next quarterly audit to catch a gap that's been running for weeks, instead of a daily number that would have caught it in days.
They fix the dashboard's mistake with a training deck on fairness, instead of a second number the system checks automatically, every day, on its own.
How to use it live. Say the reframe before naming a single tile: "a leading indicator is really just a promise about who's using the product right now. The day that changes, the same number keeps running, but it stops measuring what you think it measures, and nothing tells you unless you built something to watch the crowd itself." That buys you room to give the real design, instead of reciting "watch for bias" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if a cohort genuinely has less real risk, and the lower flag rate is correct?" Response: that's exactly why the fix isn't "flag clinic patients more," it's "stop letting one population's baseline stand in for a group it was never built from." The rebuild doesn't assume the new cohort needs more calls, it makes sure the number is actually measuring them at all.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #5 What early signal predicts churn from an AI feature?
- #6 How do you build an early warning system for silent quality degradation?