ConceptAdvancedQuality, Cost & Token Economics / Leading vs lagging indicators for AI / #17

Which leading indicators break when your user mix changes?

Steadyline's early-warning score for a mental health crisis went quiet on one whole group of patients, and nothing on the dashboard said why.

The direct answer
Rebuild your leading indicator's baseline per user cohort, not once for everyone, and mark a cohort's baseline stale the moment its share of your weekly traffic shifts past a set point, falling back to a wider default cutoff until a fresh baseline is built and checked. A single global threshold is not a safety rule. It is a description of who was using the product when someone last looked, and it stops being true the day the crowd changes shape.
Do this, in order
  1. Rebuild the leading indicator's baseline per cohort, not once for the whole population.Why: a threshold trained on one crowd quietly stops describing any group whose mix has moved away from that crowd.
  2. Flag a cohort's baseline stale the moment its share of weekly traffic shifts past a set point, and fall back to a wider default cutoff until a fresh one ships.Why: a stale baseline that keeps quietly running is worse than a broken one, because nobody thinks to check it.
  3. Track the composition of who is signing up as its own daily number, not just the indicator's own output.Why: the mix shift always happens weeks before the indicator visibly breaks, so watching the crowd catches it before watching the outcome does.
  4. Give partners and referral sources their own cohort's numbers, not just your aggregate dashboard.Why: the people closest to the affected group are the ones most likely to notice a gap first, if they can actually see it.
  5. Leave any absolute, rule-based safety trigger alone.Why: a fixed rule doesn't drift with the population the way a relative score does, so it needs no cohort baseline at all.
  6. Validate every fresh cohort baseline against a held-out sample before trusting it, never against the same check-ins it's about to score.Why: a baseline checked on the data it will later grade always looks calibrated, whether or not it actually is.

How to answer this, stage by stage

Nobody is grading whether you can say the words "distribution shift." They are grading whether you can turn it into a tile someone actually rebuilds before it costs a real person a phone call. Seven moves get you there.

1
Ground it in one real product, one real owner
Say it like this
"Let's ground this. Steadyline is a daily mood check-in app with personalized coping suggestions. Katya Vronsky runs the small team that works Steadyline's crisis-response queue every morning."
Why this works
Grounds the design in a real product and a real owner before naming a single number.
2
Name both groups before naming a single number
Say it like this
"Two groups touch this score. Katya's team reads a flagged list every morning. Patients like Farhana Chowdhury read one blank check-in screen and never learn a score exists at all."
Why this works
This is GUARD's first move, and skipping it is the most common way this kind of question turns generic.
3
Say plainly where the average stops describing anyone
Say it like this
"Steadyline's distress-drift score compares each check-in to a population baseline. That baseline was built back when every signup came from the same app-store funnel. Once patients referred through community clinics reached twenty two percent of weekly check-ins, the baseline kept describing last year's crowd, not this year's."
Why this works
This is the reframe the whole answer turns on. A candidate who skips it just says "watch for drift" and moves on.
4
Name who has no way to flag it
Say it like this
"Farhana never sees a score. She never learns a threshold exists. If her drift never crosses the pooled cutoff, nothing happens, nobody logs a miss, and the clinic that referred her has no dashboard of its own that would show the gap either."
Why this works
This is GUARD's hardest step, and the one that actually answers a risk question instead of describing one.
5
Give the fix, in one sentence
Say it like this
"So here's what I'd build. A baseline per referral cohort instead of one for everyone, marked stale the moment that cohort's share of weekly check-ins shifts past a set point, with a wider default cutoff standing in until a fresh baseline ships."
Why this works
Matches the direct answer. A named mechanism beats "we'd monitor for bias" every time.
6
Prove it with the shift that already happened
Say it like this
"Here's what happens without it. Ten weeks after the clinic partnership launched, clinic-referred patients were flagged for a counselor callback at zero point six percent a week, against three point one percent for the original cohort. Nobody noticed until a clinic counselor asked how many of her own patients had actually been called back."
Why this works
A four-sentence failure beats a paragraph about the importance of fairness monitoring.
7
Close on the tradeoff you're accepting
Say it like this
"We looked at just reviewing every clinic-referred patient's first thirty check-ins by hand, and ruled it out. It doesn't scale past a few hundred referrals a month, and it treats the whole cohort as suspect instead of fixing the score itself. The real cost of the fix is a slower rollout for every new cohort, a wider, less precise cutoff for a few weeks while its own baseline gets built, and one more baseline to watch for every channel we add. That's the trade I'd take over a quiet miss nobody can even count."
Why this works
Naming a rejected option and a real cost turns a safety instinct into a defensible decision.
If you remember one thing A leading indicator is a promise about who "everyone" is right now. The day the crowd changes shape, the same number keeps running, and it quietly stops being about anyone in particular.

Let's learn

What happens when an early-warning score keeps reporting on a crowd that already left?

Steadyline is a check-in app. Every day it asks how you're doing, offers a small coping suggestion, and watches for signs that someone might be heading toward a real crisis. Before the app had any early-warning score at all, the only sign the team had was silence: no check-in for ten straight days sent a wellness email. That caught real trouble, but late, about nine days after things had already started going wrong.

Hand sketched comparison titled Same morning, two very different views. Left, a lever and dial icon labeled Renata's queue, caption a lever, a daily flagged list, a number she can act on. Right, a person icon labeled Marisol's check-in, caption one blank screen, no score shown, no lever at all.
Same app, same morning. One side gets a number and a lever. The other gets one blank screen.

Then the team built the distress-drift score. It reads each day's check-in, compares the words and tone against that person's own last thirty days, and turns the drift into a number. If the number climbs past a cutoff, a counselor calls within twenty four hours. It cut the time to catch real trouble from about nine days down to about two. For two years it worked exactly the way it was supposed to.

Weekly counselor-callback flag rate, ten weeks after the clinic partnership
5% 0% 2.6% 3.1% 0.6% Aggregate Original cohort Clinic-referred
The one number Katya's team watched every morning sat near two point six percent for weeks. It never showed that clinic-referred patients were being flagged at a fifth the rate of everyone else.

Here is the turn. Say it plainly: the aggregate number was not lying. It was doing exactly what an average does, blending everyone into one crowd. Once Steadyline partnered with three community clinics, clinic-referred patients, mostly first-time app users, mostly writing shorter and more guarded check-ins than the original cohort, made up twenty two percent of weekly volume. Their real distress still moved, week over week, the same way anyone's does. It just never crossed a cutoff built for a different crowd's normal.

The score wasn't wrong about "everyone." It just stopped being about anyone new.
Knowledge spark: what is a population baseline? A model turns each check-in into a number for how far it drifts from "normal." Normal isn't just that one person's own history, it's also compared against everyone who has ever used the app. A cutoff, here set at two standard deviations, says how far from that shared "everyone" number counts as worth a phone call. Change who "everyone" is, and the same cutoff means something else.

At its worst, this doesn't look like a scandal. It looks like nothing. Farhana Chowdhury was referred to Steadyline through her local clinic in her third week using the app. Over three weeks her check-ins got shorter. Flatter. The kind of change that, inside her own history, was real. But her cohort's check-ins already read as terse and guarded against the pooled baseline, so nothing about her drift ever looked unusual next to "everyone." No call happened. Nobody logged a miss, because the system has no way to know what it never caught.

Hand sketched decision tree titled What happens after the drift score runs. Root box reads Steadyline scores today's check-in. One branch, crosses the cutoff, leads to counselor calls within 24 hours. Other branch, stays under the cutoff, leads to no flag, no callback, no record anywhere.
Only one of these two paths ends in a phone call. Steadyline can't tell which path a given patient is on, because it never asks who "the cutoff" was actually built for.
The decision that mattered Building the distress-drift score against one global population baseline, made back when literally every signup came from the same app-store funnel. It was the right call at launch. Nobody ever came back to ask whether that was still true.

What I would leave alone: Steadyline's explicit self-harm keyword trigger, the one that fires an immediate hotline banner and an always-logged escalation the moment specific crisis language shows up in a check-in, doesn't need a cohort baseline at all. It's a fixed rule, not a score built by comparing you to everyone else, so a shift in who signs up doesn't touch it.

The lesson: a leading indicator that's secretly a statement about "everyone" quietly stops being about anyone in particular the day everyone changes shape. And nothing you built tells you the day that happened, unless you built something to watch the shape itself.

Now here is the same thing as a story

The short version sits above. Read on for the phone call where a clinic counselor's question was the only reason anyone found out.

Katya Vronsky has run Steadyline's crisis-response operations for two years. She was in the room when the distress-drift score first shipped, back when Steadyline's whole user base was people who found the app themselves, downloaded it, signed up, and mostly already knew the language of therapy.

For most of two years, the flagged list on Katya's screen each morning ran between two and four percent of active users. A boring number, and that was the point. It meant the score was catching today what it caught yesterday.

The fading happened in three beats, and none of them looked like a mistake at the time. First, after the clinic partnership launched, the daily flagged list stayed roughly the same size and shape, so nobody thought to break it out by referral channel. Second, the weekly cohort report that existed at launch got quietly folded into a general accuracy dashboard during a reorg, because it kept agreeing with the aggregate anyway. Third, engineering shipped a faster version of the drift model to cut the cost of scoring every check-in, and re-checked it only against the same aggregate numbers everyone already trusted.

The trigger was a routine call. Ten weeks after the clinic partnership launched, one of the partner clinic's on-site counselors asked Steadyline's partnerships lead a plain question, almost in passing: "How many of our patients has your callback team actually reached so far?"

Katya pulled the real number that afternoon. Clinic-referred patients were being flagged for a callback at zero point six percent a week. The original cohort sat at three point one percent, five times higher. The aggregate flag rate had barely moved the whole time, sitting near two point six percent, because clinic-referred patients were still a minority of total check-ins.

Here is the part that actually cost something. Nobody at Steadyline could say how many patients like Farhana Chowdhury had real distress that never crossed the pooled cutoff over those ten weeks, because a check-in that doesn't trigger a call leaves no record anywhere that anyone got missed. When Katya's director asked how long this had been happening, the honest answer was ten weeks, and nobody had known for nine of them.

It was never the score lying about everyone. It just stopped being about anyone new.

Two years earlier, building that very first version of the score, the decision took one afternoon with three engineers in a room. One baseline, built from the whole population, updated automatically as more check-ins came in. It was the right call. The user base really was one crowd back then, and slicing it by referral channel would have shown identical charts saying the same thing six times over. Nobody ever put a date on that decision. Nobody ever came back to ask whether the population was still one crowd.

Run the same ten weeks through the fixed design. The referral-mix tracker crosses its stale-baseline trigger by week four, not week ten, and it crosses because clinic referrals had already climbed past the shift threshold, not because of one noisy day. The clinic cohort's baseline gets flagged stale automatically, engineering rebuilds it from a held-out sample of clinic check-ins within days, and by week five Farhana's real drift crosses the new, cohort-specific cutoff. The counselor calls her that same week.

One design watched a number that used to describe everyone. The other watches whether it still does.

What I would tell myself, back in that one-afternoon meeting: the day you build a score around the word everyone, write down what everyone means today. Someday it will stop meaning that, and nothing you built will tell you when, unless you go build that too.

GUARD, and the group the average stopped seeing

This reads like a metrics question, but the honest brief is a fairness check that runs every day, so GUARD does the work here, not a metric-only framework.

G
Groups. Who is affected.
Katya's team, who read a flagged list every morning and trust it to be complete. Patients like Farhana Chowdhury, who read one blank check-in screen and never learn a score, a threshold, or a callback queue exists.
In this answer: not patients in general. Farhana, specifically, and the roughly one in five weekly check-ins that now come from her referral cohort.
U
Unequal. Where the harm lands.
Clinic-referred patients get flagged for a callback at zero point six percent a week. The original cohort gets flagged at three point one percent. The aggregate, two point six percent, hides the gap because the harmed group is a minority of total volume.
This is the whole reason one global baseline fails here. A number built from everyone can look healthy while it is actively wrong for someone.
A
Ability to contest. Who never gets to push back.
A check-in that scores under the cutoff triggers nothing: no flag, no callback, no note anywhere. Farhana has no way to learn the score exists, let alone that it was built for a different crowd than hers. The clinic that referred her sees Steadyline's aggregate marketing numbers, not a cohort-level flag rate.
Nobody decided on purpose that clinic patients wouldn't get a lever. It fell out of treating one population baseline as a permanent fact instead of a snapshot of who signed up first.
R
Reduce. The specific product change.
A distress-drift baseline built per referral cohort, not once for the whole population, with a wider, more conservative default cutoff applied to any cohort whose baseline hasn't been rebuilt and checked yet.
Not a fairness policy. A rebuild trigger and a fallback threshold that Katya's system already runs, the same way it already runs the aggregate score.
D
Detect. How you'd know before a partner has to ask.
Track each referral channel's share of weekly check-ins as its own number. When a channel's share shifts more than eight percentage points inside any rolling four weeks, mark that channel's baseline stale automatically and switch it to the wider default cutoff until a fresh baseline clears a held-out check.
Four weeks of exposure instead of ten, caught by a number moving, not by a person asking a hard question after the fact.

Three things worth naming directly here, since this is where the real judgment lives. The easy alternative on offer was to have a person manually review every clinic-referred patient's first thirty check-ins by hand, the way a cautious team might handle any new, unfamiliar population. That got ruled out on purpose: it doesn't scale past a few hundred referrals a month without a triage backlog, and it treats an entire cohort as suspect instead of fixing what's actually broken, which is the baseline itself. The failure worth naming by name is distribution shift: the drift score was never taught a rule that says "trust clinic patients less," it simply started scoring a population whose normal check-in style looks different from the population its baseline was built on, and nobody evaluated it on that split before the clinic partnership scaled up. The guardrail is a cohort-sliced held-out check, a batch of past check-ins from each referral channel with known outcomes, that any new or rebuilt baseline has to clear before it's trusted, the same way the original baseline was checked once, years ago, and then never checked again. And the shift trigger itself isn't a hard alarm that fires on one odd day. It fires on a sustained four-week move in a channel's share of traffic, because one strange week is often just one strange week, and a trigger that fires on noise gets ignored within a month. That protection costs something real too: every new referral channel now runs on a wider, less precise cutoff for the first few weeks while its own baseline gets built and checked, and the team carries one more baseline, one more dashboard, one more thing to watch, for every channel it adds. That's the trade. A slower, less precise start for every new group of patients, in exchange for catching a real miss in weeks instead of months.

And if you want to be sure it really works, try it somewhere else

Same five letters, a car insurance claims queue instead of a check-in app, and the mix shift runs the harm the other direction this time, proof the method isn't a fluke of one industry or one kind of miss.

Ridgeway Mutual runs TrueClaim, a model that reads the written narrative on every auto claim and scores how much it reads like the claims Ridgeway's fraud team has caught before. A high score routes the claim to manual investigation and delays the payout. Sindre Bakke runs Ridgeway's claims fraud operations team.

G, groups. Sindre's team, who work the flagged-claims queue every day. Policyholders like Ximena Manrique, who file a claim and wait to hear whether it clears automatically or gets pulled aside.
U, unequal. Twelve weeks after Ridgeway added a new underwriting region through a regional broker partnership, claims from that region were flagged for manual review at fifteen point eight percent, against four point two percent for the original region. TrueClaim's anomaly score compares each claim's language against a baseline built almost entirely on the original region's narrative style, and the new region's claimants simply describe accidents differently.
A, ability to contest. A flagged policyholder gets a letter saying "additional review needed," with no mention of a language score. The payout waits weeks. Ridgeway's formal dispute process takes months and never reaches the team that tuned the model.
R, reduce. The same fix. A claims-anomaly baseline built per underwriting region, with a wider default cutoff for any region whose baseline hasn't been rebuilt yet.
D, detect. Track each region's share of weekly new claims. A sustained shift past the same eight-point trigger marks that region's baseline stale automatically, instead of waiting for a broker's operations manager to ask why so many of her clients are stuck in review.

New region's share of Ridgeway's weekly claims, first twelve weeks
30% 0% stale-baseline trigger, week 5 wk 1 wk 5 wk 9 wk 12
Ridgeway's old dashboard had no version of this line at all. By the time anyone asked a question, the new region was already close to a third of total claims.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever leading indicator your product already trusts, put that same indicator's per-cohort version, and a stale flag, right next to it.
Cost: engineering says a real per-cohort baseline pipeline is two months out. Don't read the aggregate alone as a stopgap in the meantime, pull the cohort split by hand from raw logs once a week until the real pipeline ships.
The model got better, for real: say Steadyline's overall drift-model accuracy genuinely improves next quarter. That still isn't the same claim as every cohort doing fine. A better aggregate model can make a cohort gap easier to miss, because the aggregate number gets healthier at the same time one cohort's real risk keeps climbing underneath it.

Where people run it wrong.
They build the cohort split, then leave it as a chart nobody's on the hook for reading, instead of a number with a trigger and an owner.
They wait for the next quarterly audit to catch a gap that's been running for weeks, instead of a daily number that would have caught it in days.
They fix the dashboard's mistake with a training deck on fairness, instead of a second number the system checks automatically, every day, on its own.

How to use it live. Say the reframe before naming a single tile: "a leading indicator is really just a promise about who's using the product right now. The day that changes, the same number keeps running, but it stops measuring what you think it measures, and nothing tells you unless you built something to watch the crowd itself." That buys you room to give the real design, instead of reciting "watch for bias" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a question about who a leading indicator is quietly failing?
Tap to flip
ANSWER
GUARD: name the groups, find where harm lands unevenly, ask who can't push back, design the reduce, build the daily detect.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Katya Vronsky, who runs Steadyline's crisis-response operations team, and Farhana Chowdhury, a clinic-referred patient whose real drift never crossed the pooled cutoff.
3 · WHAT BROKE
What actually broke when Steadyline's user mix changed?
Tap to flip
ANSWER
The distress-drift score's population baseline, built on the original app-store cohort, silently stopped describing clinic-referred patients once they became 22 percent of weekly check-ins.
4 · THE TWO GROUPS
Name the two groups this answer is built around.
Tap to flip
ANSWER
Katya's team, who read a flagged list every morning and can act on it. Patients like Farhana, who get one blank check-in screen with no score shown and no way to know a threshold ever existed.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building the distress-drift score against one global population baseline. It was right when every signup came from the same app-store funnel; nobody revisited it once clinic referrals changed the mix.
6 · THE NUMBER
Fill in the blank: clinic-referred patients were flagged for a callback at ___ percent a week, against ___ percent for the original cohort.
Tap to flip
ANSWER
0.6 percent; 3.1 percent. The aggregate sat near 2.6 percent the whole time and never showed the gap.
7 · THE REPLAY
Same ten weeks, new design, what changes?
Tap to flip
ANSWER
The referral-mix tracker flags the clinic cohort's baseline stale by week four. A fresh baseline clears its held-out check within days, and Farhana's real drift crosses the new cutoff by week five.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
TrueClaim, Ridgeway Mutual's auto-claims anomaly score. Same GUARD steps, opposite direction: a new region gets over-flagged instead of under-flagged.

Check yourself Score: 0 / 0

Fill in the blank
1. Before the fix, clinic-referred patients were flagged for a counselor callback at ___ percent a week, against ___ percent for Steadyline's original cohort.
Show hint
Check the chart in "Let's learn."
Show answer
0.6 percent; 3.1 percent. A five-times gap, hidden under an aggregate that sat near 2.6 percent the whole time.
Multiple choice
2. Why not just widen the global cutoff a little instead of building cohort-aware baselines?
  • A. A wider global cutoff would flag everyone more often, and the ops team can't staff that many calls.
  • B. A single global cutoff can't tell whether a low score means someone is fine, or means the baseline no longer describes their cohort at all.
  • C. Steadyline's legal team requires a cohort-specific design by policy.
  • D. Widening a threshold slows down the app's push notifications.
Show hint
Think about what a "low score" actually proves, and what it doesn't.
Show answer
B. Widening the cutoff is still one number for everyone. It doesn't fix the real problem, which is that the number is being compared to the wrong crowd.
True or false
3. True or false: the flag-rate gap between the two cohorts showed up clearly on Steadyline's aggregate weekly dashboard.
  • True
  • False
Show hint
Check what share of total check-ins the clinic-referred cohort actually was.
Show answer
False. The aggregate stayed near 2.6 percent the whole time, because clinic-referred patients were still a minority of total weekly check-ins.
Multiple choice
4. What actually broke when Steadyline's user mix changed?
  • A. The check-in app crashed for new users.
  • B. The distress-drift score's population baseline no longer described the new cohort's normal check-in language, so real drift under-fired for them.
  • C. The counselor callback phone line stopped working.
  • D. Steadyline's servers couldn't handle the extra signups.
Show hint
Nothing technically crashed. Think about what the number was actually being compared against.
Show answer
B. The app worked fine. The baseline the score compares each check-in to had simply stopped describing a growing share of the people using it.
Short answer, apply it yourself
5. Pick a product you use that scores you against "typical" behavior of other users. What's one group whose normal behavior might already look like a red flag, or look perfectly fine, only because it's measured against everyone else instead of people like them?
Show hint
Look for a place where the product compares you to an average built from a crowd you might not actually belong to.
Show answer
Model answer: A food delivery app that predicts "late" orders by comparing a rider's delivery times to the platform-wide average. A rider working a rural route with longer roads always looks slow next to a baseline built mostly from dense-city riders, so the app keeps nudging them toward a warning that has nothing to do with how well they're actually doing their job.
Short answer, name what wouldn't break
6. Name a place inside Steadyline where this same user-mix shift would NOT break a safety check, and explain why.
Show hint
Look at "what I would leave alone" in "Let's learn."
Show answer
Model answer: The explicit self-harm keyword trigger. It fires on an absolute rule, specific words typed into a check-in, not on a score computed relative to a population baseline. It means the same thing no matter who signs up next.
Before you close the answer
Why this works
Tests whether you'll trust a leading indicator's own output, or go check what crowd it was actually built to describe. Most candidates answer a metric-safety question by describing a threshold. Few describe what makes a threshold stop meaning what it used to.
Follow-up traps
"Isn't an eight-point shift trigger just a different threshold that could also be wrong?" Response: that's why it forces a rebuild and a wider fallback cutoff, not an automatic shutoff. A cohort crossing it doesn't stop Steadyline from working, it puts a conservative default in place the same day, the way a low reading sends a nurse to check on someone instead of writing a prescription by itself.

"What if a cohort genuinely has less real risk, and the lower flag rate is correct?" Response: that's exactly why the fix isn't "flag clinic patients more," it's "stop letting one population's baseline stand in for a group it was never built from." The rebuild doesn't assume the new cohort needs more calls, it makes sure the number is actually measuring them at all.
If pressed
The rebuild itself works by holding out the most recent four weeks of a cohort's check-ins, at least 500 of them, computing the same drift feature within that cohort alone, and setting a new cutoff calibrated to catch real escalations at roughly the same rate the original cohort's cutoff did, checked against a held-out clinical-review sample of flagged and unflagged check-ins before it ever goes live. That check is what makes it a real baseline and not just a guess dressed up as one.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more