ConceptAdvancedQuality, Cost & Token Economics / Leading vs lagging indicators for AI / #9

Explain why thumbs-down rate is a lagging indicator disguised as a leading one.

A number that reacts the second someone taps it feels like the fastest thing on the dashboard. It can still be the slowest thing to tell you the truth.

The direct answer
Thumbs-down rate feels immediate because the click happens the second the model answers, but three separate filters sit between a bad suggestion and that click: a cook has to notice it's wrong, care enough to stop and act, then find a small button. Most upset cooks never clear all three, so the rate stays flat while the real problem grows underneath it. Watch it next to a behavior proxy, like recipe completion or return rate, on the same group of users, and trust that instead.
Do this, in order
  1. Watch thumbs-down rate next to a behavior proxy on the same users, never on its own.Why: this is the real fix. A click count with three filters built in cannot be read alone.
  2. Rule out the boring explanation first: check the button still fires and traffic didn't drop.Why: don't call a metric structurally broken before you've ruled out an honest accident.
  3. Break the click into its three filters, notice it, care about it, find it, before trusting the raw number.Why: this is the recut that shows the rate is measuring who survives three steps, not how bad the output was.
  4. Name the real cause candidates by name: silent abandonment, trust erosion, or a button nobody sees.Why: each one needs a different fix. Guessing "quality" broadly wastes the fix.
  5. Gate any model swap behind a hand-checked set of examples, never behind thumbs-down staying under a threshold.Why: this is the decision to take back. The click was never built to catch a quiet drop like this.
  6. Leave thumbs-down alone as a per-recipe fix-it channel. Just stop using it as the health gate.Why: it still does its original job well. It was never the right job for a launch decision.

How to answer this, stage by stage

Nobody is grading whether you can spot that a metric is misleading. They're grading whether you can show your work getting there, one check at a time, instead of just announcing the twist. Seven moves.

1
Scope it to one real product
Say it like this
"Let's ground this. Coalridge runs ScrapChef, an app that looks at what's in your fridge and suggests a recipe, swaps included, like 'out of buttermilk, use milk and a spoon of lemon juice.' Cosmin Fassbinder runs recipe quality there, and the metric on trial is the thumbs-down button under every suggestion."
Why this works
Grounds the abstract question in a real product and a real owner before touching the metric itself.
2
Say your structure out loud
Say it like this
"Here's how I'd take this apart. I'm going to rule out the boring explanation first, then recut the click itself before I trust the rate it produces."
Why this works
Two seconds of structure tells the interviewer you have a plan before you say a word about the metric.
3
Lay out the timeline
Say it like this
"Ten weeks ago, Coalridge swapped in a cheaper model for suggesting substitutions, about a third of the cost per call. Thumbs-down rate sat at two point three percent the week it shipped, and two point six percent ten weeks later. Basically flat."
Why this works
Anchors the diagnosis to a real date and a real number instead of a vague sense that something's off.
4
Rule out the innocent read
Say it like this
"Before I say the metric is lying, I check the easy stuff. Did traffic drop? No. Is the button still firing? Yes, I tested it myself. So a flat two point six percent isn't an accident. Something else is holding it down."
Why this works
Ruling out a tracking accident first is what separates a real diagnosis from a hunch dressed up as one.
5
Recut the click into what it actually requires
Say it like this
"So I recut it. Out of four hundred sessions where a food scientist confirmed the suggestion was genuinely bad, fifty eight percent of cooks even noticed. Of those, about a third cared enough to stop and act instead of just winging it. Of that group, only about a quarter found the button. Multiply it out and you get under five in a hundred bad suggestions ever producing a click."
Why this works
This is the load-bearing move. It shows the rate isn't measuring how bad the output was, it's measuring who survives three filters.
6
Name the real suspects, then the one test that settles it
Say it like this
"Three things could sit under that flat number: cooks who close the app with no click at all, cooks who've quietly stopped trusting it enough to bother correcting it, or the button itself, buried behind a small icon nobody taps with flour on their hands. The test that settles it: compare thumbs-down rate against seven-day return rate, same group, swapped-model users against a control still on the old model. Thumbs-down stayed flat. Return rate dropped from sixty four percent to forty one."
Why this works
Naming specific candidates, then one test that separates them, beats a paragraph of general worry.
7
Close on the fix and the option you turned down
Say it like this
"So here's what I'd do. Gate any model swap behind a hand-checked set of examples before it ships, and watch return rate alongside thumbs-down for the first few weeks after. We talked about just making the button bigger instead. I turned that down, it only fixes 'find the button,' not 'notice' or 'care,' and a bigger button would inflate the number for reasons that have nothing to do with quality. The real cost is time, the check adds about two weeks before a cheaper model reaches everyone, against the savings we're trying to bank."
Why this works
Naming a rejected option and its real cost turns "track something else" into a defensible decision, not a wish list.
If you remember one thing A click that fires instantly is not the same as a click that fires reliably. Three filters sit between a bad answer and a thumbs-down. Most people never clear all three.

Let's learn

What happens when the number a team built to catch bad answers fast turns out to be the slowest number on the whole dashboard?

ScrapChef is an app inside Coalridge. Point your phone at what's left in the fridge, and it suggests a recipe, with a swap for anything you're missing, out of buttermilk, use milk and a spoon of lemon juice.

Before the swap, checking a substitution used to mean stopping mid-recipe to look it up online, a couple of extra minutes hunting on a phone with sticky hands. ScrapChef cut that to zero. The suggestion just appeared, and thumbs-down rate sat around two percent, an honest number, because the model was good enough that almost nobody needed to push back.

Ten weeks ago, Coalridge swapped in a cheaper model for generating substitutions, about a third the cost per call, a real saving at their volume. Thumbs-down rate today reads two point six percent. Basically the same as before.

Ten weeks, and the dashboard never moved
70% 35% 0% thumbs-down rate 7-day return rate week 0, model swap ships week 10
Thumbs-down rate barely moves, 2.3% to 2.6%. Seven-day return rate for the same swapped-model group falls from 64% to 41%. Same ten weeks, same users.

Here is the turn. The flat line is not proof nothing changed. It is proof the number was never built to catch what happens when a suggestion goes wrong mid-recipe. Most cooks don't click anything. They wing it, or they quietly close the app and order takeout, and neither one shows up as a thumbs-down.

The button did not stay quiet because nothing broke. It stayed quiet because breaking never looked like a click.
What a thumbs-down actually needs from a cook, out of 400 confirmed-bad suggestions
58% 19% 4.5% Noticed it Cared, acted Found the button
Each bar is a filter, not a separate group. A cook has to clear all three to produce a click. Under five in a hundred genuinely bad suggestions ever do.
Knowledge spark: what is a behavior proxy? A number built from what people actually do, not what they choose to report. Recipe completion, or whether someone comes back next week, moves whether or not anyone clicks a button. A click needs a person to decide to file a complaint. A behavior proxy doesn't ask permission.

At its worst, this costs the exact users a subscription app can't afford to lose. Seven-day return rate for cooks who hit a bad suggestion in the swapped-model group fell from sixty four percent to forty one, and not one of them left a review or tapped a thing. The dashboard stayed green the whole time.

The decision that mattered Coalridge gated the model swap on thumbs-down rate staying under three percent, with no hand-checked set of substitutions and no behavior check before rollout.

The choice I would take back is that gate. It made sense while everyone believed thumbs-down was a real-time gauge, a click, fired instantly, right there under the recipe. Nobody had worked out that a real-time click and a reliable one are not the same thing.

What I would leave alone: for a low-stakes swap, like table salt for kosher salt, thumbs-down really is a fine, honest signal. Nobody gets that wrong even distracted at the stove, so a flat near-zero rate there genuinely does mean fine. Not every metric here needs a behavior proxy standing next to it.

The lesson: a number that's easy to click isn't the same as a number that's easy to catch. The three steps between a bad answer and a thumbs-down are exactly the three steps most people skip when they're standing at a stove with wet hands and a timer running.

Now here is the same thing as a story

The short version is above. Read on for the Tuesday sync where a one-line remark turned into a two-week audit.

Cosmin Fassbinder has run recipe quality at Coalridge for three years. He built the golden set the old substitution model was checked against, four hundred pantry combinations, hand-scored by a food scientist named Reya who used to test recipes for a cookbook publisher. He trusts numbers. He also trusts Reya more than any dashboard.

When the cheaper substitution model shipped ten weeks ago, Cosmin watched thumbs-down rate the way you'd watch a smoke alarm, not because he expected it to go off, but because if it stayed quiet, that meant the room was fine. Two point three percent the first week. Two point four the second. He stopped checking it daily around week four. It kept being fine, so he had other things to look at.

Hand sketched comparison of three panels titled Three ways a bad recipe never becomes a click. Left, a person icon labeled Never noticed, closes the app mid-cook with no click of any kind. Middle, a question mark box labeled Stopped trusting it, checks everything by hand now, feedback feels pointless. Right, a plain box labeled Never found the button, small icon after the last step, flour on the screen.
Three separate ways a genuinely bad suggestion can happen and still never reach the dashboard as a click.

The remark came on a Tuesday, in a weekly sync, from Zinaida Trevino, who runs support. She wasn't presenting anything. She just said it, almost as an aside, on her way to the next agenda item: "App Store reviews mentioning 'wasted ingredients' or 'recipe didn't work' have nearly tripled this quarter. Nobody's filing a support ticket, they're just leaving a star rating and moving on." Then she moved on too, to a question about return labels.

Cosmin almost let it go. The dashboard was green. Thumbs-down rate had not moved. But he pulled a sample that afternoon anyway, four hundred sessions from the swapped-model group, and had Reya score them against the same golden set she'd used for the old model.

Two hundred and thirty two of the four hundred suggestions were genuinely bad, ratios that didn't work, pairings that ruined a dish. That's fifty eight percent, roughly a quarter of all sessions overall. Cosmin pulled the session recordings for those two hundred and thirty two. Most cooks didn't stop. They adjusted on the fly, tasted, added more of something, or just plated whatever came out and moved on with their evening. A handful closed the app outright, mid-recipe, and never came back that night. Eighteen tapped the thumbs-down button. Eighteen, out of two hundred and thirty two confirmed-bad suggestions.

He ran the return-rate check the same afternoon, comparing that group against a small control group Coalridge had quietly kept on the old model for exactly this kind of comparison. Sixty four percent of the control group opened the app again within seven days. Forty one percent of the swapped-model group did. That gap had been sitting there the entire ten weeks, visible the moment anyone thought to look at it next to thumbs-down instead of looking at thumbs-down alone.

It was never that the model got a little worse and nobody minded. It was that almost nobody who minded ever told the dashboard.

A year before any of this, when the thumbs-down button was first designed, the team debated where to put it, prominent under every step, or tucked under the finished card so it didn't clutter a recipe someone was actively cooking from. They chose tucked away, a fair call at the time, because the model was good enough that almost nobody needed it. That call was never revisited once the model changed underneath it.

Run the same ten weeks through the fixed design. The model swap still ships, same savings. But it ships behind Reya's golden set first, which would have caught the ratio problem in a two-day check instead of a ten-week silence, and return rate rides alongside thumbs-down on the same dashboard from day one, not discovered after a hallway remark. The gap between sixty four and forty one gets caught in week one, not week ten, and Coalridge fixes the substitution logic before a single subscriber quietly stops coming back.

What I would tell myself, before any of this: a metric that reacts instantly to a click is not the same as a metric that reacts quickly to a problem. Those are two different kinds of fast, and only one of them is the one you actually need.

TRACE, and the letter that actually cracked it

This is a diagnosis question dressed as a critique. TRACE is built for exactly this: rule out the boring answer, then narrow to the real one, one letter at a time.

T
Timeline. When did it actually start, and what shipped near that date.
The cheaper substitution model shipped week 0. Thumbs-down rate moved from 2.3% to 2.6% over the following ten weeks, effectively flat.
The gap between what shipped and when anyone noticed is the whole story.
R
Recut. Slice the number by what it actually requires, not just by segment.
Notice it's wrong, care enough to act, find the button. Three filters. Of 232 confirmed-bad suggestions, 58% got noticed, roughly a third of those got acted on, and about a quarter of that group found the button.
This is the strongest move in the whole answer. It shows the metric was never wired to the thing it claims to measure.
A
Assume nothing. Rule out a tracking accident before diagnosing behavior.
Checked traffic (stable) and the click event itself (firing correctly) before concluding the flat rate meant something structural, not broken tracking.
Skip this step and you risk chasing a ghost that's really a tracking bug.
C
Cause candidates. Three named, not a shrug at "quality."
Silent abandonment, no click of any kind. Trust erosion, cooks who've stopped bothering to correct a tool they no longer expect to fix itself. Button placement suppressing the click independent of quality.
Each candidate points at a different fix, so naming them separately matters.
E
Evidence test. The one check that separates "fine" from "lagging."
Seven-day return rate, same group, swapped-model users against a held-out control. 64% versus 41%, while thumbs-down rate held flat between the two groups.
This is the proof, not an opinion about which cause candidate is most likely.

Two things worth naming directly, since this is where the real judgment sits. First, the rejected fix: make the thumbs-down button bigger and easier to find. Turned down on purpose, that only widens the "find the button" filter, leaves "notice" and "care" untouched, and would inflate the click count for reasons unrelated to quality, muddying the exact signal you're trying to clean up. Second, the AI-specific failure worth naming by name: a cheaper model swap can quietly change behavior on the long tail of pantry combinations it saw less of in training, a form of distribution shift the golden set exists to catch before launch, not after. The guardrail is that same golden set, a batch of hand-checked substitutions a food scientist scores, required to clear a near-total pass rate before any model swap ships broadly, with a behavior-proxy dashboard running in parallel for the first few weeks regardless of what thumbs-down says. That guardrail has a real cost: the cheaper model still saves roughly two-thirds of the cost per suggestion, but the eval-and-monitor step adds about two weeks before the savings reach every user. That delay is the price of not finding out the hard way, from a support lead's aside in a Tuesday meeting, that a quiet number was never quiet because things were fine.

And if you want to be sure it really works, try it somewhere else

Same five letters, a completely different job, a technician's "not helpful" tap instead of a cook's thumbs-down, and this time the pressure comes from speed, not cost.

Larchdon runs FixFlow, a tool that points an HVAC technician's phone at a unit and an error code and suggests a diagnostic step. Idony Northrup runs product for it.

T, timeline. Two months ago, Larchdon shortened FixFlow's prompt template to cut response time from six seconds to two, so a technician isn't standing there waiting with a wrench in hand. "Not helpful" rate has stayed under one percent the whole time.
R, recut. A technician has to notice the fix was wrong, sometimes not until a callback the next week. Then care enough to log it instead of just moving to the next job, techs are paid per job, not per feedback. Then find the icon, buried behind a settings menu on a rugged tablet, gloves on.
A, assume nothing. Job volume and tablet sessions stayed flat over the same two months, ruling out "fewer sessions" as the reason the rate looks steady.
C, cause candidates. Silent workaround, techs diagnose it themselves and never reopen the tablet for that job. Trust erosion, senior techs skip the suggestion entirely and call a colleague instead. A UI regression from the last app update that moved the icon behind a menu, suppressing the click independent of the fix's quality.
E, evidence test. Compare "not helpful" rate against fourteen-day callback rate, same technicians, shortened-prompt group against a control still getting the longer prompt.

Hand sketched flow diagram titled What a not-helpful tap needs from a technician. Three connected boxes in sequence: Notices it's wrong, Stops to log it, Finds the tiny icon, the last box emphasized in blue to show it is the narrowest filter.
Same three filters as ScrapChef's cooks, a different tool, a different job, gloves instead of flour.
Same shape, different lever ScrapChef's flat number hid behind a cost-driven model swap. FixFlow's hides behind a latency-driven prompt change. Different reason to ship fast, same three filters standing between a bad answer and a click.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the recut: whatever "negative feedback" number you're watching, ask what three things a person has to do before that click fires, then check whether most of your users clear all three.
Cost: the behavior-proxy check can't ship for a month. Don't read the raw negative-feedback rate alone as a stopgap in the meantime, hold off on trusting any "still healthy" read from it until the real check exists.
The model got better, for real: say FixFlow's suggestions genuinely improve next quarter. That still isn't proof the "not helpful" rate would have caught it if they hadn't. A better model just makes the next quiet regression harder to spot, because the same three filters are still standing there.

Where people run it wrong.
They treat a flat, low rate as proof of health, on reflex, without asking what a person had to do to produce it.
They respond to a hunch that something's off by watching the same metric harder, instead of building a second one that doesn't depend on a click at all.
They wait for a support lead's aside in a meeting to go looking, instead of running the behavior-proxy comparison on a fixed schedule whether anyone remembers to ask or not.

How to use it live. Say the reframe before naming a single fix: "A click-based number can only be as fast as the person clicking it, and most people never clear all three things it takes to click it at all." That buys you room to give the real answer instead of reaching for "it's a lagging indicator" as a label with nothing under it.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question like this, and what's its one-line job?
Tap to flip
ANSWER
TRACE. Rule out the boring explanation, then narrow to the real cause with one evidence test.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Cosmin Fassbinder, who has run recipe quality at Coalridge for three years and built the golden set ScrapChef's models get checked against.
3 · THE FALSE BELIEF
What did the team believe about thumbs-down rate that turned out to be wrong?
Tap to flip
ANSWER
That it was a real-time gauge, because the click fires instantly. Really, it only counts the small slice of users who clear three separate filters first.
4 · THE RECUT
What are the three filters a click has to clear here?
Tap to flip
ANSWER
Notice the output is wrong, care enough to stop and act, and find the button. Most dissatisfied cooks never clear all three.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Gating the model swap on thumbs-down staying under three percent, with no hand-checked golden set. It made sense while the team believed the click was a real-time gauge.
6 · THE NUMBER
Fill in the blank: of 232 confirmed-bad suggestions, only ___ produced a thumbs-down click.
Tap to flip
ANSWER
18, about 4.5% of them. That gap between "actually bad" and "actually clicked" is the whole answer in one number.
7 · THE EVIDENCE TEST
What's the one check that proves the rate was lagging, not honest?
Tap to flip
ANSWER
Seven-day return rate on the same group, swapped-model users against a control. It fell from 64% to 41% while thumbs-down rate stayed flat.
8 · CROSS PRODUCT TRANSFER
Section 4 runs TRACE on a different product. Which one, and what's the flat number this time?
Tap to flip
ANSWER
Larchdon's FixFlow, an HVAC diagnostic tool. The "not helpful" tap technicians rarely bother to give, latency-driven instead of cost-driven.

Check yourself Score: 0 / 0

True or false
1. True or false: ScrapChef's thumbs-down rate stayed low because the cheaper substitution model's suggestions were actually fine.
  • True
  • False
Show hint
Check the return rate chart, not just the thumbs-down line.
Show answer
False. Return rate for the same group dropped from 64% to 41% over the same ten weeks, while thumbs-down rate barely moved.
Multiple choice
2. What's the real problem with treating thumbs-down rate as a leading indicator here?
  • A. It costs too much engineering time to track accurately.
  • B. It structurally undercounts, since most dissatisfied users never clear all three filters to click it.
  • C. It only works for apps with a free tier.
  • D. It measures the model's confidence score instead of the user's reaction.
Show hint
Think about what has to happen between a bad answer and a click.
Show answer
B. Notice, care, find the button, three filters most upset users never clear, so the raw rate reflects a small, biased slice.
Fill in the blank
3. Of the 232 confirmed-bad suggestions Cosmin reviewed, only about ___ percent produced a thumbs-down click.
Show hint
Check the recut chart under "Let's learn."
Show answer
4.5 percent (18 out of 232). Multiplying the three filters together is what turns "58% noticed" into a number that small.
Multiple choice
4. Which of these is NOT one of the three AI-specific cause candidates named in this answer?
  • A. Silent abandonment, no click of any kind.
  • B. Trust erosion, users who stop bothering to correct the model.
  • C. A server outage that took the app offline for a day.
  • D. A UI placement issue that suppresses the click rate on its own.
Show hint
Two of the four are about the user's behavior, one is about the button, and one doesn't appear in the answer at all.
Show answer
C. A server outage never comes up in this answer. It's a plausible-sounding distractor, not one of the three named candidates.
Short answer, apply it yourself
5. Think of an app or tool you use where you'd never bother clicking "report a problem," even if something genuinely went wrong. What's the real reason you wouldn't click it?
Show hint
Look for one of the three filters: you didn't notice, you didn't care enough, or you couldn't find it.
Show answer
Model answer: A GPS app that gives a slightly wrong turn-by-turn instruction. Most drivers just correct course by eye and keep driving, they don't pull over to report it. The app's "report an issue" rate would stay near zero even during a genuinely bad patch of bad directions, because reporting takes more effort than just working around it.
Short answer
6. Why wouldn't just lowering the alert threshold, say flagging concern at 2% instead of 3%, fix the problem with thumbs-down rate here?
Show hint
Ask whether the fix changes what the number is actually made of.
Show answer
Model answer: It's a dial, not a fix. The number is still built from the same three filters, so a lower threshold just moves the same blind spot to trip a little earlier. It doesn't change the fact that under five in a hundred bad suggestions ever reach a click.
Before you close the answer
Why this works
Tests whether you'll take a flat, healthy-looking number at face value or ask what it takes to actually produce that number. Most candidates stop at "vanity metric" without showing the recut that proves it.
Follow-up traps
"Couldn't the drop in return rate just be seasonal, nothing to do with the model swap?" Response: that's what the control group is for. The held-out group on the old model didn't show the same drop over the same ten weeks, same season, different model.

"Isn't a behavior proxy just as gameable as thumbs-down, teams could juice return rate with notifications?" Response: yes, which is why it's paired with thumbs-down, not swapped in for it. A number nudged up by notifications and a click rate that still won't move together tell a different story than either one alone.
If pressed
The golden set isn't static either. Coalridge rotates about ten percent of it back to Reya each quarter, because a set built to catch last year's common substitutions starts missing new pantry combinations the moment the ingredient list changes, and a stale golden set would let the next quiet regression through the same way this one got through.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more