ConceptFoundationalDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #2

Explain the difference between explicit and implicit feedback signals.

PICK the product is Kestrel Loop's Daily Mix, a playlist rebuilt every morning for each listener

Kestrel Loop builds Daily Mix, a playlist every listener gets fresh each morning, tuned from what they've played, skipped, replayed, and saved before. Marisol Cantu has run the recommendations team for four years.

The direct answer
Build the quality read around implicit signals, skips, replays, saves, since almost every listener leaves one whether they mean to or not. Keep the rare explicit rating too, but only as a small check on what an implicit signal actually means, never as the main read. A system tuned only on people who bother to rate is tuned on a tiny, biased slice of everyone who uses it.
Do this, in order
  1. Build the quality read around implicit signals first, not explicit ratings.Why: implicit signals are the only ones with enough volume to trust at scale.
  2. Keep a small explicit-rating sample running, purely to calibrate the implicit read.Why: explicit ratings tell you what an implicit action actually meant, which implicit data alone can't.
  3. Never read one implicit action alone as proof something was bad.Why: a skip can mean the song was bad, or just that the moment was wrong for it.
  4. Re-check the calibration regularly, not once at launch.Why: what a skip or a save means can drift as listening habits and the catalog both change.
  5. Watch for a shrinking explicit sample as an early warning sign on its own.Why: a thinning calibration sample means your checks on the implicit read are getting shakier too.
  6. Leave the ranking model's own math alone.Why: this decision is about which feedback to trust, not how the model scores a song.

How to answer this, stage by stage

Nobody is grading whether you can define two terms. They're grading whether you'd actually commit to building a system around one of them.

Stage 1
Scope it to one real system
Say it like this
"I'll answer this for Kestrel Loop's Daily Mix, a playlist a listener gets fresh every morning."
Why this works
Grounds a textbook-sounding question in one real, checkable system.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my pick before any reasoning. Impact, who feels each kind of error. Cost asymmetry, which error actually costs more. Kill criteria, what would change my mind."
Why this works
Signals a method for a question that sounds like a definition but is really a commitment.
Stage 3
Give the position, before any reasoning
Say it like this
"I'd build the quality read around implicit signals, skips, replays, saves, not the explicit thumbs rating. Explicit stays, but only as a small check on what implicit actually means."
Why this works
This is the direct answer, said as a real commitment instead of a list of pros and cons.
Stage 4
Name who feels each kind of error
Say it like this
"Trusting only explicit ratings tunes the whole system off the two percent of people who bother to rate, mostly the pickiest ones. Trusting implicit signals with no check risks reading an ordinary mood-switch skip as 'this song was bad' and quietly burying a good track."
Why this works
Shows both sides have a real cost, not just one obvious villain.
Stage 5
Name the cost asymmetry
Say it like this
"A shortage of explicit ratings is easy to spot, you just see a small number and know to be careful. A misread implicit signal is invisible. It looks exactly like real data and drives the whole model before anyone checks it."
Why this works
This is PICK's hardest step, and the one that actually decides the pick.
Stage 6
Give the kill criteria
Say it like this
"If a calibration check ever shows implicit and explicit signals disagreeing on the same songs more than about one time in five, I stop trusting the implicit read until we find out why."
Why this works
Shows the pick is a real bet, not a stubborn stance with no way to be proven wrong.
Stage 7
Close on the one line
Say it like this
"The real skill isn't picking implicit over explicit. It's never letting either one run the whole show alone."
Why this works
Restates the direct answer in one breath, ready for a follow-up push.

Let's learn

Four years ago, Marisol Cantu watched roughly one in three of Daily Mix's earliest listeners tap a thumbs-up or thumbs-down after almost every song.

Kestrel Loop's Daily Mix is a playlist built fresh each morning for every listener, out of what they've played, skipped, and saved before.

Share of Daily Mix sessions producing each kind of signal, today
100% 50% 0% 2% Explicit rating 100% Implicit signal
Explicit ratings show up in 2 out of every 100 sessions. Implicit signal is sitting there in all of them, whether anyone taps anything or not.

In Daily Mix's first weeks, when only a small group of beta listeners used it, that rating rate held near thirty-four percent. Marisol's team tuned the model against those ratings, and it looked like it was working, because the people who rated were the same people who kept using it.

The extra silence was never the problem. It was Marisol's team, quietly deciding month after month that an old rating habit still meant what it used to mean, without checking whether the ninety-eight percent who never rated anything were being served just as well as the two percent who did.
Weekly explicit-rating rate since Daily Mix's national launch
40% 20% 0% 5% reliable-sample threshold Wk 1: 34% Wk 5: 19% Wk 9: 9% Wk 14: 4% Wk 20: 2%
Somewhere around week eleven, the rating rate slid under the level the team could still trust for a per-song check. Nobody marked the date.

At its worst: the model spends months confidently serving indie and alternative fans well, since that community rates more than most, while it quietly gets worse for casual pop listeners who never touched the thumbs button at all. Nobody notices until churn creeps up on a group with no complaint anywhere in the data to point to.

The decision I would take back We built the thumbs button as the trust signal and left it there as the default, because it was cheap to build and unambiguous to read, back when Daily Mix had a few thousand beta listeners who nearly all clicked it. That made sense when almost everyone rated. It stopped making sense once rating became the rare exception instead of the norm.

What I would leave alone: for a brand-new song with no play history yet, the rare explicit rating still matters a lot, on purpose, since there's no implicit pattern yet for anything else to read.

The lesson: explicit and implicit signals aren't a choice between a good one and a bad one. Implicit is what you build the whole system on, since it's the only one that shows up for everyone. Explicit is what tells you whether you're reading implicit correctly, and that job only works if you keep checking it.

Now here is the same thing as a story

The short version above is what you'd say defending this design to Kestrel Loop's leadership. Read this one for how the gap actually got found.

Marisol Cantu's laptop has had a cracked hinge for two years; she still opens it every night after her kids are asleep to check Daily Mix's numbers before bed.

For most of those four years, the numbers looked fine. The thumbs-rating rate was never huge, but whenever Marisol pulled up the ratio of thumbs-up to thumbs-down, it stayed comfortably positive, and the model kept getting better at pleasing the people who rated.

Hand sketched comparison diagram titled The asymmetry, drawn. Left panel, a document icon labeled Explicit rating, caption rare but clear. Right panel, a gauge icon labeled Implicit signal, caption constant but noisy.
One kind of signal is honest and rare. The other is constant and needs a translator.

The habit thinned in three beats nobody clocked at the time. First, only the pickiest listeners bothered to rate a song they merely liked. Then, only the sharpest dislikes got a thumbs-down, since a shrug-worthy song wasn't worth reaching for the button. Then, on most mornings, Daily Mix ran its full course for a listener and produced not one single rating, and that stopped being unusual at all.

The trigger was small. A data analyst, reviewing an unrelated dashboard, mentioned to Marisol over coffee, "You know only two percent of sessions ever produce a rating, right? We're tuning a model for millions of people off the opinions of a few thousand who happen to tap a button."

Knowledge spark: what makes an implicit signal harder to read than an explicit one? An explicit rating tells you what someone meant. An implicit action, like a skip, only tells you what someone did, and a skip can mean the song was bad, or that a call came in, or that their mood just shifted. The action is easy to log. The meaning behind it takes real work to pin down.

Marisol had heard some version of that line before, in passing, and let it go. This time she pulled the actual numbers, and found the rating rate had drifted from thirty-four percent at launch down to two percent at scale, while implicit signals, skip, replay, save, had sat there with full coverage the entire time, mostly unused for anything but raw ranking.

Hand sketched flow diagram titled How a skip becomes a signal. Five boxes: track plays, listener skips early, system logs it, checks session pattern highlighted, scored maybe bad.
The step everyone skipped for years is the one in the middle: checking what the skip actually meant.

Run the same drift forward under the old design: the ratings dashboard would have kept looking calm for months, since a small, steady trickle of ratings from the same loyal fans doesn't look like a crisis on a line chart. Nothing about a shrinking sample size sets off an alarm by itself.

Hand sketched quadrant titled Sorting signals by volume and clarity. Axes how often it happens from rare to constant, and how clear the meaning is from ambiguous to obvious. Explicit thumbs sits top left, rare and obvious. Replay same song sits top right, constant and obvious. Skip in 3 seconds sits middle right. Skip mid song sits bottom right, constant and ambiguous.
Only one signal is both constant and obvious on its own. Everything else needs a second signal to read it right.

With the redesign, Daily Mix's quality read now runs primarily off implicit signals, weighted by session context, and a rolling calibration sample checks those implicit reads against whatever explicit ratings still come in, refreshed every week instead of trusted forever.

Hand sketched labeled parts diagram titled Whats inside a weekly calibration check. Center document icon labeled Calibration sample, with four callouts: sessions with both signals, agreement rate, disagreement flagged, weekly refresh.
Four parts, and the last one, the weekly refresh, is what keeps the whole check from going stale the same way the old rating habit did.

Marisol reran the same three-month window that had gone unnoticed before. This time, the calibration check flagged a drift in how "skip within three seconds" was being read for a newer genre of upbeat, short-intro tracks, catching it in the second week instead of never.

The old design asked a shrinking rating count to keep doing a job it was never built to do at this size. The new one asks the rating for something smaller and more honest: not "is this song good," but "does our reading of a skip still match what a real person meant by it."

I built the thumbs button first because it was the cheapest thing to ship, and for a long while, cheap and correct looked like the same thing. It took a colleague's offhand remark over coffee to see that the button hadn't stopped working. It had just stopped being anyone's.

PICK, in one pageNot a debate over which signal is "better." PICK is what forces you to commit to one and defend the asymmetry.

P
Position. The pick, before any reasoning.
Build the quality read around implicit signals. Keep explicit ratings, but only to calibrate what implicit actually means.
Committing first is what separates a real answer from "it depends."
I
Impact. Who feels each kind of error.
Explicit-only tunes the system off the pickiest two percent. Implicit-only risks misreading an ordinary mood-switch skip as a bad song.
Both sides get named, so the pick isn't a strawman.
C
Cost asymmetry. Which error actually costs more.
A shortage of explicit ratings is easy to spot. A misread implicit signal is invisible, and it drives the whole model before anyone checks it.
The hardest step, and the one the entire pick rests on.
K
Kill criteria. What would change the pick.
If implicit and explicit disagree on the same songs more than one time in five in the weekly check, stop trusting the implicit read until it's understood.
Turns a confident answer into a testable bet, not a stubborn one.

The recap, one line per letter: position is implicit-first with explicit as a check, impact is the two groups each kind of error hits, cost asymmetry is the invisible misread beating the visible shortage, and kill criteria is the one-in-five disagreement rate that would flip the pick.

And if you want to be sure it really works, try it somewhere elseSame four letters, a veterinary clinic's AI visit summaries instead of a playlist. A different field, the same asymmetry.

Hemlow Veterinary Partners uses an AI tool to draft the after-visit summary a pet owner receives once their appointment ends. Junot Alvear is the product owner for that tool.

Mapped onto PICK: position is the same, build the quality read around implicit signals, specifically how often a vet edits the AI draft before signing it, and how often an owner calls back with a question about something the summary should have already answered, while keeping the rare owner star-rating as a calibration check. Impact: trusting only the star rating tunes the system off the small share of owners who bother to rate at all, usually the ones who had either an unusually great or unusually bad visit. Trusting edits and callbacks alone, with no check, risks reading a vet's stylistic tweak as proof the AI draft was medically wrong, when it might have just been a wording preference. Cost asymmetry: a thin star-rating sample is easy to notice and account for. A misread edit pattern, treating harmless rewording as a real accuracy problem, is invisible and can quietly push the drafting tool in the wrong direction. Kill criteria restates the same one-in-five disagreement rule, checked against a monthly sample of visits with both signals present.

Hand sketched decision tree titled Explicit or implicit, for a vet visit summary. Root, new AI visit summary drafted. Branches: owner leaves a star rating leads to explicit rare, vet edits before signing leads to implicit common, owner calls with a question leads to implicit common, summary sent unedited leads to no signal at all.
Three of the four paths through this tree produce a signal. Only one of them is explicit, and it's the rarest branch on the page.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "trust implicit signals for coverage, use the rare explicit rating only to calibrate what implicit actually means," and stop.
Cost: there's no budget this quarter for a full calibration pipeline. Say so honestly, and ship a simple weekly spot-check on a small sample first, since a thin calibration still beats none at all.
The model gets better, for real: if skip-prediction accuracy improves, that's still not a reason to drop the explicit check, a better model can still be confidently wrong on the exact case nobody weighed in on explicitly.

Where people run it wrong.
They treat every explicit rating as gospel and every implicit signal as noise, when at scale it's often closer to the reverse.
They read a single implicit action alone as certain, with no session context to check it against.
They build the calibration sample once at launch and never refresh it, so the very thing it's supposed to protect against drifts too.

How to use it live. When someone asks "explicit or implicit," ask yourself which one has enough volume to trust at scale first. That's usually implicit. Then ask what would make you doubt it. That's your calibration sample. Both questions buy a few seconds of thinking time and land on a real, defensible answer.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an explicit-vs-implicit tradeoff question?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. The cost asymmetry step decides which signal actually carries the weight.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Marisol Cantu, who has run Kestrel Loop's recommendations team for four years and tunes the Daily Mix model.
3 · THE POSITION
What's the pick here, before any reasoning?
Tap to flip
ANSWER
Build the quality read around implicit signals, and keep the rare explicit rating only as a calibration check, never as the main signal.
4 · THE COST ASYMMETRY
Which kind of error costs more, and why?
Tap to flip
ANSWER
A misread implicit signal, since it looks exactly like real data and drives the whole model before anyone checks it. A shortage of explicit ratings is easy to spot by comparison.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Leaving the thumbs button as the default trust signal long after most listeners stopped using it, since it made sense only while nearly everyone rated.
6 · THE NUMBER
Fill in the blank: the explicit-rating rate fell from 34% at launch to ___% at scale.
Tap to flip
ANSWER
2%. Implicit signal coverage stayed at 100% of sessions the entire time, mostly unused for calibration.
7 · THE REPLAY
Same three-month drift, redesigned system. What changes?
Tap to flip
ANSWER
The weekly calibration check flags the drift in how a fast skip on short-intro tracks gets read in its second week, instead of never catching it at all.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the pick there?
Tap to flip
ANSWER
Hemlow Veterinary Partners' AI visit summaries. Same pick: vet edits and owner callbacks as the primary signal, the rare star rating as the calibration check.

Check yourself Score: 0 / 0

Multiple choice
1. Why does this answer build the quality read around implicit signals instead of explicit ratings?
  • A. Because explicit ratings are technically harder to store in a database.
  • B. Because implicit signals show up in nearly every session, while explicit ratings come from a tiny, biased slice of listeners who bother to rate.
  • C. Because listeners find the thumbs button visually unappealing.
  • D. Because implicit signals are always more accurate than explicit ones.
Show hint
Look at the "share of sessions" bar chart.
Show answer
B. Coverage is the real reason, not accuracy. Implicit signals aren't automatically more correct, they're just present far more often, which is why they need a calibration check rather than blind trust.
True or false
2. True or false: a single skip early in a song is, by itself, reliable proof that the song was a bad recommendation.
  • True
  • False
Show hint
Look at the knowledge spark about what makes implicit signals harder to read.
Show answer
False. A skip could mean the song was bad, or it could mean a call came in or the listener's mood shifted. That's exactly why implicit signals need session context and a calibration check, not blind trust.
Fill in the blank
3. Fill in the blank: by week twenty, Daily Mix's explicit-rating rate had fallen to about ___ percent of sessions.
Show hint
Look at the line chart tracking the weekly rating rate.
Show answer
2 percent. It crossed below the team's 5 percent reliable-sample threshold somewhere around week eleven, without anyone marking the date.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Leaving the thumbs button as the default trust signal. It made sense when nearly every beta listener rated, and stopped making sense once rating became the rare exception rather than the norm.
Short answer, where it wouldn't matter
5. Name a case in Daily Mix where an explicit rating still matters more than an implicit one.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A brand-new song with no play history yet. There's no implicit pattern to read yet, so the rare explicit rating carries real weight there.
Short answer, apply it yourself
6. Pick a product you use yourself. What's one implicit signal it could read from your behavior that you'd never bother to say out loud as a rating?
Show hint
Think about something you do without thinking, like re-reading a text or closing an app fast.
Show answer
Model answer: Many apps could read re-opening the same screen, abandoning a cart, or re-reading a message twice as a signal, all things a person would never bother rating explicitly.
Before you close the answer
Why this works
Tests whether you can hold two feedback types in tension instead of picking a "best" one and moving on, and whether you know which kind of error a design should actually be optimized against.
Follow-up traps
"Isn't explicit feedback obviously more trustworthy, since a person said it on purpose?" Response: on purpose doesn't mean unbiased. Only people who bother to rate show up in that data, and that's a very specific, self-selecting slice of everyone using the product.

"What if implicit signals get gamed too, like repeated skips aimed at burying a competitor's song?" Response: they can, which is exactly why implicit signals need their own guardrail too, checked against the small explicit sample and account-level rate limits, never trusted blind either.
If pressed
Kestrel Loop's real calibration sample only counts a rating and a nearby implicit action as "agreeing" when both happen in the same session, not any implicit action from that listener ever, since an older session may reflect a mood that's already changed by the time the rating comes in.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more