ConceptAdvancedResponsible AI & Advanced Practice / Responsible AI as a product requirement / #19

What responsible AI metric would you put on your team dashboard?

LEAD the product is Northlight Stream's recommendation engine, the "up next" rail on every title page

Northlight Stream is a video streaming service. Its recommendation engine builds the "up next" rail that appears the second a title ends. Tomasz Adamik is the product manager who owns that engine, and the one Friday he pulled a metric nobody had ever put on a dashboard before.

The direct answer
Put source diversity on the dashboard: the number of distinct shows or channels a viewer's recommendations actually pull from each week, not just how much they watched. Watch time and satisfaction scores stay flat while a viewer gets quietly boxed into the same five shows, and by the time complaints show up, the narrowing has been running for weeks.
Do this, in order
  1. Track source diversity per viewer, per rolling week, as its own line on the dashboard.Why: it's the number that moves first, while watch time and satisfaction still look completely healthy.
  2. Split it by cohort, not just as one company-wide average.Why: a shrinking average can hide one heavily-affected group inside a company-wide number that still looks fine.
  3. Set a real alert threshold, not just a number someone glances at monthly.Why: a metric with no threshold and no owner is a chart, not an early warning system.
  4. Watch for the metric being gamed by counting technically-different-but-similar content as "diverse."Why: any metric can be satisfied on paper without the real thing, a viewer's own sense of variety, actually improving.
  5. Pair it with the lagging outcome, like 1-star reviews mentioning repetition, to prove the leading signal was real.Why: a leading indicator nobody's checked against the outcome it's supposed to predict is just a guess with a graph.
  6. Don't force every viewer toward maximum diversity regardless of what they actually want.Why: some viewers genuinely want to rewatch the same five comfort shows on a loop, and that's a preference, not a failure.

How to answer this, stage by stageSeven moves. Say your structure out loud so the interviewer knows you're hunting for a leading signal, not a vanity number.

Stage 1
Ground it in one real system
Say it like this
"I'll answer this for a streaming service's recommendation engine, the 'up next' rail that shows the moment a title ends."
Why this works
"Responsible AI metric" is abstract until it's tied to one real screen a viewer actually looks at.
Stage 2
Name the framework
Say it like this
"I'll use LEAD. Link to the real outcome, early signal that moves first, abuse, how it gets gamed, and decision, what we'd actually do at each threshold."
Why this works
Signals a search for a leading indicator, not just a list of nice-sounding metrics.
Stage 3
Name the real outcome first
Say it like this
"The outcome that actually matters isn't watch time. It's whether a viewer feels like the platform still knows them, week after week, instead of narrowing them into a rut."
Why this works
Without naming the real outcome, any metric proposal is just a number in search of a reason.
Stage 4
Give the leading signal, concretely
Say it like this
"Source diversity: how many distinct shows or channels a viewer's recommendations actually pull from in a rolling week. It's the number that would have started dropping weeks before anyone complained."
Why this works
This is the actual answer to the question, a specific, measurable metric, not a vague value statement.
Stage 5
Prove it with the real timeline
Say it like this
"At Northlight, source diversity started falling in week one of a recommendation tweak. 1-star reviews mentioning repetitive content didn't climb until week nine. Watch time never dropped at all."
Why this works
A concrete lead time, in weeks, is what makes "leading indicator" a real claim instead of a slogan.
Stage 6
Say how it gets gamed
Say it like this
"A team could satisfy this on paper by counting near-identical shows under different titles as 'diverse' sources. That's why the metric needs a real content-similarity check behind it, not just a distinct-ID count."
Why this works
Every metric has a cheap way to be hit without doing the real work. Naming it shows you'd actually watch for it.
Stage 7
Say what you wouldn't force, and close
Say it like this
"I wouldn't force diversity onto a viewer who's rewatching the same five comfort shows on purpose. The metric flags a company-wide narrowing pattern, it doesn't override one person's actual preference."
Why this works
Shows judgment: the metric targets a real risk, not a blanket rule applied regardless of what people actually want.

Let's learn

Say a streaming service builds a rail of "up next" picks that appears the instant a show ends, so a viewer never has to go back to a menu at all.

For most viewers, most of the time, that rail works exactly as intended: it saves the two minutes of browsing they used to spend hunting for what to watch next.

Knowledge spark: what's a leading indicator? A number that starts moving before the thing you actually care about does. A lagging indicator, like a customer complaint or a churn number, tells you something already went wrong. A leading indicator tells you it's starting to go wrong, while there's still time to fix it quietly.

Northlight's team watched two numbers closely at launch: watch time, and an overall satisfaction score gathered through occasional surveys. Both looked healthy for months.

Hand sketched flow diagram titled From raw signal to dashboard. Four boxes: sessions logged, diversity scored highlighted, weekly rollup, dashboard.
This pipeline didn't exist yet. Sessions were logged, but nothing scored how many distinct sources they actually pulled from.

Then a change meant to boost short-term watch time, weighting recommendations more heavily toward whatever a viewer had watched most recently, quietly began narrowing what most viewers were shown week over week.

Source diversity, weekly, ten weeks after the recommendation change
14 7 0 14 5 Week 1 Week 10
This number started dropping the same week the recommendation change shipped. Nobody was watching it, because it didn't exist on a dashboard yet.

At its worst: a viewer who used to get a genuine mix of documentaries, comedy, and drama found her rail narrowing, week by week, down to five true-crime titles that all looked nearly identical, without her ever changing a single setting.

The decision I would take back We tracked watch time and an occasional satisfaction survey because those were the numbers the whole company already watched, and they were easy to defend in a planning review. That made sense while the recommendation engine changed slowly. It stopped making sense the moment a single tuning change could quietly narrow what people saw for weeks before either of those numbers so much as flinched.

What I would leave alone: a viewer who's deliberately rewatching the same five comfort shows on a rainy weekend doesn't need this metric flagging their account. The signal is meant to catch a company-wide narrowing pattern, not override one person's actual, chosen preference.

Watch time never told us anything was wrong. It couldn't. A viewer who's shown the same five shows on repeat often watches just as much, sometimes more, right up until the week she finally gets tired of it and says so.

The lesson: a metric that only moves after someone's already upset isn't measuring the problem. It's measuring how long people tolerate a problem before they say something.

Now here is the same thing as a story

The short version above is what you'd say defending this dashboard addition to Northlight's leadership. Read this one for how the gap actually got found.

For most of two years, Tomasz Adamik's job ran off a shared dashboard TV mounted in the team's open-plan area, two numbers front and center: watch time, and a quarterly satisfaction score.

Hand sketched comparison diagram titled Satisfied on paper, not in the room. Left panel, gauge icon labeled Metric says fine, caption average diversity looks stable. Right panel, person icon labeled Viewer walks away, caption same five channels every night.
The dashboard TV showed the left panel for two full months. The right panel was happening the entire time, just not on any screen anyone was watching.

Both numbers looked exactly the same, month after month, even after a tuning change shipped that was meant to boost short-term engagement by weighting recommendations toward a viewer's most recent watches.

Hand sketched timeline titled When each signal would have rung. Four milestones: diversity starts falling week 1 highlighted, session chains lengthen week 4, 1-star reviews climb week 9, press picks it up week 11.
Eight weeks separate the moment the real problem started and the moment anyone official noticed.

A new data scientist joined the team in week nine and, looking for something to explore in her first month, asked a question nobody else had thought to ask out loud: "do we track how many different things we actually recommend to the same person, or just how much they watch?"

Hand sketched labeled parts diagram titled What the metric definition needs. Center gauge icon labeled Source Diversity, with four callouts: unique sources per session, rolling 7 day window, per cohort split, alert threshold.
Four parts, and none of them had existed before that question got asked out loud.

Nobody had an answer. So she built one, pulling two months of session logs to count distinct sources per viewer per week, and the chart she brought back showed a decline that had been running since the very week of the tuning change.

1-star reviews mentioning "repetitive" or "same content," before vs. after diversity started dropping
60 30 0 4 Before decline (month 1) 61 After decline (month 3)
By the time this lagging number moved, the leading signal had already been dropping for two months.

With source diversity on the dashboard now, the same kind of tuning change gets caught inside its first week: a per-cohort alert fires the moment average distinct sources drops below a set floor, well before any viewer thinks to write a review about it.

Hand sketched quadrant titled Leading vs lagging, for a rec system. Axes: how early it warns from late to early, how easy to measure from hard to easy. Source diversity sits far right and mid height. 1-star reviews and watch time sit upper left. Session chain length sits lower right.
The two easiest numbers to measure, watch time and reviews, both sit in the "warns late" half. The leading signal was harder to build, which is exactly why nobody had built it yet.

The old dashboard asked "are people still watching." The new one asks "are people still being shown something new," which is the question that was actually breaking, weeks before the first one showed a single crack.

I kept watch time and the satisfaction survey front and center because those were the numbers everyone already trusted and could defend in a planning review. It took a new hire's plain question, one nobody senior had thought to ask, to see that trusting a number doesn't make it the right one to watch.

LEAD, the signal that rings firstNot a KPI wishlist. LEAD is what forces you to find the number that moves before the damage does, not after.

L
Link. The real outcome, not the model's score.
Whether a viewer still feels like the platform knows them, not raw watch-time totals.
Every metric answer needs a named outcome or it's just a number in search of a reason.
E
Early signal. What moves first.
Source diversity, distinct shows or channels per viewer per rolling week, started dropping in week one, eight weeks before anything else moved.
The hardest step, and the actual answer to the question.
A
Abuse. How it gets gamed.
Counting near-identical shows under different titles as separate "sources" would satisfy the metric without giving viewers real variety.
Every metric has a cheap way to be hit; naming it shows you'd actually watch for it.
D
Decision. What we'd do at each threshold.
Below the floor for one cohort, roll back the specific tuning change for that cohort first, not the whole platform.
A metric nobody acts on is a dashboard decoration, not a real early-warning system.
Hand sketched icon list titled What makes a good raw metric. Three items: a gauge icon labeled moves weeks before the outcome does, a person icon labeled tied to a real viewer habit, a box icon labeled hard to satisfy without doing the real work.
Watch time and the satisfaction survey both fail the first test on this list. Source diversity is the one that passes all three.

The recap, one line per letter: link is a viewer still feeling known week to week, early signal is source diversity dropping eight weeks before anything else, abuse is near-identical shows counted as fake variety, and decision is a per-cohort rollback the moment the floor is crossed.

And if you want to be sure it really works, try it somewhere elseSame four letters, a waste management company instead of a streaming service. A completely different industry, and here the leading signal isn't about variety at all, it's about a truck's own confidence.

Cindergale Waste Recovery uses an AI system that plans daily pickup routes and flags bins likely to be contaminated with the wrong material. Consuelo Reyes manages the routing tool, and drivers get their flags through a two-way radio at the start of each shift.

Mapped onto LEAD: link is whether the contamination-sorting facility downstream actually runs efficiently, not just how many stops the routing tool completes on schedule. Early signal is the rate at which the model flags a bin as "uncertain" instead of confidently right or wrong, since a rising uncertain-flag rate means the model is quietly seeing more borderline cases than its training ever covered, weeks before the sorting facility's own contamination rate climbs. Abuse is a team tuning the model to just call more borderline cases "confidently clean" to make the uncertain rate look lower, which would hide the real problem instead of fixing it. Decision is that crossing an uncertain-rate threshold for a route triggers a manual spot-check of that route's next five pickups, not a company-wide retraining effort.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "source diversity per viewer per week, it drops eight weeks before anything else does," and stop.
Cost: there's no engineering time this sprint to build the per-cohort alert. Say so honestly, and start with a monthly manual pull of the same number, even without automated alerting, since a slow signal beats no signal.
The model gets better, for real: if overall recommendation accuracy improves, that's still not proof diversity is healthy, a model can get better at predicting what someone will click while narrowing what it ever shows them.

Where people run it wrong.
They put watch time or a satisfaction survey front and center because those are the numbers everyone already trusts, not because they warn early.
They track one company-wide average and miss a cohort quietly cratering underneath a stable-looking overall number.
They build the metric and never pair it with the lagging outcome, so nobody can actually prove the leading signal predicted anything real.

How to use it live. When someone asks what metric belongs on your dashboard, don't answer with the number that's easiest to explain. Ask yourself: what would have looked perfectly healthy right up until the morning everything went wrong. That number, however inconvenient to build, is the one that actually belongs there.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what responsible AI metric would you put on your team dashboard"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Early signal is the metric that would have looked healthy right up until things broke.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Tomasz Adamik, PM for Northlight Stream's recommendation engine, whose team dashboard TV showed only watch time and a quarterly survey.
3 · THE HABIT
What did the team stop doing while watch time looked healthy?
Tap to flip
ANSWER
They stopped asking whether recommendations were actually getting more varied or narrower, since the only two numbers they watched couldn't tell them either way.
4 · THE EARLY SIGNAL
What's the leading metric this answer proposes, in one line?
Tap to flip
ANSWER
Source diversity: the number of distinct shows or channels a viewer's recommendations pull from in a rolling week.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Tracking only watch time and an occasional satisfaction survey, since those were the numbers everyone already trusted and could defend in a planning review.
6 · THE NUMBER
Fill in the blank: source diversity dropped from 14 to ___ distinct sources per viewer over ten weeks.
Tap to flip
ANSWER
5. It started falling the same week a recommendation tuning change shipped, eight weeks before 1-star reviews began climbing.
7 · THE REPLAY
Same tuning change, with source diversity on the dashboard. What changes?
Tap to flip
ANSWER
A per-cohort alert fires within the first week the average drops below the set floor, instead of waiting eight weeks for reviews to catch it.
8 · CROSS PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
Cindergale Waste Recovery's contamination-flagging routing tool. The early signal there is the model's own rising "uncertain" flag rate, not a lagging contamination number.

Check yourself Score: 0 / 0

Multiple choice
1. Why does this answer reject watch time as the metric to put on the dashboard?
  • A. Watch time is too hard to measure accurately.
  • B. Watch time can stay flat even while recommendations quietly narrow, so it never warns early.
  • C. Watch time only applies to free-tier viewers.
  • D. Watch time always goes up when a model gets worse.
Show hint
Look at the highlight block about watch time and the quadrant comparing leading and lagging signals.
Show answer
B. A viewer narrowed to the same five shows can watch just as much, so watch time tells you nothing is wrong until she finally says so.
True or false
2. True or false: this answer recommends forcing every viewer's recommendations toward maximum diversity, even ones who prefer to rewatch familiar shows.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. A viewer deliberately rewatching comfort shows isn't flagged by this metric. It's meant to catch a company-wide narrowing pattern, not override individual preference.
Fill in the blank
3. Fill in the blank: 1-star reviews mentioning repetitive content climbed from 4 per month to ___ per month, two months after diversity started dropping.
Show hint
Look at the bar chart comparing before and after the diversity decline.
Show answer
61. By the time this lagging number moved, the leading signal had already been dropping for two full months.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Tracking only watch time and an occasional satisfaction survey. It made sense while the recommendation engine changed slowly enough that neither number needed to move fast.
Short answer, apply it yourself
5. Pick an app that recommends things to you. What's one leading signal it could track that would warn it's narrowing your options, before you'd ever notice or complain?
Show hint
Think about a music app, a shopping site, or a news feed, and what "narrowing" would actually look like there.
Show answer
Model answer: Most people land on something like "number of distinct artists, sellers, or sources shown per week," the same shape as source diversity, just renamed for their own app.
Before you close the answer
Why this works
Tests whether you can find the number that would have looked completely fine right up until the moment things broke, instead of defaulting to whichever metric is easiest to explain in a meeting.
Follow-up traps
"Isn't source diversity just going to hurt engagement if you force it?" Response: the metric triggers a review, not an automatic override, and it never applies to a viewer who's deliberately rewatching familiar shows on purpose.

"Couldn't a team just game this by recommending random unrelated content to boost the diversity number?" Response: that's exactly the abuse case this answer names, which is why the metric needs a real content-similarity and relevance check behind the raw source count, not just a distinct-ID tally.
If pressed
Northlight's real fix ties the per-cohort alert threshold to each cohort's own historical baseline rather than one company-wide number, since a documentary-heavy viewer and a sitcom-rewatcher have very different normal diversity levels to begin with.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more