ConceptAdvancedResponsible AI & Advanced Practice / Responsible AI as a product requirement / #10

What does responsible AI require in the eval spec specifically?

LEAD the product is the Aldergate Advisor, a robo-advisor that recommends how customers should invest

Aldergate Wealth runs the Aldergate Advisor, a robo-advisor that recommends how a customer's savings should be split between stocks, bonds, and cash. Renata Kowalska is the product manager who owns the eval spec every new model version has to clear before it ships.

The direct answer
A responsible-AI eval spec needs a golden set stratified by the segments most likely to be harmed, a separate pass bar for each named harm category instead of one blended score, a named owner accountable for keeping both current, and a version pin recording exactly which model produced which result. A single overall accuracy number, no matter how high, is not a responsible-AI eval spec. It's a number that can hide the one group it was supposed to protect.
Do this, in order
  1. Stratify the golden set by the segments most likely to be harmed, not by convenience.Why: a randomly sampled set quietly under-represents exactly the group a blended score can't protect.
  2. Set a separate pass bar per harm category instead of one blended score.Why: a 96% blended average can sit on top of a 71% failure rate in the one segment that matters most.
  3. Name an owner accountable for the spec, not just the model.Why: a spec nobody owns quietly goes stale while the product keeps changing underneath it.
  4. Pin every result to an exact model version and date.Why: a passing score with no version pin can't be reproduced, defended, or trusted later.
  5. Re-check the segment-level bars every release, not just at launch.Why: a segment that passed once can quietly slip as later versions optimize for the average case.
  6. Report the segment scores next to the blended score, always, even when they're ugly.Why: a number nobody has to look at stops getting looked at.

How to answer this, stage by stage

Seven moves. Practice these until the words come without reaching for them.

Stage 1
Ground the question in one real spec
Say it like this
"I'll answer this for a robo-advisor's eval spec, since 'responsible AI' means something concrete there: whether the advice is suitable for who's actually asking."
Why this works
Keeps the answer from turning into a list of buzzwords with nothing underneath them.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link to the real outcome, find the early signal, name how it gets gamed, then say what decision each threshold triggers."
Why this works
Signals a real method instead of a list recited from memory.
Stage 3
Link to the outcome that actually matters
Say it like this
"The outcome isn't 'model accuracy.' It's whether a customer near retirement gets an allocation suited to someone who can't afford a bad year."
Why this works
Grounds "responsible AI" in a real person's real risk, not an abstraction.
Stage 4
Name the early signal
Say it like this
"The signal that moves first is the pass rate for the near-retirement segment specifically, tracked on its own. It can be sliding for months while the blended score still looks fine."
Why this works
This is the actual answer to "what does the eval spec require": a segment-level number, not just a blended one.
Stage 5
Name how it gets gamed
Say it like this
"A blended pass rate gets hit by over-sampling the easy cases. If the golden set is 96% young savers, of course the overall number looks great."
Why this works
Shows you understand a metric can be true and still be misleading.
Stage 6
Say what each threshold triggers
Say it like this
"If any single segment's pass rate drops below 90%, the release is blocked, no matter what the blended score says. That's a hard rule, not a suggestion."
Why this works
Turns a metric into a real decision instead of a dashboard nobody acts on.
Stage 7
Close on the one line
Say it like this
"So: a stratified golden set, a bar per harm category, a named owner, a version pin. A single blended score is not a responsible-AI eval spec, it's the thing a responsible one replaces."
Why this works
Leaves the interviewer with the actual answer, said plainly, one more time.

Let's learn

Say a robo-advisor recommends how to split a customer's savings between stocks, bonds, and cash. Every quarter, a new model version ships, and every version has to clear an eval spec first.

For a long time, that spec was one number: run the new model against a golden set of 2,000 real customer conversations, check how often the recommended allocation matched what a licensed advisor would have suggested, and if the match rate cleared 95%, ship it. Version after version cleared comfortably, usually around 96%.

Knowledge spark: what's a golden set? A fixed collection of example cases with a known correct answer, used to check a model before it ships. If the golden set doesn't look like the real customers, a high score on it doesn't mean much.

Nobody had asked how those 2,000 conversations were chosen. They were pulled at random from real traffic, and real traffic skewed young: most Aldergate customers are in their 30s and 40s, still adding to their savings every month. Only about 4% of the golden set represented someone within five years of retirement, starting to draw money out instead of putting it in, even though that group made up roughly 20% of actual customers.

Here is the turn. The 96% pass rate was never wrong, exactly. It was just answering a question nobody needed answered as urgently: how good is the model at the easy majority. It was silent on the one group where a bad recommendation does the most damage in the least recoverable way.

Pass rate: blended score vs the near-retirement segment, by model version
100% 50% 0 Jan Feb Mar Apr May Jun Blended Near retirement
The blended line barely moves all year. The near-retirement line falls for four straight months before anyone outside the eval team saw it.

At its worst: a 61-year-old customer, five years from retiring, gets recommended an allocation built for someone decades younger, loses a chunk of it in a downturn with no time left to recover, and the eval dashboard says everything passed.

The decision I would take back We built one golden set, sampled at random from real conversations, and kept no separate record of how each customer segment performed inside it. That was fine while the product mostly served people decades from retiring, since a wrong recommendation there has years to correct itself. It stopped being fine the moment a meaningfully different, higher-stakes segment existed and nobody was tracking it on its own.

What I would leave alone: the blended score itself isn't the problem, and dropping it would be a mistake too. It's still useful for catching a broad regression across the whole product. The fix is adding segment bars next to it, not replacing it.

The problem was never that the model got a near-retirement case wrong. It's that nobody had a number whose job was to notice.

The lesson: "responsible AI" in an eval spec isn't a separate checklist item bolted onto the real one. It's the decision to measure the group most at risk on its own, instead of letting it disappear into an average.

Now here is the same thing as a story

The short version above is what you'd say in the room. Read this one for how Renata actually found the gap.

Renata Kowalska has spent four years building eval infrastructure for financial products, the last two of them at Aldergate. She is the person other PMs call when a metric looks too good to be true.

Hand sketched flow diagram titled How a new model version gets checked. Five boxes: new version, stratify by segment highlighted, check each bar, ship, or block.
The second box is the one that didn't exist for the first six model versions Aldergate ever shipped.

For those first six versions, the process worked exactly as designed: run the golden set, check the blended score, ship at 95% or above. Every version cleared. Nobody had a reason to look closer, because nothing on the dashboard was telling them to.

Then a regulatory inquiry landed on Renata's desk. A 61-year-old customer had filed a complaint after a downturn wiped out nearly a fifth of his retirement savings, invested at an allocation the Advisor had recommended only four months before he planned to retire. The allocation would have made sense for someone with thirty years to recover. He had less than one.

Hand sketched comparison diagram titled What the dashboard showed vs what was true. Left panel, a gauge icon labeled Blended score, caption 96 percent looks fine. Right panel, a person icon labeled Near retirement, caption 71 percent nobody counted it.
Both numbers were true at the same time. Only one of them was being watched.

Renata pulled the golden set apart by segment for the first time since it was built. What she found was not a bug in the model. It was a hole in the eval spec: near-retirement cases made up a fifth of real customers and a fraction of that in the golden set, and inside that small slice, the real pass rate had been sliding for months, invisible under a blended average that had barely moved.

Hand sketched quadrant titled Which cases the golden set forgot. Axes share of the golden set and harm if the advice is wrong. Near retirement and sole survivor sit top left, severe harm and barely covered. Young saver sits bottom right, small harm and well covered.
The two highest-harm segments sit in exactly the corner the golden set almost never reached.

We did not build a model that failed near-retirement customers. We built a spec that never had to notice if it did.

Renata rebuilt the spec around four pieces instead of one number: a golden set deliberately stratified so every harm-prone segment gets real weight regardless of how rare it is in daily traffic, a separate pass bar for each of those segments, herself named as the spec's owner with a standing quarterly review on her calendar, and every passing result pinned to the exact model version and date that earned it.

Hand sketched labeled parts diagram titled What's in the eval spec. Center document icon labeled Eval Spec, with four callouts: stratified set, bar per harm, named owner, version pin.
Four pieces. The old spec only ever had the first half of the first one.
Hand sketched metaphor scene titled The signal that moves first. Left, a gauge icon labeled Segment Bar, caption sliding for months. Right, a document icon labeled Complaints, caption still flat for now.
The segment bar was already telling the story. The complaint count just hadn't caught up to it yet.

The next model version failed the new spec on its first run, on the near-retirement bar specifically, at 74%. It stayed home two extra weeks while the team retrained specifically on decumulation cases. The version after that cleared every segment bar at 90% or above, blended score included.

What I would tell myself a year earlier: a 96% pass rate was never proof of anything on its own. It was an average, and an average's whole job is to hide exactly the thing you actually need to see.

LEAD, the spec Renata actually builtNot "measure accuracy." LEAD forces you to find the number that would have warned you first.

L
Link. The outcome that actually matters.
Not model accuracy. Whether a customer near retirement gets an allocation suited to someone who can't absorb a bad year.
Grounds the whole spec in a real risk, not a model score.
E
Early signal. The reason LEAD exists.
The near-retirement segment's own pass rate, tracked apart from the blend. It fell for four months while the blended score barely moved.
The actual answer to "what does the spec need": a number that moves before the harm does.
A
Abuse. How it gets gamed.
A golden set sampled at random from traffic quietly under-represents the highest-harm, lowest-volume segment, inflating the blended score without anyone doing anything dishonest.
Shows the metric can be gamed by convenience, not just by bad intent.
D
Decision. What each threshold triggers.
Any single segment bar under 90% blocks the release, regardless of the blended score.
Turns the metric into a real gate instead of a dashboard nobody acts on.
Customer complaints about unsuitable allocations, by quarter
14 7 0 Q1 Q2 Q3 Q4 3 5 14 2
Complaints stayed low for two full quarters after the segment pass rate had already started sliding. The lagging number only caught up once the damage had compounded.

The recap, one line per letter: link is a suitable allocation, not a model score, early signal is the segment pass rate falling months before any complaint, abuse is a random golden set quietly under-weighting the highest-harm group, and decision is a hard per-segment bar that blocks a release the blended score alone would have passed.

And if you want to be sure it really works, try it somewhere elseSame four letters, a staffing agency's resume screener instead of a wealth app. A completely different field.

Fenwick Staffing runs Fenwick Screen, an AI tool that ranks job applicants for recruiters before a human reviews the shortlist. Dmitri Kessler owns the eval spec that gates every new ranking model.

Mapped onto LEAD: link is not "ranking accuracy," it's whether qualified candidates from every demographic group actually reach the human-reviewed shortlist at a similar rate. Early signal is the shortlist rate broken out by protected group, tracked on its own, since it can drift for months while the overall "top candidates match hiring manager picks" score stays flat. Abuse: a golden set built from the agency's own historical placements quietly bakes in whatever imbalance existed in who got hired before, so a high match score can mean "matches old bias" as easily as "found the best candidate." Decision: if any group's shortlist rate falls more than 15 points below the overall rate, the release is blocked and the ranking model is re-audited before it ships to any client.

The same labeled parts pattern reused for a second product: a golden set stratified by group, a bar per harm category, a named owner, and a version pin, this time for a resume screening tool.
Same four pieces again. A hiring tool needs them for exactly the same reason a robo-advisor does.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a stratified golden set, a bar per harm category, a named owner, a version pin, not one blended score," and stop.
Cost: there's no budget to build a fully custom stratified set this quarter. Say so honestly, and start by re-weighting the existing set so the highest-harm segment isn't under 5% of it, even before better data collection catches up.
The model gets better, for real: if the blended score climbs to 99%, that's still not a reason to drop the segment bars. A near-perfect average model can still be quietly wrong for the one group whose mistakes cost the most.

Where people run it wrong.
They chase one blended number because it's simpler to report to leadership, and simplicity quietly wins over honesty.
They build the golden set from whatever data is easiest to pull, instead of the data that actually represents who could get hurt.
They check the spec once at launch and never again, so a segment that passed at launch can fail two versions later with nobody watching.

How to use it live. If someone asks "isn't 96% good enough," ask back: 96% of what set, built how. That question alone usually tells you whether the eval spec in front of you is a real one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what does responsible AI require in the eval spec specifically"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Early signal is the step that answers the question directly.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Renata Kowalska, the product manager who owns the Aldergate Advisor's eval spec.
3 · THE HABIT
What did the team stop doing once the blended score kept passing?
Tap to flip
ANSWER
They stopped checking whether the golden set actually represented every customer segment, since the one overall number never gave them a reason to look.
4 · THE EARLY SIGNAL
What's the leading indicator in this story, and what does it predict?
Tap to flip
ANSWER
The near-retirement segment's own pass rate. It fell for four months before customer complaints ever rose.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building one randomly sampled golden set with no segment-level tracking, since it made sense while nearly every customer was decades from retiring.
6 · THE NUMBER
Fill in the blank: at its lowest point, the near-retirement segment's real pass rate was ___%, while the blended score still read about 96%.
Tap to flip
ANSWER
71%. The gap between those two numbers is the entire argument for stratifying the golden set.
7 · THE REPLAY
Same next model version, new eval spec in place. What changes?
Tap to flip
ANSWER
It fails on the near-retirement bar at 74% and stays home two weeks for retraining, instead of shipping on a blended pass. The version after clears every segment bar at 90% or above.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's its early signal?
Tap to flip
ANSWER
Fenwick Screen, a resume-ranking tool. Its early signal is the shortlist rate broken out by demographic group, tracked apart from the overall match score.

Check yourself Score: 0 / 0

Multiple choice
1. Why did the blended pass rate stay near 96% while the near-retirement segment's real pass rate fell to 71%?
  • A. The model treated near-retirement cases better than average.
  • B. Near-retirement cases were a tiny share of the golden set, so their failures barely moved the average.
  • C. The blended score was calculated incorrectly.
  • D. Regulators had already approved the near-retirement recommendations.
Show hint
Look at the line chart and the quadrant diagram.
Show answer
B. Near-retirement cases were only about 4% of the golden set, so even a steep decline there barely dented the overall average.
Fill in the blank
2. Fill in the blank: complaints stayed low through two full quarters before jumping to ___ in Q3, when the regulatory inquiry became public.
Show hint
Look at the quarterly complaints bar chart.
Show answer
14. That's the lagging outcome. The leading indicator, the segment pass rate, had already been falling for months before this jump.
True or false
3. True or false: this answer recommends dropping the blended pass rate entirely once segment bars exist.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The blended score stays, since it's still useful for catching a broad regression. Segment bars get added next to it, not instead of it.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Building one randomly sampled golden set with no segment-level tracking. It made sense while nearly every customer was decades from retiring and the stakes of a wrong answer were low across the board.
Short answer, where it wouldn't matter
5. Name a customer segment in this same product where a single blended pass rate would genuinely be fine on its own.
Show hint
Think about who has time to recover from a bad recommendation.
Show answer
Model answer: A 28-year-old just starting to save. A wrong allocation there has decades to correct itself, so the blended score isn't hiding much real risk for that group.
Short answer, apply it yourself
6. Pick a product you use yourself that shows you one overall rating or score. What group of users might that score be quietly failing?
Show hint
Think about who uses the product differently from the typical, most common user.
Show answer
Model answer: A common one: a navigation app's "on-time rate" can look great overall while quietly being much worse for wheelchair users, since routing for stairs and curbs is a small share of all trips.
Before you close the answer
Why this works
Tests whether you understand that "responsible AI" in an eval spec is a structural requirement, a stratified set and a per-segment bar, and not an extra checklist item layered on top of a normal accuracy check.
Follow-up traps
"Isn't stratifying the golden set just adding more work for a small group of customers?" Response: the group is small in volume but carries the most harm per case, and the whole point of the early signal is that it's cheap to check relative to what a missed case costs.

"Couldn't the segment bars just get gamed the same way the blended score was?" Response: less easily, since each bar has its own named owner and its own version pin, so a slipping bar is visible on its own instead of hiding inside an average.
If pressed
Aldergate's real spec also requires the golden set's segment weights to be reviewed against actual customer demographics every two quarters, so a segment that grows in real life gets more golden-set weight automatically, not just when someone remembers to update it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more