The direct answer
Split the eval in two. Let quality stay a weighted score that can slide up or down. Make safety and policy compliance a fixed, pass or fail gate, a short list of things the model can never do, like naming a protected characteristic, and run that gate before quality is even scored. A fluent, well-written answer that breaks the gate should never be able to buy its way past it just by reading well everywhere else.
Do this, in order
Split the rubric in two: quality stays a score, safety becomes a pass or fail gate that quality can't outvote.Why: a rule worth points inside a bigger score can be outvoted by everything else in that score.
Write the safety criteria as a fixed, testable list, not a general "be fair" line.Why: a vague rule can't be run as a pass or fail test, so it never actually blocks anything.
Run the safety gate before quality gets scored, so one hit blocks release no matter how the rest of the draft reads.Why: the real gap here showed up in a draft that scored well everywhere else.
Track flag and reject rates by the affected group, not only the average quality score.Why: an average that looks fine can hide a rate that isn't, once you cut it by who it's happening to.
Leave tone and formatting checks inside the quality rubric.Why: how warm a line reads isn't a safety risk, and a gate there wastes the gate's own credibility.
Re-scan analysis that's already shipped, on a rolling basis, not only at the next launch.Why: a group-level gap doesn't show up in a one-time launch check, it shows up over months.
How to answer this, stage by stage
Seven moves. The middle three, naming who each eval serves, what a good score would let slide, and who can't push back, are where the real answer lives.
1
Ground it in one real feature before naming a framework
Say it like this
"Say a company sells HireLens. It records a job interview, writes it out word for word, then drafts a short analysis for the recruiter: what the candidate did well, what's worth a second look, a recommendation. Before I touch a framework, I want the actual eval in front of me. Right now it's one hundred-point rubric, and 'never names a protected characteristic' is one line inside it, worth ten of the hundred points."
Why this works
Grounds an abstract "how do you split two evals" question in one real rubric before naming a method.
2
Say your structure out loud
Say it like this
"I'd use GUARD here, because this is a risk question wearing a spec-writing question's clothes. Who each eval actually serves, where a quality-only score would let a problem slide, who can't push back if it slips through, the actual spec change, and how I'd catch it in production."
Why this works
Two seconds that prove there's a plan before the interviewer hears you improvise a fix.
3
Name who each eval actually serves
Say it like this
"The quality eval serves the recruiter. It's asking, is this summary clear, useful, worth reading. The safety eval serves someone who's never in the room: the candidate the summary is about. Those are two different people with two different questions, and right now one eval is answering for both of them."
Why this works
This is GUARD's G step, naming both sides before a fix gets proposed.
4
Say exactly what a quality-only score would let slide
Say it like this
"A line like 'candidate may have caregiving responsibilities that could limit on-call availability' reads clean. It's well-formatted, relevant to the hiring call, calmly worded. Score it on quality alone and it does well, ninety-six out of a hundred in our case, because quality never asks whether the sentence should exist at all."
Why this works
This is U. It shows, with a real line, how fluent writing can carry a policy violation straight past a quality bar.
5
Name who can't push back if it slips through
Say it like this
"The candidate never reads that line. She gets an email saying the role went to someone else. Nobody tells her a sentence about her kids sat inside an internal report and nudged a scheduling flag against her. There's no button she can press to see it, let alone argue with it."
Why this works
This is A, GUARD's hardest step, and the one most answers skip.
6
Write the actual spec change, not a policy sentence
Say it like this
"Instead of one hundred-point rubric with a safety line worth ten points, split it. Quality stays a weighted score. Safety becomes a separate pass or fail check: a fixed list of protected characteristics the model can never name or imply, run against every draft before quality is even scored. One hit blocks the draft, no matter what it scores on everything else."
Why this works
This is R, and it's the whole question. A safety rule that shares a scoreboard with quality can always be outvoted by it.
7
Say how you'd catch it in production, then close
Say it like this
"I wouldn't just watch the average quality score, because averages hide exactly this. I'd track the flag rate by group, how often candidates who mention family, health, or age in small talk get a 'concern' flag, against everyone else. If that gap opens, the gate needs a new pattern in it, and I'd re-scan the last quarter's reports the same week I find it. That's the whole answer: two scores, one that can slide, one that can't be bought."
Why this works
This is D, plus a close that restates the decision in one breath.
Let's learn
Here's what happens when an answer sounds right and breaks a rule at the same time, and only one of those two things shows up on the scoreboard.
Say a company builds HireLens. It records a job interview, writes it out word for word, then drafts a short analysis for the recruiter: what the candidate did well, what's worth a second look, a recommendation.
Before the tool, a recruiter wrote those notes by hand. Twenty minutes after every interview, five interviews a day, that's about an hour and forty minutes just writing up what already happened.
Knowledge spark: what's a protected characteristic?
A fact about a person that's against the law to use in a hiring decision: age, disability, family or marital status, religion, and more. A hiring tool isn't supposed to notice these, let alone write them into a report.
Now HireLens drafts the same notes in under a minute. The recruiter reads it, edits a line or two, and moves on. The team built a hundred-point rubric to grade every draft: is it accurate, is it clear, is it useful, and one line, worth ten of the hundred points, checking that it never names a protected characteristic.
For over a year the average score sat at ninety-two out of a hundred, so the rubric looked like it was working. Here's the turn. Those ten points were never the real problem. The real problem is that ten points can get outvoted. A draft that infers a candidate's caregiving status can still write that inference in clean, confident, well-formatted prose, and clean, confident, well-formatted prose is exactly what the other ninety points reward.
We didn't build a tool that broke the rule. We built a scorecard where breaking the rule couldn't fail you.
At its worst, this lands on a candidate who never gets a callback, filtered out by an inference nobody meant to make and nobody checked, while the dashboard upstairs still reads ninety-two out of a hundred and calls the quarter clean.
The decision I would take back
We put the safety check inside the same hundred-point rubric as quality, worth ten points out of a hundred. That was fine back when violations were rare and mild. It stopped being fine the day a fluent one showed up and the other ninety points paid its way through.
What I would leave alone. Tone and formatting checks: does the summary read warm enough, is it the right length, does it follow the house style. None of that touches a protected characteristic. Folding those into the same forgiving quality score is fine, they were never the problem.
The lesson. A rule that shares a scoreboard with something else can be outvoted by it. If you want a rule that never loses, don't let it compete.
Now here is the same thing as a story
Read the short version above if you want it fast. Read this one when you want to feel why the fix matters, not just know what it is.
Miriam Solano can read a hundred-point rubric and tell you inside a minute which line is doing real work and which one is decoration. She's the eval lead for HireLens, Solvane's interview-analysis product, and she wrote the original rubric herself during the pilot, eighteen months ago, with three early customers watching every score.
For the first several months, that rubric was the whole job. Miriam read every flagged "concern" draft, line by line, every single week, checking that nothing had drifted. The average score climbed to the low nineties and held there. She kept reading anyway, out of habit more than anything.
Then she started skimming instead. The score held. Weekly checks became monthly. By the time the average had sat at ninety-two for a year, she was glancing at the dashboard number and moving on to the next fire.
Nobody told her this was risky. It looked, every single week, exactly like it was working.
Then a new recruiter, three weeks into the job, asked her a plain question in the hallway. "Why did HireLens flag this candidate's availability as a 'moderate concern' when he never once said anything about his schedule?" Miriam pulled the transcript. Near the top of the call, making small talk, the candidate had mentioned he had two young kids. HireLens's analysis, in the scheduling section, had written: "Candidate indicated potential caregiving responsibilities; may affect flexibility for early or on-call shifts. Moderate concern." Ninety-six out of a hundred on the quality rubric. Clear. Well-formatted. Relevant to the hiring call, on paper.
Same report, same outcome. Only one of them can change the setting.
We didn't build a tool that broke the rule. We built a scorecard where breaking the rule couldn't fail you.
Miriam pulled the last quarter's flagged drafts that weekend, fifty of them, out of four hundred interviews HireLens had processed. Six of the fifty contained something like it: a caregiving inference, an age assumption dressed up as "recent graduate, may need more ramp time," a comment on "communication clarity" that tracked suspiciously well with non-native accents. All six had scored above ninety on the quality rubric. Their average was ninety-three, a full point above the rubric's own yearlong average. The violations weren't dragging the score down. They were hiding inside it.
I would take back the decision to put the safety line inside that same hundred-point score at all. Not the rule itself, the rule was right. The scoreboard it lived on was wrong. Ten points is a discount a good sentence can always afford to pay.
Run the quarter again with a pass or fail gate sitting in front of the quality score instead of inside it. All six of the backlog's violating drafts get caught before they'd have reached a recruiter's inbox. Going forward, the gate catches about one draft in forty-five before release, each one rewritten before anyone sees it, and the flag-rate gap between candidates who mentioned family or health topics and everyone else starts closing toward the baseline instead of sitting there unmeasured.
The thing I'd tell myself, eighteen months back: a score that averages two questions into one number can only ever answer the easier one.
GUARD, aimed at one merged score
This is a risk question, so the framework is GUARD. "How do you split two evals" reads like a spec-writing question, which is exactly why nobody flagged the merged rubric as the actual gap until a new hire's plain question found it.
G, groups. The quality eval serves the recruiter, is the draft clear and useful. The safety eval serves the candidate, someone who never reads the draft at all, only lives with what it led to.
U, unequal. A violation doesn't cost the recruiter anything, the draft still reads well. It lands on the candidate, whose transcript happened to include a family, health, or age detail in ordinary small talk.
Concern flag rate, by whether the transcript mentioned family or health topics
Four hundred interviews processed last quarter, fifty flagged "concern" drafts examined.
Mentioned family or health topics (80 candidates)
34%
Didn't mention either (320 candidates)
9%
Twenty-seven of eighty, against twenty-nine of three-twenty. The rubric's own average score never showed this, because the average was never cut by who the candidate was.
A, ability to contest. The candidate never sees the analysis, the rubric, or the reasoning behind either one. She sees a rejection email, two steps removed from the sentence that helped shape it.
The box that should sit here, and doesn't.
R, reduce. A separate pass or fail gate, a fixed list of protected characteristics the model can never name or imply, run against every draft before quality is scored. One hit blocks release, no matter what the rest of the draft scores.
D, detect. A flag-rate dashboard cut by whether a transcript mentioned family, health, or age topics, not just the overall quality average. A gap that opens there triggers a fresh scan of already-shipped drafts, on a rolling basis.
Where this answer would fail
If the fix here is telling recruiters to read the transcript more carefully, or raising the safety line's point value from ten to twenty, none of it counts. A pass or fail gate and a by-group dashboard are both build items with an owner and a cost. Somebody can fund them this quarter, and you can check whether they did.
And if you want to be sure it really works, try it somewhere else
A city permits office uses an AI tool, PermitScope, to review building-permit applications and draft a risk note for the inspector. Different building entirely, same five letters, same trap.
G, groups. Inspectors, and whoever owns PermitScope's rubric, enforce the rule that a risk note can never cite an applicant's neighborhood, income level, or the look of their existing structure as a reason for extra scrutiny. Applicants, especially first-time homeowners and small contractors, are who it protects.
U, unequal. A note reading "modest existing structure suggests owner-builder work, recommend closer review" doesn't land evenly. It lands hardest on lower-income applicants whose homes already look more improvised, and it slows their permit exactly where it should move fastest.
A, ability to contest. An applicant reading a "needs review" notice has no way to see the risk note behind it, let alone the line that triggered it. By the time she could ask, the permit has already lost a filing season.
R, reduce. Fifty adversarial applications, a modest home, an unusual-sounding name, a low-income zip code, and the model has to write its risk note using only the actual code violation, never the property's appearance or the applicant's background. Zero tolerance, same shape as the hiring gate.
D, detect. Every risk note gets checked against a by-neighborhood flag-rate dashboard, not just the office's average review time. A gap opening between neighborhoods with similar violation rates blocks the rubric until it's fixed.
Swap the trigger and it still runs
- Speed: PermitScope rolls out to a second, busier district within a month, and the same untested gate ships with it, because the pilot only checked the district it launched in.
- Cost: the office trims a reviewer's hours because the average review time looks great, which is exactly the number that hides which neighborhood is waiting longest.
- The model gets better: accuracy on citing real code sections climbs to ninety-eight percent, and that's exactly when nobody proposes building the by-neighborhood dashboard, because the top-line number already looks finished.
Where people run it wrong
- Writing the never-rule once, in a policy memo, and calling it done instead of a test that has to keep growing.
- Testing a fix against the same application types that already passed, instead of the ones that actually failed.
- Building a fairness dashboard that only runs at launch, not on every batch that ships after it.
How to use it live
Ask, out loud, "does the safety check share a score with anything else?" If the honest answer is yes, say plainly that a shared score can always be outvoted, then fix it live by naming the one item on the list that should never be allowed to average out.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits a question about splitting evals for safety from evals for quality, and why?
Tap to flip
ANSWER
GUARD, for risk and safety. The real question isn't how to word two rubrics, it's who each eval serves, who can't push back if a bad one slips through, and how you'd catch it in production, exactly what GUARD is built to find.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Miriam Solano, eval lead for HireLens at Solvane. She built the tool's original hundred-point rubric during the pilot and used to read every flagged draft herself.
3 · THE HABIT
What did Miriam stop doing once the rubric's score held steady?
Tap to flip
ANSWER
Reading every flagged "concern" draft line by line. She moved to a monthly spot-check, then stopped, because the average score sat at ninety-two for over a year.
4 · THE GAP
What's the two-number gap this whole answer turns on?
Tap to flip
ANSWER
Concern flags hit thirty-four percent of candidates who mentioned family or health topics in small talk, against nine percent of everyone else. A merged, quality-weighted score never showed that gap.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Putting the safety check inside the same hundred-point rubric as quality, worth ten of the hundred points. It made sense because violations were rare and mild when the rubric was built, so nobody expected one fluent enough to outvote it.
6 · THE NUMBER
Fill in: of the last quarter's fifty flagged drafts, ______ contained a protected-characteristic inference, and their average quality score was ______ out of 100.
Tap to flip
ANSWER
Six of fifty, average quality score 93. The violating drafts scored slightly above the rubric's own yearlong average of 92.
7 · THE REPLAY
Same next quarter, new pass or fail gate in place. What changes, and by how much?
Tap to flip
ANSWER
The gate catches all six of the fifty backlog drafts before they'd have reached a recruiter. Going forward, it catches roughly one draft in forty-five before release, and the family/health flag-rate gap starts closing toward the nine percent baseline.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the reduce step become?
Tap to flip
ANSWER
A city permits office's risk-note tool, PermitScope. Reduce: fifty adversarial applications testing that risk notes cite only real code violations, never a property's appearance or an applicant's background, zero tolerance, same pass or fail shape as the hiring gate.
Check yourself Score: 0 / 0
Short answer
1. Why isn't putting the "no protected characteristic" rule inside the same hundred-point quality rubric, even worth real points, enough to catch it?
Show hint
Ask what happens to those ten points once the other ninety are written well.
Show answer
Model answer: "Because a violation only costs ten of a hundred points, so a well-written draft can lose those ten and still clear the bar. Averaging a safety rule into a quality score means a fluent answer can buy its way past the rule with good writing everywhere else. A rule that's meant to never happen needs its own gate, not a slice of somebody else's score."
Fill in the blank
2. In the last quarter's sample, ______ of the fifty flagged drafts contained a protected-characteristic inference, and they scored ______ on average, against the rubric's overall average of 92.
Show hint
It's the pair of numbers behind "the old decision" flashcard.
Show answer
Six of fifty, and 93 on average. The violating drafts scored slightly higher than the yearlong average, because the writing itself read clean. The score never asked whether the sentence should exist.
True or false
3. True or false: because a recruiter still reads the draft before deciding, the eval spec itself doesn't need a hard safety gate.
Show hint
Ask how long a recruiter actually spends reading each draft, and what a fluent line looks like to someone moving fast.
Show answer
False. Recruiters read a summary in under a minute and rarely see the underlying transcript. A confident, well-formatted line about "caregiving responsibilities" reads as normal shorthand, not a red flag. The eval spec is the only place actually built to catch it.
Multiple choice
4. Which old decision does this answer take back?
- A. Putting the safety check inside the same hundred-point rubric as quality, worth ten of the hundred points.
- B. Building a pass or fail safety gate that runs before quality is scored.
- C. Building HireLens's analysis feature at all.
- D. Telling recruiters to read the transcript more carefully before deciding.
Show hint
Look for the decision made when the rubric was first built, not the fix proposed after.
Show answer
A. B is the fix, not the reversal. C is the answer that gives up instead of designing something. D is a dial turned up ("be more careful"), not a decision taken back. Only A names the actual choice, one merged rubric, that the answer undoes.
Short answer, apply it yourself
5. Pick a product you've used yourself that scores both "is this good" and "is this safe" in some way. How would you check whether the safety part could get outvoted by a good quality score?
Show hint
Look for a single overall rating or score, then ask what it's actually made of underneath.
Show answer
Model answer: "A ride-share app's driver rating blends punctuality, cleanliness, and safety complaints into one number out of five. I'd check whether a driver with one serious safety complaint but otherwise great ratings can still land at 4.8, the same as a driver with zero complaints. If the two look identical from the score alone, the safety complaint got averaged away instead of gating anything."
Multiple choice
6. In this story, who holds the setting on the eval gate, and who only ever sees the outcome?
- A. Miriam holds the setting; the candidate only ever sees the outcome.
- B. The candidate holds the setting; Miriam only ever sees the outcome.
- C. Both hold the setting equally.
- D. Neither holds it; the rubric sets itself with no owner.
Show hint
Ask who could actually rewrite the rubric, and who could only ever receive a rejection email.
Show answer
A. Miriam owns and can change the eval spec. The candidate never sees the report, the rubric, or the reasoning behind either one, only the decision that comes out the other end.