The direct answer
Don't let a pre-checked box at signup stand in for consent. Ask again, in plain words, right after the chat ends, whether that specific conversation can go into the golden set. Send every "yes" through a person who rewrites the identifying specifics into paraphrase before it's stored, not just an automated scrubber that only catches names. Then plant a few unique, made-up phrases inside the set so you can search for a leak in production before a customer ever finds one themselves.
Do this, in order
Ask separately, after the chat, in plain words, instead of a box pre-checked at signup.Why: a box someone didn't read isn't consent, and it's the decision the whole leak traces back to.
Route every contributed chat through a human reviewer who paraphrases identifying specifics, not an automated scrubber alone.Why: a scrubber catches a name; it doesn't catch a manager's first name, a street, or a dose.
Cap how much of any one person's real words survive unchanged into the golden set.Why: even with the name gone, someone's exact sentence can still point straight back to them.
Seed a handful of entries with a unique canary phrase and watch production output for it.Why: it turns "we hope nothing leaked" into something you can actually search for.
Track the review backlog against the contribution rate as its own number.Why: the promise of hand review quietly stops being true the moment volume outruns it, and nobody notices without a number to watch.
Leave aggregate labels, like the urgency tier a chat got sorted into, freely logged with no per-entry consent needed.Why: a tier number carries nobody's actual words, so gating it the same way just slows the team down for no protection gained.
How to answer this, stage by stage
Seven moves. The middle three, naming who each version of the golden set actually serves, which chats put a person at the most risk, and who can never check whether theirs is one of them, are where the real answer lives.
1
Ground it in one real product before naming a framework
Say it like this
"Say we build Haven Line. You open it on a bad night, type what's going on, and in a few messages it decides whether you need a breathing exercise, a counselor next week, or a crisis line right now. To test that decision well, the eval team wants real chats, not just scripts three clinicians wrote, in a golden set they can grade the model against."
Why this works
Grounds an abstract "how do customers contribute" question in one real product before naming a method.
2
Say your structure out loud
Say it like this
"I'd use GUARD, because this reads like a data-collection question wearing a spec question's clothes. Who's affected, where the harm lands hardest, who can't tell it happened to them, the actual design change, and how you'd catch it before they do."
Why this works
Two seconds that prove there's a plan before the interviewer hears you improvise a policy.
3
Name who each version of the golden set actually serves
Say it like this
"There's the person who typed the chat, someone having a hard night who never expected a stranger to read it back later. And there's everyone else who uses Haven Line, who's better off if the model has actually seen real crisis language instead of three clinicians' best guess at it. Those two groups both have a real claim, and right now nobody's written down which one wins when they conflict."
Why this works
This is GUARD's G step, naming both sides before proposing a fix.
4
Say which chats put a person at the most risk
Say it like this
"A chat that just says 'I've been really anxious lately' is safe to store almost as is, nobody could trace it back. But the chats we actually need most for the golden set are the specific ones. Someone naming their manager, their exact clinic, their medication and dose. Take the name off one of those and the rest still points straight at one person."
Why this works
This is U. It shows, with a real example, that the harm doesn't land evenly, and that the most valuable chats and the most dangerous ones are often the same chats.
5
Name who can't tell it happened to them
Say it like this
"She checked a box, or didn't uncheck one, months ago at signup, on a night when reading settings was the last thing on her mind. She has no way to look up whether her exact chat got pulled into the set, whether a contractor read it while labeling severity levels, or whether it's sitting in a shared file somewhere. There's no 'your chat was used' notice. By the time anyone would find out, it's already out."
Why this works
This is A, GUARD's hardest step, and the one most answers skip.
6
Write the actual design change, not a policy line
Say it like this
"Two changes. First, ask again after the chat ends, in plain words, not a box ticked at signup. Second, every 'yes' goes to a human reviewer who rewrites the specific details, the manager's name, the exact clinic, the dose, into a close paraphrase that keeps the severity and the shape of what was said, but drops what makes it locatable. No chat reaches the golden set on the scrubber's word alone."
Why this works
This is R, and it's the whole question. A default isn't a decision, and a name-only scrubber isn't a review.
7
Say how you'd catch a leak before the customer does, then close
Say it like this
"I'd seed a few golden-set entries with a made-up phrase nobody would type by accident, then search production outputs and any place the set gets shared for that exact string. I'd also log who opens each entry, and once a quarter try to match ten anonymized entries back to a real account. If I can do it once, the review bar tightens. That's the whole answer: ask twice, review by hand, and build a way to catch yourself before someone else does."
Why this works
This is D, plus a close that restates the decision in one breath.
Let's learn
Here's what happens when the chats you most need are the same chats that give away the most about the person who wrote them.
Say we build Haven Line. You open it on a hard night, describe what's going on, and it decides in a few messages whether you need a breathing exercise now, a counselor this week, or a crisis line right now.
Knowledge spark: what's a golden set?
A pile of real examples you use to grade the model, before it ever reaches a user. Not training data. A test the model has to pass every time something about it changes.
Before real chats went in, the eval team tested Haven Line against pretend conversations three clinicians wrote by hand. Against those scripts, the model caught 91 chats out of 100 where someone was clearly describing self-harm. Against a set of real, already-anonymized crisis-line transcripts borrowed for a research check, it caught only 72. Real people don't type the way three clinicians imagined they would.
So the team started asking customers to let their real chats join the golden set. The ask lived as a box at signup, pre-checked by default, worded as "help us improve the model." Here's the turn. The extra crisis language they picked up wasn't the problem. The problem is what a pre-checked box does at scale: 89 out of every 100 sessions ended up marked shareable, because most people never noticed the box, let alone unchecked it. Volume went from about 40 shared chats a week to 600 a week inside four months. Nobody had time to hand-review 600 a week, so the team leaned on an automated scrubber that stripped names and only names.
The transcripts we needed most were the ones that gave away the most.
At its worst, this lands on the one person whose chat named a manager, a clinic, and a medication dose, all still sitting in a "de-identified" file a contracted annotator can open, a file that could just as easily end up pasted into a bug report, a shared slide, or a future prompt. A leak like that doesn't just cost a person their privacy. It costs a mental health product the one thing it's actually selling: that you can say the specific, ugly, true thing and it stays where you put it.
The decision I would take back
We made "share this chat" a box pre-checked at signup instead of a real question asked after each session. That was fine when contribution was low enough for someone to read every entry by hand. It stopped being fine the moment volume outran review and a scrubber became the only real gate left standing.
What I would leave alone. The urgency tier the model assigns to a session, logged in bulk with no chat text attached, self-help, book this week, crisis now. That number carries nobody's words. Tracking it in aggregate needs no per-entry consent at all, it was never the risk.
The lesson. A box nobody read isn't consent, it's a legal fig leaf. And once a team can't review everything by hand, the risky detail has to be designed out before it ever reaches the set, not caught by a filter after the fact.
Now here is the same thing as a story
Read the short version above if you want it fast. Read this one when you want to feel why the fix matters, not just know what it is.
Chinelo Eze can read a spec document and find the one line doing all the real work in about a minute. She's the product manager who owns eval-driven specification for Haven Line, and she wrote the original golden-set plan herself, back when the team was still testing against three clinicians' pretend conversations.
For the first few months after real chats started coming in, Chinelo read every single one that got marked shareable. Forty or so a week, over coffee, before the rest of her day started. She was looking for exactly one thing: a detail specific enough to point back at a real person even with the name removed.
Then the number climbed. Fifty a week. A hundred and twenty. By the time it hit six hundred, she'd stopped reading them herself and started trusting the automated scrubber to catch what mattered. It always found the names. The dashboard kept saying "processed," and "processed" started to feel the same as "safe."
Nobody told her this was risky. It looked, every single week, exactly like it was working.
Then a contracted annotator, three weeks into labeling severity levels for the golden set, sent her a short message. "This one's supposed to be de-identified, but it still has a name and a street in it." Chinelo pulled the entry. The scrubber had removed the writer's own name. What was left read: "I can't keep hiding this from David, my manager, and I already told my therapist at the Fifth Street clinic that I upped my sertraline to 150 without telling my doctor." No name. Three ways to find her anyway.
Same golden set, same outcome. Only one of them can pull the lever.
The transcripts we needed most were the ones that gave away the most.
Chinelo pulled a sample that weekend, two hundred entries the scrubber alone had processed from the backlog. Twenty-eight of them, 14 out of every 100, still carried a detail beyond a name: a workplace, an exact clinic, a dose, a method. She pulled a second sample, a hundred and fifty entries a human had reviewed properly before the volume outran the team. Three of those, 2 out of every 100, still had something. The gap wasn't the scrubber failing occasionally. It was the scrubber doing exactly the one job it was built for, and nothing else.
I would take back the decision to make "share this chat" a box pre-checked at signup. Not the idea of asking, the idea was right. The moment we let a default stand in for a real answer, we also signed up to review everything that default produced, and we never built the capacity to keep that promise once it started working too well.
Run the same six months again with two changes: the ask happens after the chat, in plain words, and every "yes" goes to a person who paraphrases the specifics before storage. Contribution drops from 89 shareable chats out of 100 to about 61, a real number from people who actually meant it. Review load drops with it, because a smaller, informed pool is one a team can actually keep up with. And the twenty-eight identifying entries in that backlog sample never reach the golden set with David's name, the clinic, or the dose still readable in them at all.
The thing I'd tell myself, back when I wrote that first plan: a scrubber that only removes names isn't privacy, it's the one part of privacy that was easy to automate.
GUARD, aimed at a box nobody read
This is a risk question, so the framework is GUARD. "How do customers contribute safely" reads like a data-collection question, which is exactly why the pre-checked box didn't look like the actual decision until an annotator's message found it.
G, groups. The person who wrote the chat, on a hard night, never expecting a stranger to read it back. And the wider base of Haven Line users, who are safer if the model has actually seen real crisis language instead of a clinician's best guess at it.
U, unequal. A vague, anxious chat is nearly impossible to trace back to anyone. The specific ones, naming a manager, a clinic, a dose, are the ones the golden set needs most, and the ones that stay identifiable long after the name is gone.
Entries still carrying an identifying detail after processing
Two backlog samples pulled the same weekend, before the redesign shipped.
Automated scrubber only (200 entries)
14%
Human-reviewed before volume outran the team (150 entries)
2%
Twenty-eight of two hundred, against three of a hundred and fifty. The dashboard's "processed" count never showed this, because it never asked what survived processing.
A, ability to contest. The customer who checked, or didn't uncheck, a box at signup has no later way to see whether her exact chat is in the set, who has read it, or where it's been shared. There's no notice, and no undo.
The step that should sit before the last box, and doesn't.
R, reduce. A separate, plain-language ask after the chat ends, not a default at signup. Every "yes" goes through a human reviewer who paraphrases identifying specifics before storage, with a cap on how much verbatim text survives.
D, detect. A handful of golden-set entries seeded with a unique canary phrase, searched for in production output and anywhere the set gets shared. A quarterly re-identification test, ten anonymized entries, matched back to a real account by hand, tightens the review bar the day it succeeds even once.
Where this answer would fail
If the fix is a longer terms-of-service paragraph, or telling the scrubber to "be more thorough," none of it counts. A separate consent ask, a human review step, and a canary search are all build items with an owner and a cost. Somebody can fund them this quarter, and you can check whether they did.
And if you want to be sure it really works, try it somewhere else
A company runs an anonymous workplace misconduct line, ReportWell, where an AI triages incoming reports for severity and routes the urgent ones straight to HR. Different building entirely, same five letters, same trap.
G, groups. The employee who filed the report, trusting "anonymous" to mean something. And the wider workforce, who's safer if the triage model has actually seen real report language instead of HR's best guess at how people phrase a complaint.
U, unequal. A vague report, "there's a culture problem on my team," is nearly untraceable. A specific one, naming a project, a date, and a manager's habit of closing the door during one-on-ones, points straight at one desk even with the reporter's name stripped out.
A, ability to contest. The reporter has no way to check whether their report, or a close paraphrase of it, ended up in a training set an engineer or the accused manager's own leadership chain might eventually see.
R, reduce. A report only joins the golden set after a second, explicit ask, once the case is closed, and a reviewer rewrites any detail specific enough to name a desk before it's stored, not just the reporter's name.
D, detect. A canary detail seeded into a few closed-case entries, checked against anything the model later surfaces to a manager during an unrelated review, catching a leak before a reporter has to wonder if it was them.
Swap the trigger and it still runs
- Speed: the golden set needs to triple before a big model swap next quarter, and the fastest way to get there is to loosen the consent ask, exactly the move that broke it the first time.
- Cost: a cheaper scrubber ships because it's "good enough on the benchmark," and the benchmark was built the same lax way the real backlog was.
- The model gets better: catch rate on real crisis phrasing climbs from 72 to 94, and that's exactly the moment nobody wants to slow down and add a canary search, because the top-line number already looks finished.
Where people run it wrong
- Writing the consent language once, in a policy document, and treating it as done instead of a design that has to survive volume.
- Trusting a scrubber's pass rate on a benchmark built from the same easy cases it already catches.
- Running a canary search once, at launch, instead of on every batch that joins the set after it.
How to use it live
Ask, out loud, "would the person who wrote this recognize it if they saw it again?" If the honest answer is yes, the review step isn't finished. Say that plainly, then name the one specific detail in the example you were just given that should never have survived.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits a question about customers contributing real data to a golden set, and why?
Tap to flip
ANSWER
GUARD, for risk and safety. The real question isn't how to ask for data, it's who each version of the set serves, whose data is hardest to make anonymous, and who can never check whether theirs leaked, exactly what GUARD is built to find.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Chinelo Eze, the product manager who owns eval-driven specification for Haven Line. She wrote the original golden-set plan and used to read every shareable chat herself.
3 · THE HABIT
What did Chinelo stop doing once contribution volume grew?
Tap to flip
ANSWER
Reading every shareable chat by hand. Once volume passed a hundred a week, she started trusting the automated scrubber alone, and treated "processed" as the same thing as "safe."
4 · THE GAP
What's the two-number gap this whole answer turns on?
Tap to flip
ANSWER
Fourteen out of every hundred scrubber-only entries still carried an identifying detail, against two out of every hundred that a person had actually reviewed. The dashboard's "processed" count never showed the difference.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Making "share this chat" a box pre-checked at signup instead of a real question asked after each session. It made sense when contribution was low enough that someone could still read every entry by hand.
6 · THE NUMBER
Fill in: with a pre-checked box, ______ out of every 100 sessions were marked shareable. With a real, separate ask, that dropped to about ______.
Tap to flip
ANSWER
89 out of 100 with the pre-checked box, dropping to about 61 with a real, separate ask. Fewer chats, but every one of them from a person who actually meant it.
7 · THE REPLAY
Same six months, new consent and review design in place. What changes, and by how much?
Tap to flip
ANSWER
Contribution drops from 89 to about 61 shareable chats per 100 sessions, review load drops with it, and the twenty-eight identifying entries in the sample backlog never reach the golden set with a manager's name, a clinic, or a dose still readable.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the reduce step become?
Tap to flip
ANSWER
ReportWell, a workplace misconduct triage line. Reduce: a report only joins the golden set after a second, explicit ask once the case is closed, and a reviewer rewrites any detail specific enough to name a desk, not just the reporter's name.
Check yourself Score: 0 / 0
Short answer
1. Why isn't a "share your chat to help us improve" checkbox at signup, on its own, enough consent for a golden set built from mental-health conversations?
Show hint
Ask when the checkbox is shown, and what state of mind someone is usually in at signup versus right after a hard chat.
Show answer
Model answer: "Because most people never read it, and being pre-checked means silence counts as yes. Someone signing up isn't thinking about a future chat they haven't had yet. Real consent has to be asked about the specific thing, right after it happens, in words plain enough that saying yes means something."
Fill in the blank
2. In the backlog sample, ______ out of 200 scrubber-only entries still carried an identifying detail, against ______ out of 150 entries a person had reviewed.
Show hint
It's the pair of numbers behind "the gap" flashcard.
Show answer
28 out of 200, and 3 out of 150. Fourteen percent against two percent. The scrubber caught names every time; it never caught a manager's name, a street, or a dose, because it was never built to look for those.
True or false
3. True or false: once a chat's writer's name is removed, the chat is safe to store in a shared golden set.
Show hint
Ask what else, besides a name, could point back to exactly one person.
Show answer
False. A manager's name, an exact clinic, a medication and dose, or any combination of specific details can still narrow a chat down to one real person, even with the writer's own name gone. Name removal alone isn't the same as making something unidentifiable.
Multiple choice
4. Which old decision does this answer take back?
- A. Making "share this chat" a box pre-checked at signup instead of a real question asked after each session.
- B. Building a golden set from real customer chats at all.
- C. Adding a human reviewer who paraphrases identifying specifics before storage.
- D. Telling the scrubber to run twice instead of once.
Show hint
Look for the decision made when the ask was first designed, not the fix proposed after.
Show answer
A. C is the fix, not the reversal. B is the answer that gives up instead of designing something. D is a dial turned up ("run it again"), not a decision taken back. Only A names the actual choice, a default standing in for consent, that this answer undoes.
Short answer, apply it yourself
5. Pick a product you've used yourself that asks to use your data to "improve the product." How would you check whether that consent was a real, separate ask, or a default that just happened to be in your favor to leave alone?
Show hint
Look for whether the option was already turned on when you first saw it, and whether it was asked about the specific thing or buried in a general settings page.
Show answer
Model answer: "A fitness app asks to use my workout data to 'improve recommendations for everyone,' and the toggle was already on when I checked. I'd look for whether it was asked right when I logged something specific, in plain words about that specific data, versus buried in a settings menu I'd have to go find. If it's the second one, it's a default wearing consent's clothes."
Multiple choice
6. In this story, who holds the decision over whether a specific chat joins the golden set, and who only finds out after the fact, if ever?
- A. Chinelo and the eval team hold the decision; the customer who wrote the chat only finds out after the fact, if ever.
- B. The customer holds the decision; Chinelo only finds out after the fact.
- C. Both hold the decision equally, at the same time.
- D. Neither holds it; the scrubber decides on its own.
Show hint
Ask who could actually change the review process, and who could only ever receive silence.
Show answer
A. Chinelo's team owns the intake, the review step, and the golden set itself. The customer who wrote the chat has no dashboard, no notice, and no way to ask later whether her exact words are in it.