How do you turn a responsible AI principle into a testable product requirement?
Verve is a short-video app. Its comment-ranking model decides which replies show up first under a video, including which ones climb into the Top 10 a creator actually reads. Dax Okonkwo runs trust and safety for that ranking system, and keeps a laminated one-page checklist taped to the edge of his monitor.
- Turn the principle into one number tied to a check, not a value statement.Why: a principle nobody can test never actually gets tested, no matter how often it's repeated in a deck.
- Build the held-out eval set before you write the threshold.Why: a threshold with no real set behind it is a guess wearing a number's clothes.
- Pick a leading signal, not the harm itself.Why: if you wait for real harm to show up in a support ticket, the harm already happened.
- Make the check block the launch automatically.Why: a check a person can wave through under deadline pressure isn't a check, it's a suggestion.
- Name one way the number gets gamed, and watch for that specifically.Why: every metric has a cheap way to look good without anyone doing the real work.
- Re-test the eval set itself as the product's comment shapes change.Why: a set built for last year's harassment patterns goes stale quietly, and a stale set is worse than no set because it looks like proof.
How to answer this, stage by stage
The interviewer isn't grading whether you can recite "responsible AI" as a phrase. They're grading whether you can turn it into something an engineer could actually fail a build against.
Let's learn
Say a company writes a value into its charter: we don't amplify harassment. Everyone nods. Nobody can tell you what would make that sentence false.
Verve's comment-ranking model used to get spot-checked by hand. Once a quarter, someone on the trust and safety team would pull ten ranked comment threads and read through them, looking for anything that felt like a pile-on had climbed too high. It always looked fine. So the checks got shorter, then less frequent, then folded into a slide nobody presented anymore.
There was no single bad quarter that started this. The number that mattered, the share of reported comments still sitting in a video's Top 10 an hour later, drifted upward for six straight weeks while nobody was watching it, because nobody had ever named it as a number worth watching.
At its worst: a coordinated pile-on comment climbed to third place under a creator's video, ahead of hundreds of ordinary replies, and stayed there for forty minutes before a moderator who happened to be scrolling that day caught it.
What I would leave alone: a single blandly negative comment sitting low in the rankings doesn't need this kind of gate. Not every unkind reply is a pile-on, and treating every low-rated comment as a harassment signal would bury real, if harsh, feedback along with the actual danger.
The lesson: a principle with no number behind it isn't a value the company holds, it's a sentence the company hopes stays true.
Now here is the same thing as a story
The short version above is what you'd say cold in an interview. Read this one for how the near miss actually happened.
Dax has run trust and safety for Verve's comment system for four years. He's the only one at the company whose whole job is watching what the ranking model does to a thread, not just whether creators are posting.
When the ranking model first shipped, harassment complaints were rare enough that a quarterly hand-check felt like plenty. Dax would pull ten threads, read them over coffee, and email a one-line summary to the team: looked fine, nothing to flag.
It built up slowly, over about six weeks, and nobody noticed because nobody had a number pointed at the right thing. The ranking model had gotten better at predicting which replies would get more responses, which meant it had also gotten better, completely by accident, at surfacing the exact kind of comment a coordinated pile-on produces: short, reactive, and dozens of people replying to each other in a chain.
Then came a Tuesday. A creator posted a video about a local news story, and within twenty minutes a coordinated group of accounts had replied to each other with the same accusatory comment, worded just differently enough each time to avoid Verve's word-filter. The ranking model, reading all that reply-to-reply activity as high engagement, pushed the comment to third place under the video.
Dax happened to be scrolling that creator's page for an unrelated reason. He saw the comment sitting at third place, checked the report queue, and found three people had already flagged it. Nobody had told him. Nothing had told him. He just happened to be looking.
He pulled the ranking logs for the past six weeks that same afternoon, and built the number for the first time: the share of reported comments still sitting in Top 10 an hour after posting had gone from 2 percent to 11 percent, quietly, while every dashboard the team actually watched stayed flat.
We did not almost lose a launch that day. We almost let a coordinated pile-on outrank hundreds of real replies under a stranger's video, and the only reason it didn't ship to more people was luck, not design.
The team wrote "no more than 5 percent, checked every launch" that same week, and built the eval set behind it from six months of past reported comments, tagged for whether they'd stayed in Top 10.
I want to say the model got worse. It didn't, not really. It got better at exactly the thing it was told to optimize, engagement, and a pile-on is engagement wearing a bad costume. That's the part a value statement can never catch, because a value statement doesn't know what the model is actually being scored on underneath it.
We wrote "we don't amplify harassment" into a launch doc two years ago because it was true, and true felt like enough. It took a stranger's pile-on almost reaching third place under a video to see that a sentence nobody can fail isn't a promise. It's a hope with good intentions.
LEAD, in one screenNot "how sure is the model." How early would this number have told us, and what do we do at each level of it.
The recap, one line per letter: link is creator retention behind the harassment promise, early signal is the reported-comment-visibility rate that moved six weeks ahead of anything else, abuse is a coded pile-on that dodges a slur filter, and decision is the three-tier response from ship to auto-block.
And if you want to be sure it really works, try it somewhere elseSame four letters, a handmade-goods marketplace instead of a comments feed. A completely different harm, and this time the input mix is the trick.
Almscroft Market lets independent makers list handmade goods, and an AI tool drafts each listing's description from a few uploaded photos and a short seller note. The company's principle is "we don't let AI listings mislead a buyer." Priya, who owns that feature, mapped it onto LEAD like this. Link: buyer trust in the marketplace, since a buyer burned once by a misleading AI description stops buying from anyone on the platform, not just the one seller. Early signal: the return rate specifically on AI-drafted listings for items over 100 dollars, since that's where sellers started rationing which items they'd trust the AI tool with. Abuse: a seller could pass this check by only ever using the AI tool on cheap, low-risk items, which makes the eval set look great while telling you nothing about the expensive listings buyers actually complain about. Decision: below a 4 percent return-for-mismatch rate, ship as-is; above it, require a seller to confirm three AI-drafted claims by hand before the listing goes live.
Swap the trigger and it still runs.
Speed: an interviewer gives you thirty seconds. Say "turn the value into a number, build the eval set first, and make the number block the launch," and stop there.
Cost: there's no headcount to build a full held-out set this quarter. Say so honestly, and start with the twenty worst historical cases as a rough set rather than none at all, since a rough gate beats a hope.
The model gets better, for real: if the ranking model's overall relevance score improves, that's exactly when this check matters most, because a smarter model is also a smarter pile-on amplifier, not a safer one by default.
Where people run it wrong.
They write the value statement and call it done, since it feels complete on the page.
They pick the harm itself as the metric, which means by definition the harm already happened before the number moved.
They build a threshold with no eval set behind it, so the number is really just a guess with a percent sign on it.
How to use it live. When someone hands you a values statement and asks you to make it real, ask yourself one question first: what would move, quietly, weeks before anyone could actually see the harm. Name that thing out loud before you touch a single threshold.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Doesn't auto-blocking the launch just slow the team down constantly?" Response: only when the number is actually bad, and a launch that would push a pile-on into Top 3 is exactly the launch that should be slowed down.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Responsible AI as a product requirement
- #2 What safety requirements belong in every AI PRD regardless of feature?
- #3 Describe how you would assess a feature for potential harm before building it.
- #4 Explain the difference between a safety issue and a quality issue.
- #5 How would you handle a feature that works well overall but poorly for one demographic?
- #6 What is a content policy and who should own it in a product organization?
- #7 Design the guardrails for an AI feature aimed at teenagers.