CaseAdvancedResponsible AI & Advanced Practice / Responsible AI as a product requirement / #17
How would you handle a user population using your product in an unintended harmful way?
GUARD the product is Reelhaven's auto-clip tool, which stitches a user's video moments into a highlight reel
Reelhaven is a short-video app. Its auto-clip tool watches a user's uploads and stitches the most-replayed moments into a shareable highlight reel, no editing skill required. Marguerite Solano is the trust and safety product manager who found out what a small slice of uploaders were actually stitching together.
The direct answer
Attach a visible, un-strippable watermark to every auto-clip tying it back to the original video, and add a rate limit that flags when the same person keeps appearing, uploaded by different accounts, across many clips. Name a real contact path for anyone who appears in a clip they never agreed to be in, not just the uploader who made it.
Do this, in order
Keep a visible, hard-to-strip watermark on every auto-generated clip.Why: once a clip is untraceable to its source, there's no way to act on a pattern of misuse at all.
Flag when the same filmed person shows up across many clips from different uploaders.Why: one clip of someone is a moment. Dozens, stitched by strangers, is a pattern the person being filmed can't see or stop.
Give the person who was filmed, not just the uploader, a real way to ask for a takedown.Why: the uploader chose to use the tool. The person filmed never got a choice at all.
Block re-download of a clip once it's taken down.Why: a takedown that still leaves a saved copy in twenty other phones isn't a takedown.
Track the misuse rate by account age, not just by clip volume.Why: brand-new accounts built around a single target look nothing like an ordinary creator's upload pattern.
Leave the ordinary meme-remix and highlight-reel use case exactly as fast as it is today.Why: the overwhelming majority of clips are someone celebrating their own dog or their own goal, and slowing that down to catch a rare pattern punishes everyone for a few.
How to answer this, stage by stageEight moves, since this question stretches wider than most: it needs both a response plan and a defense of why that response is the right one.
Stage 1
Pick one real product and one real misuse
Say it like this
"I'll answer this for Reelhaven's auto-clip tool, which stitches a user's uploaded moments into a shareable highlight reel. The unintended harmful use is a small group of accounts using it to build harassment compilations targeting one person."
Why this works
"Unintended harmful use" is too broad to answer until it's one real tool and one real pattern.
Stage 2
Name the framework
Say it like this
"I'll use GUARD. Groups affected, where the harm lands unequally, who can actually contest it, the specific reduction, and how we'd detect it."
Why this works
Signals a structured answer about power and harm, not a vague promise to "take this seriously."
Stage 3
Name both people, on purpose
Say it like this
"There are two people in this story. The uploader, who chose to use the tool. And the person appearing in the clip, who never chose anything and doesn't even know the clip exists until someone sends it to them."
Why this works
Most weak answers only ever picture the uploader, since the uploader is the one with an account and a login.
Stage 4
Ask who never gets to push back
Say it like this
"The person filmed can't inspect the clip before it ships, can't appeal a decision they were never told was being made, and can't opt out of appearing in someone else's upload in the first place."
Why this works
This is the load-bearing question in a fairness answer: not "is this bad," but "who has zero ability to stop it."
Stage 5
Give the concrete reduction
Say it like this
"Concretely: an un-strippable watermark on every clip, a flag when the same person appears across many clips from different accounts, and a real contact path for the person filmed, not just the uploader."
Why this works
This is the actual answer, a specific product change, not a review board or a policy update.
Stage 6
Prove it with the real pattern
Say it like this
"When Reelhaven's team finally looked, forty-two percent of flagged harmful clips came from accounts under seven days old, versus three percent from accounts over a year old. That's not ordinary creators drifting into bad behavior. That's accounts built for this specific purpose."
Why this works
A real, checkable number is what separates "we're concerned" from "we found the actual pattern."
Stage 7
Say how you'd detect it going forward
Say it like this
"We'd track weekly reports of the same target appearing across multiple clips as its own metric, separate from total flagged content, since that's the number that would have shown the pattern building for months before the news story."
Why this works
Shows you think about ongoing detection, not just a one-time cleanup after the damage is public.
Stage 8
Say what you'd leave alone, and close
Say it like this
"The ordinary use, a parent stitching their kid's soccer season, a friend group's trip highlights, stays exactly as fast and frictionless as it is today. The reduction targets the pattern, not the tool."
Why this works
Shows judgment: the fix is aimed at a specific, named harm, not a blanket slowdown that punishes everyone.
Let's learn
Here is what happens when a tool built to celebrate a user's own best moments gets pointed, by a handful of people, at someone else's worst ones.
Reelhaven's auto-clip tool watches everything a user uploads over a week and stitches the most-replayed seconds into a single shareable reel. Most people use it exactly as intended: a highlight of their own gym progress, their own cooking wins, their own dog being ridiculous.
Knowledge spark: what does "unintended use" actually mean here?
It means the tool is working exactly as designed. Nothing is broken. The auto-clip feature correctly finds the most-replayed moments and stitches them together. The harm isn't a bug in the stitching, it's in what a small group of people chose to feed it.
A small number of accounts began uploading clips of the same private individual, pulled from public livestreams and old videos, feeding them into the tool to auto-generate a compilation built to humiliate, not celebrate.
The uploader has an account, a login, and a choice. The person filmed has none of the three.
At its worst, one of these compilations reached a local news story about online harassment, five months after the first version had already been quietly flagged and taken down once, only to reappear under a new account within a week.
Share of flagged harmful clips, by account age
Brand-new accounts made up a wildly disproportionate share of the harm. Ordinary long-time users almost never triggered a flag.
The person targeted found out the same way everyone else did, a friend forwarded her the clip. She had no account on Reelhaven at all, and no way to have known any of it existed before it did.
The decision I would take back
We shipped auto-clips with no watermark, because a clean, unbranded video felt more shareable and less like an ad for our own app. That made sense when clips were mostly being shared inside small friend groups. It stopped making sense the moment a clip needed to be traced back to its source across a dozen re-uploads on other platforms, and there was nothing on it to trace.
What I would leave alone: the ordinary case, a parent's kid's soccer season, a friend group's weekend trip, gets none of this friction. The fix targets a specific, measurable pattern, not the tool's basic speed for everyone else.
The exact same clip, the exact same harm. One version can be traced back to where it started. The other can't be traced at all.
The problem was never that people were being cruel with the tool, cruelty is old. The problem was a clip designed to travel with nothing attached to it, so once it left the app, nobody could tell where it came from or stop it from coming back under a new name.
The lesson: a tool that works exactly as designed can still need a guardrail, because the design never asked who the tool's output might be used against, only who was using it.
Now here is the same thing as a story
The short version above is what you'd say defending this response to Reelhaven's leadership. Read this one for how the pattern actually surfaced.
Before any of this, Marguerite Solano's job looked like a fairly ordinary content-moderation queue: copyright flags, spam accounts, the occasional graphic-content report. Nothing that kept her up at night.
The moderation queue had a review step. It never had a step for the person in the clip to say anything at all.
The first flagged compilation showed up in week six of the feature's life, reported by a random viewer, not the person in it. Marguerite's team took it down under a general harassment policy and moved on, treating it as a one-off.
Eight weeks sit between the first flagged clip and anyone internally noticing it wasn't a one-off.
It came back under a different account name within the week. Then again. Each individual takedown looked, on its own, like the moderation system doing its job correctly.
Weekly reports of harassment-compilation clips, eight weeks before the news story
Nobody had a metric for "same target, many accounts." Once someone built one, the climb had already been running two months.
By week fourteen, Marguerite finally cross-referenced the takedown log by target instead of by account, and the pattern was immediate: the same person, dozens of clips, a rotating cast of week-old accounts.
The worst two categories sit in the hardest corner to see. A per-account moderation queue was built to catch the easy, obvious corner instead.
With the watermark and the cross-account flag in place, the same compilation attempt now gets caught before it ever fully renders: the system sees the same face flagged across five new accounts inside a week and freezes the pattern for review, instead of waiting for a viewer to report it after the fact.
Four pieces. The old system had none of them; a per-account queue only ever looks at one account at a time.
The old design asked "is this one account behaving badly." The new one asks "is the same person, across many accounts, being targeted the same way," which is the actual shape the harm took.
I skipped the watermark because an unbranded clip felt more shareable, and shareability was the whole metric everyone was watching that quarter. It took watching the same clip resurface under a new account name, every single week, to see that a clip nobody could trace was a clip nobody could actually stop.
GUARD, who has the lever and who doesn'tNot a moderation flowchart. GUARD is what makes sure the person who never logged in still gets a way to push back.
G
Groups. Who is affected, named.
The uploader, who chose the tool. The person filmed, who never chose anything and often never even knows.
Naming both, not just the account holder, is where a real fairness answer starts.
U
Unequal. Where the harm lands unevenly.
Forty-two percent of flagged clips came from accounts under a week old, built for exactly this, versus three percent from established users.
A real, measurable skew, not a hunch about "some bad actors."
A
Ability to contest. Who never gets to push back.
The person filmed can't inspect, appeal, or opt out, since she doesn't have a Reelhaven account and never knew the clip was being made.
The hardest step, and the one most moderation systems never ask.
R
Reduce. The specific design change.
A visible, un-strippable watermark, a cross-account same-target flag, and a real contact path for the person filmed, not just the uploader.
A product decision, not a policy memo or a review board.
D
Detect. How you'd know before it's public.
A weekly count of same-target reports across accounts, tracked as its own metric, would have shown this climbing two months before the news story.
Turns a hindsight realization into an ongoing, checkable number.
The recap, one line per letter: groups is the uploader versus the person filmed, unequal is the 42 percent versus 3 percent account-age skew, ability to contest is the person filmed having no account and no appeal, reduce is the watermark plus the cross-account flag, and detect is the weekly same-target report count.
And if you want to be sure it really works, try it somewhere elseSame five letters, a restaurant supply wholesaler instead of a video app. A completely different industry, and here the person who can't push back isn't a stranger in a clip, it's a small supplier undercut by their own order history.
Larder & Line runs an AI ordering assistant that suggests reorder quantities to restaurant buyers based on past purchase patterns. Boyd Achterberg manages the assistant, and a printed weekly order sheet is what most buyers actually annotate by hand before submitting.
Mapped onto GUARD: groups are the restaurant buyer using the assistant, and the small independent food supplier whose product the assistant quietly stops recommending. Unequal is that the assistant, trained mostly on high-volume purchase data, began systematically under-recommending small suppliers' items in favor of larger distributors with denser order histories, regardless of quality or price. Ability to contest is the small supplier having no visibility into the recommendation at all, no dashboard, no account on the buyer's system, nothing to appeal, since the recommendation happens entirely inside the buyer's private ordering assistant. Reduce is a rule that guarantees a minimum visibility floor for suppliers below a certain order-volume threshold, so a real quality product doesn't vanish from suggestions purely for having less historical data. Detect is a quarterly audit comparing recommendation share against actual product ratings, watching for exactly this kind of volume-based crowding out.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "watermark every clip, flag the same target across accounts, and give the person filmed a real contact path," and stop.
Cost: there's no engineering budget this quarter for an automated cross-account flag. Say so honestly, and start with a manual weekly review of takedowns grouped by target instead of by account, even done by hand.
The model gets better, for real: if Reelhaven's overall clip quality improves, that's still not a reason to drop the watermark, a rarer misuse case can still be the one that reaches a news camera.
Where people run it wrong.
They review every flagged clip in isolation, one account at a time, and never notice the same target being hit by a rotating cast of new accounts.
They design the appeal path only for the uploader, since the uploader is the one who's logged in, and forget the person who never had an account to begin with.
They wait for a news story to build the detection metric that could have shown the pattern two months earlier.
How to use it live. When someone asks how you'd handle a user population misusing your product, don't start with moderation policy. Start by naming the second person in the story, the one with no account, no login, and no way to see what's being built out of their own life, and ask what lever you'd actually give them.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "how would you handle a user population using your product in an unintended harmful way"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Ability to contest is the step that asks who never gets to push back.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Marguerite Solano, trust and safety product manager at Reelhaven, whose ordinary moderation queue turned into tracking a rotating-account harassment pattern.
3 · THE HABIT
What did the moderation team do, one clip at a time, that hid the real pattern?
Tap to flip
ANSWER
They reviewed every flagged clip by account, one at a time, instead of by who was being targeted, so each takedown looked like the system working, not a repeating pattern.
4 · WHO CAN'T PUSH BACK
Name the two people in this story. Which one never gets a lever?
Tap to flip
ANSWER
The uploader, who chose the tool, and the person filmed, who never chose anything. The person filmed has no account, no appeal, and no way to know the clip exists.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Shipping auto-clips with no watermark, since an unbranded clip felt more shareable. It made sense while clips mostly stayed inside small friend groups.
6 · THE NUMBER
Fill in the blank: accounts under 7 days old made up ___ percent of flagged harmful clips, versus 3 percent from accounts over a year old.
Tap to flip
ANSWER
42 percent. A wildly disproportionate share, pointing at accounts built specifically for this misuse, not ordinary users drifting into it.
7 · THE REPLAY
Same compilation attempt, redesigned tool. What changes?
Tap to flip
ANSWER
The system catches the same face flagged across five new accounts within a week and freezes the pattern for review, instead of waiting for a viewer to report it after it's already live.
8 · CROSS PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which product, and who can't push back there?
Tap to flip
ANSWER
Larder & Line's restaurant ordering assistant. The small independent supplier being quietly under-recommended has no visibility or appeal into the buyer's private assistant at all.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: by week 8, weekly reports of harassment-compilation clips had climbed to ___.
Show hint
Look at the line chart tracking weekly reports before the news story.
Show answer
47. It started at 3 in week 1, and nobody had a metric tracking "same target, many accounts" until week 14.
True or false
2. True or false: this answer recommends adding review friction to every clip Reelhaven's auto-clip tool generates.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. Ordinary clips stay exactly as fast as before. The reduction targets the measurable misuse pattern, a watermark and a cross-account flag, not a blanket slowdown.
Multiple choice
3. Why did reviewing flagged clips one account at a time fail to catch the real pattern?
A. The moderation team wasn't trained properly.
B. The auto-clip tool had a bug in how it stitched videos.
C. Each individual takedown looked like the system working correctly, since the pattern only appears when you compare across accounts by target, not within one account.
D. Reelhaven didn't have a moderation team at all.
Show hint
Look at the flow diagram and the quadrant showing hidden versus obvious harm.
Show answer
C. The harassment compilation and doxxing-adjacent categories sit in the hardest-to-spot corner, since a per-account queue only ever looks at one account at a time.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Shipping clips with no watermark, since an unbranded clip felt more shareable. It made sense while clips mostly stayed inside small friend groups, before anyone needed to trace one back to its source.
Short answer, where it wouldn't matter
5. Name a use of the auto-clip tool where none of this new friction applies.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A parent stitching a highlight reel of their kid's soccer season, or a friend group's trip clips. Ordinary use stays exactly as fast as before.
Short answer, apply it yourself
6. Pick a product you use that lets one person generate content involving another person. Who in that setup has no account and no way to push back?
Show hint
Think about tagging, photo-sharing, or clip tools where the subject of the content isn't the one who uploaded it.
Show answer
Model answer: Most people land on photo tagging or clip-sharing tools, where the person tagged or filmed often has no account on the app at all and no way to see or contest what's been made.
Before you close the answer
Why this works
Tests whether you can name the person with zero power in the room, the one with no account and no appeal, instead of only designing for the account holder who's easy to picture because they're the one who's logged in.
Follow-up traps
"Won't a visible watermark just make the clip less shareable and hurt engagement?" Response: yes, slightly, and that's the accepted cost. A clip nobody can trace back to its source is a clip nobody can act on when it's used to harm someone.
"Isn't a rate limit on same-target clips going to catch legitimate reaction videos and remixes too?" Response: the flag is for review, not automatic takedown, so a legitimate remix gets a human look, while a rotating-account harassment pattern gets caught instead of resurfacing under a new name every week.
If pressed
Reelhaven's real cross-account flag uses a perceptual hash of the filmed face, not just the account or the clip file, so a new account re-uploading the same underlying footage still trips the same flag even with a different username and a re-encoded video.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.