CaseAdvancedDesigning for Uncertainty & Trust / Designing for failure and graceful degradation / #7

How should the product behave when the model produces output that fails a safety filter?

GUARD a blocked listing has an operator who can flip the switch, and a seller who can't

Thriftloop is a resale marketplace. Its listing assistant turns a seller's photos and rough notes into a finished description, then a separate safety filter scans that description before it goes live. Marisol Kwan sells curated estate-sale finds, including antique scientific and medical instruments, and her filter story is the one worth telling.

The direct answer
Never let a safety block be a dead end. Give every blocked listing a specific reason, hold it in a reviewable "needs a look" state instead of deleting it, and give the seller a real appeal path. Track false blocks by category, not just true catches, so a filter that quietly punishes one kind of seller gets caught by a dashboard instead of by that seller quitting.
Do this, in order
  1. Give every block a specific, written reason, never a form notice.Why: "doesn't meet our guidelines" gives a seller nothing to fix and nothing to argue with.
  2. Hold blocked listings in review, not deleted, with a real appeal button.Why: a seller with no lever at all can only guess, quit, or start hiding words.
  3. Track the false-block rate by category, not just the catch rate.Why: an average false-block rate hides one small seller group getting blocked at ten times everyone else's rate.
  4. Set the block threshold per category instead of one number for the whole site.Why: the same word means something dangerous in one listing and something ordinary in another.
  5. Route repeat false blocks in a category to a person, not another automatic pass.Why: a pattern of wrong calls in the same category is a bug in the rule, not one unlucky listing.
  6. Leave the filter's true catches, weapons, scams, counterfeits, exactly as strict as they are today.Why: this fix is about the wrongly blocked, not about loosening the real protection.

How to answer this, stage by stage

Nobody is grading whether you can describe a safety filter. They're grading whether you noticed it can be wrong in a direction nobody's watching.

Stage 1
Scope it to one platform, one blocked listing
Say it like this
"I'll use Thriftloop, a resale marketplace, and a seller named Marisol whose antique instrument listings kept getting blocked by the safety filter."
Why this works
Stops the answer from turning into a generic content-moderation essay.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups affected, where the harm lands unequally, ability to contest, the design change to reduce it, and how you'd detect it in production."
Why this works
Shows the interviewer you have a repeatable way to reason about risk, not just a strong opinion.
Stage 3
Name both people the filter touches
Say it like this
"Thriftloop's trust team can always override a block. Marisol, on the other end of it, gets a form notice and nothing else. She's the one with no lever."
Why this works
This is GUARD's sharpest move: naming who can push back and who genuinely can't.
Stage 4
Give the one decision
Say it like this
"A block should never be silent. Give a specific reason, hold the listing instead of deleting it, and add a real appeal. Then track false blocks by category so a whole seller group getting hit isn't invisible."
Why this works
This is the direct answer, said as a concrete build decision, not a value statement about safety.
Stage 5
Name the trade-off out loud
Say it like this
"Loosening the threshold for a category means a genuinely bad listing in that category takes a little longer to catch, since it now waits for a person instead of an instant block. That's a real cost, and it's still worth it against silently punishing an honest seller group for months."
Why this works
Shows you're not pretending the fix is free.
Stage 6
Close on the one line
Say it like this
"A safety filter that can't explain itself and can't be appealed isn't safer, it's just unaccountable. The fix isn't a softer filter, it's a filter that has to show its work."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Before Marisol ever opened a laptop, she spent about six hours a week writing her own listing copy for the estate-sale pieces she found: apothecary bottles, brass loupes, a boxed set of 1920s surgical instruments sold for display and collecting, not use. She wrote maybe eighteen listings a week that way, and every one of them went live the moment she posted it.

Now Thriftloop's assistant drafts the copy for her in seconds from a few photos, and she posts thirty or more listings a week instead of eighteen.

Hand sketched two-panel scene titled Two people, one lever. Left figure in blue labeled Trust Team, caption holds the lever. Right figure in red labeled Marisol, caption empty hands, no appeal.
Same block, same day. One side of it can always flip the outcome. The other side can only guess.

Here's the turn: the extra listings were never the problem. The problem is that Thriftloop's safety filter, built to catch weapon and drug-paraphernalia listings, treats words like "scalpel" and "surgical" as red flags no matter what they're attached to. Marisol's antique-instrument listings get blocked at a rate wildly higher than the rest of the site, and every block arrives as the same four words: "doesn't meet our guidelines."

False-block rate by seller category, last quarter
70% 35 0 Site avg, 2% Clothing, 3% Electronics, 5% Instruments, 34% Weapon replicas, 61%
Weapon replicas belong up there, most of those blocks are correct. Antique instruments don't, and nothing on the dashboard told anyone the difference.

At its worst, an honest seller gets treated as a repeat offender by a filter that never learns the difference between a scalpel in a listing and a scalpel in a threat, and quits the category, or the platform, without a single support ticket to explain why.

Hand sketched quadrant titled Which categories get blocked without cause. Axes: real violation rate low to high, block rate low to high. Vintage clothing and vintage electronics sit low on both. Antique medical instruments sits low on real violations but high on block rate. Weapon replicas sit high on both.
Only the top-left corner is a real problem. The top-right corner is the filter doing its job.
The decision I would take back Thriftloop's early trust and safety team built the filter to reject instantly and show a single, generic notice for every category, on the theory that a fast, uniform block was safer than a slow, explained one. That made sense when the filter mostly caught obvious cases. It stopped making sense the day an honest seller category started absorbing the same blunt block meant for actual threats.

What I would leave alone: the filter's true catches don't need softening. A listing that's actually selling a weapon or a counterfeit should still get blocked instantly, no appeal delay, no second-guessing.

The lesson: a safety filter is a guess, not a verdict. The moment a guess is delivered with no reason and no way to push back, the person on the receiving end has no choice left but to disappear.

Now here is the same thing as a story

The short version above is what you'd say defending this redesign to Thriftloop's trust and safety lead. Read this one for how quietly it wore Marisol down.

Marisol Kwan is good at exactly one unglamorous thing: knowing what an estate sale is actually worth before anyone else in the room does. Nine years of Saturday mornings taught her which box of "junk" has a hand-cranked centrifuge worth four hundred dollars sitting at the bottom of it.

Her first few months on Thriftloop were good ones. She'd photograph a find, let the assistant draft the copy, tweak a line or two, and post. Instruments, bottles, brass tools, all of it went up clean.

Knowledge spark: why would a safety filter block something harmless? Most safety filters are trained to catch a pattern, not to understand a sentence. A model that's learned "scalpel" shows up near weapon and self-harm listings will flag it wherever it appears, including a 1920s display set a collector is proud of. It isn't reading intent. It's matching a word to a risk it was trained to fear.

Then a listing for a boxed surgical set got blocked. She rewrote it without the word "scalpel," swapped in "blade," and it got blocked again. She rewrote it a third time with no medical words at all, just "vintage tool set," and it finally went live, worth less to a buyer searching for exactly what it was.

Hand sketched flow diagram titled Where the appeal should be, and isn't. Five boxes: Listing drafted, Filter scans it, Blocked no reason highlighted, Appeal path none, Seller guesses why.
The fourth box is empty on purpose. That's the actual bug.
The filter did not lie. It just never said which word it was afraid of, so Marisol had to guess with her own inventory.

Over two months, three more listings got the same treatment. She started stripping every precise word from her copy before the assistant even touched it, and buyers who searched for "surgical" or "apothecary" stopped finding her at all. Then, at a regional collectors' meetup, another dealer said, half-joking, "you're still fighting that filter? I just stopped listing anything with 'surgical' in the title months ago."

Hand sketched timeline titled Marisol's two months with the filter. Five milestones: First block scalpel set listing, Rewrites blocked again swapped in blade, A dealer's remark just drop the word highlighted, Escalates to support reaches a human, Redesign ships reason codes and appeal.
The remark didn't teach her anything new. It just told her she wasn't the only one who'd quietly given up.

That was the moment she stopped assuming it was her own bad luck. She filed a support ticket demanding an actual human look at the pattern, not just her one listing, and Thriftloop's trust team finally pulled the numbers: her category was being blocked seventeen times more often than the site average, with no matching rise in actual violations.

Hand sketched labeled parts diagram titled What a good block notice needs. A document icon at the center labeled Block Notice, with four callouts around it: reason code, what to fix, appeal link, review time.
The old notice had none of these four. Marisol built her own folk theory instead, one blocked listing at a time.

With the redesign, a blocked listing now says exactly which phrase tripped the filter, offers a specific edit or an appeal button, and repeat false blocks in the same category get routed to a person within a day instead of bouncing off the same rule again. Run the same two months forward: Marisol's first block comes back within a day with "flagged for 'scalpel,' likely a false match, appeal reviewed by a person," and she never has to strip a single accurate word from her own listings.

Hand sketched icon list titled What the redesign adds. Four items: a document icon labeled A real reason not a form notice, a gauge icon labeled Block limit set per category, a person icon labeled A person reviews repeat misses, a funnel icon labeled Wrong blocks counted by category.
None of these four existed before. All four came out of one seller's meetup complaint.

The old filter asked Marisol to trust a silence. The new one shows her the actual word it flagged.

I signed off on the instant, unexplained block because it felt like the safe default, better to over-block than under-block. It took watching one honest seller category get quietly ground down to see that "safe" and "silent" were never the same thing.

GUARD, in one screenNot a lecture on content policy. GUARD is what tells you whose morning gets ruined by a wrong call nobody explains.

G
Groups. Who is affected.
Thriftloop's trust team, who set the rule and can always override it, and Marisol, who receives the block with no say in it.
Names both sides before assigning blame to either.
U
Unequal. Where the harm lands hardest.
Antique instrument sellers, a small niche, absorb a false-block rate seventeen times the site average, with no matching rise in real violations.
Shows the harm isn't spread evenly, it's concentrated on one group nobody was watching.
A
Ability to contest. Who never gets to push back.
A blocked seller gets a form notice with no reason and no appeal, so she can't inspect the call, argue it, or even know what to change.
This is the hardest step and the heart of the answer: naming the person with no lever at all.
R
Reduce. The specific design change.
Show the exact phrase that tripped the filter, hold the listing for review instead of deleting it, and give a real appeal button.
A concrete product decision, not a policy memo about being more careful.
D
Detect. How you'd know in production.
Track false-block rate by category on a standing dashboard, not just overall catch rate, so a group like Marisol's shows up as a number before it shows up as a quiet exodus.
Turns "we found out from a support ticket" into "we found out from our own data."

The recap, one line per letter: groups is the trust team and the seller, unequal is the instrument category's absurd false-block rate, ability to contest is the empty appeal box, reduce is a real reason plus a real appeal, and detect is a false-block dashboard split by category.

Hand sketched decision tree titled When a blocked reply gets a human look. Root: auto-reply flagged by filter. Three branches: first time low-risk group leads to hold for edit, repeat wrong block group leads to send to a person, real policy break leads to confirm block.
This same tree runs the identical redesign inside a completely different product, a support inbox instead of a marketplace.

And if you want to be sure it really works, try it somewhere elseSame five letters, a B2B help desk instead of a resale marketplace. The blocked thing is a support reply, not a listing.

HelpCrate is a support-ticketing platform where an AI drafts reply suggestions for agents, and a separate safety filter blocks any draft that looks like it might promise a refund or a legal commitment the company hasn't approved. Dabir Farrow is a support agent at a small hardware retailer who uses HelpCrate daily. Mapped onto GUARD: groups is Dabir and the policy team that wrote the refund-language filter; unequal is that agents on the warranty-claims queue get blocked far more than agents on the general queue, since warranty replies naturally use words like "replace" and "reimburse."

The ability-to-contest gap here is structurally identical: a blocked draft just vanishes from Dabir's suggestion box with no note, so he assumes the assistant "doesn't understand warranty cases" and stops trusting its drafts for that entire queue, typing everything from scratch again. The reduce step: show the specific phrase the filter caught, and let a supervisor clear a whole approved phrase, like "we'll replace the unit under warranty," so it stops tripping the filter for every agent, not just Dabir, the next time it's used correctly.

Share of HelpCrate's AI-drafted replies agents actually used, by queue, before and after the phrase-clearing fix
100% 50 0 71% 74% General queue 22% 68% Warranty queue
The general queue barely moved. The warranty queue, the one absorbing nearly all the false blocks, nearly tripled once cleared phrases stopped being re-flagged.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "give every block a reason, hold instead of delete, and add an appeal, then track false blocks by category," and stop.
Cost: there's no time to review every historical block right now. Say so honestly, and start with whichever category shows the widest gap between its block rate and its actual violation rate.
The model gets better, for real: if the filter's overall accuracy genuinely improves, that's still not a reason to skip the appeal path, a rarer false block is still a false block, and it's the one the improved model will be most confidently wrong about.

Where people run it wrong.
They watch the filter's overall catch rate and call it healthy, without ever splitting it by category.
They treat a silent block as neutral, when a person on the other end always builds a theory to fill the silence, usually the wrong one.
They wait for a support ticket or a public complaint to notice a pattern, instead of asking upfront which group has no way to push back at all.

How to use it live. When someone asks how a product should behave after a safety block, ask yourself one question first: if this exact call is wrong, does the person it happened to have any way to find out, or fix it? Design for the answer being no.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a safety, risk, or fairness question?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Name who can push back and who can't before designing the fix.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Kwan, a Thriftloop seller of antique scientific and medical instruments, whose honest listings kept tripping a filter built for weapon and drug-paraphernalia language.
3 · UNEQUAL HARM
Where did the harm land hardest, and why?
Tap to flip
ANSWER
Antique instrument sellers, a small niche whose honest vocabulary happened to match the filter's weapon-language pattern, absorbed a false-block rate seventeen times the site average.
4 · ABILITY TO CONTEST
Who never got to push back here?
Tap to flip
ANSWER
Marisol. A blocked listing gave her a form notice with no reason and no appeal, so she had no way to inspect, argue, or fix the call.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building the filter to reject instantly with one generic notice for every category, reasonable when it mostly caught obvious cases, wrong once an honest category started absorbing the same blunt block.
6 · THE NUMBER
Fill in the blank: the antique instrument category's false-block rate was ___ percent, against a site average of 2 percent.
Tap to flip
ANSWER
34 percent, seventeen times the site average, with no matching rise in real violations in that category.
7 · THE REDESIGN
Same two months, redesigned filter. What changes?
Tap to flip
ANSWER
Marisol's first block returns within a day naming the exact flagged phrase and an appeal reviewed by a person, and she never has to strip an accurate word from a listing again.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the reduce step there?
Tap to flip
ANSWER
HelpCrate, a support-ticketing platform. The reduce step lets a supervisor clear an approved warranty phrase so it stops tripping the filter for every agent, not just the one who first hit it.

Check yourself Score: 0 / 0

True or false
1. True or false: in this answer, the safety filter's overall catch rate for real violations should be loosened.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The fix targets false blocks in one category. Real catches, like weapon replica listings, stay exactly as strict.
Multiple choice
2. Why does this count as a GUARD-shaped question rather than a plain design question?
  • A. Because the product involves a marketplace.
  • B. Because one group is affected far more than others, and that group has no way to contest a wrong call.
  • C. Because Thriftloop's engineering team is small.
  • D. Because the filter runs instantly instead of overnight.
Show hint
Look at the Unequal and Ability to Contest steps.
Show answer
B. A concentrated, unequal harm with no lever for the affected group is exactly what GUARD is built to surface.
Fill in the blank
3. Fill in the blank: HelpCrate's warranty queue saw its AI-draft usage rise from 22 percent to ___ percent after the phrase-clearing fix.
Show hint
Look at the grouped bar chart in Section 4.
Show answer
68 percent. The general queue barely moved, which is what shows the fix targeted the right group.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Rejecting instantly with one generic notice for every category. It made sense while the filter mostly caught obvious, clearly bad listings, and stopped making sense once an honest category absorbed the same blunt treatment.
Short answer, apply it yourself
5. Think of a time an app rejected or flagged something you did with no explanation. What did you assume the reason was, and were you right?
Show hint
Ask whether you ever found out the real reason, or just built your own guess and worked around it.
Show answer
Model answer: Most people can recall guessing wrong at least once, exactly the folk-theory problem Marisol's story shows.
Short answer, where it wouldn't matter
6. Name a case on Thriftloop where this same appeal-and-reason redesign genuinely doesn't need to slow anything down.
Show hint
Look at the quadrant diagram's bottom-right, weapon replicas.
Show answer
Model answer: A listing that's an actual weapon or counterfeit. Those blocks are usually correct, so they should still be instant with no appeal delay.
Before you close the answer
Why this works
Tests whether you'll design for the person a safety system silently punishes, or just describe the filter itself and call the job done.
Follow-up traps
"Won't showing the exact flagged phrase teach bad actors how to dodge the filter?" Response: the same information already leaks out through trial and error, like Marisol rewriting her listing three times. Showing it to everyone just removes the unequal cost of learning it the hard way.

"Isn't a human review queue just slower and more expensive?" Response: yes, and it's scoped to repeat false blocks in one category, not every listing, so the added cost lands only where the filter is actually wrong often.
If pressed
Thriftloop's category-aware thresholds are reviewed by the trust and safety team on a set schedule, not silently auto-tuned, since loosening a threshold for one category is a judgment call about real-world risk, not something a model should decide for itself.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more