How should the product behave when the model produces output that fails a safety filter?
Thriftloop is a resale marketplace. Its listing assistant turns a seller's photos and rough notes into a finished description, then a separate safety filter scans that description before it goes live. Marisol Kwan sells curated estate-sale finds, including antique scientific and medical instruments, and her filter story is the one worth telling.
- Give every block a specific, written reason, never a form notice.Why: "doesn't meet our guidelines" gives a seller nothing to fix and nothing to argue with.
- Hold blocked listings in review, not deleted, with a real appeal button.Why: a seller with no lever at all can only guess, quit, or start hiding words.
- Track the false-block rate by category, not just the catch rate.Why: an average false-block rate hides one small seller group getting blocked at ten times everyone else's rate.
- Set the block threshold per category instead of one number for the whole site.Why: the same word means something dangerous in one listing and something ordinary in another.
- Route repeat false blocks in a category to a person, not another automatic pass.Why: a pattern of wrong calls in the same category is a bug in the rule, not one unlucky listing.
- Leave the filter's true catches, weapons, scams, counterfeits, exactly as strict as they are today.Why: this fix is about the wrongly blocked, not about loosening the real protection.
How to answer this, stage by stage
Nobody is grading whether you can describe a safety filter. They're grading whether you noticed it can be wrong in a direction nobody's watching.
Let's learn
Before Marisol ever opened a laptop, she spent about six hours a week writing her own listing copy for the estate-sale pieces she found: apothecary bottles, brass loupes, a boxed set of 1920s surgical instruments sold for display and collecting, not use. She wrote maybe eighteen listings a week that way, and every one of them went live the moment she posted it.
Now Thriftloop's assistant drafts the copy for her in seconds from a few photos, and she posts thirty or more listings a week instead of eighteen.
Here's the turn: the extra listings were never the problem. The problem is that Thriftloop's safety filter, built to catch weapon and drug-paraphernalia listings, treats words like "scalpel" and "surgical" as red flags no matter what they're attached to. Marisol's antique-instrument listings get blocked at a rate wildly higher than the rest of the site, and every block arrives as the same four words: "doesn't meet our guidelines."
At its worst, an honest seller gets treated as a repeat offender by a filter that never learns the difference between a scalpel in a listing and a scalpel in a threat, and quits the category, or the platform, without a single support ticket to explain why.
What I would leave alone: the filter's true catches don't need softening. A listing that's actually selling a weapon or a counterfeit should still get blocked instantly, no appeal delay, no second-guessing.
The lesson: a safety filter is a guess, not a verdict. The moment a guess is delivered with no reason and no way to push back, the person on the receiving end has no choice left but to disappear.
Now here is the same thing as a story
The short version above is what you'd say defending this redesign to Thriftloop's trust and safety lead. Read this one for how quietly it wore Marisol down.
Marisol Kwan is good at exactly one unglamorous thing: knowing what an estate sale is actually worth before anyone else in the room does. Nine years of Saturday mornings taught her which box of "junk" has a hand-cranked centrifuge worth four hundred dollars sitting at the bottom of it.
Her first few months on Thriftloop were good ones. She'd photograph a find, let the assistant draft the copy, tweak a line or two, and post. Instruments, bottles, brass tools, all of it went up clean.
Then a listing for a boxed surgical set got blocked. She rewrote it without the word "scalpel," swapped in "blade," and it got blocked again. She rewrote it a third time with no medical words at all, just "vintage tool set," and it finally went live, worth less to a buyer searching for exactly what it was.
Over two months, three more listings got the same treatment. She started stripping every precise word from her copy before the assistant even touched it, and buyers who searched for "surgical" or "apothecary" stopped finding her at all. Then, at a regional collectors' meetup, another dealer said, half-joking, "you're still fighting that filter? I just stopped listing anything with 'surgical' in the title months ago."
That was the moment she stopped assuming it was her own bad luck. She filed a support ticket demanding an actual human look at the pattern, not just her one listing, and Thriftloop's trust team finally pulled the numbers: her category was being blocked seventeen times more often than the site average, with no matching rise in actual violations.
With the redesign, a blocked listing now says exactly which phrase tripped the filter, offers a specific edit or an appeal button, and repeat false blocks in the same category get routed to a person within a day instead of bouncing off the same rule again. Run the same two months forward: Marisol's first block comes back within a day with "flagged for 'scalpel,' likely a false match, appeal reviewed by a person," and she never has to strip a single accurate word from her own listings.
The old filter asked Marisol to trust a silence. The new one shows her the actual word it flagged.
I signed off on the instant, unexplained block because it felt like the safe default, better to over-block than under-block. It took watching one honest seller category get quietly ground down to see that "safe" and "silent" were never the same thing.
GUARD, in one screenNot a lecture on content policy. GUARD is what tells you whose morning gets ruined by a wrong call nobody explains.
The recap, one line per letter: groups is the trust team and the seller, unequal is the instrument category's absurd false-block rate, ability to contest is the empty appeal box, reduce is a real reason plus a real appeal, and detect is a false-block dashboard split by category.
And if you want to be sure it really works, try it somewhere elseSame five letters, a B2B help desk instead of a resale marketplace. The blocked thing is a support reply, not a listing.
HelpCrate is a support-ticketing platform where an AI drafts reply suggestions for agents, and a separate safety filter blocks any draft that looks like it might promise a refund or a legal commitment the company hasn't approved. Dabir Farrow is a support agent at a small hardware retailer who uses HelpCrate daily. Mapped onto GUARD: groups is Dabir and the policy team that wrote the refund-language filter; unequal is that agents on the warranty-claims queue get blocked far more than agents on the general queue, since warranty replies naturally use words like "replace" and "reimburse."
The ability-to-contest gap here is structurally identical: a blocked draft just vanishes from Dabir's suggestion box with no note, so he assumes the assistant "doesn't understand warranty cases" and stops trusting its drafts for that entire queue, typing everything from scratch again. The reduce step: show the specific phrase the filter caught, and let a supervisor clear a whole approved phrase, like "we'll replace the unit under warranty," so it stops tripping the filter for every agent, not just Dabir, the next time it's used correctly.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "give every block a reason, hold instead of delete, and add an appeal, then track false blocks by category," and stop.
Cost: there's no time to review every historical block right now. Say so honestly, and start with whichever category shows the widest gap between its block rate and its actual violation rate.
The model gets better, for real: if the filter's overall accuracy genuinely improves, that's still not a reason to skip the appeal path, a rarer false block is still a false block, and it's the one the improved model will be most confidently wrong about.
Where people run it wrong.
They watch the filter's overall catch rate and call it healthy, without ever splitting it by category.
They treat a silent block as neutral, when a person on the other end always builds a theory to fill the silence, usually the wrong one.
They wait for a support ticket or a public complaint to notice a pattern, instead of asking upfront which group has no way to push back at all.
How to use it live. When someone asks how a product should behave after a safety block, ask yourself one question first: if this exact call is wrong, does the person it happened to have any way to find out, or fix it? Design for the answer being no.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a human review queue just slower and more expensive?" Response: yes, and it's scoped to repeat false blocks in one category, not every listing, so the added cost lands only where the filter is actually wrong often.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Designing for failure and graceful degradation
- #1 What should happen in the UI when the model returns nothing usable?
- #2 Design the fallback experience for an AI feature when the provider is down.
- #3 Explain the difference between failing loudly and failing silently, and which you prefer.
- #4 How do you design a feature that degrades to a non-AI version rather than breaking?
- #5 Describe three failure modes to design for before launch.
- #6 What error message would you write for a model timeout, and what would you avoid saying?