What is the product cost of a guardrail that is too aggressive?
Solace Telecom is a phone carrier. Solace Assistant handles billing questions and disputes in chat, for consumers and small businesses alike. Ezinne Okoro runs a catering business and has used Solace Assistant to fix billing mistakes for four years.
- Replace the blanket keyword refusal with a routing decision: flagged cases go to stricter review, not a dead end.Why: refusing outright teaches people to hide the details a reviewer would actually need.
- Watch how vague incoming messages get, not just how many get refused.Why: a refusal rate can look stable while the honesty of what people say to you quietly collapses.
- Separate "sounds like fraud" from "sounds like a genuine dispute with specific numbers," since they overlap in language but not in intent.Why: the same words, an amount, a date, a request to reverse a charge, show up in both, and only one of them is the threat.
- Measure round trips per resolved dispute, before and after any guardrail change.Why: a customer who has to restate their issue three extra times has already paid the guardrail's real cost.
- Keep the guardrail strict on account-changing actions, like adding a new payment method.Why: that's genuinely higher-stakes territory, and pre-editing a request there is a much smaller loss than pre-editing a billing complaint.
- Tell customers plainly when something is under stricter review, instead of going silent.Why: silence is what taught Ezinne to guess at what was safe to say, instead of just being told.
How to answer this, stage by stage
Six moves, since this is a focused concept question, not a sprawling one.
Let's learn
Say a phone carrier builds a chat assistant to handle billing disputes. A customer types out what went wrong, and the assistant checks the account and, in most cases, fixes it on the spot.
For a long time, this worked well. A typical dispute, described in full detail, took about one round trip to resolve: the customer explained the specific charge, and the assistant fixed it. Average resolution time sat close to a single day.
Then one attempted fraud got through: a caller falsely claimed a duplicate charge and asked for a refund to a new account. In response, the team added a filter that flagged messages combining a dollar amount, a request to reverse a charge, and specific dispute language, refusing them outright and routing the customer to a generic "please call support" message.
Here is the turn. The filter caught the fraud pattern. It also caught nearly every real, honest billing dispute, since a real dispute needs exactly those things: an amount, a date, a request to fix it. The extra refusals were never really the problem on their own. The problem is what customers did next: they learned which words got them blocked, and started leaving them out.
At its worst: a genuine dispute takes four or five exchanges to resolve, each one a little vaguer than a real complaint should need to be, while the assistant asks clarifying questions it wouldn't have needed if the customer had just been able to say the truth plainly the first time.
What I would leave alone: the guardrail around account-changing actions, like adding a new payment method mid-chat, should stay just as strict. Nobody's stripping honest detail out of "please add a card," so the same aggressive rule costs much less there than it does on a billing complaint.
The lesson: a guardrail's cost doesn't show up as a refusal count. It shows up as a quiet change in what people are willing to tell you, and that number doesn't appear on any dashboard unless someone goes looking for it.
Now here is the same thing as a story
The short version above is what you'd say in the room. Read this one for how Ezinne actually found the workaround.
Ezinne Okoro has run her catering business for six years and used Solace's small-business line for most of that time. She knows her monthly bill down to the line item, because she's the one who reconciles it against her own receipts every month.
For most of those six years, fixing a billing mistake through Solace Assistant took one message. She'd type exactly what was wrong: "I was charged 340 dollars twice for July's international add-on, please refund the duplicate." The assistant checked, confirmed, and refunded it, usually within the same conversation.
Then Solace had a scare: someone had used the chat to falsely claim a duplicate charge and get a refund routed to a different account entirely. The response, shipped within a week, was a filter that flagged any message combining a specific dollar amount, a request to reverse a charge, and dispute language, and refused it outright with a generic message asking the customer to call in instead.
The next time Ezinne had a real billing dispute, an actual duplicate charge, she wrote it exactly the way she always had. The chat refused her and pointed her to a phone line with a forty-minute wait. She tried again, worded slightly differently. Refused again.
A friend, another small-business owner on the same carrier, told her what had worked for her: don't name the amount, don't say "duplicate" or "refund," just ask "can you check my July bill" and let the assistant ask its own follow-up questions from there.
It worked, in the sense that the message got through. It also meant every real dispute she had from then on took three or four extra exchanges, since the assistant now had to ask the very questions she used to just answer up front.
We did not make her more careful. We made her vaguer, and vaguer is not the same thing as safer.
Solace eventually redesigned the guardrail: instead of an outright refusal, a flagged message now routes to a stricter, faster-turnaround human review lane, with the customer told plainly that their request is being checked more carefully, rather than told nothing at all. Ezinne can say the whole truth again. Her last two disputes each took one message and a same-day fix.
What I would tell myself, back when the filter first shipped: refusing a message doesn't make the underlying request go away. It just makes the next version of that request harder to read.
The five steps, if you want to remember itFLIPS finds the behavior change a guardrail actually causes, not just the refusal count it produces.
The recap, one line per letter: find is Ezinne, a specific small-business customer; locate is her habit of full, honest detail; identify is the flip between full truth and safe vagueness; pinpoint is one merged filter built right after a fraud scare; show is resolution time returning from four days to about one.
And if you want to be sure it really works, try it somewhere elseSame five letters, an airline's claims chatbot instead of a phone carrier. A different flip family entirely.
Windrow Airlines built Windrow Assistant to handle delayed-baggage compensation claims by chat. After one attempted double-claim, the team added a guardrail that flagged any message mentioning a specific bag, a dollar amount, and a resubmission in the same conversation.
Mapped onto FLIPS: find is Teo Alvear, a frequent business traveler who's filed a dozen legitimate claims over the years. Locate is his habit of just replying directly in the same chat thread when a claim needed a correction. Identify is a workaround flip, not a pre-editing one: instead of rewording his message, he starts keeping his own screenshots and re-filing entirely new claims by email whenever the chat stalls, building a private process the airline can't see or support. Pinpoint is the decision to flag and stall any resubmission-sounding message in-thread, rather than giving a clear, visible path to amend an existing claim. Show is that once a visible "amend this claim" button replaces the guardrail's silent stall, Teo's corrections go back to taking one message instead of a parallel email thread nobody on the airline's side ever sees.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "an aggressive guardrail doesn't reduce risky requests, it strips the honest detail out of genuine ones," and stop.
Cost: there's no engineering time this quarter to build a full review-routing lane. Say so honestly, and start with a visible message telling flagged customers what tripped the filter, even without rebuilding the routing yet.
The model gets better, for real: if the guardrail's fraud detection accuracy improves, that's still not proof the cost is gone. A more accurate filter can still strip honest detail from genuine disputes if it refuses instead of routing them.
Where people run it wrong.
They measure the guardrail's success by refusal count, not by what happens to the requests that get refused.
They build one rule to catch both fraud and genuine edge cases, because it's faster than building two.
They let a refusal go silent, with no explanation, which teaches the customer to guess instead of simply telling them what's happening.
How to use it live. If someone asks "isn't a strict guardrail always safer," answer with the actual number: how many extra exchanges, or how much extra time, a real customer now needs. That's the concrete cost a strict guardrail rarely gets billed for.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"How do you know the vaguer wording is actually caused by the guardrail, and not something else?" Response: the round-trip and resolution-time numbers move in lockstep with the guardrail's tightening and its later redesign, both up and back down, which is the strongest evidence available short of asking every customer directly.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Responsible AI as a product requirement
- #1 How do you turn a responsible AI principle into a testable product requirement?
- #2 What safety requirements belong in every AI PRD regardless of feature?
- #3 Describe how you would assess a feature for potential harm before building it.
- #4 Explain the difference between a safety issue and a quality issue.
- #5 How would you handle a feature that works well overall but poorly for one demographic?
- #6 What is a content policy and who should own it in a product organization?