ConceptAdvancedResponsible AI & Advanced Practice / Responsible AI as a product requirement / #13

What is the product cost of a guardrail that is too aggressive?

FLIPS the product is Solace Assistant, Solace Telecom's billing chat for small-business customers

Solace Telecom is a phone carrier. Solace Assistant handles billing questions and disputes in chat, for consumers and small businesses alike. Ezinne Okoro runs a catering business and has used Solace Assistant to fix billing mistakes for four years.

The direct answer
An overly aggressive guardrail costs you the honest detail in what customers tell you. Once specific words trip a refusal, people learn to strip those words out, and the request that reaches you is vaguer, not safer. The guardrail was built to catch fraud. Its real cost shows up in genuine disputes, which get slower and harder to resolve because the very details that would prove them true are exactly what the guardrail flags.
Do this, in order
  1. Replace the blanket keyword refusal with a routing decision: flagged cases go to stricter review, not a dead end.Why: refusing outright teaches people to hide the details a reviewer would actually need.
  2. Watch how vague incoming messages get, not just how many get refused.Why: a refusal rate can look stable while the honesty of what people say to you quietly collapses.
  3. Separate "sounds like fraud" from "sounds like a genuine dispute with specific numbers," since they overlap in language but not in intent.Why: the same words, an amount, a date, a request to reverse a charge, show up in both, and only one of them is the threat.
  4. Measure round trips per resolved dispute, before and after any guardrail change.Why: a customer who has to restate their issue three extra times has already paid the guardrail's real cost.
  5. Keep the guardrail strict on account-changing actions, like adding a new payment method.Why: that's genuinely higher-stakes territory, and pre-editing a request there is a much smaller loss than pre-editing a billing complaint.
  6. Tell customers plainly when something is under stricter review, instead of going silent.Why: silence is what taught Ezinne to guess at what was safe to say, instead of just being told.

How to answer this, stage by stage

Six moves, since this is a focused concept question, not a sprawling one.

Stage 1
Ground the question in one real chat feature
Say it like this
"I'll answer this for a billing dispute chatbot, since 'guardrail too aggressive' has a very specific cost there: it doesn't just refuse people, it changes what they're willing to tell you."
Why this works
Keeps the answer from staying an abstract debate about safety settings.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals a repeatable method for finding the real cost, not a guess.
Stage 3
Locate the habit that used to work
Say it like this
"Before the guardrail tightened, she described her billing problems in full, honest detail, the charge, the date, the amount, and the assistant handled most of it directly."
Why this works
Shows the thing that was working before naming what broke it.
Stage 4
Name the flip
Say it like this
"She went from describing her issue fully to stripping out anything that might trip the guardrail. Full honest detail, or vague and safe. No in-between."
Why this works
This is the specific product cost the question is asking about, named as a real behavior change.
Stage 5
Pinpoint the old decision
Say it like this
"After one fraud attempt got through, we built a single keyword filter that refused anything that sounded like a dispute, instead of routing it to a stricter human check."
Why this works
Names a real, reasonable-at-the-time decision, not a vague "we were too cautious."
Stage 6
Show the replay, and close
Say it like this
"With flagged cases routed to review instead of refused, she can say the whole truth again. Her resolution time drops back from four days to about a day. That's the cost we get back."
Why this works
Ends on something countable, which is what makes the cost real instead of theoretical.

Let's learn

Say a phone carrier builds a chat assistant to handle billing disputes. A customer types out what went wrong, and the assistant checks the account and, in most cases, fixes it on the spot.

For a long time, this worked well. A typical dispute, described in full detail, took about one round trip to resolve: the customer explained the specific charge, and the assistant fixed it. Average resolution time sat close to a single day.

Knowledge spark: what's a pre-editing flip? It's what happens when someone learns a tool's failure shape and starts grooming their own words to avoid it. They don't complain less. They just stop giving you the detail that would have let you actually help them.

Then one attempted fraud got through: a caller falsely claimed a duplicate charge and asked for a refund to a new account. In response, the team added a filter that flagged messages combining a dollar amount, a request to reverse a charge, and specific dispute language, refusing them outright and routing the customer to a generic "please call support" message.

Here is the turn. The filter caught the fraud pattern. It also caught nearly every real, honest billing dispute, since a real dispute needs exactly those things: an amount, a date, a request to fix it. The extra refusals were never really the problem on their own. The problem is what customers did next: they learned which words got them blocked, and started leaving them out.

Average round trips to resolve a billing dispute, before and after the guardrail change
5 2.5 0 Before After 1.4 4.1
Nearly three extra round trips per dispute, on average, once customers started leaving out the details the guardrail was scanning for.

At its worst: a genuine dispute takes four or five exchanges to resolve, each one a little vaguer than a real complaint should need to be, while the assistant asks clarifying questions it wouldn't have needed if the customer had just been able to say the truth plainly the first time.

The decision I would take back After the fraud attempt, we merged two separate jobs into one filter: catching fraud, and handling genuine disputes. One aggressive rule did both, because building a single rule was faster than building a routing decision. That made sense in the scramble right after the incident. It stopped making sense the moment "customer names an amount and asks for a fix" became the very pattern the rule was built to block.

What I would leave alone: the guardrail around account-changing actions, like adding a new payment method mid-chat, should stay just as strict. Nobody's stripping honest detail out of "please add a card," so the same aggressive rule costs much less there than it does on a billing complaint.

We did not just refuse her more often. We taught her that the truth was the thing getting her blocked.

The lesson: a guardrail's cost doesn't show up as a refusal count. It shows up as a quiet change in what people are willing to tell you, and that number doesn't appear on any dashboard unless someone goes looking for it.

Now here is the same thing as a story

The short version above is what you'd say in the room. Read this one for how Ezinne actually found the workaround.

Ezinne Okoro has run her catering business for six years and used Solace's small-business line for most of that time. She knows her monthly bill down to the line item, because she's the one who reconciles it against her own receipts every month.

Hand sketched timeline titled How the guardrail got here. Five milestones: launch with disputes handled directly, a scare where one fraud attempt gets through, tightened where dispute language now gets blocked, workaround where she learns to say less, and redesigned where flagged cases get reviewed instead of refused.
The middle milestone changed everything downstream of it, for every honest customer along with the one dishonest one it was built to stop.

For most of those six years, fixing a billing mistake through Solace Assistant took one message. She'd type exactly what was wrong: "I was charged 340 dollars twice for July's international add-on, please refund the duplicate." The assistant checked, confirmed, and refunded it, usually within the same conversation.

Then Solace had a scare: someone had used the chat to falsely claim a duplicate charge and get a refund routed to a different account entirely. The response, shipped within a week, was a filter that flagged any message combining a specific dollar amount, a request to reverse a charge, and dispute language, and refused it outright with a generic message asking the customer to call in instead.

Hand sketched flow diagram titled Where her honest message gets stopped. Five boxes: describes charge, names the amount, guardrail fires highlighted, chat goes silent, she starts over.
Every honest dispute runs through the same three middle boxes as the one fraud attempt ever did.

The next time Ezinne had a real billing dispute, an actual duplicate charge, she wrote it exactly the way she always had. The chat refused her and pointed her to a phone line with a forty-minute wait. She tried again, worded slightly differently. Refused again.

A friend, another small-business owner on the same carrier, told her what had worked for her: don't name the amount, don't say "duplicate" or "refund," just ask "can you check my July bill" and let the assistant ask its own follow-up questions from there.

Hand sketched comparison diagram titled A small move a big snap. Left panel, a gauge icon labeled Guardrail, caption a little stricter. Right panel, a question mark box icon labeled Her wording, caption suddenly stripped bare.
The guardrail moved a little. What she was willing to type moved all the way to the other setting.

It worked, in the sense that the message got through. It also meant every real dispute she had from then on took three or four extra exchanges, since the assistant now had to ask the very questions she used to just answer up front.

We did not make her more careful. We made her vaguer, and vaguer is not the same thing as safer.

Hand sketched icon list titled FLIPS the five letters. Five rows: F find the person, L locate the habit, I identify the flip, P pinpoint the old decision, S show the replay.
The method that found this cost, laid out in the same five steps every time.
Hand sketched metaphor scene titled What we assumed she had. Left, a gauge icon labeled A Dial, caption check a little more. Right, a box icon labeled A Switch, caption say it plain or say nothing true.
We built the filter as if honesty could be dialed down gradually. It never had a middle setting to dial.

Solace eventually redesigned the guardrail: instead of an outright refusal, a flagged message now routes to a stricter, faster-turnaround human review lane, with the customer told plainly that their request is being checked more carefully, rather than told nothing at all. Ezinne can say the whole truth again. Her last two disputes each took one message and a same-day fix.

What I would tell myself, back when the filter first shipped: refusing a message doesn't make the underlying request go away. It just makes the next version of that request harder to read.

The five steps, if you want to remember itFLIPS finds the behavior change a guardrail actually causes, not just the refusal count it produces.

F
Find the person.
Ezinne Okoro, small catering business owner, four years on Solace's small-business line.
Grounds the cost in one real customer's actual words, not an abstract complaint rate.
L
Locate the habit.
She used to describe disputes in full, specific detail, and the assistant resolved most of them in one message.
Shows the working habit before naming what broke it.
I
Identify the flip. The hard step.
Full honest detail, or stripped, vague, guardrail-safe wording. No middle setting.
This flip is the actual product cost the question is asking about.
P
Pinpoint the old decision.
Merging "catch fraud" and "handle genuine disputes" into one aggressive refusal filter, right after a real fraud scare.
Names a specific, reasonable-at-the-time choice, not a vague overcaution.
S
Show the replay.
With flagged cases routed to review instead of refused, her resolution time drops from four days back to about one.
Ends on something countable, proving the fix actually restores what was lost.
Median resolution time for a billing dispute, by month, in days
5 2.5 0 Jan Feb Mar Apr May Jun Tightened Redesigned
Resolution time nearly quadrupled after the guardrail tightened, then dropped back below its starting point once flagged cases got routed instead of refused.

The recap, one line per letter: find is Ezinne, a specific small-business customer; locate is her habit of full, honest detail; identify is the flip between full truth and safe vagueness; pinpoint is one merged filter built right after a fraud scare; show is resolution time returning from four days to about one.

And if you want to be sure it really works, try it somewhere elseSame five letters, an airline's claims chatbot instead of a phone carrier. A different flip family entirely.

Windrow Airlines built Windrow Assistant to handle delayed-baggage compensation claims by chat. After one attempted double-claim, the team added a guardrail that flagged any message mentioning a specific bag, a dollar amount, and a resubmission in the same conversation.

Mapped onto FLIPS: find is Teo Alvear, a frequent business traveler who's filed a dozen legitimate claims over the years. Locate is his habit of just replying directly in the same chat thread when a claim needed a correction. Identify is a workaround flip, not a pre-editing one: instead of rewording his message, he starts keeping his own screenshots and re-filing entirely new claims by email whenever the chat stalls, building a private process the airline can't see or support. Pinpoint is the decision to flag and stall any resubmission-sounding message in-thread, rather than giving a clear, visible path to amend an existing claim. Show is that once a visible "amend this claim" button replaces the guardrail's silent stall, Teo's corrections go back to taking one message instead of a parallel email thread nobody on the airline's side ever sees.

The same timeline pattern reused for a second product: launch, a scare, a guardrail tightening, a workaround, and a redesign, this time for an airline's baggage claims chatbot.
Same shape, different flip. This time the cost isn't vaguer wording, it's a whole shadow process the airline can't see.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "an aggressive guardrail doesn't reduce risky requests, it strips the honest detail out of genuine ones," and stop.
Cost: there's no engineering time this quarter to build a full review-routing lane. Say so honestly, and start with a visible message telling flagged customers what tripped the filter, even without rebuilding the routing yet.
The model gets better, for real: if the guardrail's fraud detection accuracy improves, that's still not proof the cost is gone. A more accurate filter can still strip honest detail from genuine disputes if it refuses instead of routing them.

Where people run it wrong.
They measure the guardrail's success by refusal count, not by what happens to the requests that get refused.
They build one rule to catch both fraud and genuine edge cases, because it's faster than building two.
They let a refusal go silent, with no explanation, which teaches the customer to guess instead of simply telling them what's happening.

How to use it live. If someone asks "isn't a strict guardrail always safer," answer with the actual number: how many extra exchanges, or how much extra time, a real customer now needs. That's the concrete cost a strict guardrail rarely gets billed for.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Pre-editing flip: the person learns the tool's failure shape and starts sanitizing their own input before the tool ever sees the real version.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ezinne Okoro, a catering business owner who has used Solace Assistant for billing disputes for four years.
3 · THE HABIT
What did Ezinne stop doing because it worked?
Tap to flip
ANSWER
She stopped describing her billing disputes in full, honest detail, since that specific detail was what got her messages refused.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Full, honest, specific detail, versus stripped, vague, guardrail-safe wording. No setting in between.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging fraud detection and genuine dispute handling into one aggressive refusal filter, right after a real fraud scare.
6 · THE NUMBER
Fill in the blank: average round trips per dispute rose from 1.4 to ___ after the guardrail tightened.
Tap to flip
ANSWER
4.1. Median resolution time rose alongside it, from about a day to over four days at its worst.
7 · THE REPLAY
Same bad month, redesigned guardrail. What changes?
Tap to flip
ANSWER
Flagged disputes route to a fast human review lane instead of an outright refusal. Ezinne's resolution time drops to about 1.3 days, and she can say the whole truth again.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and which family?
Tap to flip
ANSWER
Windrow Airlines' baggage claims chatbot. There the flip is a workaround: the customer builds a private email-based re-filing process instead of pre-editing his wording.

Check yourself Score: 0 / 0

Multiple choice
1. Why does the aggressive guardrail cost more than it looks like on a simple refusal-count dashboard?
  • A. Refusal counts are always undercounted by the system.
  • B. It changes what customers are willing to tell you, which a refusal count alone never captures.
  • C. Refused messages are automatically deleted from the log.
  • D. Customers stop using the chat entirely once refused.
Show hint
Look at the highlight line and the round-trips chart.
Show answer
B. The real cost is the honesty of the input, tracked by round trips and resolution time, not the refusal count itself.
True or false
2. True or false: this answer recommends loosening the guardrail on account-changing actions, like adding a new payment method.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. That guardrail stays strict, since honest detail isn't being stripped out of those requests the way it is for billing disputes.
Fill in the blank
3. Fill in the blank: median resolution time peaked at about ___ days in May before the redesign shipped.
Show hint
Look at the line chart's highest point.
Show answer
4.2 days. It dropped to 1.3 days the month after the guardrail was redesigned to route instead of refuse.
Short answer, why no middle setting
4. Why couldn't Ezinne have just been "a little more careful with her wording" instead of stripping the details out completely?
Show hint
Think about what the guardrail actually scans for.
Show answer
Model answer: The filter flagged the exact combination a real dispute needs: an amount, a date, a request to fix it. Any version specific enough to be useful was specific enough to trip the filter, so there was no safe middle ground.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse?
Show hint
Think about a tool you've learned to phrase things carefully around.
Show answer
Model answer: A common one: learning to avoid certain words in a customer-service chatbot's search box, because those exact words used to trigger an unhelpful canned response instead of a real answer.
Short answer, the number
6. If round trips per dispute had only risen from 1.4 to 2.0 instead of 4.1, would the same fix still be worth building? Why or why not?
Show hint
Think about whether the fix depends on the size of the increase.
Show answer
Model answer: Yes, though with less urgency. Any sustained rise means genuine disputes are quietly getting harder to resolve, which is worth fixing before it grows into a bigger gap.
Before you close the answer
Why this works
Tests whether you can name a cost that never shows up on the metric a team is already watching, and whether you understand that refusing an input doesn't remove the underlying need, it just changes how that need gets expressed.
Follow-up traps
"Isn't some friction on financial disputes a good thing, to slow down fraud?" Response: friction that routes to review is fine; friction that refuses outright and teaches people to hide detail is a different thing, and it makes fraud harder to spot too, since real signal gets stripped out along with the noise.

"How do you know the vaguer wording is actually caused by the guardrail, and not something else?" Response: the round-trip and resolution-time numbers move in lockstep with the guardrail's tightening and its later redesign, both up and back down, which is the strongest evidence available short of asking every customer directly.
If pressed
Solace's real redesign also lets a customer attach a screenshot of their bill directly, which gives the review lane the same specific detail without requiring the customer to type the flagged words at all.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more