CaseIntermediateResponsible AI & Advanced Practice / Internal AI tooling and enablement products / #10

Describe an internal knowledge assistant and its most common failure mode.

FLIPS the tool: Ironclad Builders, and CodeCompass, its internal assistant for building and safety code questions

Interviewer's question: "Describe an internal knowledge assistant and its most common failure mode." Ironclad Builders runs mid-rise commercial construction sites across one metro region. Dale Whitcombe has supervised sites there for nine years.

The direct answer
An internal knowledge assistant's most common failure mode isn't a wrong answer, it's a confidently wrong answer that a site supervisor has quietly stopped double-checking, because the tool kept being right for long enough that checking stopped feeling necessary. The fix isn't a smarter model. It's showing the source citation and how current it is on every answer, so trust has somewhere to look instead of somewhere to hide.
Do this, in order
  1. Show the source citation and its staleness on every answer, always.Why: without it, a confident wrong answer and a confident right one are indistinguishable to the person reading them.
  2. Flag rare or edge-case questions differently from routine ones.Why: a rare code interaction is exactly where a model is most likely to be confidently wrong and least likely to get checked.
  3. Never remove a trust signal just because people don't fully understand it yet.Why: an unclear confidence indicator can be redesigned; a removed one leaves nothing behind at all.
  4. Expect trust to erode fastest exactly when the model is improving.Why: good news is still a perturbation; people relax their checking as accuracy climbs, not just as it falls.
  5. Leave routine, well-established answers alone.Why: not every citation needs the same scrutiny, and treating all of them as equally risky trains people to ignore the flag that matters.

How to answer this, stage by stage

Nobody is grading whether you can describe a chatbot that answers code questions. They're grading whether you can name the exact moment trust in it becomes dangerous.

Stage 1
Scope it to one tool, one person
Say it like this
"I'll answer this for CodeCompass, Ironclad Builders' internal assistant for building and safety code questions, and Dale, a site supervisor who uses it daily."
Why this works
Turns an abstract "describe a tool" prompt into one real relationship the interviewer can inspect.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Signals a method before making a single claim about failure modes.
Stage 3
Ground it in the person and the habit
Say it like this
"Dale used to double-check CodeCompass's citation against the physical code book on anything unfamiliar. He's nine years in, he knows most of it from memory anyway."
Why this works
Shows the habit was real, specific, and rational, not carelessness.
Stage 4
Name the flip
Say it like this
"As the model got measurably better, Dale checked less and less, until he stopped checking almost entirely. He didn't decide to trust it completely. It just became the case."
Why this works
This is the direct answer's core: a confident wrong answer treated as ground truth, with no middle setting between checking and not.
Stage 5
Give the old decision that made it reasonable
Say it like this
"Early on, CodeCompass showed a confidence percentage on every answer. It got pulled a few months in because a survey said people didn't understand what '73 percent confident' meant."
Why this works
Names a specific, sensible-at-the-time decision, not a vague "the model got worse."
Stage 6
Show the replay, and close
Say it like this
"Put the citation and a staleness flag back, in plain language instead of a percentage, and Dale catches the same misread fire-code amendment before signing off, not after."
Why this works
Ends on something countable: a stop-work order avoided, not just "trust restored."

Let's learn

CodeCompass answers a site supervisor's question about a building or safety code clause, citing the relevant section, in place of digging through a physical code book or a PDF.

Before it, Dale looked up unfamiliar code sections by hand, about ten minutes per lookup on something outside his usual memory. With CodeCompass, an answer comes back in seconds.

CodeCompass's own error rate, month over month
10% 5% 0% Month 1 Month 18 9% 2%
This is genuinely good news. It's also the reason Dale's checking habit had nothing left to hold onto.

The turn: the model getting better wasn't the problem. The problem is that Dale's checking dropped faster than the error rate did, so the gap between how often CodeCompass was wrong and how often anyone would catch it kept widening in exactly the wrong direction.

Hand sketched comparison diagram titled Small move, big snap. Left panel, a gauge icon labeled The model, caption error rate falls gradually, month over month. Right panel, a box icon labeled Dale's trust, caption stays high, then drops to near zero checking almost overnight.
One line moves in small steps. The other one snaps, and there's no in-between setting once it does.
The decision that mattered CodeCompass launched showing a confidence percentage on every answer. It was removed a few months in, after a survey found most users didn't know what "73 percent confident" meant and thought it made the tool look unfinished. Every answer since has looked equally certain, whether it actually was or not.

At its worst: CodeCompass confidently misreads an updated fire-code amendment on a rare mezzanine-platform configuration, Dale signs off without a second look, because nothing on screen distinguished this rare case from the routine ones it handles perfectly every day.

Hand sketched labeled parts diagram titled What CodeCompass's answer needs now. Center document icon labeled Code Answer, with four callouts: cited source clause, staleness flag shown, edge case warning, link to full text.
Two of these four existed at launch. The other two are what the near miss actually forced back in.

What I would leave alone: CodeCompass's answers on routine, high-frequency questions, standard fire-exit spacing, common load ratings, don't need the same scrutiny. Those are exactly where nine years of Dale's own memory would catch a wrong answer instantly anyway.

The lesson: a knowledge assistant doesn't fail by getting worse. It fails by getting good enough that nobody's still looking when it isn't.

Now here is the same thing as a story

The short version above is what you'd say to Ironclad's safety director. Read this one for how the near miss actually unfolded.

Dale Whitcombe has supervised commercial builds at Ironclad for nine years. Ask him the fire-exit spacing requirement for a standard warehouse floor and he'll answer before you finish the question.

For CodeCompass's first several months, Dale treated it the way he'd treat a sharp new hire: useful, but worth checking. On anything outside his own memory, he'd pull the actual code book and confirm the citation before signing off on a decision.

Knowledge spark: why would a model getting better make trust more dangerous, not less? A model that's often wrong keeps people checking out of necessity. A model that's rarely wrong lets people stop checking out of habit. The rare mistakes don't go away, they just get harder to catch, because nobody's still looking for them.

Over the next year, CodeCompass's own error rate kept falling, genuinely, month over month. Dale noticed, the way anyone would, and slowly checked less. Not a decision, exactly. Just a habit that quietly wore down as the tool kept being right.

Hand sketched timeline titled Dale's trust arc. Four milestones: launch double-checks everything, model improves fewer errors month over month, checking nearly stops nobody decided this highlighted, near miss fire code amendment misread.
The third milestone never got announced anywhere. It just happened, the same way habits usually fade.

The near miss came on a mezzanine platform design, a configuration Ironclad had built maybe twice before. A recent fire-code amendment had changed the required clearance for that specific configuration. CodeCompass, still working off an index that hadn't caught the amendment's edge-case interaction, confidently cited the old clearance requirement, worded exactly like every other answer it gave.

Dale signed off. A city inspector, doing a routine walkthrough two days later, caught the discrepancy before any real harm occurred, and a stop-work order got issued while the platform was corrected.

Nobody at Ironclad ever decided that Dale should stop checking CodeCompass. The tool just kept being right long enough that checking stopped feeling like a decision at all.

Here's the decision I'd take back: pulling the confidence percentage a few months into launch. It made sense at the time, the number confused people and made the product look unfinished. It stopped making sense the moment every answer, confident or shaky, started looking exactly the same.

Replayed with a plain-language flag instead of a percentage, "current as of this month" versus "not recently verified, rare case": the same mezzanine question comes back flagged as an edge case with an unverified citation. Dale checks it against the actual amendment before signing anything, catches the clearance change himself, and the stop-work order never gets issued.

I approved removing the confidence number because it tested poorly and felt like the responsible, user-friendly call. It took a stop-work order on a real platform to see that the number wasn't the problem. Showing no signal instead of a confusing one was.

The five steps, run against one near missNot a story about carelessness. FLIPS is what tells you exactly where trust snapped, and why nobody would have noticed it happening.

F
Find the person. Whose morning is this?
Dale Whitcombe, nine years supervising commercial builds, who can answer most code questions from memory alone.
Grounds the failure mode in a real, competent person, not a generic "user."
L
Locate the habit. What did he stop doing?
He stopped pulling the physical code book to double-check CodeCompass's citation on unfamiliar questions.
The habit was rational: checking kept confirming the tool was right, so checking less felt earned.
I
Identify the flip. The hardest step.
A confident wrong answer gets treated as ground truth, the same way a confident right one always was. No middle setting.
This is the direct answer: the failure isn't the wrong answer, it's the indistinguishability from a right one.
P
Pinpoint the old decision. Which choice only made sense before?
Removing the confidence percentage a few months into launch, because it tested poorly and looked unfinished.
A specific, reasonable-at-the-time decision, not a vague "the model wasn't good enough."
S
Show the replay. Same bad day, new design.
A plain-language staleness flag catches the same misread amendment before sign-off, and the stop-work order never happens.
Ends on something countable: an avoided stop-work order, not just "more trust."
Hand sketched icon list titled The five letters. Five items: F find the person, Dale nine years on site. L locate the habit, he stopped checking the code book. I identify the flip, a confident wrong answer trusted as ground truth. P pinpoint the old decision, the confidence flag got removed. S show the replay, citations and a staleness flag restore the check.
The most screenshot-able page in this whole answer. Five lines, one story.
Hand sketched comparison diagram titled Switch, not dial. Left panel, a gauge icon labeled The dial, assumed, caption many settings, always adjustable. Right panel, a box icon labeled The switch, real, caption two positions only, no smooth middle.
The whole answer in one image. Nobody designed for a switch, because everyone assumed they were building a dial.

The recap, one line per letter: find the person is Dale, nine years in; locate the habit is checking the code book on unfamiliar questions; identify the flip is confident-wrong treated as ground truth; pinpoint the old decision is removing the confidence percentage; show the replay is a staleness flag catching the same amendment before sign-off.

Rate of seeded test errors actually caught by a human, before vs. after trust eroded
100% 50% 0% 95% Month 2 10% Month 16
The model's own error rate fell across these same months. The human catch rate fell even faster, which is the actual risk this whole answer is about.

And if you want to be sure it really works, try it somewhere elseA different flip family, a port authority's customs desk instead of a construction site. This is what proves FLIPS isn't a one-off story.

Harrow Point Port Authority gives customs officers an assistant that classifies incoming cargo against tariff and trade-agreement rules. Ingrid Falkner is a customs compliance officer there.

Mapped onto FLIPS, with a different flip family: F, Ingrid, five years classifying cargo, sharp on routine tariff codes. L, she used to spot-check the assistant's classification against the actual tariff schedule on a sample of shipments each week. I, this time it's a verification flip, not over-trust: after two misclassifications in one week on shipments under a newly signed trade agreement, she swung from spot-checking a sample to manually re-verifying every single classification, which is nearly as slow as not having the tool at all. P, the old decision: the assistant showed no distinction between a well-established tariff code and one from a trade agreement signed only weeks earlier, so nothing separated an ordinary answer from a genuinely untested one. S, the replay: flagging classifications tied to agreements less than ninety days old lets Ingrid verify only the small, genuinely new slice instead of everything, restoring the speed the tool was built to provide.

Hand sketched quadrant titled Sorting customs rulings by verification risk. Axes how often this ruling type comes up and cost of a wrong ruling. Standard tariff classification and routine duty calculation sit common low cost, bottom right. New trade agreement clause and restricted goods exemption sit rare high cost, top left.
A different shape of picture than Section 2 used: a sort by risk instead of a switch metaphor. The top-left corner is where Ingrid's full re-check should have stayed focused.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the failure mode is a confident wrong answer nobody's still checking for; fix it by showing the citation and how current it is," and stop.
Cost: if a full staleness system is too expensive to build immediately, start with a simple age-of-source-data flag, even a rough one beats no signal at all.
The model gets better, for real: this is the twist most candidates miss. Improvement is a perturbation too, and it's exactly when checking habits erode fastest, because good news feels like permission to relax.

Where people run it wrong.
They treat a falling error rate as proof the tool needs less oversight, without checking whether human verification is falling even faster.
They remove a confidence signal because it tested poorly, instead of redesigning it in plainer language.
They apply the same scrutiny to every answer, which trains people to tune out the flag on the rare case that actually needed it.

How to use it live. When someone asks you to describe a knowledge assistant's failure mode, ask yourself first: what happens the day it's confidently wrong, and would anyone even notice? If the honest answer is no, that's the failure mode, not a hypothetical one.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip. Dale checked sometimes, then stopped checking almost entirely, right as the model itself was genuinely improving.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dale Whitcombe, a nine-year site supervisor at Ironclad Builders, who can answer most routine code questions from memory alone.
3 · THE SITUATION
How did code lookups happen before CodeCompass?
Tap to flip
ANSWER
Dale looked up unfamiliar code sections by hand in a physical code book or PDF, about ten minutes per lookup.
4 · THE ANCHOR
What exactly turns a confident answer into a dangerous one?
Tap to flip
ANSWER
Nothing on screen distinguishes a routine, well-established answer from a rare, unverified one. Both look equally certain.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Removing CodeCompass's confidence percentage a few months into launch, because it tested poorly and looked unfinished.
6 · THE NUMBER
Fill in the blank: by month 16, only ___ percent of seeded test errors were still being caught before sign-off.
Tap to flip
ANSWER
10 percent, down from 95 percent in month 2, even as the model's own error rate was falling over the same period.
7 · THE REPLAY
Same mezzanine question, redesigned staleness flag. What changes?
Tap to flip
ANSWER
The answer comes back flagged as an unverified edge case. Dale checks it against the real amendment, catches the clearance change, and the stop-work order never happens.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Harrow Point Port Authority's cargo-classification assistant. Verification flip: two errors in a row swung the officer from spot-checking to re-verifying everything.

Check yourself Score: 0 / 0

True or false
1. True or false: CodeCompass's near miss happened because the model was getting worse.
  • True
  • False
Show hint
Look at the line chart of CodeCompass's own error rate.
Show answer
False. The model's own error rate was genuinely falling. The near miss happened because human checking fell even faster than the model improved.
Multiple choice
2. Why did removing the confidence percentage make the failure mode worse specifically?
  • A. It made CodeCompass slower to respond.
  • B. It required Dale to log in more often.
  • C. Every answer, confident or shaky, started looking exactly the same, removing the one signal that could have prompted a check.
  • D. It increased CodeCompass's actual error rate.
Show hint
Look at the "pinpoint the old decision" step.
Show answer
C. The flip depends on a confident wrong answer being indistinguishable from a confident right one, and removing the confidence signal is exactly what made that true.
Fill in the blank
3. Fill in the blank: CodeCompass's error rate fell from 9 percent in month 1 to ___ percent by month 18.
Show hint
Look at the line chart.
Show answer
2 percent. A genuinely strong improvement, and the exact reason Dale's checking habit had nothing left to hold onto.
Short answer, where it wouldn't matter
4. Name a CodeCompass question type where this failure mode barely applies.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine, high-frequency question like standard fire-exit spacing, where nine years of Dale's own memory would catch a wrong answer instantly anyway.
Short answer, apply it yourself
5. Pick a tool you use yourself. What's one habit it built in you that you'd stop doing if it got a little better, not worse?
Show hint
Think about a tool where getting more reliable made you check its output less closely.
Show answer
Model answer: Many people stop proofreading autocorrect's suggestions once it's been right often enough, which is exactly the over-trust pattern this answer describes.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "here's the decision I'd take back."
Show answer
Model answer: Removing the confidence percentage. It made sense because it tested poorly and looked unfinished, and it broke the moment every answer started looking equally certain.
Before you close the answer
Why this works
Tests whether you can name a failure mode tied to trust and confidence signaling, specific to how a model communicates uncertainty, rather than a generic "sometimes it's wrong" answer.
Follow-up traps
"Isn't this just Dale being careless?" Response: no, he did the sensible thing every step of the way; a tool that keeps being right rationally earns less scrutiny, that's not negligence, that's how trust normally works.

"Wouldn't a confidence percentage have prevented all of this?" Response: not on its own, since it got removed for being confusing; the actual fix is a clearer signal, in plain language, not necessarily the same number brought back unchanged.
If pressed
Ironclad's actual fix ties the staleness flag to the age of the underlying code index specifically, not a general confidence score, so it flags "this section hasn't been re-verified since the last code cycle" rather than a number nobody found meaningful the first time.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more