Describe an internal knowledge assistant and its most common failure mode.
Interviewer's question: "Describe an internal knowledge assistant and its most common failure mode." Ironclad Builders runs mid-rise commercial construction sites across one metro region. Dale Whitcombe has supervised sites there for nine years.
- Show the source citation and its staleness on every answer, always.Why: without it, a confident wrong answer and a confident right one are indistinguishable to the person reading them.
- Flag rare or edge-case questions differently from routine ones.Why: a rare code interaction is exactly where a model is most likely to be confidently wrong and least likely to get checked.
- Never remove a trust signal just because people don't fully understand it yet.Why: an unclear confidence indicator can be redesigned; a removed one leaves nothing behind at all.
- Expect trust to erode fastest exactly when the model is improving.Why: good news is still a perturbation; people relax their checking as accuracy climbs, not just as it falls.
- Leave routine, well-established answers alone.Why: not every citation needs the same scrutiny, and treating all of them as equally risky trains people to ignore the flag that matters.
How to answer this, stage by stage
Nobody is grading whether you can describe a chatbot that answers code questions. They're grading whether you can name the exact moment trust in it becomes dangerous.
Let's learn
CodeCompass answers a site supervisor's question about a building or safety code clause, citing the relevant section, in place of digging through a physical code book or a PDF.
Before it, Dale looked up unfamiliar code sections by hand, about ten minutes per lookup on something outside his usual memory. With CodeCompass, an answer comes back in seconds.
The turn: the model getting better wasn't the problem. The problem is that Dale's checking dropped faster than the error rate did, so the gap between how often CodeCompass was wrong and how often anyone would catch it kept widening in exactly the wrong direction.
At its worst: CodeCompass confidently misreads an updated fire-code amendment on a rare mezzanine-platform configuration, Dale signs off without a second look, because nothing on screen distinguished this rare case from the routine ones it handles perfectly every day.
What I would leave alone: CodeCompass's answers on routine, high-frequency questions, standard fire-exit spacing, common load ratings, don't need the same scrutiny. Those are exactly where nine years of Dale's own memory would catch a wrong answer instantly anyway.
The lesson: a knowledge assistant doesn't fail by getting worse. It fails by getting good enough that nobody's still looking when it isn't.
Now here is the same thing as a story
The short version above is what you'd say to Ironclad's safety director. Read this one for how the near miss actually unfolded.
Dale Whitcombe has supervised commercial builds at Ironclad for nine years. Ask him the fire-exit spacing requirement for a standard warehouse floor and he'll answer before you finish the question.
For CodeCompass's first several months, Dale treated it the way he'd treat a sharp new hire: useful, but worth checking. On anything outside his own memory, he'd pull the actual code book and confirm the citation before signing off on a decision.
Over the next year, CodeCompass's own error rate kept falling, genuinely, month over month. Dale noticed, the way anyone would, and slowly checked less. Not a decision, exactly. Just a habit that quietly wore down as the tool kept being right.
The near miss came on a mezzanine platform design, a configuration Ironclad had built maybe twice before. A recent fire-code amendment had changed the required clearance for that specific configuration. CodeCompass, still working off an index that hadn't caught the amendment's edge-case interaction, confidently cited the old clearance requirement, worded exactly like every other answer it gave.
Dale signed off. A city inspector, doing a routine walkthrough two days later, caught the discrepancy before any real harm occurred, and a stop-work order got issued while the platform was corrected.
Here's the decision I'd take back: pulling the confidence percentage a few months into launch. It made sense at the time, the number confused people and made the product look unfinished. It stopped making sense the moment every answer, confident or shaky, started looking exactly the same.
Replayed with a plain-language flag instead of a percentage, "current as of this month" versus "not recently verified, rare case": the same mezzanine question comes back flagged as an edge case with an unverified citation. Dale checks it against the actual amendment before signing anything, catches the clearance change himself, and the stop-work order never gets issued.
I approved removing the confidence number because it tested poorly and felt like the responsible, user-friendly call. It took a stop-work order on a real platform to see that the number wasn't the problem. Showing no signal instead of a confusing one was.
The five steps, run against one near missNot a story about carelessness. FLIPS is what tells you exactly where trust snapped, and why nobody would have noticed it happening.
The recap, one line per letter: find the person is Dale, nine years in; locate the habit is checking the code book on unfamiliar questions; identify the flip is confident-wrong treated as ground truth; pinpoint the old decision is removing the confidence percentage; show the replay is a staleness flag catching the same amendment before sign-off.
And if you want to be sure it really works, try it somewhere elseA different flip family, a port authority's customs desk instead of a construction site. This is what proves FLIPS isn't a one-off story.
Harrow Point Port Authority gives customs officers an assistant that classifies incoming cargo against tariff and trade-agreement rules. Ingrid Falkner is a customs compliance officer there.
Mapped onto FLIPS, with a different flip family: F, Ingrid, five years classifying cargo, sharp on routine tariff codes. L, she used to spot-check the assistant's classification against the actual tariff schedule on a sample of shipments each week. I, this time it's a verification flip, not over-trust: after two misclassifications in one week on shipments under a newly signed trade agreement, she swung from spot-checking a sample to manually re-verifying every single classification, which is nearly as slow as not having the tool at all. P, the old decision: the assistant showed no distinction between a well-established tariff code and one from a trade agreement signed only weeks earlier, so nothing separated an ordinary answer from a genuinely untested one. S, the replay: flagging classifications tied to agreements less than ninety days old lets Ingrid verify only the small, genuinely new slice instead of everything, restoring the speed the tool was built to provide.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the failure mode is a confident wrong answer nobody's still checking for; fix it by showing the citation and how current it is," and stop.
Cost: if a full staleness system is too expensive to build immediately, start with a simple age-of-source-data flag, even a rough one beats no signal at all.
The model gets better, for real: this is the twist most candidates miss. Improvement is a perturbation too, and it's exactly when checking habits erode fastest, because good news feels like permission to relax.
Where people run it wrong.
They treat a falling error rate as proof the tool needs less oversight, without checking whether human verification is falling even faster.
They remove a confidence signal because it tested poorly, instead of redesigning it in plainer language.
They apply the same scrutiny to every answer, which trains people to tune out the flag on the rare case that actually needed it.
How to use it live. When someone asks you to describe a knowledge assistant's failure mode, ask yourself first: what happens the day it's confidently wrong, and would anyone even notice? If the honest answer is no, that's the failure mode, not a hypothetical one.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Wouldn't a confidence percentage have prevented all of this?" Response: not on its own, since it got removed for being confusing; the actual fix is a clearer signal, in plain language, not necessarily the same number brought back unchanged.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Internal AI tooling and enablement products
- #1 Why do companies underinvest in internal AI tooling, and what does it cost them?
- #2 Describe the first internal AI tool you would build at a 500-person company.
- #3 How do you measure adoption of an internal AI tool?
- #4 What is different about PMing for internal users who cannot churn?
- #5 Design an internal prompt library and explain who maintains it.
- #6 How would you build an internal eval platform that other teams actually use?