CaseFoundationalAI Opportunity & Model Strategy / When NOT to use AI / #12
When is a search box the correct answer to a request for a chatbot?
PICK · Eighty six percent of FindDesk's real questions already had one right document. AskFennelworth wrote that document wrong once in every five tries
Fennelworth Transit Authority runs the buses and light rail for a mid-size metro area, about nineteen hundred employees between operators, station agents, dispatchers, and HR staff. FindDesk is the internal search tool that holds every one of their policies, safety bulletins, and maintenance manuals. Ianto Colbrun owns it. When Selamawit Girma, the VP of People Experience, asked him to swap FindDesk's search bar for a chatbot called AskFennelworth before the next town hall, he had six weeks to work out whether that was a good idea, or just a good demo.
The direct answer
Build the search box for any question that already has one right, findable document behind it, people just need to land on it fast. Save the chatbot for the fraction of questions that genuinely need two or more documents pulled together, or that are too fuzzy to point at one clean answer. At Fennelworth, that fuzzy fraction was 14 percent. Writing a full sentence for the other 86 percent didn't get anyone to the answer any faster. It got some of them a wrong one, dressed up to sound sure of itself.
Do this, in order
Send any question with one right, findable document straight to the search box. Save the chatbot for the ones that genuinely need pulling two or more documents together.Why: this is the whole call, restated as the thing you'd actually do.
Tag a real sample of questions as one-answer or needs-pulling-together before you build either interface.Why: guessing which kind of question you're facing is how Fennelworth nearly built the wrong front door for 86 percent of its own traffic.
Never let the chatbot write its own sentence for a question that already has one exact document behind it. Quote it and link it instead.Why: this is the choice Ianto's team took back, and it's the one change that stopped the wrong answers cold.
Treat a wrong, confident-sounding chatbot answer as a bigger problem than a slow search.Why: it gets believed and acted on. Two operators forfeited real paid time off believing a sentence FindDesk's own document never said.
Recheck the one-answer share on a real schedule, not once and done.Why: the whole call rests on that share staying high. If it ever drops, the right front door changes too.
Leave the truly low-stakes lookups, like a potluck signup date, as chatbot-friendly ground.Why: shows the rule is about what a wrong answer costs, not about chat itself being bad.
How to answer this, stage by stage
Nobody in the room is grading whether you sound excited or suspicious about chatbots. They're grading whether you can tell a lookup from a real question, on the spot, and say why it matters.
1
Put one real help-desk decision on the table
Say it like this
"Let's make this concrete. Fennelworth Transit Authority runs buses and light rail for a mid-size metro area, about nineteen hundred employees. FindDesk is the internal search tool that holds every HR policy, safety bulletin, and maintenance manual they've got. Ianto Colbrun owns it, and the ask on the table is to replace its search bar with a chatbot called AskFennelworth before the next town hall."
Why this works
Keeps the interviewer grading one real call, not a general opinion about chatbots.
2
Name the method before you use it
Say it like this
"I'll run this through PICK. Position: where I land, no hedging. Impact: what each option actually costs, in real units. Cost asymmetry: which mistake is cheap and which one is expensive. Kill criteria: the one test that tells me which kind of question I'm actually looking at."
Why this works
Two seconds of structure tells the room a method is coming, not a mood.
3
Give the position, no hedging
Say it like this
"My position: build the search box for any question that already has one right, findable document behind it. Save the chatbot for the fraction that genuinely needs two or more documents pulled together, or that's too fuzzy to point at one clean answer."
Why this works
This is the direct answer, said early enough that the story underneath can't blur it.
4
Anchor it to the one thing only a model creates
Say it like this
"This only matters because a model sits underneath the chat answer. A search box points at a document, it can't misquote it. A chatbot reads that same document and writes its own sentence, and sometimes the sentence it writes is confident, well formatted, and wrong."
Why this works
Keeps the answer anchored in what a model actually does, not general advice about picking a tool.
5
Bring the numbers that back the search side
Say it like this
"We tagged a real six-week sample of FindDesk's traffic, about a thousand queries. Eighty six percent had exactly one document with the right answer in it, a PTO cap, a form number, a route's radio channel. Search already got people to that document in about thirty five seconds. There was nothing there to fix."
Why this works
A real, checkable number the "search already works" claim would fall apart without.
6
Bring the numbers that back, and limit, the chatbot
Say it like this
"The other fourteen percent were genuinely fuzzy, questions that needed two or three policies stitched together. There, the chatbot got it right eighty nine percent of the time, in forty five seconds against six minutes of a person flipping between documents by hand. But when we replayed the known-answer questions through it, just to check, it paraphrased the source wrong one time in five. Two operators got told their unused time off had no cap. It does. They forfeited twelve hours between them before payroll's year-end sweep caught it."
Why this works
Shows exactly where the chatbot earns its keep, and exactly where it doesn't.
7
Hand over the test and close on it
Say it like this
"So here's the test I use. Does this question already have one document with the right answer sitting in it? If yes, that's search, every time, because a written sentence can't be more right than the document it's copying from, it can only be slower or wrong. If no, if it genuinely needs pulling from more than one place, that's where the chatbot is actually worth the risk."
Why this works
Ends on a test instead of a preference, the part a candidate can repeat under pressure.
Let's learn
What does a company's own help desk actually need: a box you type into, or someone to talk to?
FindDesk is the search bar Fennelworth's employees use to find a policy answer instead of emailing HR and waiting. Before FindDesk existed, three years ago, finding out something like a PTO rule meant emailing HR and waiting about a day for someone to look it up by hand. Now FindDesk answers most of those same questions itself, in well under a minute, across about fourteen thousand searches a month.
Almost every real question FindDesk gets is one of these two shapes. The shape decides which tool should answer it.
Ianto Colbrun owns FindDesk. Selamawit Girma, the VP of People Experience, came back from a conference having seen a slick chatbot demo and wanted FindDesk's search bar replaced with one, called AskFennelworth, in time for the next town hall, six weeks out. Her reasoning: fewer HR tickets, and a tool that looked like it belonged in the same decade as everything else the company had just bought.
Knowledge spark: what's a RAG chatbot?
AskFennelworth reads whichever real documents seem to match a question, then writes a brand new sentence based on them, instead of just showing the document itself. That's called retrieval augmented generation. The finding part is reliable, it's usually good at locating the right document. The writing part is where a fact can quietly slip from what the document actually said.
Ianto's team built a pilot and opened it to forty North Line operators. It answered every question the exact same way: find whatever documents matched, then write a full sentence back. Friendly, confident, done.
The extra mistakes were never really the problem. The problem was that a wrong-sounding answer and a right one read exactly the same on a screen.
Here's the part that's easy to miss. Getting a few more answers wrong wasn't the real issue on its own. The real issue was what people did with a wrong answer that sounded exactly as sure of itself as a right one. Nobody double checks a sentence that reads like policy. They act on it.
At its worst, that's two people believing FindDesk told them their extra time off never expired. It does. Forty hours, used by March thirty first, in the same document FindDesk had linked correctly for two years running. Both of them lost hours of paid time off they thought they still had, and it took Ianto's team three full workdays combined to trace both cases back to the same chatbot session and make it right. FindDesk had never once gotten that question wrong before AskFennelworth started answering it in its own words.
The choice I would take back
AskFennelworth's default was to write a full answer for every question it got handed, lookup or not, instead of routing lookup-shaped questions straight to the document and saving its own sentences for the fraction that actually needed one. I would take that default back. Quote the exact document for anything with one right answer. Only let it write when the question genuinely needs pulling two or three documents together.
What I would leave alone: FindDesk also fields things like "when's the potluck signup close" or "what's on the cafeteria menu this week." Those stay chatty and machine-written, on purpose. A wrong answer there costs someone a shrug, not twelve hours of pay. The rule isn't that a chatbot is bad at lookups in general. It's that a wrong lookup answer has to be cheap before you let a model write it in its own words.
The lesson: a chatbot and a search box can answer the exact same question. The difference that matters is what a wrong answer costs, not which one feels more modern in a town hall demo. Check what the real questions look like before deciding which one gets built.
Now here is the same thing as a story
Use this version when you've got the extra two minutes. A number moves your head. A person's actual Tuesday moves your gut.
Ianto Colbrun has run internal tools at Fennelworth for five years, and he has one rule he doesn't break: never change something everybody depends on without pulling a real sample of what people actually ask it first. It's saved him from shipping at least two bad ideas that nobody else caught in time.
FindDesk itself has been steady for three years, the search bar nobody thinks about because it just works. Then Selamawit Girma came back from a vendor conference in March with a demo on her phone and a deadline in her calendar: swap it for a chatbot, AskFennelworth, before the next town hall, six weeks out.
Ianto started the way he always does, pulling a sample of real questions before touching anything. Six weeks isn't much runway, and Selamawit wanted a pilot ready for a demo, not a research project. He told himself he'd run the full two-week sample after the beta shipped, since it was "just forty people," and pulled a two-day sample instead, enough to sanity check the demo, not enough to catch anything rare. The pilot went out answering every question the same way, lookup or not, because that was the simplest thing to build in six weeks. Nobody flagged it. It felt done.
The two-day sample that felt like enough never got the chance to catch what a six-week one later found.
Then, in week five, two operators on the North Line mentioned to their union rep, in the same week, that they thought their unused time off just carried over on its own. Neither one filed a complaint. They just mentioned it, the way you'd mention the weather.
Payroll's year-end sweep runs automatically. It zeroed out anything over the forty-hour cap, same as it does every year. Both operators lost time off they thought they still had, twelve hours between them, and this time the union rep asked where they'd gotten the idea it didn't expire. Both of them said the same thing: they'd asked AskFennelworth.
Every step up to the fourth one worked exactly as designed. The fourth step is where a fact quietly stopped making it into the sentence.
We didn't lose twelve hours of somebody's vacation. We lost the one thing FindDesk had that a chatbot demo never could: the fact that when it told you something, it was actually true.
I want to say the problem was the twenty percent wrong rate. It isn't really that either. Nobody at Fennelworth had a number in their head for how often AskFennelworth should be allowed to get it wrong. They just knew that a question about their own paycheck should never come from something that sometimes makes a detail up and never says so.
Back when the pilot got scoped, in the meeting where six weeks became the whole plan, somebody asked whether the bot should answer every question type the same way, or handle some differently. The answer, at the time, was that splitting it into two paths sounded like twice the work for a demo that only needed to look good once. So it wrote a sentence for everything. That was the choice.
Here's the replay. Same six weeks, same beta group. This time, any question that matches one document by itself gets the document quoted back with a link, no written sentence at all. The PTO question comes back reading exactly what the policy says: forty hours, used by March thirty first. It takes about the same thirty five seconds plain search always took. The fourteen percent of questions that really do need pulling two or three documents together still get a written answer, sources shown underneath, and it's right about eighty nine times out of a hundred. Six weeks later, replayed the same way: zero wrong answers on the lookup questions. Zero.
One design let the model answer every question the same confident way. The other design let it answer only the questions that actually needed it to think.
What I'd tell myself, in the meeting where six weeks became the whole plan: the two-path build was never twice the work. It was maybe a week more, for a bot that couldn't get a paycheck question wrong. I skipped the two easy days of sampling that would have told me that, because a demo felt more urgent than a sample. It cost twelve hours of somebody else's vacation to learn it back.
PICK: the test that decided FindDesk's front door
Not permission to keep every chatbot out, and not a reason to build one everywhere either. PICK only earns its place here if it turns "chat feels modern" into a real test.
One mistake costs a second try. The other costs real pay and a grievance, and it never looks like a mistake while it's happening.
PPosition. Where the answer actually lands.
Build the search box for any question that already has one right, findable document behind it. Save the chatbot for the fraction that genuinely needs two or more documents pulled together, or that's too fuzzy to point at one clean answer.
This isn't a case against chatbots. AskFennelworth was genuinely better than search for the fuzzy fourteen percent, and it stayed that way even after everything else got fixed.
State the position before any story, so it doesn't read as invented after the fact to fit what already went wrong.
IImpact. What's lost each way.
Build the chatbot as the front door for everything, and the fraction of questions with one right answer stop getting that answer. They get a written paraphrase instead, and one time in five it drops or changes a number that mattered.
Keep everything on search only, and the fuzzy fourteen percent pay for it. Those questions took six minutes of flipping between three or four documents by hand, against forty five seconds and an eighty nine percent hit rate with a written answer and its sources shown.
Naming both losses stops the answer from collapsing into "chatbots are unsafe" or "search is outdated," neither of which is a real decision.
CCost asymmetry. The heart of it.
A slow search costs someone thirty extra seconds, at worst, and it's visible the moment it happens, they just try a different search term. A wrong chatbot answer costs nothing visible at the moment it happens. It reads exactly like a right one, gets believed, and gets acted on. That's the twelve hours of paid time off two operators forfeited without anyone catching it until payroll's sweep ran weeks later. Building search first for the lookup-shaped questions is cheap and works the day it ships. Building a chatbot as the front door for those same questions adds a real chance of a confident, wrong, hard-to-catch answer, in exchange for an interface that only looks more modern.
KKill criteria. The one test.
Do most real questions already have one correct, findable existing answer? Then search is the right tool. Or do they genuinely require pulling together or reasoning across more than one source, with real ambiguity in what's actually being asked? Then a chatbot has real value there. Ianto's team considered one shortcut instead of routing by question type: switch AskFennelworth to quote the source instead of paraphrasing it, and leave it answering everything. It didn't hold up. Retrieval sometimes pulled in two similar policy documents at once, a general leave policy and a department-specific carryover addendum, and even in quote mode the bot stitched together quotes from both into one answer that was technically sourced and still wrong. They dropped that fix and routed by question type instead.
Two shortcuts got tried and dropped before the real fix. Neither one was lazy, both just solved the wrong half of the problem.
Cost, by the numbers: guessing wrong, both directions, over the six-week pilot
Cost of chat answering lookupsCost of search alone on fuzzy questions
Both bars are real hours. Only the left one comes with a policy grievance and a payroll correction attached, which the hours alone don't show.
The kill line, charted: share of FindDesk queries that are exact-answer lookups, six straight weeks
Weekly lookup shareThe week of the PTO incident
Ianto set sixty percent as the line where he'd flip and make the chatbot the default front door instead. The share never dropped below eighty two. Search stays the front door, not because chat failed, but because most real questions never needed it to answer them.
The PTO cap and the cafeteria menu are both one-document lookups. Only one of them belongs on the danger side of the chart.
The trade worth saying out loud: routing lookup questions to search means FindDesk's front door won't feel like one smooth conversation. Some questions come back as a quoted line and a link, not a written paragraph. That's a real loss for how modern the tool feels in a town hall demo, and it's worth paying, because the alternative is a fast, friendly-sounding answer that's wrong on exactly the questions people trust it most to get right.
And if you want to be sure it really works, try it somewhere else
Same four letters, a library system instead of a transit authority, and the AI-specific risk moves from a PTO cap to a fine policy that gets copied into thirty four branches' training notes.
Four branches, one root question: does this question already have one right answer sitting in a document somewhere.
Thackwood County Library System runs StackWise, the internal portal its staff use across thirty four branches for circulation rules, interlibrary loan policy, and cataloging standards. Kaarlo Virtanen is the systems librarian who owns it. About 2,200 staff searches a month go through StackWise. A checked sample found 79 percent were single-answer lookups, like "how many renewals can a patron get on an interlibrary loan," and 21 percent genuinely needed pulling two or more policies together, like whether a fine can be waived for a patron who's also part of a school partnership account. Position: build search for the 79 percent, and save a chatbot for the 21 percent that really does need stitching sources together. Impact: a chatbot paraphrasing the fine cap for a lost children's book wrong, even once, risks getting copied straight into a training note that all thirty four branches follow, since staff trust a clean written answer more than they trust digging through the policy binder themselves. Cost asymmetry: a missed search costs a staffer a second try with different words. A wrong fee policy, copied into branch training notes, costs a corrected memo that has to go out to every branch and get read again. Building search first for the lookup-shaped 79 percent is cheap and works immediately. Letting a chatbot answer fee and policy questions with a written paraphrase risks a wrong number reaching thirty four branches before anyone catches it. Kill criteria: does the question have one document with the answer in it, or does it genuinely need two or more policies read together with real judgment involved? For the interlibrary loan renewal count, yes, one document, so it's search. For the fine waiver tied to a partnership account, no, it needs two policies and a judgment call, so a chatbot can draft it, sources shown, a person still signs off before it becomes a branch-wide answer.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the test: does this question already have one document with the answer in it, or not.
Cost: no time to tag a thousand real queries before a decision is due. Fine, ask whoever's requesting the chatbot for their own best guess at the split, and make them defend that number out loud.
The model got better, for real: say a newer version of the retrieval step starts citing the exact sentence, not just the document, and testing shows it stops dropping numbers and dates. That's the moment lookup questions can start moving to the chatbot too, not before.
Where people run it wrong.
They build the chatbot because it demos well, not because they checked what real questions actually look like.
They treat "the chatbot named its source" as the same as "the chatbot got the source right," when it can point at the correct document and still write the wrong number from it.
They fix one wrong answer by hand and call the whole thing safe, instead of checking whether the same mistake happens again on the next hundred tries.
How to use it live. If you're ever asked whether to build search or chat, buy yourself a second with one plain question, said out loud: "do most of these questions already have one right answer sitting in a document somewhere?" That question is the whole method, asked instead of assumed.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question about when to reach for a search box instead of a chatbot?
Tap to flip
ANSWER
PICK: state the real position, name what each side actually costs, find which mistake is cheap versus expensive, then give the one test that decides it.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Ianto Colbrun, who owns FindDesk, Fennelworth Transit Authority's internal search tool, and Selamawit Girma, the VP of People Experience who pushed for a chatbot front door.
3 · THE POSITION
What's the real rule for choosing search over chat, or chat over search?
Tap to flip
ANSWER
Build the search box for any question with one right, findable document behind it. Save the chatbot for questions that genuinely need two or more documents pulled together, or that are too fuzzy to point at one answer.
4 · THE TWO SHAPES
What are the two shapes a FindDesk question can take?
Tap to flip
ANSWER
A lookup, one document already has the exact answer, like a PTO cap or a form number. Or a synthesis question, it genuinely needs two or three policies pulled together, like a seniority question spanning two job titles.
5 · THE CHOICE I'D TAKE BACK
What old decision would Ianto's team take back?
Tap to flip
ANSWER
AskFennelworth's default was to write a full sentence for every question, lookups included. They'd take that back and route lookup questions straight to the sourced document, saving written answers for the questions that actually needed one.
6 · THE NUMBER
Fill in the blank: 86 percent of FindDesk's real queries were single-answer lookups. The chatbot paraphrased those wrong ___ percent of the time. Two operators forfeited ___ hours of paid time off because of it.
Tap to flip
ANSWER
20 percent, one time in five. 12 hours.
7 · THE KILL TEST
What's the one test for whether a question belongs to search or to chat?
Tap to flip
ANSWER
Does it already have one document with the right answer in it? If yes, that's search, every time. If it genuinely needs pulling from more than one place, that's where a chatbot is worth the risk.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of FindDesk's PTO cap?
Tap to flip
ANSWER
Thackwood County Library System's StackWise. The role goes to the fine-cap policy for a lost children's book, a single-document lookup that a wrong paraphrase could get copied into training notes at all thirty four branches.
Check yourself Score: 0 / 0
True or false
1. True or false: because AskFennelworth got 89 percent of the fuzzy, multi-document questions right, Ianto should trust it to write answers for the lookup questions too, since its overall accuracy is high.
True
False
Show hint
Ask what a single overall accuracy number hides when you average two very different rates together.
Show answer
False. One overall accuracy number would have hidden the split. The lookup questions were wrong one time in five, a far worse and far more dangerous rate than the fuzzy questions' own 89 percent, and blending the two together would have hidden the exact number that mattered.
Multiple choice
2. What actually explains why two operators forfeited paid time off?
A. FindDesk's search engine was down that week.
B. AskFennelworth paraphrased the PTO carryover policy wrong, dropping the cap.
C. Payroll changed the carryover policy without telling HR.
D. The operators never read the PTO policy at all.
Show hint
Think about which tool the two operators actually asked, and what it told them.
Show answer
B. The policy never changed, and the document had always been correct. The chatbot's written sentence dropped or changed the cap, and both operators believed it because it read exactly like a real policy answer.
Fill in the blank
3. The six-week audit tagged ___ real FindDesk queries. ___ percent had one document with the exact right answer. The other ___ percent genuinely needed pulling two or more documents together.
Show hint
Look at stage five of the walkthrough, and the opening of Let's learn.
Show answer
1,050. 86 percent. 14 percent. That split is the whole reason search stayed the default front door instead of AskFennelworth.
Multiple choice
4. Why didn't switching the chatbot to quote-only mode, instead of paraphrasing, fully fix the wrong-answer problem?
A. Quoting used too much computing time and had to be turned off.
B. Retrieval sometimes pulled in two similar policy documents at once, so even a quoted answer blended the wrong one in.
C. Employees trusted a full written sentence more than a quote.
D. Quoting only worked for safety bulletins, not HR policies.
Show hint
Look at the K step, cost asymmetry section, for the shortcut that got ruled out.
Show answer
B. Quoting fixes the writing step, but not a retrieval mistake made before the writing even starts. If the wrong second document gets pulled in, quoting it faithfully just quotes the wrong thing correctly.
Short answer, apply it yourself
5. Think of a search tool or help center you use yourself, at work or anywhere else. Name one question you ask it that already has exactly one right answer sitting in a document somewhere. Would you rather see that document, or trust a chatbot to reword it for you?
Show hint
Pick a real question with a real number or date in it, then ask what a paraphrase could quietly get wrong.
Show answer
Model answer: My health insurer's help center chatbot can summarize what a specific procedure costs under my plan. That number sits in one document, my plan's summary of benefits. I'd rather see the actual line from that document, because a chatbot rewording a dollar figure is exactly the kind of question where a small paraphrase error would matter, and I'd have no way to notice it had happened.
Fill in the blank
6. If AskFennelworth had fully replaced FindDesk's search bar for every question, the pilot's own numbers put fixing wrong lookup answers at about ___ hours over six weeks, against about ___ hours saved on the fuzzy questions over the same period. Net, that's a real ___.
Show hint
Look at the "Cost, by the numbers" chart in the framework recap.
Show answer
24 hours. 13 hours. Loss. A net loss of about 11 hours, and that's before counting the policy grievance and the trust it cost, which the hours alone don't capture.
Before you close the answer
Why this works
Tests whether you default to "chatbots feel more modern, people prefer them" or actually check what real questions look like before deciding, and whether you understand why a wrong lookup answer is a different, worse kind of risk than a slow one.
Follow-up traps
"Couldn't you just show a confidence score next to the chatbot's answer instead of restricting it to search for lookups?" Response: a confidence score measures how sure the model is that it found the right document, not whether the sentence it wrote drifted from what that document said. Both wrong PTO answers came back with high confidence, the model was sure it had the right policy, it just wrote the number wrong.
"What if leadership just wants one clean interface for the demo, are you saying no to that?" Response: not no. The front door can still look like one chat box. Behind it, lookup-shaped questions get answered by finding and quoting the document instead of writing a new sentence. The demo still looks like one tool.
If pressed
The twenty percent wrong-paraphrase rate wasn't spread evenly across every document. It clustered hard on policies that had a number and a date in the same sentence, the retrieval step split some of those sentences at a chunk boundary, so the model saw the cap without the deadline, or the reverse, in close to forty percent of the documents built that way.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.