Your support team handles 8,000 tickets a month. Structure a discovery process to find the AI opportunity.
Three weeks after finance flags support cost per ticket at the quarterly review, Dilnoza Karimova is handed all 8,000 of Loamstone's monthly tickets and one line: reduce this with AI. Loamstone backs up the files on your laptop to the cloud automatically, so a dead hard drive never costs you a lost photo folder or a client's project files. She almost scopes the build around what the ticket subject lines suggest, before she's read what the tickets actually are.
- Segment the 8,000 tickets by category and by how repeatable the fix actually is, before naming what to build.Why: this is the actual reversal; skip it and "reduce ticket volume with AI" quietly turns into "ship a chatbot," the first idea anyone had.
- Pull a real, stratified sample across a full month, and tag it from what agents actually did, not from ticket subject lines.Why: subject lines describe the complaint, not the fix; two very differently worded tickets can share one repeatable resolution, and two similar-sounding ones can not.
- Check every high-scoring category against a fix that already exists before scoping a build.Why: password reset scores higher on repeatability than restores, but a self-serve flow already clears most of it before it becomes a ticket; building there mostly duplicates what's already working.
- Scope the first build narrowly around the one category that's both large and genuinely repeatable, not "support" broadly.Why: restoring a deleted file is 23 percent of volume and 71 percent solved by the same four macros; that combination, not the raw ticket count, is what makes it buildable.
- Set a real target and a held-out check before anything reaches a customer directly.Why: a tool that drafts steps for an agent to check carries a different risk than a bot answering a customer alone, and a wrong guess in the second case has no human standing between it and them.
- Leave the genuinely bespoke categories, like corrupted backups and billing disputes, to people, on purpose.Why: skip this and every future AI request gets treated as "which category's next," when some categories were never going to clear a repeatability bar and don't need re-litigating every quarter.
How to answer this, stage by stage
Nobody is grading whether you can name a chatbot vendor. They are grading whether you'll treat "reduce ticket volume with AI" as a scope with the actual solution missing, and go find out what 8,000 tickets a month actually are before naming what gets built.
Let's learn
The AI project already had a name before anyone had read a single ticket: "AI Support Chatbot."
Loamstone backs up the files on your laptop to the cloud automatically, so a dead hard drive never costs you a lost photo folder or a client's project files. About 150,000 people pay for it. Thirty-eight support agents clear roughly 8,000 tickets a month between them, oldest first, no sorting beyond that.
Before Dilnoza pulled a single real ticket, all she had was a mandate and that name already sitting on the roadmap, penciled in three quarters back because the roadmap deck needed one AI line item and a chatbot was the fastest thing to draw. So she starts where the mandate points: write the chatbot's first ten answers, pulled from the ten most common words in ticket subject lines. Password. Billing. Sync. Restore. She opens a folder of real "restore" tickets to write that one, expecting something shaped like a question she can answer in two lines.
It isn't a question. It's three panicked sentences and a screenshot. No two restore tickets read alike. But every single one has the same three-click fix behind it, because the agent working it always does the same thing: find the file, check the snapshot date, restore it, confirm with the customer.
At its worst, that gap reaches a customer directly. Say the chatbot had shipped anyway, built off subject lines the way the first draft was headed. Jonty, a wedding photographer, types "my wedding photos are gone" into it a little after midnight, three days before he's due to deliver the album. The bot's best match is the password-reset flow's cousin: check your trash folder, check your sync settings. He tries both. Neither is the actual fix, which is a snapshot restore from four days back, a fix an agent could do in under five minutes. He spends forty minutes fighting a script that was never built for what he actually needed, at one point nearly overwriting the correct backup snapshot by following a setting the bot suggested, before he finally reaches a person.
Here's what the sample actually showed, once Dilnoza stopped guessing from subject lines and pulled 450 real tickets, spread across a full month and every plan tier, and tagged each one by category.
What I would leave alone: corrupted or incomplete backup complaints don't get an AI build, not now and probably not ever. Only 9 percent of them share a fix, because each one is closer to a small forensic investigation than a repeated task, and the stakes, someone's only copy of something, are too high for a canned pattern to guess at. Billing gets left alone too, for a related but different reason: it's the second-biggest category by volume, but only 12 percent of it is macro-driven, because every case involves a negotiated discount, a refund judgment, or an account-specific exception. Big and repeatable are two separate questions, and billing only answers one of them.
The lesson: "reduce ticket volume with AI" sounds like an instruction, but it's actually a question with the real work still ahead of it. A mandate can tell you the size of the problem. It can never tell you the shape of the fix, and the shape of the fix is the only thing that decides whether a model can actually help.
Now here is the same thing as a story
The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what a rushed name on a roadmap slide costs, three quarters later.
Twenty months into the job, Dilnoza Karimova had one real skill people asked her for by name: handed a big, vague ask, she always found the smallest true thing worth building instead of chasing the whole wishlist. Before support experience, she'd run onboarding at a smaller company, the kind of job where you learn fast that "make it easier" means nothing until you can point at the one step people actually get stuck on.
Loamstone's support team had been steady for a long time. Thirty-eight agents, 8,000 tickets a month, a queue that never quite emptied but never quite drowned anyone either. Then, at a quarterly review, someone from finance put a number on a slide: support cost per ticket, up 14 percent over the year, next to a line asking what the AI roadmap was doing about it. Petronio Drennan, who runs support, walked out of that meeting and handed Dilnoza the whole queue with four words. Reduce this with AI.
There was already a name for the project. It had been sitting on the roadmap since the planning cycle before, "AI Support Chatbot," penciled in because the deck needed one AI line item that quarter and nobody had time to argue about what it should actually be. Dilnoza did what the name suggested. She pulled a week of ticket subject lines, sorted them by the words that showed up most, and started drafting the chatbot's first ten answers. Password reset. Easy. Billing question. Doable, mostly. Restore a file.
She opened the ticket to write that answer and stopped.
The ticket wasn't a question at all. It was three run-on sentences from someone whose laptop had just eaten a semester's worth of thesis drafts, and a screenshot of an empty folder. She pulled two more. Same thing: messy, urgent, nothing FAQ-shaped about any of them. But underneath all three, the agent's actual reply was almost identical: find the file, check which backup snapshot it last appeared in, restore that snapshot, confirm with the customer. Ingaborg Havlicek, a six-year agent who handled more restore tickets than anyone, had four saved macros she reused for nearly all of them.
We considered the obvious fix first: rank every category by size alone and build for the biggest one that wasn't restore, which was billing. We rejected that. Billing's fixes are negotiated case by case, a discount here, a refund judgment there, not looked up from a pattern, and a model guessing at a refund amount is guessing in exactly the place a wrong guess costs the most.
So Dilnoza put the chatbot draft down and ran the process properly. A real sample, 450 tickets, stratified across every day of a full month and every plan tier, so a quiet Tuesday and a chaotic Monday both counted. Each ticket tagged by category, and by whether Ingaborg's team solved it from one of a small set of saved macros or wrote something new every time. Then a check most people skip: did the category already have a fix that wasn't AI at all. Password reset scored highest on repeatability, 88 percent macro-driven, but 90 percent of resets never became tickets in the first place, caught by the existing forgot-password flow before an agent ever saw them. Building a model for the leftover 10 percent would have meant re-solving a problem Loamstone had already mostly solved.
Restore a deleted file cleared both bars. Twenty-three percent of the queue, and 71 percent of it solvable from four macros that already existed in Ingaborg's saved replies. Dilnoza scoped the first build around exactly that, and nothing wider: a tool that reads a restore ticket, drafts the snapshot-selection steps from the pattern in thousands of past agent replies, and hands the draft to an agent to check and send. It never touches the actual restore. A wrong guess drafted for a human to catch costs a few seconds. A wrong guess executed straight against someone's only backup doesn't.
Average handle time on restore tickets had sat at 14 minutes for as long as anyone could remember, nobody had ever asked whether the steps behind that number ever changed. They didn't, which is exactly what made the tool possible. By week five it was under 4.
What I'd tell the version of me who wrote "AI Support Chatbot" on that roadmap slide three quarters early: a name on a deck is not a decision, even though it feels like one the moment it's typed. The real decision was always going to be made by 8,000 tickets nobody had read yet, and the only choice that mattered was whether to read them before or after the build started.
FLIPS, or the three weeks that decided what got built
Not a trick to sound structured. FLIPS is the difference between a chatbot that fits the ten most common words in a ticket queue, and a tool that fits the one category actually shaped like a pattern.
The AI-specific failure worth naming plainly is a labeling mistake, not a support-ops one: grouping tickets by wording instead of by resolution pattern. A quick topic model on raw ticket text could have merged "restore a deleted file" in with "restore my password" because both contain the word restore, or split the real restore category into ten wording-based clusters that look unrelated when the fix behind every one of them is identical. The guardrail is tagging the sample from the agent's own applied macro, the ground truth of what actually happened, not from a fresh re-clustering of the raw text. There's a real cost accepted on purpose here too: three weeks of a PM and an analyst reading tickets by hand, with nothing that looks like "AI shipped" to show that quarter, in exchange for a tool that clears a real repeatability bar instead of a guess dressed up as a roadmap line.
And if you want to be sure it really works, try it somewhere else
Same five letters, a completely different flip family this time. Nobody reaches for a familiar chatbot here. The trap is quieter: the discovery process itself gets cleaned up before anyone notices.
A city's building permits department fields around 5,000 applicant inquiries a month through its citizen portal. Fadzai Herzfeld, the department's digital services lead, got the same kind of ask Dilnoza did: find the AI opportunity in that queue. She didn't reach for a generic chatbot. She went straight to running a discovery process, sampling real inquiries and tagging them by issue, exactly the move that worked at Loamstone. The trap here was quieter, and it lived inside the tagging itself.
F · Fadzai Herzfeld, digital services lead for the city's permits department, handling about 5,000 applicant inquiries a month.
L · Her small tagging team's habit: when a real applicant message rambled across two or three issues at once, whoever was tagging it picked the one that seemed most important and recorded a single tidy label, because the spreadsheet template only had a field for "primary issue" and messy input slowed the count down.
I · The pre-editing flip, a different shape from Dilnoza's. Old setting: the raw applicant message gets tagged as it actually arrived, bundled issues and all. New setting: a messy, multi-issue message gets sanitized into one clean label before it's ever counted, so the tagged sample looks like a tidy, addressable population when the real inbox never was. Nothing in between: either the count reflects the real messages, or it reflects a cleaned-up stand-in for them.
P · The tagging guide, built early on, told taggers to record "the primary issue," one label per inquiry, because the original spreadsheet only had room for one and the team just wanted a rough headcount by topic at the time.
S · Fadzai adds a second field, "bundled issue, if any," and requires the raw first two sentences recorded exactly as written, not paraphrased. A fresh 300-inquiry sample comes back showing 41 percent of "permit status" inquiries carry a second, hidden issue, usually a missed callback or a document that never got acknowledged, a number the tidy sample had erased entirely. The AI opportunity gets rescoped to draft the callback-and-document response alongside the status answer, not the status question alone, and it holds up against the real, messy inbox instead of only the clean sample used to find it.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: a "reduce X with AI" mandate never tells you what to build on its own, because it names the size of the problem and never the shape of the fix.
Cost: no budget for a full 450-ticket read. A smaller stratified sample, even 150 tickets read properly across a real spread of days, beats a "no time for this" chatbot scoped off subject lines alone.
The model got better, for real: say a general chatbot's off-the-shelf accuracy climbs on its own. Doesn't matter, maybe it matters more. A better model still can't learn a fix from tickets whose resolution was never a repeatable pattern to begin with.
Where people run it wrong.
They rank categories by size alone, and mistake "biggest" for "most buildable," which is exactly the billing trap.
They tag a sample from re-read ticket text instead of from what agents actually did, and quietly reintroduce a subject-line guess wearing a spreadsheet's clothes.
They treat "we ran a discovery process" as proof the sample is honest, instead of checking whether the tagging itself quietly cleaned up the mess that made the real inbox hard in the first place.
How to use it live. If an interviewer hands you a raw ticket count and asks where the AI opportunity is, ask yourself one thing before answering: could I name the one category that's both big and provably repeatable, right now, from real evidence. If the honest answer is "I'd have to guess," that's the whole question, answered.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a three-week discovery process just slow-walking a mandate finance wants solved this quarter?" Response: those three weeks buy a tool that actually clears a repeatability bar. A chatbot scoped off subject lines would have shipped faster and changed nothing, since restore tickets were never shaped like questions a canned FAQ bot could answer.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Opportunity identification for AI
- #1 What characteristics make a workflow a good candidate for AI? List five.
- #2 Describe a method for finding AI opportunities inside an existing product without starting from the technology.
- #3 How do you distinguish a problem AI solves from a problem AI merely touches?
- #4 Rank these by AI suitability and justify: expense approval, contract review, invoice matching, hiring decisions.
- #5 Explain why high-volume, low-stakes, tolerant-of-error tasks are the best first targets.
- #7 What signals in user research suggest an AI solution rather than a better interface?