CaseIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #6

Your support team handles 8,000 tickets a month. Structure a discovery process to find the AI opportunity.

FLIPS · scoping the first AI build inside Loamstone's 8,000-ticket support queue

Three weeks after finance flags support cost per ticket at the quarterly review, Dilnoza Karimova is handed all 8,000 of Loamstone's monthly tickets and one line: reduce this with AI. Loamstone backs up the files on your laptop to the cloud automatically, so a dead hard drive never costs you a lost photo folder or a client's project files. She almost scopes the build around what the ticket subject lines suggest, before she's read what the tickets actually are.

The direct answer
Before building anything, pull a real sample of the 8,000 tickets and tag each one two ways: its category, and whether the fix behind it is a repeatable pattern or a one-off judgment call. Scope the first AI build narrowly around the single category that scores high on both, not around "support" as a whole. For Loamstone that's tickets asking to restore a deleted file, not a general chatbot.
Do this, in order
  1. Segment the 8,000 tickets by category and by how repeatable the fix actually is, before naming what to build.Why: this is the actual reversal; skip it and "reduce ticket volume with AI" quietly turns into "ship a chatbot," the first idea anyone had.
  2. Pull a real, stratified sample across a full month, and tag it from what agents actually did, not from ticket subject lines.Why: subject lines describe the complaint, not the fix; two very differently worded tickets can share one repeatable resolution, and two similar-sounding ones can not.
  3. Check every high-scoring category against a fix that already exists before scoping a build.Why: password reset scores higher on repeatability than restores, but a self-serve flow already clears most of it before it becomes a ticket; building there mostly duplicates what's already working.
  4. Scope the first build narrowly around the one category that's both large and genuinely repeatable, not "support" broadly.Why: restoring a deleted file is 23 percent of volume and 71 percent solved by the same four macros; that combination, not the raw ticket count, is what makes it buildable.
  5. Set a real target and a held-out check before anything reaches a customer directly.Why: a tool that drafts steps for an agent to check carries a different risk than a bot answering a customer alone, and a wrong guess in the second case has no human standing between it and them.
  6. Leave the genuinely bespoke categories, like corrupted backups and billing disputes, to people, on purpose.Why: skip this and every future AI request gets treated as "which category's next," when some categories were never going to clear a repeatability bar and don't need re-litigating every quarter.

How to answer this, stage by stage

Nobody is grading whether you can name a chatbot vendor. They are grading whether you'll treat "reduce ticket volume with AI" as a scope with the actual solution missing, and go find out what 8,000 tickets a month actually are before naming what gets built.

1
Scope it to one product, one person, one ask
Say it like this
"Let's ground this. Loamstone backs up files to the cloud for freelancers and small studios. Dilnoza Karimova owns support experience there. Finance flags support cost per ticket at the quarterly review, and she gets handed all 8,000 monthly tickets with one line: reduce this with AI."
Why this works
Naming the product, the person, and the mandate stops the answer from staying a vague statement about handling a support backlog.
2
Name the structure before the story starts
Say it like this
"I'll run this as FLIPS. Find the person whose scoping habit is at risk. Locate what she reaches for by default. Identify the flip, the exact question that changes. Pinpoint the old decision that only made sense before anyone had the data. Show the replay with a real discovery process in place."
Why this works
Two sentences of structure buy you a plan before a single detail arrives.
3
Reframe what's actually being tested
Say it like this
"This isn't really asking me to design a support chatbot. It's asking whether I'll treat 'reduce ticket volume with AI' as already telling me what to build, or whether I'll go find out what the 8,000 tickets actually are first."
Why this works
Compresses the whole answer into one breath before a detail can bury it.
4
Give the one decision
Say it like this
"Here's what I'd actually do. Pull a real sample of the tickets, something like 450, spread across a full month. Tag each one by category and by whether the fix is a repeatable pattern or a judgment call. Rank by volume and repeatability together, check what already has a non-AI fix, and scope the first build around the one category that clears both bars."
Why this works
This is the direct answer, said plainly, before the story arrives to earn it.
5
Prove it with a compressed failure
Say it like this
"Here's what happens without it. The chatbot ships, built off ticket subject lines. It handles password questions fine. Then a customer named Jonty types 'my wedding photos are gone' into it at midnight, gets a canned answer about checking his trash folder, and spends forty minutes fighting a script before anyone can pull the right backup snapshot."
Why this works
Three sentences carry a whole incident that a full retelling would take a page to earn.
6
Close on the one line
Say it like this
"So: 'reduce ticket volume with AI' never tells you what to build on its own. Somebody has to go find the one category that's both big and actually learnable from what agents already did, and that costs about three weeks before a line of code gets written, which is exactly the three weeks that keeps you from shipping a chatbot nobody's ticket text was ever shaped for."
Why this works
Leaves the interviewer with the actual decision, not just a well-told story about one queue.

Let's learn

The AI project already had a name before anyone had read a single ticket: "AI Support Chatbot."

Loamstone backs up the files on your laptop to the cloud automatically, so a dead hard drive never costs you a lost photo folder or a client's project files. About 150,000 people pay for it. Thirty-eight support agents clear roughly 8,000 tickets a month between them, oldest first, no sorting beyond that.

Before Dilnoza pulled a single real ticket, all she had was a mandate and that name already sitting on the roadmap, penciled in three quarters back because the roadmap deck needed one AI line item and a chatbot was the fastest thing to draw. So she starts where the mandate points: write the chatbot's first ten answers, pulled from the ten most common words in ticket subject lines. Password. Billing. Sync. Restore. She opens a folder of real "restore" tickets to write that one, expecting something shaped like a question she can answer in two lines.

Hand sketched two panel comparison titled The I step, in one picture. Left panel a document icon labeled The small move, caption Dilnoza drafts chatbot FAQ answers straight from ticket subject lines, same as always. Right panel a question mark box icon in terracotta labeled The big snap, caption the ticket text is a mess but the resolution steps repeat every time, so she scopes by real pattern, not subject line.
Nothing about the mandate changed. What counted as evidence before she'd scope anything changed completely.

It isn't a question. It's three panicked sentences and a screenshot. No two restore tickets read alike. But every single one has the same three-click fix behind it, because the agent working it always does the same thing: find the file, check the snapshot date, restore it, confirm with the customer.

The tickets weren't shaped like questions. The fixes behind them were shaped exactly like a pattern.

At its worst, that gap reaches a customer directly. Say the chatbot had shipped anyway, built off subject lines the way the first draft was headed. Jonty, a wedding photographer, types "my wedding photos are gone" into it a little after midnight, three days before he's due to deliver the album. The bot's best match is the password-reset flow's cousin: check your trash folder, check your sync settings. He tries both. Neither is the actual fix, which is a snapshot restore from four days back, a fix an agent could do in under five minutes. He spends forty minutes fighting a script that was never built for what he actually needed, at one point nearly overwriting the correct backup snapshot by following a setting the bot suggested, before he finally reaches a person.

The choice I would take back Loamstone's roadmap already carried a line called "AI Support Chatbot" before anyone had segmented the 8,000 tickets, because the last planning cycle needed a nameable AI project and a chatbot was the fastest thing to put on a slide. That made sense back when nobody had pulled the data yet and a name was all the moment needed. Nobody ever went back and checked whether the name still fit once the data existed to check it against.
Knowledge spark: what makes a fix "repeatable"? It means the steps an agent takes to solve it barely change from ticket to ticket, even when the words the customer used are completely different. A model can learn a repeatable fix from what agents already did. A fix that's a fresh judgment call every time, a refund amount, a legal question, has nothing steady in it for a model to learn.

Here's what the sample actually showed, once Dilnoza stopped guessing from subject lines and pulled 450 real tickets, spread across a full month and every plan tier, and tagged each one by category.

Loamstone's 8,000 monthly tickets, by category
Restore a deleted file 1,840 (23%) Password reset / login 1,360 (17%) Billing, plan, refunds 1,120 (14%) Sync errors 960 (12%) Storage quota / upgrade 800 (10%) App crashes / install 640 (8%) Corrupted backup 560 (7%) Account transfer 400 (5%) Everything else 320 (4%)
The category scoped forEvery other category
Restore was the biggest category in the queue. Size alone is only half the case for it, the other half is on the next chart.

What I would leave alone: corrupted or incomplete backup complaints don't get an AI build, not now and probably not ever. Only 9 percent of them share a fix, because each one is closer to a small forensic investigation than a repeated task, and the stakes, someone's only copy of something, are too high for a canned pattern to guess at. Billing gets left alone too, for a related but different reason: it's the second-biggest category by volume, but only 12 percent of it is macro-driven, because every case involves a negotiated discount, a refund judgment, or an account-specific exception. Big and repeatable are two separate questions, and billing only answers one of them.

The lesson: "reduce ticket volume with AI" sounds like an instruction, but it's actually a question with the real work still ahead of it. A mandate can tell you the size of the problem. It can never tell you the shape of the fix, and the shape of the fix is the only thing that decides whether a model can actually help.

Now here is the same thing as a story

The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what a rushed name on a roadmap slide costs, three quarters later.

Twenty months into the job, Dilnoza Karimova had one real skill people asked her for by name: handed a big, vague ask, she always found the smallest true thing worth building instead of chasing the whole wishlist. Before support experience, she'd run onboarding at a smaller company, the kind of job where you learn fast that "make it easier" means nothing until you can point at the one step people actually get stuck on.

Loamstone's support team had been steady for a long time. Thirty-eight agents, 8,000 tickets a month, a queue that never quite emptied but never quite drowned anyone either. Then, at a quarterly review, someone from finance put a number on a slide: support cost per ticket, up 14 percent over the year, next to a line asking what the AI roadmap was doing about it. Petronio Drennan, who runs support, walked out of that meeting and handed Dilnoza the whole queue with four words. Reduce this with AI.

There was already a name for the project. It had been sitting on the roadmap since the planning cycle before, "AI Support Chatbot," penciled in because the deck needed one AI line item that quarter and nobody had time to argue about what it should actually be. Dilnoza did what the name suggested. She pulled a week of ticket subject lines, sorted them by the words that showed up most, and started drafting the chatbot's first ten answers. Password reset. Easy. Billing question. Doable, mostly. Restore a file.

She opened the ticket to write that answer and stopped.

Hand sketched horizontal timeline titled Dilnoza's habit, and the moment it flipped. Four milestones: The good months, caption scopes the one real thing. The mandate lands, caption 8,000 tickets, reduce with AI. The FAQ draft, caption not FAQ-shaped at all, this milestone emphasized in terracotta. The replay, caption one slice, measured.
Nobody announced a change of plan. It showed up in the middle of writing ten answers that were supposed to be easy.

The ticket wasn't a question at all. It was three run-on sentences from someone whose laptop had just eaten a semester's worth of thesis drafts, and a screenshot of an empty folder. She pulled two more. Same thing: messy, urgent, nothing FAQ-shaped about any of them. But underneath all three, the agent's actual reply was almost identical: find the file, check which backup snapshot it last appeared in, restore that snapshot, confirm with the customer. Ingaborg Havlicek, a six-year agent who handled more restore tickets than anyone, had four saved macros she reused for nearly all of them.

We didn't need a smarter chatbot. We needed to know what the eight thousand actually were.

We considered the obvious fix first: rank every category by size alone and build for the biggest one that wasn't restore, which was billing. We rejected that. Billing's fixes are negotiated case by case, a discount here, a refund judgment there, not looked up from a pattern, and a model guessing at a refund amount is guessing in exactly the place a wrong guess costs the most.

So Dilnoza put the chatbot draft down and ran the process properly. A real sample, 450 tickets, stratified across every day of a full month and every plan tier, so a quiet Tuesday and a chaotic Monday both counted. Each ticket tagged by category, and by whether Ingaborg's team solved it from one of a small set of saved macros or wrote something new every time. Then a check most people skip: did the category already have a fix that wasn't AI at all. Password reset scored highest on repeatability, 88 percent macro-driven, but 90 percent of resets never became tickets in the first place, caught by the existing forgot-password flow before an agent ever saw them. Building a model for the leftover 10 percent would have meant re-solving a problem Loamstone had already mostly solved.

Hand sketched left to right flow diagram titled The discovery process, five steps before a build gets named. Five boxes connected by arrows: Pull a sample. Tag by pattern. Rank by repeat, this step emphasized. Check for a fix. Scope one slice.
Three weeks, mostly spent reading. Nobody wrote a line of code until step five.

Restore a deleted file cleared both bars. Twenty-three percent of the queue, and 71 percent of it solvable from four macros that already existed in Ingaborg's saved replies. Dilnoza scoped the first build around exactly that, and nothing wider: a tool that reads a restore ticket, drafts the snapshot-selection steps from the pattern in thousands of past agent replies, and hands the draft to an agent to check and send. It never touches the actual restore. A wrong guess drafted for a human to catch costs a few seconds. A wrong guess executed straight against someone's only backup doesn't.

Average handle time on restore tickets had sat at 14 minutes for as long as anyone could remember, nobody had ever asked whether the steps behind that number ever changed. They didn't, which is exactly what made the tool possible. By week five it was under 4.

What I'd tell the version of me who wrote "AI Support Chatbot" on that roadmap slide three quarters early: a name on a deck is not a decision, even though it feels like one the moment it's typed. The real decision was always going to be made by 8,000 tickets nobody had read yet, and the only choice that mattered was whether to read them before or after the build started.

FLIPS, or the three weeks that decided what got built

Not a trick to sound structured. FLIPS is the difference between a chatbot that fits the ten most common words in a ticket queue, and a tool that fits the one category actually shaped like a pattern.

Hand sketched numbered list titled FLIPS, five questions before a ticket queue becomes a build. Five rows: F, find the person, Dilnoza, scoping Loamstone's AI ask. L, locate the habit, reach for the familiar chatbot first. I, identify the flip, subject lines to real resolution pattern, this row in terracotta. P, pinpoint the old decision, chatbot named before data pulled. S, show the replay, one category, scoped and measured.
Four setup and payoff letters, and one question that only had to be asked once the data existed to ask it of.
FFind the person. Whose habit is this?
Dilnoza Karimova, the PM who owns support and post-purchase experience at Loamstone, the person whose first move shapes what "reduce ticket volume with AI" turns into before anyone else weighs in.
The flip belongs to whoever decides what counts as evidence before a build gets scoped, not whoever eventually writes the code.
LLocate the habit. What did she reach for by default?
Whenever a "reduce X with AI" mandate landed, she scoped the build around what ticket subject lines suggested, the fastest, most familiar move, and the one that matched what leadership already had in their head as "the AI project."
That habit cost nothing when nobody was checking. It stops being free the moment the fix behind the words turns out not to match the words at all.
IIdentify the flip. What verb snaps?
Old setting: "how do we reduce ticket volume with AI" is treated as the whole brief, and a chatbot scoped from subject-line keywords counts as an answer. New setting: that question needs a second one first, what does a real sample's category-and-repeatability breakdown say is worth automating, worth assisting, and worth leaving alone, and the real data picks the opportunity instead of a generic mandate. Nothing in between: there's no version where a mandate alone tells you what to build, because the fix's shape was never in the mandate to begin with.
This is the answer to the question in one line. Ticket volume tells you the size of the problem. It never tells you the shape of the fix, and only the second thing decides whether AI can help.
PPinpoint the old decision. Which choice made sense before?
Loamstone's roadmap named "AI Support Chatbot" as a single line item before anyone had segmented the 8,000 tickets, because the planning deck needed one nameable AI project and a chatbot was the fastest thing to draw.
"Read the subject lines more carefully" would be a new dial on the same broken input. Naming the category before naming the solution shape is the reversal actually taken back.
SShow the replay. Same mandate, better ending?
A similar mandate returns, finance flags cost per ticket again the following year. This time Dilnoza runs the sample, tag, rank, and check process first, about three weeks, and scopes narrowly around restore tickets: 1,840 a month, 71 percent solvable from four existing macros, no competing self-serve fix already covering it.
The replay ends in a count: average handle time on restore tickets drops from 14 minutes to under 4 by week five, saving roughly 300 agent-hours a month, while nine other categories get looked at honestly and most of them get left alone on purpose.
Hand sketched full page metaphor titled What the roadmap assumed, and what was true. Left panel a gauge icon labeled DIAL, caption we assumed a bigger mandate just needs a bigger chatbot. Right panel a box icon in terracotta labeled SWITCH, caption either the sample says a category is worth building, or it doesn't, nothing between.
The whole answer, in one picture. Nobody designed a dial. Everybody got a switch, and for three quarters nobody had to throw it themselves.
Average handle time on "restore a deleted file" tickets, week by week
16 min 8 min 0 14.2 tool ships 3.9 min, wk 5 Wk -2 Wk 0 Wk 5
Average handle time, that weekTool shipsStabilizes
Fourteen minutes was normal for so long that nobody thought to ask whether the steps behind it ever changed. They didn't, and that's exactly what made this buildable.
Every category, plotted by volume and by how repeatable its fix actually is
roughly the bar 100% 50% 0 Tickets a month 0 1,000 2,000 Restore a file the pick Password reset already self-serve Billing big, not learnable Sync errors Storage quota App crashes Corrupted backup Account transfer
Scoped forBig, rejectedEverything else
Restore sits alone in the corner that matters: big enough to earn a build and steady enough for a model to actually learn. Billing is just as big and nowhere near steady.

The AI-specific failure worth naming plainly is a labeling mistake, not a support-ops one: grouping tickets by wording instead of by resolution pattern. A quick topic model on raw ticket text could have merged "restore a deleted file" in with "restore my password" because both contain the word restore, or split the real restore category into ten wording-based clusters that look unrelated when the fix behind every one of them is identical. The guardrail is tagging the sample from the agent's own applied macro, the ground truth of what actually happened, not from a fresh re-clustering of the raw text. There's a real cost accepted on purpose here too: three weeks of a PM and an analyst reading tickets by hand, with nothing that looks like "AI shipped" to show that quarter, in exchange for a tool that clears a real repeatability bar instead of a guess dressed up as a roadmap line.

And if you want to be sure it really works, try it somewhere else

Same five letters, a completely different flip family this time. Nobody reaches for a familiar chatbot here. The trap is quieter: the discovery process itself gets cleaned up before anyone notices.

A city's building permits department fields around 5,000 applicant inquiries a month through its citizen portal. Fadzai Herzfeld, the department's digital services lead, got the same kind of ask Dilnoza did: find the AI opportunity in that queue. She didn't reach for a generic chatbot. She went straight to running a discovery process, sampling real inquiries and tagging them by issue, exactly the move that worked at Loamstone. The trap here was quieter, and it lived inside the tagging itself.

Hand sketched labeled parts diagram titled How a tidy sample hides its own gap. Central document icon labeled Tagged sample, with four radiating labels: Messy message arrives. Tagger picks one clean label. Bundled issue dropped. Sample looks cleaner than the real inbox.
Same five questions, a completely different way the flip hides. This time the discovery process cleaned up the very thing it was supposed to measure.

F · Fadzai Herzfeld, digital services lead for the city's permits department, handling about 5,000 applicant inquiries a month.
L · Her small tagging team's habit: when a real applicant message rambled across two or three issues at once, whoever was tagging it picked the one that seemed most important and recorded a single tidy label, because the spreadsheet template only had a field for "primary issue" and messy input slowed the count down.
I · The pre-editing flip, a different shape from Dilnoza's. Old setting: the raw applicant message gets tagged as it actually arrived, bundled issues and all. New setting: a messy, multi-issue message gets sanitized into one clean label before it's ever counted, so the tagged sample looks like a tidy, addressable population when the real inbox never was. Nothing in between: either the count reflects the real messages, or it reflects a cleaned-up stand-in for them.
P · The tagging guide, built early on, told taggers to record "the primary issue," one label per inquiry, because the original spreadsheet only had room for one and the team just wanted a rough headcount by topic at the time.
S · Fadzai adds a second field, "bundled issue, if any," and requires the raw first two sentences recorded exactly as written, not paraphrased. A fresh 300-inquiry sample comes back showing 41 percent of "permit status" inquiries carry a second, hidden issue, usually a missed callback or a document that never got acknowledged, a number the tidy sample had erased entirely. The AI opportunity gets rescoped to draft the callback-and-document response alongside the status answer, not the status question alone, and it holds up against the real, messy inbox instead of only the clean sample used to find it.

What finally surfaced it A new hire tagging inquiries kept a private side note for messages that didn't fit one label cleanly, because it bothered her to throw that detail away. Three weeks in, she asked Fadzai in passing why so many "status" inquiries also mentioned a callback nobody had returned. Nobody had an answer, because the tagging sheet had no field where that detail could have shown up at all.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: a "reduce X with AI" mandate never tells you what to build on its own, because it names the size of the problem and never the shape of the fix.
Cost: no budget for a full 450-ticket read. A smaller stratified sample, even 150 tickets read properly across a real spread of days, beats a "no time for this" chatbot scoped off subject lines alone.
The model got better, for real: say a general chatbot's off-the-shelf accuracy climbs on its own. Doesn't matter, maybe it matters more. A better model still can't learn a fix from tickets whose resolution was never a repeatable pattern to begin with.

Where people run it wrong.
They rank categories by size alone, and mistake "biggest" for "most buildable," which is exactly the billing trap.
They tag a sample from re-read ticket text instead of from what agents actually did, and quietly reintroduce a subject-line guess wearing a spreadsheet's clothes.
They treat "we ran a discovery process" as proof the sample is honest, instead of checking whether the tagging itself quietly cleaned up the mess that made the real inbox hard in the first place.

How to use it live. If an interviewer hands you a raw ticket count and asks where the AI opportunity is, ask yourself one thing before answering: could I name the one category that's both big and provably repeatable, right now, from real evidence. If the honest answer is "I'd have to guess," that's the whole question, answered.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
A scope flip. Old setting: scope the AI build around all 8,000 tickets at once, one general chatbot. New setting: shrink the unit down to the single category that's both large and genuinely repeatable. The atom went from "all of support" to "restoring a deleted file."
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dilnoza Karimova, the PM who owns support and post-purchase experience at Loamstone, a cloud storage and backup company. Known for finding the one real thing worth building instead of chasing the whole wishlist.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Whenever a "reduce X with AI" mandate landed, she scoped the build around what ticket subject lines suggested, the fastest, most familiar move, instead of reading a real sample of full tickets first.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Old: "how do we reduce ticket volume with AI" is the whole brief, and a chatbot built off subject-line keywords counts as an answer. New: that question needs a second one first, what does the sample's real category-and-repeatability breakdown say is worth automating, assisting, or leaving alone.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Loamstone's roadmap already carried a line called "AI Support Chatbot" before anyone had segmented the 8,000 tickets, because the roadmap needed a nameable AI project and a chatbot was the fastest thing to put on a slide.
6 · THE NUMBER
Fill in the blank: restore-a-file tickets are ___ a month, ___% of the queue, and ___% of them are solved with one of just four saved macros.
Tap to flip
ANSWER
1,840 a month; 23%; 71%. That combination, size and repeatability together, is what made the category buildable, not the raw ticket count on its own.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The same kind of mandate lands again. This time Dilnoza runs the sample, tag, rank, and check process first, about three weeks, and scopes an assist tool around restore tickets alone. Average handle time drops from 14 minutes to under 4 by week 5, saving roughly 300 agent-hours a month.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs FLIPS again on a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
A city permits department (Fadzai Herzfeld). The pre-editing flip: the team tagging applicant inquiries for the same kind of discovery process quietly cleaned up messy, multi-issue messages into one tidy label each before tagging them, hiding how often issues actually bundled together.

Check yourself Score: 0 / 0

Fill in the blank
1. Before the flip, Dilnoza scoped the chatbot around what the ticket ___ suggested. After the flip, she scoped it around what the sample's real ___ actually showed.
Show hint
Look at the I step in the framework recap.
Show answer
Subject lines; category-and-repeatability breakdown. The words in a ticket describe the complaint. Only a real sample, tagged from what agents actually did, shows whether the fix behind those words is learnable.
True or false
2. True or false: Dilnoza's fix was to read ticket subject lines a little more carefully before writing the chatbot's FAQ answers.
  • True
  • False
Show hint
Look at what changed as evidence, not how carefully she read.
Show answer
False. "Read more carefully" is a dial. The real flip changed what counted as evidence at all: a real sample tagged by resolution pattern, not a closer read of the same subject lines.
Short answer, name the reversal
3. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Loamstone's roadmap already named "AI Support Chatbot" as a line item before anyone had segmented the 8,000 tickets. It made sense at the time because the planning deck needed one nameable AI project and nobody had pulled the ticket data yet to check the name against.
Multiple choice
4. Which category did Dilnoza deliberately leave for a human to handle, and why?
  • A. Password reset, because it's too easy for AI to bother with.
  • B. Corrupted or incomplete backup complaints, because each case is close to a small forensic investigation and the stakes are too high for a canned pattern.
  • C. Sync errors, because customers prefer solving those themselves.
  • D. Restore a deleted file, because it happens too rarely to be worth building for.
Show hint
Look at "What I would leave alone" in Let's learn.
Show answer
B. Only 9 percent of corrupted backup complaints share a fix, and the stakes, someone's only copy of something, are too high to hand to a pattern-matching guess.
Short answer, apply it yourself
5. Think of a broad "do something with AI" mandate you've seen, or can imagine landing on a team. What's the one number you'd want to see before agreeing on what to build?
Show hint
Look for a number that separates "big" from "learnable," not just a total count.
Show answer
Model answer: For a "let's add AI to our onboarding emails" ask: what share of new-user questions in the first week are actually answered the same way every time, versus questions that depend on that specific person's account, because only the first kind has a real pattern for a model to learn.
Multiple choice
6. If restore-a-file tickets had scored a repeatability of only 20 percent instead of 71 percent, would the same reversal, scoping narrowly around that category, still make sense?
  • A. Yes, because it's still the biggest category in the queue.
  • B. No, because a large category with a low repeatability score is exactly the billing trap: big enough to look tempting, but not actually learnable from what agents did.
  • C. Yes, because ticket volume alone always determines the right AI opportunity.
  • D. No, because categories under 2,000 tickets a month are never worth building for.
Show hint
Check the scatter chart and where billing sits on it.
Show answer
B. Size and repeatability are two separate questions. A category that's big but not repeatable is billing all over again, no matter how large the volume looks on its own.
Before you close the answer
Why this works
Tests whether you'll treat "reduce ticket volume with AI" as already telling you what to build, or as a scope with the actual solution shape still missing. Most candidates jump straight to a chatbot. Fewer notice that ticket volume answers a different question than whether the fix behind those tickets is something a model could actually learn.
Follow-up traps
"What if the second-largest category, billing, is actually where the real savings are?" Response: size it against repeatability first; billing's fixes are negotiated case by case, only 12 percent macro-driven, so most of its cost is judgment calls a model can't safely make, not steps it could learn.

"Isn't a three-week discovery process just slow-walking a mandate finance wants solved this quarter?" Response: those three weeks buy a tool that actually clears a repeatability bar. A chatbot scoped off subject lines would have shipped faster and changed nothing, since restore tickets were never shaped like questions a canned FAQ bot could answer.
If pressed
The restore-assist tool never executes a restore itself. It only drafts the snapshot-selection steps for an agent to check and send, because letting a model act directly on a customer's only backup is the one place a wrong guess can't be undone.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more