ConceptIntermediateAI Opportunity & Model Strategy / When NOT to use AI / #6

Describe a case where adding AI increased user effort rather than reducing it.

FLIPS · a reply drafting tool for live chat agents at Larkstone Outfitters, and the one ticket type it made slower

Larkstone Outfitters sells camping and hiking gear through live chat support. Cerelia Culloway has worked the damaged item and refund dispute queue for four years. Once Larkstone's drafting tool started writing her replies for her, she built on top of its suggestions like everyone else, until one Thursday she caught a wrong number seconds before sending it, and never let the tool's numbers near a reply again.

The direct answer
At Larkstone Outfitters, a live chat drafting tool wrote the refund amount and the eligibility decision as ordinary sentences, mixed in with the apology and the sign off, for damaged item dispute tickets. It got that one fact wrong about 1 in 6 times, confidently, in fluent prose that read exactly like a correct answer. Once an agent got burned once, she stopped trusting any of the draft's numbers and rebuilt every fact from the order record herself before sending, on top of reading the draft first. Real handle time for that ticket type rose from 9 minutes to 11.5, even while the team's own launch dashboard, watching only how fast a draft appeared, still looked healthy. The fix wasn't a better model. It was pulling the refund amount and the eligibility line out of the model's free text and showing them as a verified field, so the model only ever writes the sentence around a fact, never the fact itself.
Do this, in order
  1. Pull the refund amount and the eligibility decision out of the model's free generated sentence and show them as a verified field, pulled straight from the order system.Why: this is the actual reversal. Skip it and the fix becomes "ask agents to check more," which is exactly the extra effort this question is about.
  2. Rank ticket types by whether a wrong fact inside them costs real money or trust, not by how often the whole draft reads correctly.Why: damaged item disputes and simple shipping questions had a similar overall draft accuracy, but only one of them had a fact worth pulling out of the sentence.
  3. Track real, end to end resolution time by ticket type, not just time to first draft or how many words an agent changed.Why: Larkstone's own dashboard stayed healthy for nine weeks on this exact ticket type, because it was only ever watching how a draft got made, not how a ticket got finished.
  4. Give the eligibility decision a confidence line, and route anything under it to a person instead of a guess.Why: a model's certainty can't be made perfect, only measured, so the product has to decide what happens below the line, not just hope it's rare.
  5. Leave the rest of the draft, the greeting, the apology, the sign off, as ordinary AI generated text.Why: a clumsy sentence costs a small rewrite. A wrong dollar figure costs a customer's trust and, in this case, 150 dollars nobody meant to give away.
  6. Skip this whole rebuild for ticket types where the draft's only fact is already a straight lookup, like a shipping date pulled from a tracking number.Why: there's nothing to separate from the sentence if the fact was never generated in the first place.

How to answer this, stage by stage

Nobody is grading whether you can describe a tool that backfired. They are grading whether you'll notice that "we generated a draft fast" quietly got treated as "we made the job faster," when nobody ever measured the second one.

1
Scope it to one product, one agent, one ticket type
Say it like this
"Let's ground this. Larkstone Outfitters sells outdoor and camping gear through live chat. Cerelia Culloway has worked the damaged item and refund dispute queue for four years. The team built a tool that reads a ticket and drafts a reply for her to send, and on this one ticket type, it ended up costing her more time than typing from scratch."
Why this works
Naming the product, the person, and the exact ticket type stops the answer from turning into a general complaint about AI adding friction.
2
Say your structure out loud
Say it like this
"I'll run this as FLIPS. Find the person whose habit is at stake. Locate what she reached for without thinking. Identify the flip, the exact thing that snaps. Pinpoint the old decision that only worked before anyone measured the real time. Show the replay with the fact pulled out of the sentence."
Why this works
Two sentences of structure signal a plan before a single detail lands.
3
Reframe what's actually being tested
Say it like this
"This isn't really asking me to describe a broken feature. It's asking whether I'll notice that 'we generated a draft fast' quietly got treated as 'we made the job faster,' when nobody ever measured whether the second one was true."
Why this works
Compresses the whole answer into one breath before a detail can bury it.
4
Give the one decision
Say it like this
"Here's what I'd actually do. Pull the refund amount and the eligibility line out of the model's own sentence, and show them as a verified field, pulled straight from the order system. The model still writes the words around them. It just doesn't get to write the numbers anymore."
Why this works
This is the direct answer, said plainly, before the story arrives to earn it.
5
Prove it with a compressed failure
Say it like this
"Here's what happens without it. A camp stove comes back scratched. The draft tells Cerelia the customer's owed 340 dollars. She's about to send it. Then she glances at the order total, out of an old habit, and it says 190. The extra 150 was a damage credit that only applies if the customer keeps the item, and this one's a return."
Why this works
Four sentences carry a whole incident that a full retelling would take a page to earn.
6
Name the AI-specific detail you'd hold onto
Say it like this
"The real fix wasn't a smarter model. It was deciding which parts of that sentence were allowed to be a guess in the first place. A model can say a dollar figure with total confidence and still be wrong, so the number can't live inside a sentence the model is free to write however it wants."
Why this works
Shows the judgment is about where a probabilistic system gets to state a fact, not about the model needing to be better.
7
Say what you'd measure, and where the old number hid
Say it like this
"I'd track real ticket time by ticket type, not just how fast a draft appears or how little of it gets edited. Larkstone's own dashboard looked fine for nine weeks on this exact ticket type, because it was only ever watching the first two of those."
Why this works
Shows you think past launch day, and names the exact place the old metric was blind.
8
Close on the one line
Say it like this
"So: a tool that hands someone a fast, confident, sometimes wrong answer hasn't saved them a step. It's added one, reading a wrong answer before they can write the right one. Fix that by taking the fact out of the model's hands, not by asking a person to read more carefully."
Why this works
Leaves the interviewer with the actual decision, not just a well told story about one near miss.

Let's learn

Every damaged item dispute used to take Cerelia about nine minutes, start to finish, before Larkstone Outfitters ever put a model near her queue.

Larkstone sells tents, stoves, and hiking packs online, and its whole support team works over live chat, about 1,400 conversations a day. Around 95 of those, every day, are damaged item and partial refund disputes: an item came back scratched, or dented, or the wrong size, and someone has to work out what's owed. Before the drafting tool, an agent opened the order record, checked the item's history, worked out the refund amount and whether a damage credit applied, then wrote the reply herself. Cerelia was good at it. She could spot a mismatched SKU before the page finished loading.

The drafting tool reads a new ticket and writes a full reply: an apology, the refund amount, the eligibility line, and a sign off, in one smooth paragraph. Larkstone's launch review measured two things before it shipped: time to first draft, under 2 seconds, and how much of a sent reply changed from what the model wrote. On damaged item tickets in the first month, 81 percent of drafts went out with fewer than 15 words changed. That looked like a win worth telling leadership about.

Hand sketched two panel comparison titled The I step, in one picture. Left panel a gauge icon labeled The small move, caption a new ticket lands, Cerelia builds her reply on top of the draft, same as always. Right panel a question mark box icon in brick orange labeled The big snap, caption now she opens the order record first, every time, and only borrows the draft's wording.
Nothing about the ticket changed. What Cerelia required before trusting a number changed completely.

Here's the turn. The draft's overall accuracy on this ticket type never was the problem. It got the refund amount or the eligibility line wrong about 1 in 6 times, and that number barely moved the whole time. The real problem is what Cerelia had to do once she knew that. She couldn't just check a little more carefully. The moment she stopped trusting one number in a drafted reply, she couldn't trust any of them without checking, so for this ticket type, she started opening the order record and working out the refund herself, every single time, before she'd let herself read the draft at all.

Average handle time, damaged item dispute tickets: before the drafting tool vs. real total time, nine weeks after launch
12 min 6 min 0 9.0 min Before the drafting tool 11.5 min After, real time, week 9
BeforeAfter, real end to end time
The tool's own dashboard, time to first draft and words changed, looked healthy the entire nine weeks. Neither number ever measured this.
We didn't hand Cerelia a shortcut. We handed her a second job: fact checking a stranger who writes exactly like her.

At its worst, that gap isn't just a story about one agent. Multiply Cerelia's extra two and a half minutes across roughly 95 damaged item disputes a day, across a queue of agents doing the same ticket type the same way, and this one ticket type alone was costing Larkstone's support floor about four hours of agent time a day that didn't exist before the tool arrived. Worse than never building it, because the nine minute baseline was already fine.

Hand sketched left to right flow diagram titled The old pipeline, before the fix. Four boxes connected by arrows: Order record. AI drafts. Merged paragraph, this box emphasized in brick orange. Agent sends.
One box in the middle did two jobs at once: state the fact, and phrase the sentence. Nothing marked where one ended and the other began.
The choice I would take back Larkstone let the model write the refund amount and the eligibility line as ordinary sentences, mixed into the same paragraph as the apology and the sign off, instead of pulling those two facts from the order system as separate, marked fields the model could only write around. That made sense when the tool only handled simple, single item tickets, where the model's numbers were nearly always right and one smooth paragraph felt more finished than a form. It stopped making sense the moment the ticket type got messier than that.
Knowledge spark: what's a verified field? A piece of text pulled straight from a system's own records, like an order total or a ship date, instead of written fresh by a model. It can't be phrased wrong, because it was never phrased at all, only copied.

What I would leave alone: shipping status and delivery date questions on this same live chat tool don't need any of this. The draft states a date pulled straight from the tracking number, the same fact every time, nothing generated. No agent on that ticket type ever built a habit of rebuilding it by hand, because there was never a sentence hiding a guess inside it.

The lesson: a fast draft and a fast resolution are not the same claim, and the only way to find out you've been treating them as one is to actually go measure the second thing. Larkstone had a whole month of numbers that looked like proof. None of them were measuring the thing that mattered.

Now here is the same thing as a story

The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what 150 dollars, caught by luck, does to a person's whole approach to a tool.

Every shift, before the drafting tool, Cerelia opened the same second tab first: the order record. Check the item, check the return reason, check what the customer already paid, then write the reply. She'd worked the return and damage queue at Larkstone for four years, long enough that new hires got sent to her when a ticket didn't make sense. Ask her whether a scratched tent pole qualified for a damage credit or a straight refund, and she'd know before she finished reading the description.

The drafting tool arrived in the spring. For the first few months, on this ticket type, it was genuinely good. It read the ticket, pulled up what looked like the right numbers, and wrote a reply that sounded like something she'd have typed herself. She still opened the order record those first weeks, mostly out of habit, and it kept matching. So she opened it a little less. Then only for the bigger orders. Then, without ever deciding to, she stopped opening it at all for this ticket type, the same way she stopped for the easy shipping questions months earlier, because the draft kept being right and there was no reason left to check.

Hand sketched horizontal timeline titled Cerelia's habit, thinning across six weeks. Four milestones: Before the draft, caption builds every reply from the order record. The good weeks, caption the draft is right, she still glances at the record. The glance fades, caption routine tickets first, then disputes too. The near miss, this milestone emphasized in green, caption 340 shown, 190 owed.
Nobody told her to stop checking. The draft simply kept being right long enough that checking stopped feeling like part of the job.

Then came a Thursday. A customer named Vasyl Brindley had returned a camp stove with a cracked burner ring, and the draft told Cerelia he was owed 340 dollars: the item's price, plus a 150 dollar damage credit. She had her cursor over send. Out of the same old habit that had mostly gone quiet, she glanced at the order total sitting in the sidebar. It said 190.

The damage credit only applies when a customer keeps a flawed item instead of returning it. Halden was returning his. The draft had pulled the credit rule from a similar ticket earlier that week and applied it here anyway, in a sentence that read just as confidently as every correct one she'd sent that morning.

She didn't catch a typo. She caught a stranger, speaking in her own voice, about to give away 150 dollars that wasn't owed.

Here's what she did next, and it's the whole story. She didn't start reading the draft's numbers more carefully. She stopped reading them at all. From that Thursday on, for every damaged item dispute, she opened the order record first, worked out the refund and the eligibility herself, wrote the number down, and only then read the draft, purely to borrow a line or two of its wording. She never told anyone this became her process. Nothing in Larkstone's tools tracked it. The dashboard still saw a draft generated fast, and a reply sent with only a few words changed, because her rewritten numbers usually landed close to what the draft had guessed anyway.

We considered the easy fix first: put a line under every drafted reply, please confirm the amount before sending. We killed it within a day. Cerelia already knew to check. That was exactly the problem, not the fix. A reminder doesn't tell you which of the draft's four sentences is the one that's wrong, so it doesn't save a single minute, it just makes the checking official.

We also considered training the model on more damaged item examples until the error rate came down. We killed that one too, more slowly, because it took a real conversation to see why. Even a model that's right 99 times out of 100 is still, on the hundredth, confidently wrong in a sentence that looks exactly like the other 99. Cerelia can't tell which sentence is the hundredth one by reading it. Neither can anyone. A lower error rate makes the mistake rarer. It doesn't make it visible.

Here's the decision I'd take back instead, and it isn't Cerelia's, it's the one made months earlier, in the room where the drafting tool got designed. Someone decided the reply should read as one smooth, human sounding paragraph, the refund amount and the eligibility line generated in the same breath as the apology and the sign off. Reasonable, at the time. The early version only handled simple, single item tickets, where the numbers were nearly always right, and a paragraph with no visible seams felt like a better product than a form letter with blanks in it. Nobody built a second lane for the day the ticket types got messier than that.

Run the same kind of ticket again, six weeks after the fix ships. A different camp stove, a different customer, the same kind of damage credit mix up the model would have made before. This time the refund amount and the eligibility line don't come from the model at all. They're pulled straight from the order system and shown as a marked field, right in the reply composer, next to the AI written sentence around them. The eligibility rule carries its own confidence line: under 90 percent sure, and it doesn't state a value, it flags the ticket for a person instead. Cerelia reads the field once, confirms it against the order total already sitting in the sidebar, and sends. No second tab. No rebuilt reply. Handle time for this exact kind of ticket: six and a half minutes, faster than the tool ever managed, faster than the nine minutes it took before the tool existed at all.

What I'd tell myself, back in that first design meeting: making a reply sound like one smooth paragraph and making a fact something you can trust are two different jobs. We built a tool that was very good at the first one, and quietly assumed that meant it was doing the second.

FLIPS, or the fact that stopped being a sentence

Not a trick to sound structured. It's the difference between a reply that sounds finished and a reply that's actually made of the right numbers.

Hand sketched numbered list titled FLIPS, five questions before Larkstone rebuilt one reply. Five rows: F, find the person, Cerelia, rebuilding a reply by hand. L, locate the habit, trusting the draft's numbers outright. I, identify the flip, what verb snaps, this row in brick orange. P, pinpoint the old decision, facts and phrasing merged into one paragraph. S, show the replay, the amount becomes a field, not a sentence.
Four setup and payoff letters, and one question that only mattered once a model was allowed to state a dollar figure on its own.
FFind the person. Whose habit is this?
Cerelia Culloway, four years in Larkstone's damaged item and refund dispute queue, the agent who used to spot a mismatched SKU before the page finished loading.
The flip belongs to whoever actually reads a drafted number and decides whether to send it, not whoever approved the launch.
LLocate the habit. What did she stop doing because it worked?
Once the drafting tool launched, she stopped opening the order record before answering a damaged item dispute. She let the draft's stated refund amount and eligibility line go straight into her reply, because it kept being right on the tickets she happened to check.
That habit cost nothing while the draft's numbers stayed reliable. It became expensive the day one of them wasn't, and she had no way to tell which sentence was the wrong one.
IIdentify the flip. What verb snaps?
Old setting: Cerelia builds her reply on top of the draft, trusts its numbers, borrows its wording wholesale. New setting: she never lets the draft's numbers anywhere near what she sends. She opens the order record fresh, works out the refund and eligibility herself, and only keeps a phrase or two of the draft's tone. Nothing in between: the moment she can't trust one number in a drafted reply, she can't trust any of them without checking, so she checks all of them, every time, for this ticket type, and rebuilds the fact rather than merely re-reading it.
This is the answer to the question in one line. The tool didn't add a step to her old process. It added a whole second process, running quietly underneath the one anyone could see on a dashboard.
PPinpoint the old decision. Which choice made sense before?
Larkstone let the model write the refund amount and the eligibility decision as ordinary generated sentences, mixed into the apology and the sign off, instead of pulling those two facts from the order system as separate, marked fields the model could only write around.
"Add a confirm-the-amount reminder" would be a new dial bolted onto the same broken design. Separating fact from phrasing is the actual reversal being taken back.
SShow the replay. Same trigger, better ending?
A similar damage credit mix up lands again, six weeks after the fix ships. This time the refund amount and eligibility line render as a verified field pulled straight from the order system, with a confidence line that routes anything under 90 percent to a person instead of a guess.
The replay ends in a count: handle time for this ticket type drops to 6.5 minutes, beating both the 11.5 minute rebuild-everything phase and the original 9 minute baseline, because Cerelia is no longer doing two people's jobs on one reply.
Hand sketched full page metaphor titled What the launch review assumed, and what was true. Left panel a gauge icon labeled DIAL, caption we assumed a fast draft means a faster reply, by degrees. Right panel a question mark box icon in brick orange labeled SWITCH, caption either the fact came from the record or nobody could trust it, nothing between.
The whole answer, in one picture. The launch review designed for a dial. Cerelia lived inside a switch, for nine quiet weeks before anyone else noticed.
Real handle time for damaged item disputes, week by week, while the launch dashboard stayed flat
12 min 10 min 8 min 9.0 11.5, wk 9 Wk 1 Wk 5 Wk 9
Real handle time, minutesWeek the near miss surfaced it
Time to first draft and words-changed both stayed flat and healthy across all nine weeks. Neither one was built to notice a habit forming underneath it.
Hand sketched labeled parts diagram titled What the rebuilt reply is made from. Central document icon labeled Damaged item reply, with four radiating labels: Refund amount, pulled from the order system. Eligibility line, a verified field. Under 90 percent sure, routes to a person. AI only drafts the sentence around it.
Two of these four pieces are facts now. The other two are still the model's job, and always were.

The AI specific failure worth naming plainly is confident wrongness: a model stating a dollar figure or a policy exception in fluent, certain sounding prose, with no visible difference between a right answer and a wrong one. The guardrail is the verified field itself, backed by a real confidence threshold on the eligibility classifier, tuned against six months of resolved disputes rather than a guess, so anything under 90 percent routes to a person instead of asserting a value. There's a real trade off accepted here too: building the live link into the order system took about five weeks of engineering time the team didn't plan for at launch, and the tool now sometimes shows "needs a policy check" instead of a tidy, fully finished looking paragraph, on purpose, trading a little polish and shipped speed for two facts an agent can actually trust without opening a second tab.

And if you want to be sure it really works, try it somewhere else

Same five letters, a genuinely different flip family this time. Nobody here stops trusting the tool. She just starts saving it for exactly the wrong cases.

Wrenwick Veterinary Partners runs a group of clinics, and its vet techs use a drafting tool to write post visit care instruction messages that go to pet owners: medication timing, follow up scheduling, what to watch for. Kirstyn Beckendorf is known on her team for the toughest kind of message to get right, post surgical cases on two or three overlapping medications with a tapering schedule. For the tool's first two quarters, she leaned on it for every visit type, routine and complex alike, because it read to her like a fast, reasonable first pass either way.

Hand sketched two panel comparison titled Same five letters, a different I both times. Left panel a person icon in green labeled Cerelia, caption Larkstone Outfitters, workaround flip, she rebuilds every fact herself instead of trusting the draft. Right panel a person icon in brick orange labeled Kirstyn, caption Wrenwick Veterinary Partners, substitution flip, she saves her AI credits for the hardest cases, exactly where it is worst.
Same method, a different verb entirely. One agent stopped trusting the tool. The other started rationing it, straight toward the cases it's worst at.

The trap showed up after a cost review capped every tech at 200 AI assisted drafts a month, to control the clinic group's model spend. Kirstyn didn't abandon the tool. She started budgeting it. Routine wellness visit messages, the kind she can write from memory in under a minute anyway, she now writes herself and saves her credits for the complex, time consuming, post op cases, the ones she cares most about getting right and where writing from scratch takes the longest.

F · Kirstyn Beckendorf, a vet tech at Wrenwick who writes post visit care instructions, especially skilled with multi drug taper schedules.
L · She used to reach for the drafting tool on every visit type without thinking about which ones it was actually good at.
I · The substitution flip, a different shape entirely from Cerelia's. Old setting: she uses the tool for every message, easy and hard alike. New setting: with a fixed monthly quota, she rations her limited credits toward the visit types she cares most about getting right, the complex multi drug cases, which happen to be exactly the visit type the model is worst at. The model never changed. The mix of what it's being asked to draft did, and the measured error rate on AI assisted messages climbed because of it.
P · Whoever set the quota priced every visit type's AI credits the same, instead of steering the limited credits toward the visit types the model was actually reliable on.
S · Wrenwick reserves AI drafting automatically for visit types where the model's fact error rate sits under 5 percent, routine wellness and vaccine follow ups. For multi drug post op cases, techs get a structured template that pulls exact dosages straight from the medication record instead of full free text drafting. Post op message time, previously creeping toward 14 minutes as Kirstyn wrote them fully by hand to save her quota, drops to under 5.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: the model didn't get worse, the mix of what it was asked to draft got harder, and no dashboard watching an overall average would ever show you that on its own.
Cost: no budget for a full verified-field build like Larkstone's. Even a simple rule, flag any reply with a number over a set dollar threshold for a mandatory second look, catches most of the damage before a bigger fix ships.
The model got better, for real: say next quarter's accuracy climbs across the board. Doesn't matter, maybe it matters more. An improving average is exactly the kind of number that hides one ticket type still quietly getting worse underneath it.

Where people run it wrong.
They treat "drafts are going out nearly unedited" as proof the tool is working, instead of asking whether that's true on the hard cases too, or only on the easy ones nobody had to check anyway.
They fix trust with a reminder to double check, instead of asking which specific facts inside the reply are generated versus looked up.
They wait for a customer complaint to force the real comparison, instead of instrumenting total ticket time by type from the day the tool ships.

How to use it live. If an interviewer hands you a tool that's supposed to save someone time, ask yourself one thing before answering: which two or three facts in its output, if wrong, would actually cost someone money or trust, and are those facts generated by the model or looked up from somewhere real? That question alone usually is the whole answer.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
A workaround flip. Cerelia doesn't just check the draft more carefully. She builds a private process around it, pulling every fact from the order record herself and only borrowing the draft's wording, a step nobody's dashboard could see.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Cerelia Culloway, a live chat support agent at Larkstone Outfitters, an outdoor and camping gear retailer. Four years in the damaged item and refund dispute queue, known for catching a mismatched SKU before anyone else does.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Once the drafting tool launched, she stopped opening the order record before answering a damaged item dispute. She let the draft's stated refund amount and eligibility line go straight into her reply, because it kept being right on the tickets she happened to check.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Old: she builds her reply on top of the draft and trusts its numbers. New: she never lets the draft's numbers near her reply, and rebuilds the refund amount and eligibility from the order record herself, every time, keeping only the draft's wording. No middle setting.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Larkstone let the model write the refund amount and eligibility line as ordinary sentences mixed into the apology and sign off, instead of pulling those two facts from the order system as separate, marked fields the model could only write around.
6 · THE NUMBER
Fill in the blank: real handle time for damaged item disputes went from ___ minutes before AI to ___ minutes nine weeks after launch, even though the draft got the key fact wrong about 1 in ___ times the whole way through.
Tap to flip
ANSWER
9 minutes; 11.5 minutes; 1 in 6. The draft's own error rate barely moved. What changed was how much extra work each of those errors created once Cerelia stopped trusting any of the draft's numbers.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The refund amount and eligibility line render as a verified field pulled from the order system, with a confidence line that routes anything under 90 percent to a person. Handle time for this ticket type drops to 6.5 minutes, better than the tool ever managed and better than the original 9 minute baseline.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs FLIPS again on a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Wrenwick Veterinary Partners, a vet clinic group. The substitution flip: once AI-assisted drafts get capped at a monthly quota, vet tech Kirstyn Beckendorf saves her credits for the hardest post op cases, exactly the ones the tool is worst at.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Larkstone's launch dashboard still look healthy nine weeks after the near miss, while real handle time on this ticket type had climbed 28 percent?
  • A. The dashboard was broken and needed a bug fix.
  • B. The dashboard measured how fast a draft appeared and how little of it got edited, never how long the whole ticket actually took to resolve.
  • C. Cerelia had stopped reporting her real handle time.
  • D. The model's accuracy had genuinely dropped that quarter.
Show hint
Look at the two numbers Larkstone's launch review actually measured, in Let's learn.
Show answer
B. Time to first draft and words changed both stayed flat and healthy the entire nine weeks. Neither one was ever built to notice a private, invisible habit forming underneath it.
True or false
2. True or false: the drafting tool's overall accuracy on damaged item tickets dropped noticeably once Cerelia stopped trusting its numbers.
  • True
  • False
Show hint
Look at "here's the turn" in Let's learn.
Show answer
False. The model's own error rate, about 1 in 6, barely moved the whole time. What changed was Cerelia's process around it, not the model's behavior.
Fill in the blank
3. Before AI, damaged item dispute tickets took about ___ minutes. Nine weeks after launch, real handle time for the same ticket type had climbed to about ___ minutes, even though the drafting tool's own dashboard never moved.
Show hint
Check the bar chart in Let's learn.
Show answer
9 minutes; 11.5 minutes. A 28 percent increase, entirely invisible to a dashboard that only ever watched how fast a draft got made.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Larkstone let the model write the refund amount and the eligibility line as ordinary sentences mixed into the apology and sign off, instead of pulling those two facts from the order system as separate, marked fields. That made sense when the tool only handled simple tickets where the model's numbers were almost always right, and one smooth paragraph felt more finished than a form. It stopped making sense once the ticket type got messier than that.
Short answer, apply it yourself
5. Think of an AI tool you use yourself that's supposed to save time. Is there a fact inside its output you always double check by hand anyway? What would it take to make that fact something you never had to check?
Show hint
Ask which parts of the output are generated prose and which are, or could be, looked up from a real source.
Show answer
Model answer: Name the specific fact, not the whole output, a date, a total, a name spelled a certain way. Making it uncheckable usually means pulling it from the actual source system as a marked field, the same fix as Larkstone's, rather than trusting the model to state it correctly on its own.
Short answer, the number question
6. If Larkstone's drafting tool got the refund amount right 99 percent of the time instead of about 1 in 6, would pulling the amount out as a verified field still be worth building? Why or why not?
Show hint
Think about whether Cerelia could tell which sentence was the rare wrong one, at any error rate.
Show answer
Model answer: Yes, likely still worth it, because the real problem was never the error rate on its own. Even at 1 in 100, a confidently wrong sentence looks identical to a correct one, so an agent still can't trust any single reply without checking, unless the fact simply isn't generated in the first place.
Before you close the answer
Why this works
Tests whether you'll treat "the draft got made fast" as the same claim as "the job got done faster," or whether you'll go find out. Most candidates stop at describing the annoyance and never separate which specific facts inside an AI output are allowed to be a guess.
Follow-up traps
"Isn't the fix just a database lookup? What does that have to do with AI at all?" Response: the lookup only matters because the model was allowed to state that same fact freely in the first place. The real decision is which parts of an AI-generated message are allowed to be probabilistic, and which aren't.

"What if there's no clean system-of-record field to pull the fact from?" Response: then the guardrail becomes a confidence threshold that routes to a person instead of asserting a value, the same principle, just without a clean data source standing behind it.
If pressed
Pulling the order system's data into the chat tool live added about 340 milliseconds to load a ticket, well inside the 2 second budget the drafting tool already had. The fix didn't need a new latency budget. It fit inside the one that already existed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more