CaseIntermediateResponsible AI & Advanced Practice / Internal AI tooling and enablement products / #15
Design the training programme that accompanies an internal AI rollout.
SPARK the scenario: Falkirk Ground Services, an airport ground-handling company training ramp agents on an AI tool that drafts turnaround delay-cause reports
Interviewer's question: "Design the training programme that accompanies an internal AI rollout." Yolanda Prescott runs ramp operations training at Falkirk Ground Services. Her ramp agents are about to start using an AI tool that drafts the delay-cause code after every aircraft turnaround.
The direct answer
Build the training entirely out of real cases the tool has actually gotten wrong before, not a feature tour of its buttons. Trainees practice spotting and fixing about thirty real misclassified drafts before they ever touch a live one. The habit worth building isn't knowing where the confirm button is. It's noticing the one draft, in the middle of a rushed turnaround, that reads fine but is wrong in a way that would bill the wrong airline.
Do this, in order
Build the curriculum from real historical failure cases, not a feature tour.Why: an agent who's only seen a slideshow has no instinct for when to distrust a draft under time pressure.
Make trainees practice fixing wrong drafts, not just reading about them.Why: recognizing a wrong pattern under pressure is a skill you build by doing it, not by watching someone else describe it.
Test the training against a live near miss before calling it done.Why: a curriculum that hasn't been checked against a real, current failure case is still a guess about what agents actually need.
Skip full data-literacy education and every rare delay code on day one.Why: those cost training time now for a payoff most agents won't need for months, if ever.
Keep updating the curriculum with new real cases as they happen.Why: the tool's failure shape will shift over time, and a curriculum frozen at launch goes stale exactly when it matters most.
How to answer this, stage by stage
Nobody is grading whether you can list training modules. They're grading whether your curriculum survives the first time the tool is confidently wrong under real time pressure.
Stage 1
Scope it to one person and one real task
Say it like this
"I'll design this for Yolanda Prescott's ramp agents at Falkirk Ground Services, specifically the delay-cause code they draft after every turnaround."
Why this works
Grounds an abstract training question in one real, repeatable task.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, payoff, anchor, risk, keep out."
Why this works
Signals a real design method before a single curriculum detail is said.
Stage 3
Show today, without the tool
Say it like this
"Right now, an agent picks a delay code from memory in the ninety seconds after a rushed turnaround, and usually defaults to the same generic code under pressure."
Why this works
The situation step. You can't design training for a gap you haven't shown existing.
Stage 4
Name the real habit you want
Say it like this
"I don't want agents to just learn the tool's interface. I want them to catch a wrong draft before it becomes the billing record an airline sees."
Why this works
This is the payoff step, and the whole insight the question is testing for.
Stage 5
Give the one anchor decision
Say it like this
"The anchor: the whole curriculum is built from about thirty real past drafts the tool got wrong, and trainees practice spotting and fixing every one before they touch a live case."
Why this works
Matches the direct answer, concrete enough that anyone could picture the actual training session.
Stage 6
Prove it survives the day the tool is wrong
Say it like this
"Say a wrong draft slips through on a busy afternoon. An agent trained on real failure cases has actually seen this shape of mistake before, and catches it in seconds instead of rubber-stamping it."
Why this works
The risk step. Training that only works in a calm classroom isn't training, it's a feature demo.
Stage 7
Say what you're leaving out on day one
Say it like this
"Day one skips full data-literacy training and every rare delay code. Both cost real time now for a payoff most agents won't need for months."
Why this works
The keep-out step. Naming real, deliberately excluded content shows judgment about scope.
Stage 8
Close on the one line
Say it like this
"Train on real failure cases, not features, and never call the curriculum finished until it's been checked against an actual live mistake."
Why this works
Restates the direct answer, ready for whatever gets pushed on next.
Let's learn
Before any of this, a ramp agent picked a delay code from memory in about ninety seconds, right after sprinting back from an aircraft, almost always the same generic code, because the finer categories were hard to recall under real time pressure.
Four steps, and the third one is where the guessing has always lived.
The new AI tool reads turnaround data and drafts a delay-cause code in seconds, for the agent to confirm or correct. Falkirk's first training rollout was a straightforward feature tour: here's the screen, here's the confirm button, here's how to override it.
Knowledge spark: why does a wrong delay code matter more than it sounds?
Delay codes decide which airline gets billed for the extra ground time. A wrong code isn't just a paperwork error, it can send a real invoice to the wrong company, and that dispute takes far longer to untangle than the ninety seconds it took to create it.
The feature tour taught agents where every button lived. It didn't teach them what a wrong draft actually looks like, so the first time the tool confidently produced one, most agents confirmed it without a second look, exactly the way they'd been trained to.
The anchor: training built from real failure, not features
Instead of a feature tour, Yolanda rebuilt the curriculum around thirty actual past drafts the tool had gotten wrong, pulled from its first weeks of quiet internal testing. Trainees practice spotting exactly what made each one wrong, and fixing it, before they ever see a live case.
At its worst: a feature-tour-only rollout ships across the whole ramp, agents rubber-stamp wrong drafts during every rushed turnaround, and Falkirk spends months untangling billing disputes with airlines who were charged for the wrong delay.
What I would leave alone: Falkirk's actual turnaround procedures on the tarmac don't need to change at all. The fix lives entirely in how agents are trained to read the tool's output, not in how the ground work itself gets done.
The lesson: a feature tour teaches where the buttons are. Only real failure cases teach someone what wrong actually looks like under pressure.
Now here is the same thing as a story
The short version above is what you'd pitch to Falkirk's operations director. Read this one for how the near miss actually forced the redesign.
The ramp gets its worst rush around 4pm, when three aircraft turn around back to back and every agent is sprinting between gates.
Yolanda Prescott's first training rollout, the feature tour, went fine on paper. Agents could navigate the tool's screen confidently within a single afternoon session. Adoption looked strong in week one.
Same tool, same rushed afternoon. Only one of these two agents had actually seen this shape of mistake before.
In week two, during the 4pm rush, the tool drafted a delay code blaming a late inbound crew, when the real cause was a mechanical hold, an easy mistake to make when two causes happen close together. An agent trained only on the feature tour confirmed it in the same three seconds he confirmed everything else, and it went out to be billed to the wrong airline before anyone caught it.
Nobody had taught him what wrong looked like. They'd only taught him where the confirm button was, and he pressed it exactly like he'd been shown to.
The billing dispute took six weeks to untangle and cost Falkirk a real client relationship's worth of goodwill. Yolanda pulled every wrong draft the tool had produced during its quiet internal testing phase, thirty of them, and rebuilt the entire curriculum around spotting and fixing exactly those thirty patterns.
Four steps per case, repeated thirty times, and that repetition is the whole curriculum.
Wrong-draft catch rate, feature tour versus failure-case training
Same tool, same rushed afternoon shift. The gap is entirely in what the training actually practiced.
Delay-code accuracy, week by week across the training rollout
Accuracy climbed steadily as each cohort finished the failure-case curriculum, not in one jump on launch day.
Run the near miss forward under the new curriculum. The exact combination, a crew-swap and a mechanical hold landing close together, is one of the thirty cases every agent has already practiced spotting. The draft still comes out wrong. This time, it doesn't leave the gate uncorrected.
SPARK, in one screenNot a training checklist. SPARK is what forces the curriculum to survive the exact mistake that made this training necessary.
S
Situation. How the job is done today.
An agent picks a delay code from memory in ninety seconds, defaulting to the same generic code under pressure.
Grounds the design in a real, current gap instead of a hypothetical one.
P
Payoff. The habit the training should build.
Catching a wrong draft before it becomes the airline's billing record, not knowing where the buttons are.
Not "tool literacy," but a specific, testable instinct.
A
Anchor. The one design decision.
A curriculum built entirely from thirty real past failure drafts, practiced hands-on before any live case.
The hardest step, and the direct answer: train on real failure, not on features.
R
Risk. What breaks it.
A rushed 4pm turnaround, the exact moment an agent trained only on features rubber-stamps a wrong draft.
Proves the anchor survives the exact failure this question is really testing for.
K
Keep out. What waits for later.
Full data-literacy education, every rare delay code, and a formal certification exam. Real content, deliberately delayed.
Shows judgment about scope, not a longer wish list.
Three real ideas, and every one of them would have cost training time the near miss didn't allow for.
The recap, one line per letter: situation is a rushed guess with no real cause data, payoff is catching a wrong draft before it bills the wrong airline, anchor is thirty real failure cases practiced hands-on, risk is the training surviving a rushed 4pm shift, and keep out is the heavier content deliberately left for later.
And if you want to be sure it really works, try it somewhere elseSame five letters, a home-improvement retail chain instead of an airport ramp. Nothing else about the two jobs is alike.
Pikeworth Home & Hardware trains its returns-desk staff on an internal AI tool that drafts return-fraud-flag documentation for suspicious returns.
Six weeks from a built curriculum to full coverage, and the middle milestone is the one that proved it actually worked.
Mapped onto SPARK: situation is a returns clerk guessing at a fraud flag from a rushed customer interaction, often defaulting to "approve" under pressure to keep the line moving. Payoff is catching a wrongly cleared fraud flag before a repeat offender walks out with a refund, not memorizing the tool's dropdown menus. Anchor is a curriculum built from real past cases where the tool wrongly cleared or wrongly flagged a return, practiced hands-on. Risk is a busy Saturday afternoon return line, the exact moment a feature-tour-trained clerk waves through a draft they never learned to question. Keep out is full fraud-pattern statistics training and every rare return-policy edge case, both saved for a later refresher.
The generic feature tour sits alone in the least useful corner, exactly where it belongs.
Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "build training from real failure cases, not features, and check it against an actual mistake before calling it done," and stop.
Cost: if there's no time to gather thirty real cases before launch, start with the five worst ones from testing and add more each week, since five real cases still beat a feature tour of zero.
The model gets better, for real: if the tool's draft accuracy improves over time, the training still matters, because rarer mistakes are exactly the ones agents are least prepared to notice, having seen fewer of that shape by then.
Where people run it wrong.
They build training around the interface instead of around what a wrong answer actually looks like.
They treat a calm classroom session as proof the training works, without ever testing it against a real rushed shift.
They try to cover every rare case on day one and run out of time for the common ones that actually matter most.
How to use it live. When someone asks you to design training for an AI rollout, ask yourself first: has this curriculum ever been checked against something the tool actually got wrong. If the answer is no, it's still a feature tour wearing a training programme's name.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "design the training programme for an internal AI rollout"?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Built for design questions, not a metric or estimation method.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Yolanda Prescott, who runs ramp operations training at Falkirk Ground Services, redesigning the curriculum after a real billing near miss.
3 · THE SITUATION
How was the delay-cause code chosen before any AI tool existed?
Tap to flip
ANSWER
An agent guessed from memory in about ninety seconds, right after a rushed turnaround, usually defaulting to the same generic code.
4 · THE ANCHOR
What's the one design decision the whole training curriculum hangs on?
Tap to flip
ANSWER
Building it entirely from thirty real past drafts the tool got wrong, and having trainees practice spotting and fixing every one.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Rolling out a feature-tour-only curriculum first, which taught agents where the buttons were but never what a wrong draft actually looks like.
6 · THE NUMBER
Fill in the blank: in a live pilot test, feature-tour-trained agents caught only ___ percent of wrong drafts, versus 89 percent for failure-case-trained agents.
Tap to flip
ANSWER
34 percent. The same tool, the same shift, and a gap that came entirely from what the training had actually practiced.
7 · THE REPLAY
Same near miss, new curriculum. What changes?
Tap to flip
ANSWER
The exact crew-swap-plus-mechanical-hold combination is one of the thirty practiced cases, so the wrong draft gets caught before it ever leaves the gate.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different company. Which one, and what's the anchor there?
Tap to flip
ANSWER
Pikeworth Home & Hardware's returns desk. Same anchor: a curriculum built from real past fraud-flag mistakes, practiced hands-on, applied to return processing instead of delay codes.
Check yourself Score: 0 / 0
True or false
1. True or false: the agent who confirmed the wrong delay code during the 4pm rush was careless or poorly motivated.
True
False
Show hint
Look at the block-highlight line in the story.
Show answer
False. He pressed the confirm button exactly as trained. Nobody had ever taught him what a wrong draft looked like.
Multiple choice
2. Why did the feature-tour curriculum fail to prevent the billing near miss?
A. Because agents didn't pay attention during training.
B. Because it taught the interface, not what a wrong draft actually looks like under pressure.
C. Because the AI tool's model was too inaccurate to train around.
D. Because the training session ran too short.
Show hint
Look at the payoff step, P.
Show answer
B. Tool literacy and the instinct to distrust a wrong draft are two different skills, and only one of them was ever taught.
Fill in the blank
3. Fill in the blank: Yolanda rebuilt the curriculum around ___ real past drafts the tool had gotten wrong during testing.
Show hint
Look at the labeled parts diagram and its caption.
Show answer
Thirty. Each one practiced hands-on: spot the wrong draft, fix it, understand why it happened.
Short answer, where it wouldn't matter
4. Name a part of Falkirk's ramp operations this training redesign genuinely didn't need to touch.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The actual turnaround procedures on the tarmac. The fix lived entirely in how agents were trained to read the tool's output, not in how the ground work itself gets done.
Short answer, apply it yourself
5. Think of a tool you were trained on at work. Was the training built from real past mistakes, or mostly a tour of its features? What would change if it were rebuilt around real failure cases?
Show hint
Think about whether the training ever showed you what a wrong output actually looks like.
Show answer
Model answer: Most workplace tool training is a feature tour. Rebuilding it around real, specific past mistakes would teach the actual instinct to catch a wrong output, not just where the settings live.
Short answer, the number question
6. If the pilot test had used only ten real failure cases instead of thirty, would the 89 percent catch rate likely have held? Why or why not?
Show hint
Think about how many different failure shapes ten cases could realistically cover.
Show answer
Model answer: Probably not fully, since ten cases cover fewer of the tool's actual failure shapes, and an agent is less likely to have practiced the exact pattern they encounter live.
Before you close the answer
Why this works
Tests whether your training design actually builds an instinct for distrust, instead of just teaching a tool's interface and calling it done.
Follow-up traps
"Isn't thirty cases just a small sample of everything that could go wrong?" Response: yes, which is why the curriculum keeps growing with new real cases over time, rather than freezing at launch and going stale exactly when the tool's failure shape shifts.
"Won't agents just memorize the thirty cases instead of learning the underlying pattern?" Response: each case is taught with why it happened, not just what the fix was, so the training targets the reasoning, not rote memorization of thirty specific answers.
If pressed
Falkirk's actual curriculum refresh cycle pulls five new real cases every month from live corrections, and retires the oldest five once every agent has demonstrated catching them reliably, so the training set stays current without growing indefinitely.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.