The direct answer
Treat a pilot's clean numbers as proof that a small, closely-watched group's manual workarounds were working, not proof the product is ready for everyone. Before general release, write down every place a person on the pilot team quietly caught something the product should have caught on its own, and build each one into the product before that hand-holding stops.
Do this, in order
Treat the pilot's clean numbers as proof the small group's manual catches worked, not proof the product is ready.Why: the clean pilot ran on a handful of users and a specialist quietly fixing the same mistake by hand before it happened twice, and the product never had to catch it alone.
Write down every place a person on the pilot team caught something the product should have caught itself.Why: a gap nobody named is a gap nobody remembers to close before the team that was covering it moves on.
Build the specific catch back into the product, checked against every group's own real cases, before opening it past the pilot.Why: this is the actual piece of hand-holding a wider rollout removes, and it's the one that reaches a real decision.
Don't let a rollout get approved off pilot numbers alone, without anyone naming what the pilot team was doing behind them.Why: a committee that signs off on a number it never sees the manual work behind is approving a claim nobody actually tested.
Sample real output by group after launch, on a schedule, not just once before it.Why: this is the only way to catch the gap reopening before someone outside the team has to catch it instead.
Leave the low-stakes parts of the product out of the same gate.Why: forcing every draft through the same heavy check buries the one gate that actually matters.
How to answer this, stage by stage
Seven moves, from naming one real rollout to the line you'd close on.
1
Name one real rollout before you reach for a framework
Say it like this
"Say a hospital's AI tool listens to a doctor dictate a patient visit and writes the first draft of the chart note. It got tested for three months with six ICU doctors who had a direct line to the team that built it. Six weeks after it went live for all six hundred and forty doctors in the building, with none of that direct line left, I want to walk through what that gap actually contains."
Why this works
Grounds "the pilot-to-production gap" in one real artifact before any framework language shows up.
2
State your plan in one breath
Say it like this
"I'd run this through GUARD, because a pilot succeeding and a rollout being safe are two different claims, and the real question is who's exposed to the space between them. Who was protected during the pilot and who isn't once it's everyone, where that gap lands hardest, who can't tell the difference, the actual fix, and how I'd catch it if the fix never happened."
Why this works
Two seconds naming the plan, not a recited acronym before the real thinking starts.
Say it like this
"A pilot proving a product works and a pilot proving a product is ready for everyone are not the same claim. The pilot's numbers looked clean because a small team was quietly doing work by hand that never got built into the product itself. The moment that team can't keep up with the size of production, the clean number was never really about the product. It was about the team."
Why this works
This is where a checklist answer and a real answer split apart.
4
Commit to the mechanism, not the value
Say it like this
"I'd say a pilot doesn't clear a product for general release. It clears a list. Before anyone signs off on going hospital-wide, I want a written list of every place a person quietly caught something the product should have caught itself, and none of it counts as done until that catch is built into the product, not staffed by a person watching a chat window."
Why this works
A mechanism you could point to in the rollout plan, not a value everybody already agrees with.
5
Make it real with the near miss
Say it like this
"Here's what it looks like without that list. Six weeks into the rollout, a night-shift covering doctor signs a chart where the tool drafted vinBLASTine for a patient whose own doctor had actually said vinCRISTine, two different chemo drugs. Nobody was reading every flagged note anymore, there were twelve thousand notes a week instead of six hundred. A pharmacist's routine callback caught it before anything got ordered. A hand review of two hundred live notes afterward found the mix-up rate on departments the pilot never touched was six in a hundred, against one in a hundred on the ICU notes the tool had actually been checked on."
Why this works
The compressed version of the story below. Real numbers, a concrete gap, not a hypothetical one.
6
Turn detection into a habit, not a one-time audit
Say it like this
"I'd ask one question before trusting any pilot's numbers: what was a person doing by hand to make those numbers look that clean? If nobody can answer that, the numbers aren't describing the product. After launch, I'd sample live output by department and shift on a schedule, on purpose looking for the pilot-group-versus-everyone-else gap to reopen, and I'd track every catch that happened downstream, like a pharmacist's callback, as a sign of a check the product itself missed."
Why this works
Turns the D step into a repeatable check, not a one-time audit after something goes wrong.
7
Land the answer in one breath
Say it like this
"So: a pilot proves a product works for the small group a team was quietly protecting. Going wide is a different claim, and it only holds if every piece of that protection got built into the product before the team let go. Skip that step, and the clean pilot numbers were never about the product. They were about the six weeks somebody was still watching."
Why this works
Restates the decision in one breath, the line an interviewer remembers on the way out.
Let's learn
Every note used to take an ICU doctor at Brannock Regional fourteen minutes to write by hand.
Say a hospital brings in a tool called ChartEcho. It listens to a doctor talk through a patient visit and writes the first draft of the chart note, the medicines, the plan, the diagnosis, ready for the doctor to check and sign.
Knowledge spark: what's a drug-name mix-up
Some drug names look and sound almost the same but do very different things. VinCRISTine and vinBLASTine are both real chemo drugs, one letter apart when spoken fast, and mixing them up on a chart is not a small typo. Hospitals keep a running list of these look-alike, sound-alike pairs precisely because they trip up humans and machines the same way.
For three months, six ICU doctors tried it. Pennard Health, the company that built it, had a specialist read every flagged note and fix a drug-name mix-up by hand, the same day, before a doctor ever saw the same mistake twice. Across roughly six hundred notes, that specialist caught eleven of them. Documentation time on the unit dropped from fourteen minutes a note to under four. The pilot's numbers were about as clean as numbers get.
Then it went live for everyone. Six hundred and forty doctors, forty departments, night shift included. Notes went from six hundred over three months to twelve thousand a week.
Drug-name mix-up rate, six weeks into the rollout
A 200-note hand-reviewed sample, split by whether the pilot had ever touched that department.
ICU notes (the pilot's own unit)
1%
Departments the pilot never touched
6%
ChartEcho didn't get worse. The ICU rate held steady at what the specialist had quietly been catching all along. It's the departments that specialist never watched, oncology, psychiatry, cardiology, where the mix-up rate is six times higher, because nothing was ever built to catch it there.
Here is the turn. ChartEcho did not get worse. What stopped was the thing making its numbers look that clean: a person reading every flagged note and fixing the odd mix-up before it repeated. Nobody built that catch into the product. It was a person, in a chat window, for three months, and everyone downstream mistook the clean number for proof the product was ready.
The pilot's numbers were never really about ChartEcho. They were about a specialist nobody wrote down.
At its worst, that guess looks exactly like a fact. A night-shift doctor covering a patient they don't personally know signs a chart where ChartEcho drafted vinBLASTine for a patient whose own doctor had said vinCRISTine, two different chemo drugs. Nothing on the chart marks the drug name as shaky. It reads exactly as certain as every other line, right up until a pharmacist's routine callback catches it before anyone acts on it.
The decision I would take back
Brannock's rollout committee approved hospital-wide use off the pilot's clean numbers alone. Nobody in that meeting asked what a person was doing by hand to keep those numbers that clean, because nobody had written it down as work. It made sense at the time. The pilot looked done.
What I would leave alone. ChartEcho also drafts the after-visit summary a patient gets by text, the appointment reminder, the plain stuff. Get a word wrong there and the patient calls to double-check. Nobody signs a medical order off a text reminder. That part doesn't need the same gate.
The lesson. A pilot that looks finished has usually just been finished by a person nobody counted. The real work of going from pilot to production isn't retesting the parts that already worked. It's finding the parts a person was quietly holding up, and building each one into the product before that person is asked to hold up six hundred and forty people instead of six.
Now here is the same thing as a story
The short version sits above. Read this one for the six weeks nobody was watching the thing that used to be watched every day.
For three months, the best part of Katarina Boone's week was a Friday afternoon that kept finding nothing wrong.
She runs product for ChartEcho at Pennard Health. Four years in, she's the one other teams call when a rollout needs to move from one hospital to the next without anyone getting hurt on the way.
The Brannock pilot started in February. Six ICU doctors, a fourteen-bed unit, three months. Every Friday, Katarina sat down with her implementation lead and read through that week's flagged notes together, line by line, looking for anything ChartEcho had drafted that didn't match what the doctor actually said. Most weeks it was nothing. Some weeks it was a drug name, close enough to the real one that a tired doctor might not catch it, fixed that same afternoon before it could happen twice.
By April, the pilot was the easiest thing on Katarina's calendar. Documentation time on the unit was down from fourteen minutes a note to under four. The doctors liked it. The Friday review kept finding almost nothing to fix. She started trusting the dashboard's weekly summary instead of reading every flagged note herself, because the summary and her own read never disagreed.
One of them had a lever. The other one only had a signature.
Then Brannock's leadership signed off on going hospital-wide. Every doctor, every department, starting in June.
Katarina knew the volume was about to jump, and it did, from roughly six hundred notes over the whole pilot to twelve thousand a week. What she didn't clock, not really, was that the Friday review had never been a feature. It was her and one specialist, reading real notes, catching a real thing, and nobody had ever put "keep doing this at forty times the size" on a launch checklist, because nobody had ever called it a checklist item. It was just Friday.
Six weeks in, a pharmacist doing a routine callback on an oncology order noticed something off. The chart read vinBLASTine, six milligrams. The doctor who'd dictated it, covering overnight for a colleague, had actually said vinCRISTine, two milligrams, two different chemo drugs, close enough at speed to slide past a doctor who'd never met this patient before that shift. Nothing caught it. Nothing ChartEcho, or the chart, or the ordering system, marked as shaky. It read exactly as certain as the other eleven meds on the same note.
The order never went in. But Katarina spent the rest of that week pulling a sample, two hundred live notes from across the hospital, and reading them herself, the way she used to on Fridays. On ICU notes, the mix-up rate was almost exactly what her old Friday reviews used to find and fix: about one in a hundred. On notes from departments the pilot had never touched, oncology, psychiatry, cardiology, it was one in seventeen.
The step that used to be a person, and never became anything else
Back in February, nobody in the room asked what would happen to the Friday review once it wasn't six doctors anymore. If they had, the honest answer would have been: it doesn't scale, and nothing is standing in for it once it stops. That made sense at the time. The pilot only had six doctors. Reading their notes by hand on a Friday was the obviously right way to check the work.
Run the same rollout again, with one change. Before ChartEcho goes live past the ICU, the thing Katarina and her specialist were doing by hand becomes an actual feature: every drafted medicine gets checked against that department's own drug list, flagged if it's a known look-alike, sound-alike pair, before a doctor ever signs it. The same near miss happens in June. This time the chart shows "vinBLASTine, flagged, confirm dose and drug" instead of a clean line. The doctor checks it in fifteen seconds. Nobody needs a pharmacist's callback to catch it, because the product caught it first.
What I'd tell myself, looking back at those Friday afternoons: I thought I was checking the product. I was the product, and I never wrote that down anywhere anyone could see.
GUARD, once the hand-holding stops
This reads like a scaling question. The real test is whether anyone wrote down what a small pilot team was doing by hand, before the rollout removed them without replacing them.
G, groups. Katarina and her implementation lead, who could read every flagged note by hand at six doctors' worth of volume, and every doctor at Brannock once it's six hundred and forty, especially the ones covering a patient overnight who've never met them before.
U, unequal. The gap costs almost nothing on the ICU notes the pilot actually tuned itself on. It lands hardest on departments the pilot never touched, oncology, psychiatry, cardiology, where the drug names, the shorthand, the whole vocabulary are things nobody was ever reading by hand in the first place.
A, ability to contest. A night-shift doctor signing a chart for a patient they don't know has no way to tell that this specific note is coming from a system that's never actually been checked against oncology's own drug list. The chart reads exactly as certain as every note that has been checked.
R, reduce. A named list of what has to be built before general release, not an assumption that pilot success means production is ready: the drug-name cross-check the specialist ran by hand becomes a real feature, checked against every department's own list, before that department goes live.
D, detect. Ask what a person was doing by hand to keep the pilot's numbers clean, before trusting them. After launch, sample live notes by department and shift on a schedule, and treat every downstream catch, like a pharmacist's callback, as proof of a check the product itself missed.
Where this answer would fail
If the fix here is "hire more specialists to keep reading flagged notes at the new volume" or "ask doctors to double-check drug names more carefully," it doesn't count. Those are more hands doing the same manual catch, not the catch built into the product. The only version that closes the gap is the one where the check runs without a person watching a chat window for it.
And if you want to be sure it really works, try it somewhere else
A mid-size city's building department pilots an AI tool that reads a submitted permit application and flags which code sections it likely trips. Four months, three senior reviewers in the downtown office, residential permits only. The vendor's code lead watched every flagged case and patched prompt gaps for odd local rules within the day. It's set to go citywide next quarter, all nine district offices, every permit type, including the historic district the downtown office never sees.
G, groups. The downtown reviewers, including Imre Papp, who had a direct line to the vendor's code lead for four months, and every reviewer citywide, including newer staff in offices with no such line.
U, unequal. Barely matters on the plain residential permits the tool was tuned on downtown. It lands hardest on historic-district applications, an entire code chapter the pilot never triggered once, because no historic properties file through the downtown office.
A, ability to contest. An applicant in the historic district has no way to know their reviewer's sign-off leaned on a tool that has never actually seen this chapter. The permit review board approved citywide use off downtown's numbers alone.
R, reduce. A named coverage list: which code chapters the tool has actually been checked against, by real cases, not just installed for. No office gets sign-off authority for a chapter that isn't on the list yet.
D, detect. Track every signed permit by which code chapter it cited, and flag any chapter the tool was never checked against before it's used to wave one through.
Swap the trigger and it still runs
- Speed: leadership wants the rollout finished before the next budget cycle, so the gap-closing list gets treated as a nice-to-have instead of a gate.
- Cost: building the drug-name check into the product properly costs real engineering weeks, so it keeps losing to whatever else is on the roadmap for launch.
- The model gets better: a newer version of the tool handles the ICU cases even more smoothly, which makes it easier to assume it's ready everywhere, not harder, because the demo looks even more finished.
Where people run it wrong
- Treating "the pilot's numbers were clean" as proof the product is ready, instead of proof a small team was covering for it.
- Fixing the gap by adding more people to watch the same thing at a bigger scale, instead of building the catch into the product.
- Reading a steady dashboard average as evidence nothing's wrong, when the average is hiding one group's rate quietly climbing while the group the pilot tuned on looks fine.
How to use it live
Ask "what was a person doing by hand to make this pilot look this clean?" before you ask anything about the accuracy number. If nobody can answer that, you've found the gap in about five seconds.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
What framework fits a question about the pilot-to-production gap, and why?
Tap to flip
ANSWER
GUARD, for risk and fairness. The real question is who's exposed once a pilot's small, closely-supported group stops being the only ones using the tool.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Katarina Boone, a product manager at Pennard Health, who ran ChartEcho's rollout from a six-doctor ICU pilot at Brannock Regional Medical Center to all six hundred and forty physicians.
3 · THE HABIT
What did Katarina stop doing once the pilot looked finished?
Tap to flip
ANSWER
She stopped reading every flagged note herself on Fridays with her implementation lead, and started trusting the dashboard's weekly summary instead, because the two never used to disagree.
4 · THE GAP
What's the gap this answer turns on?
Tap to flip
ANSWER
The pilot's clean numbers came from a specialist quietly catching drug-name mix-ups by hand, same day, before a doctor saw one twice. That catch never became a real feature, so it disappeared the moment volume made hand-checking impossible.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
Brannock's rollout committee approved hospital-wide use off the pilot's clean numbers alone, without anyone asking what a person was doing by hand to keep them clean. It made sense because the pilot genuinely looked done, and nobody had ever written the Friday review down as a checklist item.
6 · THE NUMBER
Fill in: in the 200-note sample taken after the near miss, the drug-name mix-up rate on departments the pilot never touched was ______ percent, against 1 percent on the ICU's own notes.
Tap to flip
ANSWER
6 percent. The gap between the two numbers is the whole story: the ICU rate held steady at what the specialist used to catch by hand, and everywhere else, nothing was ever built to catch it at all.
7 · THE REPLAY
Same near miss, drug-name check built into the product this time. What changes?
Tap to flip
ANSWER
The chart shows "vinBLASTine, flagged, confirm dose and drug" instead of a clean line. The covering doctor checks it in fifteen seconds. No pharmacist callback is needed, because the product catches it before the chart is ever signed.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the "unequal" gap become?
Tap to flip
ANSWER
A city building department's AI permit-review tool, piloted downtown on residential permits only. The gap: it barely matters on the plain permits it was tuned on, and lands hardest on historic-district applications, a code chapter the downtown pilot never once triggered.
Check yourself Score: 0 / 0
Multiple choice
1. The Brannock ICU pilot's numbers looked clean, with drug-name mix-ups caught and fixed before a doctor ever saw one twice. What does that actually prove?
- A. ChartEcho is safe to use on any hospital department's drug list.
- B. A specialist's manual review was catching what the product itself couldn't yet catch.
- C. Six doctors is a large enough sample to trust the tool hospital-wide.
- D. A one-in-a-hundred error rate would hold steady no matter which department used it.
Show hint
Ask who was reading the flagged notes every Friday, and what happened to that person once the rollout was hospital-wide.
Show answer
B. A, C, and D all treat the pilot's cleanliness as a fact about the product or the sample size. The real driver was a person catching mix-ups by hand before they repeated, work that was never built into ChartEcho itself.
True or false
2. True or false: since a specialist reading flagged notes by hand caught the pilot's mix-ups, the fix for the rollout is to hire more specialists to keep reading flagged notes at the new, much bigger volume.
Show hint
Ask whether the fix changes what's checking the note, or just how many people are doing the same manual check.
Show answer
False. More people doing the same manual read at forty times the volume doesn't scale, and it's still a person, not the product, doing the catching. The fix is building the drug-name check into ChartEcho itself.
Fill in the blank
3. In the 200-note sample taken six weeks into the rollout, the drug-name mix-up rate on the ICU's own notes was ______ percent.
Show hint
It's the number that shows ChartEcho itself didn't get worse, only untested everywhere the pilot never went.
Show answer
1 percent. Almost exactly what the Friday review used to catch and fix by hand. The gap only shows up against the 6 percent figure from departments the pilot never touched.
True or false
4. True or false: the after-visit text summary ChartEcho sends to patients needs the same drug-name cross-check gate before it can be trusted.
Show hint
Ask whether a wrong word there reaches a signed medical order, or just a phone call from a patient.
Show answer
False. A wrong word in a text reminder gets caught by the patient calling to check. Nobody signs a medical order off that text, so the same heavy gate there would be process for its own sake.
Short answer, apply it yourself
5. Think of a tool at your own workplace, or one you use, that looked ready after a small pilot. What do you think a small team was quietly doing by hand during that pilot that never got built into the actual product?
Show hint
Look for the moment a tool started behaving like nobody was reading its output by hand anymore.
Show answer
Model answer: "A scheduling tool at my last job worked perfectly during its pilot with one small team, because our team lead was manually double-checking every conflict it flagged before trusting it. Once it rolled out company-wide, nobody replaced that double-check, and double-bookings started showing up in departments the pilot never used."
Short answer
6. If the audit had found the non-ICU mix-up rate was 2 percent instead of 6 percent, would the fix still be the same built-in drug-name check, or would something smaller do? Why?
Show hint
Ask whether the fix targets a specific bad number, or the process that let an unbuilt check go live at all.
Show answer
Model answer: "Still the same built-in check, just less urgently needed. Even at 2 percent, that's still one drug-name mix-up in fifty, on charts nothing marks as shaky, in a hospital that's already had one near miss. The real problem, that the check was never built into the product, would still be true regardless of the exact number."