ConceptIntermediateShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #4
Explain when a Wizard of Oz prototype beats a real model.
The direct answer
Use a Wizard of Oz prototype, a real person quietly typing the bot's replies, when the thing you still don't know is whether people will use and trust this new kind of conversation at all. Save the real model for once you know the shape of that conversation. Build the model first and you learn how to fix its prompts before you learn whether anyone wanted to type to a bot in the first place.
How to pick between a Wizard of Oz shell and a real model, in order
Run the Wizard of Oz version first, a real person behind every reply, before committing to a real model.Why: it answers whether anyone wants this conversation at all, the question a half-built model can't answer.
Name who feels each kind of miss, in real time and real days, not a feeling.Why: cheap-and-visible only outranks hidden-and-expensive once both sides have a number.
Watch for the people behind the curtain settling into a small, stable set of answers.Why: that settling, not a date on a calendar, is the real evidence the conversation is narrow enough for a model.
Scope the first real model to only the answers the pilot proved safe.Why: a model let loose on the whole conversation repeats the exact mistake the pilot was built to catch.
Set the kill line before the pilot starts, not once it stalls.Why: without one, "when do we build it" quietly turns into "whenever we get around to it."
How to answer this, stage by stage
This is a yes-or-no about which tool to build first, not a rulebook for every prototype, so PICK does the work here.
1
Anchor the question in one real decision
Say it like this
"Let's make this real. Say Corrandale Home Insurance wants customers to describe a claim in their own words in a chat window, instead of filling in a fourteen-field form. Before anyone builds a model, they call it Hearthline, and put a real claims specialist behind it, typing every reply by hand."
Why this works
Stops the answer floating at a dictionary definition of "Wizard of Oz" and gives the interviewer one real product to push on.
2
Say the shape of the answer before giving it
Say it like this
"I'll pick a side first, then say who feels each kind of mistake, then say which kind is actually worse, then say what would change my mind. Four moves, that's PICK, and I'll go in that order."
Why this works
Signals a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Reframe what the question is really testing
Say it like this
"This isn't really 'what is a Wizard of Oz prototype.' It's 'do you know that some questions only have a real answer once a person is on the other end of the chat, and a model can't give you that answer yet, only a guess at one.' So I'm going to say which question Corrandale actually needs answered first."
Why this works
Shows the interviewer you see past the surface ask to the real judgment being tested.
4
State the position, with the real mechanism in it
Say it like this
"My position: to learn whether customers will trust a guided conversation to start a claim, use a Wizard of Oz shell, a real specialist typing behind the Hearthline chat window, not a real model. Engineering estimated a first version of the real model at about fourteen weeks before a single customer could touch it. The Wizard of Oz shell was live to real customers in three."
Why this works
PICK rewards a real mechanism and a real number, not a vague promise to "test more."
5
Name who feels each kind of miss, with a real case behind it
Say it like this
"Here's the split. If a specialist writes a slow or clumsy reply, the specialist feels it, and a supervisor catches it in the nightly transcript check, same day. Now picture a model doing this instead. A customer asks if a slow pipe leak is covered. The model answers with total confidence: yes. It's actually excluded. Nobody finds out until the claim gets denied, an average of 46 days later."
Why this works
A real case turns "asymmetry" from a word into something the interviewer can picture happening to an actual person.
6
Say what you'd leave alone
Say it like this
"I wouldn't run a Wizard of Oz test on the claims math itself, working out what a payout should actually be. That number has to come from the real system. No specialist behind a curtain can safely fake it. Wizard of Oz answers the trust question. It was never built to answer the arithmetic one."
Why this works
Shows judgment instead of treating Wizard of Oz as the right tool for every part of the product.
7
Name the kill criteria and close on one sentence
Say it like this
"I'd flip to building the real model once the specialists' own answers settle into a small, stable set, at least 75 percent of conversations, held for two weeks running, with a hard rule that anything outside that set still goes to a person. Until that happens, a fourteen-week model answers a question nobody's actually asked yet."
Why this works
Ends on the line the interviewer remembers, and shows the pick isn't permanent, it's a decision that can be revisited with evidence.
One more thing before the walkthrough moves on: this pick is about order and scope, not a ban on real models. Most candidates hear "Wizard of Oz vs. a real model" and answer like it's forever. Say which specific question needs a human first, and you've shown judgment instead of reciting a definition.
Let's learn
Hearthline is a chat window on Corrandale Home Insurance's website. Instead of filling in a form, a customer types what happened, in their own words, and the window asks a few follow-up questions until a claim gets started.
Knowledge spark: what is a Wizard of Oz prototype?
A product that looks automatic to the person using it, but a real human is doing the work behind the screen, reading what comes in and typing back the reply. It lets a team test the idea before anyone builds the thing that would actually run it on its own.
Before Hearthline, starting a claim meant a fourteen-field form. Customers took about nine minutes to finish it, and 22 percent gave up partway through.
Behind Hearthline, at the very start, there was no model at all. A real claims specialist read what a customer typed and wrote every reply by hand, pretending to be the bot. Customers finished describing their claim in about three minutes, and the number who gave up fell to 6 percent.
Here is the turn. Those fast, low-drop numbers are not really the finding. They only prove that people will talk to a chat window about a claim. They say nothing about whether a real model could hold that same conversation safely, because there still wasn't one.
Days before the mistake is caught
The green bar is small on purpose: a specialist's slow or clumsy reply shows up in that night's transcript check, same day, and gets fixed with a note. The red bar is what a live model's confident wrong answer about coverage would cost: nothing flags it, because the model sounds just as sure when it's wrong as when it's right, so it sits there until a customer's claim is denied and someone finally asks why.
The fast numbers proved people would talk. They never proved a model could answer safely.
At its worst, a model gets built straight off that early success, skips the trust question, and answers a real customer's coverage question with total confidence, and gets it wrong. Nothing was ever built to catch that until a claim is denied and the customer asks why the chat told them otherwise.
The choice I would take back
Corrandale's original roadmap gated the whole pilot behind a working model: build Hearthline's model first, then test it on real customers. I would take that back. Build the chat shell first, put a person behind it, and learn whether customers want this conversation before spending fourteen weeks teaching a model to hold it.
What I would leave alone. The claims-math system, exactly as it is. Working out what a payout should actually be needs the real system. No specialist typing behind a curtain can safely fake that number. Wizard of Oz answers the trust question. It was never built to answer the arithmetic one.
The lesson. A prototype's job depends on the question. A Wizard of Oz shell tests whether people want the conversation. A real model tests whether it can hold one safely. Build the second before answering the first, and real engineering weeks get spent on a question nobody had actually asked yet.
What the customer sees, and what is actually there
The same five answers, said again and again
You don't need this to answer the question. Read it if you want to feel why the split has to happen on real evidence, not on a date picked before the pilot even started.
Thora Wexley has run product for Corrandale's digital claims team for four years, since before Hearthline was a chat window or anything else.
When the Hearthline pilot went live, three claims specialists sat behind it, invisible to the customer, writing a fresh reply to whatever someone typed. Thora read every single transcript herself each night for the first three weeks, about forty minutes a night, checking that the replies actually matched what a trained adjuster would say.
They always matched. The specialists were good at their job, and for a while Hearthline read like the best version of the old phone line, just typed.
So she stopped reading every transcript. She sampled ten a night. Then five. By week six she was skimming a daily summary a specialist wrote for her.
Then, in week eight, a new specialist joined the desk and asked the group chat a question. Why did everyone keep pasting almost the same four sentences for a water-damage claim? Wasn't the whole point to write something fresh, the way a real adjuster would?
Nobody had a good answer. So Thora pulled the real numbers instead of a feeling.
By week six, 78 percent of every conversation had already been answered using one of five saved snippets the specialists had quietly built for themselves: one for water damage, one for a cracked window, one for a late payment question, one for a burst pipe, one for storm roof damage. Nobody had told them to do this. It just turned out that most claims, most days, needed almost the same thing said back.
Same product, two very different ways to get it wrong
The specialists hadn't stopped doing their job. They'd found the shape of it.
I want to say the problem here is that the specialists got lazy. They didn't. The real problem is that nobody had set up a way to notice this on purpose. The pilot was scoped to run twelve weeks, then Thora's team would decide, model or no model, on the calendar date they'd picked before the pilot ever started.
Months earlier, in the meeting where they scoped the pilot, someone asked how they'd know it was time to build the real thing. The answer that won the room was "let's just run it for a quarter and see." Twelve weeks felt long enough to be safe and short enough to keep momentum. Nobody wrote down a number that would tell them sooner.
Once Thora had the real percentages in front of her, the pattern was obvious: 22, 38, 50, 60, 68, 78 percent, week after week, then holding above 75 for two weeks running by week seven. With a kill line marked at 75 percent held for two weeks, Thora's team would have had their answer by week seven instead of week twelve, five weeks earlier, and they'd have known exactly which five answers were safe to hand to a model first, with everything else still going to a person.
One design waits for a date on the calendar. The other watches for the moment the conversation actually stops surprising anyone.
And the thing I'd tell myself, back in that scoping meeting: we asked how long we'd run the pilot. We never asked what evidence would tell us we could stop early. On a fixed date, that's never. On real transcripts, it's usually sitting right there by week six or seven, if anyone's counting.
PICK, applied to the Hearthline pilot
This is a yes-or-no about which tool to build first, not a rule for every prototype Corrandale ever makes, so PICK is the tool.
P, position. Run the trust question, whether customers will describe a claim in their own words and follow a guided conversation, through a Wizard of Oz shell first, with a real specialist behind the chat. Save the real model for the claim types the pilot proves are safe and stable.
I, impact. A specialist's clumsy or slow reply is felt by the specialist, and caught by the nightly transcript check, same day. A model's confident wrong answer about coverage is felt by the customer, and by nobody on the team, until a denied claim brings it back, on average 46 days later.
C, cost asymmetry. The Wizard of Oz miss is cheap and visible: it shows up in a transcript that same night, and it's fixed with a note to one specialist. The real-model miss is hidden and expensive: the model sounds just as sure when it's wrong as when it's right, so nothing flags it, and the cost lands on a customer's claim, not on anyone's dashboard.
K, kill criteria. Flip to building the real model once the specialists' own answers settle into a small, stable set, at least 75 percent of conversations, held for two weeks running. Below that line, the conversation is still too varied for a model to hold safely on its own.
Knowledge spark: why not just watch what an off-the-shelf model does instead of paying specialists to type?
Because that's still a model, just one nobody at Corrandale trained or checked. It can still sound completely sure while getting a coverage detail wrong, and there's no cheap way to catch that early. A Wizard of Oz shell answers a different question first: do people even want this conversation, before anyone hands it to any model at all.
How fast the specialists' answers settled into a stable set
Weekly share on one of the five saved answers
Kill line: 75%, held two weeks running
The share climbs steadily from 22 percent in week one, crosses the 75 percent kill line in week six at 78 percent, and holds above it through week seven at 80 percent, the second week running. That two-week hold, not the pilot's planned twelve-week checkpoint, is the real signal the conversation had narrowed enough to hand a slice of it to a model.
Run PICK again, on a veterinary chain's symptom chat
Larkhollow Veterinary Group, a chain of eleven clinics, wants pet owners to describe a symptom in a chat window and get told whether it's urgent or can wait for a normal appointment. Before building a triage model, they put real vet technicians behind a chat called Vetline, typing every answer themselves.
P. Run the trust question, whether owners will describe a symptom in chat and act on what Vetline tells them, through a Wizard of Oz shell first. Keep a real triage model out until the technicians' own answers settle. I. A technician's slow reply is felt by the technician, caught the same shift, fixed with a note. A model's confident wrong answer, telling an owner a limping dog can wait until Monday when it actually needs same-day care, is felt by the dog and the owner, and nobody at Larkhollow knows until the pet arrives worse than it should have. C. The Wizard of Oz miss is loud and cheap: a supervisor hears about it before the shift ends. The real-model miss is quiet and expensive: an owner trusts a confident wrong answer, and the cost shows up as a harder case in the waiting room, not as a support ticket anyone reads in time. K. Once the technicians' answers settle into a small, stable set across enough seasons and species, fold those into a scoped model, with a hard rule that anything outside the set still goes to a person. Same kind of line Corrandale's team would use.
What I would leave alone, at Larkhollow
The technicians' actual medical judgment on ambiguous cases stays with a person, full stop. That's not a question either a Wizard of Oz shell or a model should be trusted to shortcut, no matter how it performs on the easy cases.
Swap the trigger and it still runs
Speed: if Corrandale needed something live in four weeks instead of twelve, the pick doesn't move, a Wizard of Oz shell can go live long before any model could anyway.
Cost: if the specialists' time turns out to cost more per conversation at scale than expected, the pick still doesn't move, the pilot was never meant to run at full volume forever, it's there to learn cheaply before committing engineering money.
The model got better: if a pre-built model turned out to already handle claim-intake conversations reliably, with proven low odds of a confident wrong answer on coverage, that's exactly the evidence that would flip the pick early.
Where people run it wrong
Understaffing the Wizard of Oz shell to save money, so the pilot measures patience for a slow chat instead of trust in a guided one.
Skipping the shell "to save time," so the team learns whether the model works before learning if anyone wanted to talk to it at all.
Running it forever with no kill line, so the people behind the curtain quietly become the permanent product, and nobody notices the pilot never ended.
If you are asked this cold
Say the reframe out loud before you answer with a definition. "Give me a second, I want to pick the one question a model genuinely can't answer yet, before I say where Wizard of Oz stops and a real model starts." That's true, it's already stage three of the walkthrough, and it buys you the time to find the real split instead of reciting a textbook line.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what's the hardest step to nail?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a real model's confident wrong answer costs more than a specialist's slow, honest one.
2 · THE PERSON
Who owns the Hearthline pilot, and what has she done for four years?
Tap to flip
ANSWER
Thora Wexley, the product manager who has run Corrandale's digital claims team for four years, since before Hearthline existed.
3 · THE HABIT
What did Thora stop doing once the specialists kept clearing every check?
Tap to flip
ANSWER
Reading every transcript herself each night. Then ten a night. Then five. By week six, just a daily summary someone else wrote for her.
4 · THE ASYMMETRY
Name the two kinds of miss here and what each one costs.
Tap to flip
ANSWER
A Wizard of Oz miss: a slow or clumsy reply, caught by the specialist, fixed the same day. A real-model miss: a confident wrong answer about coverage, caught only when a claim gets denied, on average 46 days later.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Run the trust question through a Wizard of Oz shell first, a real specialist behind the chat. Save the real model for the claim types the pilot proves are safe and stable.
6 · THE NUMBER
By week six, ______ percent of Hearthline conversations were being answered with one of five saved snippets.
Tap to flip
ANSWER
78. That's the number that told Thora the conversation was narrower than anyone thought, and it held above the 75 percent kill line for a second week straight by week seven.
7 · THE KILL CRITERIA
What evidence would flip this position back the other way?
Tap to flip
ANSWER
The specialists' answers settling into a small, stable set, at least 75 percent of conversations, held for two weeks running. Below that, a model isn't ready to hold the conversation alone.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the position land there?
Tap to flip
ANSWER
Vetline, a symptom-intake chat for Larkhollow Veterinary Group. The Wizard of Oz shell stays until the vet technicians' own answers settle into a stable set, the same kill line, a different waiting room.
Check yourself Score: 0 / 0
Multiple choice
1. Which of these is the actual mechanism behind this answer's pick?
A. Skip prototyping and build the real model straight away, since it will need to exist eventually anyway.
B. Put a real specialist behind the Hearthline chat window first, and only build the model once their answers settle into a small, stable set.
C. Build the model and a Wizard of Oz shell side by side, and compare their transcripts every week.
D. Ask customers to fill in a survey about whether they'd trust a chatbot, instead of building anything.
Show hint
Three of these either skip the trust question entirely or never produce evidence anyone could act on.
Show answer
B. A can't tell you if people trust the conversation. C spends the fourteen weeks anyway, on top of the shell. D never touches a real interaction. Only B tests the real question cheaply first.
Fill in the blank
2. Fill in the blank: before Hearthline, the paper claim form took customers about ______ minutes to finish, and 22 percent gave up partway through.
Show hint
It's the number from the old fourteen-field form, before any chat window existed.
Show answer
9. Nine minutes and a 22 percent drop-off were the numbers Hearthline's Wizard of Oz shell cut to three minutes and 6 percent, without a model doing any of the talking.
True or false
3. True or false: this position means Corrandale should never build a real model for Hearthline.
True
False
Show hint
Think about what the kill line is actually for.
Show answer
False. The position is about order and scope: run the Wizard of Oz shell first, then build a model once the pilot shows which answers are safe to hand off, not never build one at all.
Multiple choice
4. Why not just build the real model first and skip the Wizard of Oz shell, to save the specialists' time?
A. Because specialists are cheaper than engineers, so cost isn't really the reason.
B. Because a model's early mistakes get debugged by engineers, who learn about prompt quirks, not about whether a customer would ever want to type a claim into a chat window in the first place.
C. Because Corrandale's compliance team legally requires a human in every claim conversation.
D. Because a model technically cannot process customer text at all until it's fully trained.
Show hint
Think about who actually learns something during the fourteen weeks a model takes to build, and what they learn.
Show answer
B. Building the model first spends real engineering weeks answering a technical question before anyone's checked whether the underlying idea, people typing a claim into a chat window, is one customers even want.
Short answer
5. If the real-model miss had only cost 2 days to catch instead of 46, would the same position still hold? Walk through it.
Show hint
Think about whether the fix protects against the exact number of days lost, or against the fact that nothing in a live model catches its own confident, wrong answer before a customer acts on it.
Show answer
Yes, still worth it. The 46 days is evidence of how expensive the gap is, not the reason it exists. Even at 2 days, the real problem stands: nothing in a model's own design would catch its confident wrong answer before a customer acted on it. The day count changes how loud the alarm should ring. It doesn't decide whether the gap needs closing.
Short answer, apply it yourself
6. Pick a product you use yourself. Name one question about it that only a real person, not a model, could safely test first.
Show hint
Look for the question that's really about whether you'd want the conversation at all, not about whether a model could technically hold it.
Show answer
Model answer: "A grocery app's new voice-ordering feature. Whether shoppers actually want to say their list out loud in a kitchen, instead of typing it, is a question a real person listening on the other end could test in a week. Whether the checkout math comes out right afterward is a different question, and that one does need the real system, not a person guessing." Any answer works if it names the question whose true answer only exists once a real person is on the other end.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.