ConceptIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #16

Explain when a manual Wizard of Oz test beats a technical spike.

PICKa title decision faked by hand, before it was ever built

Mossgrove Audio is a small independent audiobook publisher. Saoirse Mullen runs editorial there, and had to decide whether to build a technical spike for an AI chapter-title generator, or fake the feature by hand first and see if anyone could even tell the difference.

The direct answer
Run a Wizard of Oz test first whenever the real question is about the quality of a judgment call a machine would have to fake convincingly, not whether a model can technically produce an output at all. Have an editor draft the titles by hand, log the pattern in their own reasoning, and only build the technical spike once a human faking the feature has shown the idea is worth automating in the first place.
Do this, in order
  1. Ask whether a human can fake this decision convincingly, fast, before building anything.Why: if a person can fake it in ten minutes, you can learn the same thing a technical spike would teach, for a tenth of the cost.
  2. Have the human log why they made each call, not just what they picked.Why: the reasoning is what you'll eventually need to hand a model, and it's invisible if nobody writes it down.
  3. Set a stopping rule for when the fake pattern is proven enough to automate.Why: without one, Wizard of Oz testing can drift on forever and never actually become a build decision.
  4. Reach for a technical spike instead when speed or scale is the actual thing in doubt.Why: a human faking real-time or high-volume behavior doesn't tell you anything true about whether a model can do it live.
  5. Never let a good Wizard of Oz result skip the technical spike that checks it holds at scale.Why: a pattern that works for ten titles by hand doesn't automatically survive ten thousand run through a model.

How to answer this, stage by stage

Nobody is scoring whether you've heard the phrase "Wizard of Oz." They're scoring whether you can tell, on the spot, which kind of doubt actually needs a model built to resolve it, and which kind doesn't.

Stage 1
Pick your position before the reasoning
Say it like this
"My pick: run the human version first whenever the doubt is about the quality of a judgment call, not the model's raw capability. I'll ground this in a real decision at an audiobook publisher."
Why this works
PICK opens on the commitment, not a tour of pros and cons. That's the first thing an interviewer is actually listening for.
Stage 2
Scope it to one real feature
Say it like this
"At Mossgrove Audio, the question was whether an AI tool could generate chapter titles good enough to replace an editor picking them by feel."
Why this works
Keeps the answer from floating into a generic debate about "testing methods" with nothing real attached.
Stage 3
Name who feels each kind of wrong bet
Say it like this
"If we build the wrong technical spike, engineering loses weeks on a version of the feature nobody wanted, and nobody notices until launch. If we run a bad Wizard of Oz afternoon, an editor wastes one session, and we know within a day."
Why this works
This is where the answer stops being a definition of Wizard of Oz testing and becomes an actual argument about cost.
Stage 4
Say which error you're optimizing against
Say it like this
"The visible, cheap error is a wasted editorial afternoon. The hidden, expensive one is an engineering team building a real pipeline around a title-generation pattern nobody ever confirmed listeners actually respond to. I'd rather risk the afternoon."
Why this works
Naming the asymmetry out loud is the whole test PICK is built around. Whichever error you'd rather eat is the actual decision.
Stage 5
Prove it with the compressed failure
Say it like this
"An imprint next door skipped this step, built a technical spike straight away, and shipped titles that scored fine on an internal rubric. Real listeners hated them, titles that gave away the twist. Nobody had ever asked a human editor to fake the decision first and notice that pattern."
Why this works
A real, specific failure someone else lived through is worth more than an abstract argument about testing philosophy.
Stage 6
Name the kill line, then close on the pick
Say it like this
"Once an editor faking the titles by hand gets a consistent, explainable pattern across twenty real books, that's the signal to build the technical spike and confirm it holds at scale. Until then: fake it by hand, log the reasoning, automate only the proven pattern."
Why this works
Closing with the exact point where you'd switch to a technical spike shows this isn't hedging forever, it's a real, bounded decision.

Let's learn

Say a listener finishes a chapter of an audiobook and sees a title pop up for the next one: something short, a little intriguing, never giving away what happens.

Before any of this, an editor at Mossgrove Audio picked every chapter title by hand, reading the full manuscript first, taking maybe fifteen minutes per book across forty chapters. Listeners rarely commented on titles at all, which the editorial team read, correctly, as a sign the titles were doing their job invisibly.

Hand sketched labeled parts diagram titled The wizard, close up. A person icon at the center labeled The Wizard, with four labeled callouts around it: Reads the manuscript, Picks the title by feel, Writes down why, Never tells the reader.
What the human faking the feature actually does, and the one part, writing down why, that most teams skip.

Here's the turn: the pitch for an AI title generator wasn't wrong. The doubt wasn't about whether a model could technically produce a plausible-sounding chapter title fast. It could, easily. The real doubt was whether it could reproduce the specific, hard-to-name judgment an editor makes about how much to reveal, and building a full technical pipeline to test that judgment call is expensive and slow compared to just faking it by hand first.

Time to a confident answer: Wizard of Oz test versus full technical spike
180 hrs 90 hrs 0 6 hrs Wizard of Oz, 1 week 160 hrs Technical spike, 5 weeks
Both paths were trying to answer the same question: is this pattern good enough to automate. One took an afternoon. The other took a month.

At its worst, engineering spends five weeks building a full pipeline, complete with retrieval and formatting logic, only to learn in week six that the underlying editorial judgment was never actually captured, because nobody faked it by hand first to find out what that judgment even was.

A technical spike answers "can a model produce this." A Wizard of Oz test answers "is this judgment call even worth automating." Building the wrong one first wastes the expensive kind of week.
The choice I would take back Mossgrove had priced editorial review time as a shared, unlimited internal resource, "free" in the way engineering time never gets treated as free. That made sense when review was light. It stopped making sense once teams started routing every new feature idea straight to engineering, since nobody's personal calendar ever showed the true cost of skipping a cheap human test first.

What I would leave alone: I wouldn't run a Wizard of Oz test for a feature where the real question is raw speed or scale, like whether a model can generate captions live during a broadcast. A human faking real-time behavior at scale doesn't tell you anything true about whether a model can actually do it.

The lesson: the cost of the wrong test was never really about hours. It was about which kind of wrong you'd rather find out about: a bad afternoon, or a bad month.

Now here is the same thing as a story

The short version above is what you'd say defending the plan in a five-minute editorial meeting. Read this one for what it felt like the week someone else's failure landed on Saoirse's desk two days before her own kickoff.

Here is what happens when a tool works fine in a demo and nobody has yet asked the question that actually matters: does it make the same call a good editor would make, for a reason nobody wrote down.

Saoirse's laptop has a cracked hinge held together with tape, three years of manuscript notes still open in half its tabs. She had scoped the AI title generator project for two weeks, ready to greenlight a technical spike, when a colleague from a sibling imprint mentioned, almost in passing, that their own title-generator pilot had just been quietly shelved.

Hand sketched comparison titled The asymmetry, drawn. Left panel, gauge icon labeled Wrong technical spike, caption weeks building the wrong thing, silent, expensive. Right panel, person icon labeled Wrong Wizard of Oz read, caption one bad afternoon, visible, cheap to redo.
The cost that decided everything: one wrong bet is loud and cheap. The other is quiet and expensive.

The sibling imprint had built the full pipeline first: real-time generation, a formatting layer, a review dashboard. It scored well against an internal rubric the engineering team had written themselves. Then it shipped to real listeners, and the titles it produced kept doing one specific thing: giving away the emotional turn of the chapter before the listener got there.

Knowledge spark: what is a Wizard of Oz test, exactly? A Wizard of Oz test fakes an automated feature using a real human doing the work behind the scenes, so users experience something that looks automated while a person is actually making every call. It answers whether the idea is worth automating, before anyone spends engineering time finding out whether a model can technically do it.

Nobody on the engineering team had done anything wrong, technically. The model produced fluent, plausible titles quickly and reliably. What nobody had ever tested was the actual editorial judgment: how much to reveal, and how much to hold back, a call a human editor makes by feel and had genuinely never written down.

Hand sketched decision tree titled Is this decision testable by hand fast. Root: can a human fake this in one afternoon. Three branches: yes one title ten minutes leads to Wizard of Oz first. No needs real speed or scale leads to technical spike first. Needs both leads to Wizard of Oz narrows it, spike confirms scale.
The one question that would have caught the sibling imprint's mistake before it shipped.

Saoirse changed her own plan that afternoon. Instead of scoping a technical spike, she asked one editor to fake the feature by hand: read each of the next twenty books' final drafts, pick a chapter title the old way, and this time write down, in one sentence, why that title and not a more revealing one.

Hand sketched flow diagram titled How a Wizard of Oz test actually runs, second step emphasized. Five steps left to right: Editor drafts by hand. Listener reacts. Editor logs the pattern. Repeat with next title. Automate the proven pattern.
Five steps, and the fourth one, logging why, is the part a straight-to-engineering plan always skips.

Across twenty books, a real pattern emerged, one the editor could finally name out loud: titles that named a place or an object worked better than titles that named a feeling. "The Locked Drawer" over "A Moment of Doubt," every time, because the object teased without revealing the emotional turn.

Hand sketched icon list titled What a good Wizard of Oz test needs. Three rows: a real reader, a real title decision. A way to log the human's hidden reasoning. A stopping rule for when to stop faking it.
Three things that make the fake version actually useful, instead of just a slower way to do the same job by hand forever.

Only once that pattern held for twenty real books did engineering start a technical spike, this time scoped narrowly around one instruction: favor concrete nouns over named emotions. The spike took nine days instead of five weeks, because it already knew exactly what to test for.

What I'd tell myself, hearing how close Mossgrove came to repeating the sibling imprint's exact mistake: the technical question was never the hard one. The hard one was a judgment call an editor had been making silently for years, and no model was ever going to guess it without someone faking it by hand first.

PICK, in one screenNot a script for faking every feature by hand forever. PICK is what tells you exactly which kind of doubt a human can resolve cheaper than a model can.

P
Position. Your pick, before the reasoning.
Fake it by hand first, whenever the doubt is about the quality of a judgment call, not raw model capability.
Commit to the pick before walking through the evidence.
I
Impact. Who feels each kind of wrong bet?
A wasted Wizard of Oz afternoon costs one editor a session. A wrongly scoped technical spike costs engineering weeks, building around a pattern nobody confirmed listeners actually wanted.
Naming who eats each cost turns a vague preference into a real comparison.
C
Cost asymmetry. Which error is hidden and expensive?
A bad Wizard of Oz afternoon is visible immediately and cheap to redo. A wrongly built technical spike is invisible until it ships, and by then real listeners are the ones who catch the mistake.
This is the step the whole answer turns on: optimizing against the hidden, expensive error, not the loud, cheap one.
K
Kill criteria. What flips the pick?
A consistent, explainable pattern across twenty real books, held by a human faking the feature by hand. That's the signal to build the technical spike.
Without this line, Wizard of Oz testing could run forever and never become a real build decision.

The recap, one line per letter: position is fake it by hand first for judgment-quality doubts, impact is a wasted afternoon against weeks of misdirected engineering, cost asymmetry favors the visible, cheap mistake over the hidden, expensive one, and kill criteria is a proven pattern across twenty real cases, which is exactly what let the eventual technical spike run in nine days instead of five weeks.

And if you want to be sure it really works, try it somewhere elseSame four letters, a real-time interpreter-matching tool instead of a chapter title. Different flip family entirely, the same discipline about which doubt to test by hand.

Oskar Lindqvist coordinates bookings at Cindermere Language Partners, weighing whether to spike an AI tool that matches incoming interpretation requests to the best available interpreter. Mapped onto PICK: position is fake the matching decision by hand first, since the doubt is about judgment, matching a client's specific dialect and context to the right interpreter, not raw model speed. Impact is a coordinator wasting an afternoon manually testing match quality, against engineers building a live matching pipeline around criteria nobody confirmed actually mattered to clients. The underlying flip here differs from Mossgrove's substitution flip: it's an abandonment flip, coordinators who get one bad automated match quietly stop trusting the tool for anything but the easiest, most obvious bookings, and nobody notices the tool going unused for the hard cases. Cost asymmetry still favors the cheap, visible test. Kill criteria is the same shape: a human coordinator's manual matching logic proven consistent across twenty real requests before any pipeline gets built.

The same hand sketched decision tree reused: is this decision testable by hand fast, applied here to an interpreter matching tool instead of a chapter title generator.
The same question, asked of a completely different booking desk: can a human fake this specific judgment call before anyone builds a pipeline around it.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "fake the judgment call by hand first, build the spike only once the pattern holds," and stop.
Cost: there's no editor available to run even a short Wizard of Oz test. Say so honestly, and scope the smallest real version, one editor, five books, rather than skipping straight to engineering.
The model demos surprisingly well, for real: even a strong early demo is worth a Wizard of Oz check first, since a model can sound fluent while still missing the specific, unwritten judgment call that actually matters to real listeners.

Where people run it wrong.
They treat "can the model technically produce this" as the same question as "is this the right thing to automate."
They skip logging the human's reasoning, so the Wizard of Oz test produces a good decision but nothing engineering can actually build from.
They let a good Wizard of Oz result skip the technical spike entirely, instead of confirming the pattern holds at real scale.

How to use it live. The moment an interviewer asks this, buy two seconds by asking yourself: is the actual doubt here about speed and scale, or about a judgment call a person makes without ever writing down why? The second one is always the Wizard of Oz answer.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Substitution flip: once editorial review time got treated as a free, unlimited resource, teams routed every new idea straight to engineering instead of the cheap human test, rationing careful review toward whatever felt urgent.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Saoirse Mullen, who runs editorial at Mossgrove Audio, deciding between a Wizard of Oz test and a technical spike for an AI chapter-title generator.
3 · THE HABIT
What did Saoirse almost stop doing, before the colleague's remark caught her?
Tap to flip
ANSWER
Testing the editorial judgment by hand first. She had scoped a technical spike straight away, the same path that had just quietly failed at a sibling imprint.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Testing a judgment call cheaply by hand versus building a full technical pipeline around a pattern nobody confirmed. No middle path once engineering time is committed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pricing editorial review time as a free, shared internal resource, which meant nobody's calendar ever showed the true cost of skipping a cheap human test before routing an idea to engineering.
6 · THE NUMBER
Fill in the blank: the Wizard of Oz test took about ___ editor hours across one week, while the full technical spike would have taken about ___ engineer hours across five weeks.
Tap to flip
ANSWER
6 editor hours versus 160 engineer hours.
7 · THE REPLAY
Same title-generation idea, Wizard of Oz test run first. What changes?
Tap to flip
ANSWER
The editor's hand-faked pattern, concrete nouns over named emotions, holds across twenty real books. The eventual technical spike, scoped around that one instruction, takes nine days instead of five weeks.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Cindermere Language Partners' interpreter-matching tool. The flip there is abandonment: coordinators who get one bad automated match quietly stop trusting the tool for hard cases, with nobody noticing.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old decision does this answer say Mossgrove would take back, and why did it make sense at the time?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Treating editorial review time as a free, unlimited resource. It made sense when review load was light, but it meant nobody felt the true cost of skipping a cheap human test before routing ideas straight to engineering.
Multiple choice
2. What actually went wrong with the sibling imprint's title generator?
  • A. The model was too slow to generate titles in real time.
  • B. It produced fluent titles that scored well internally, but gave away each chapter's emotional turn, a judgment call nobody had tested by hand first.
  • C. The engineering team ran out of budget before finishing the pipeline.
  • D. Listeners complained the titles were too short.
Show hint
Look at what real listeners reacted to, versus what the internal rubric measured.
Show answer
B. The model was technically capable; the untested part was the specific editorial judgment about how much to reveal.
True or false
3. True or false: a Wizard of Oz test is a good fit for deciding whether a model can generate live captions fast enough during a broadcast.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. A human faking real-time, high-speed behavior doesn't tell you anything true about whether a model can actually do it live; that's a technical spike question, not a judgment-quality one.
Fill in the blank
4. Fill in the blank: the pattern the editor found across twenty real books was that titles naming a ___ worked better than titles naming a ___.
Show hint
Look at the story's example, "The Locked Drawer" versus "A Moment of Doubt."
Show answer
Place or object; feeling. A concrete noun teased the reader without revealing the emotional turn the way a named feeling did.
Short answer, apply it yourself
5. Think of a feature idea you've had, for work or otherwise. Could you fake it by hand for one real user before building anything? What would you learn from watching that?
Show hint
Think about whether the doubt is about a judgment call or about raw speed and scale.
Show answer
Model answer: A strong version identifies a specific judgment call, like Mossgrove's "how much to reveal," and describes faking it for one real case to see if a consistent pattern shows up.
Short answer, work the number
6. If the Wizard of Oz test had taken 40 editor hours instead of 6, would it still have been the right call over a 160-hour technical spike?
Show hint
Think about what the real comparison is measuring, not just raw hours.
Show answer
Model answer: Yes. Even at 40 hours, it's a quarter of the technical spike's cost, and it still answers the judgment question before committing engineering time to the wrong pattern.
Before you close the answer
Why this works
Tests whether you can tell the difference between "can a model do this" and "is this worth automating," and pick the cheaper test for whichever one is actually in doubt.
Follow-up traps
"Isn't this just A/B testing with extra steps?" Response: no, because nothing is being compared to a control here, a person is standing in for the automation entirely, to learn whether the underlying judgment is even learnable before a model ever touches it.

"What if the human faking it can't actually explain their own reasoning?" Response: that's itself a real finding, it means the judgment call may not be a stable, learnable pattern at all, and a technical spike built on it would just be automating inconsistency.
If pressed
The final technical spike specifically tested the concrete-noun pattern against a held-out set of five books the editor hadn't seen during the Wizard of Oz phase, to confirm the pattern generalized rather than just fitting the twenty books it was found on.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more