ConceptFoundationalShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #1

What can you learn from a prototype that you cannot learn from a spec?

The direct answer
Build a rough, working prototype and put it in front of real customers before you commit engineering time to the spec. A spec tells you what the feature should do. Only a working prototype shows you how the model behaves when a real person interrupts it, mumbles half an item, or adds a condition onto the end of a sentence, and whether it still sounds sure of itself after it drops that condition. That confident wrongness, not word accuracy, is the thing worth catching early.
Do this, in order
  1. Build a rough, working prototype and test it with real customers before the spec gets built.Why: a spec read in a meeting room never gets interrupted, mumbled at, or spoken to with a condition tacked on the end.
  2. Watch what the model does when it drops something, not just how often it's right.Why: it sounds just as sure of itself whether it heard the whole sentence or missed half of it, and that's the part a spec can't show you.
  3. Count how many real people said something the spec never listed.Why: one odd sentence is a fluke. Nine out of fourteen is a pattern the whole design has to answer for.
  4. Keep the prototype rough on purpose. Don't build it to survive a busy Saturday yet.Why: proving it can handle real traffic is a different question from proving it understands what people actually say, and mixing the two slows both down.
  5. Feed what the prototype finds back into the spec before the six-week build starts.Why: a fix costs an afternoon before the build starts. It costs weeks after.

How to answer this, stage by stage

Seven moves. Ground it in one real spec and one real prototype week, since that's what makes the answer checkable.

1
Scope it to one feature, one store, one real spec
Say it like this
"So I'm the PM for QuickSay, the voice ordering feature inside the Kettlebridge Grocers app. I'm not going to talk about prototypes as a general idea. I'll walk through the actual spec we had signed off, and the one week prototype that found the hole in it before we spent six weeks building the wrong thing."
Why this works
A scoped example gives the interviewer something to picture and push back on.
2
Say the structure out loud
Say it like this
"Here's how I'll go: what the spec looked like and why it passed review clean, what a prototype shows you that a spec can't, the one decision I'd actually make, what happens if a team leans on prototypes too hard or not hard enough, and what I'd leave a rough prototype to prove later, not on day one."
Why this works
Two seconds of structure tells the interviewer you have a plan, so they follow instead of guessing where you're going.
3
Reframe what a spec actually can't show you
Say it like this
"Most people think a spec is 'not tested yet' and a prototype is 'tested.' That's not really the gap. A spec gets read in silence, by people who already know what it's supposed to do. A prototype gets spoken to, by someone who has no idea what it expects to hear. That's the whole difference, and it's why a spec can pass every review clean and still miss what happens the first time a real person opens their mouth."
Why this works
This is the actual insight being tested. Skip it and you're just saying "prototypes are good," which nobody disagrees with and nobody learns from.
4
Give the anchor: what the prototype is built to catch
Say it like this
"So the anchor is this: before we touched the six-week build, I had one engineer wire a rough version, real speech-to-text, a general model, running live at the pickup counter for one week. I wasn't testing whether it could hear 'bananas.' I was testing what it did when someone said 'bananas, but only if they're not too ripe,' and whether it added the item quietly and confidently without ever mentioning the part it dropped."
Why this works
Naming the specific thing you're testing for, out loud, is what separates a real design decision from a vague "let's prototype it" gesture.
5
Prove it with the near miss
Say it like this
"Here's what almost went into the spec as done. A shopper picking up her order said, 'get me bananas, but only if they're not too ripe, and my usual oat milk, not the vanilla one.' The prototype came back in about two seconds: 'Got it. One bunch of bananas, one oat milk.' It heard both nouns fine. It dropped both conditions. And it said 'got it' the exact same way it said 'got it' to every simple order that afternoon."
Why this works
A specific near miss, with the actual words said, does more work than "it might misunderstand things" ever will.
6
Name the risk, in both directions
Say it like this
"Lean on the prototype too hard, and you start treating one good week at one counter as proof it's ready for every store, every accent, every busy Saturday, which it was never built to test. Lean on it too little, and you go straight from a clean spec review to a six-week build, and the first time anyone hears what a real customer actually says is after it's already shipped."
Why this works
Naming both failure directions shows you understand the prototype's job has a limit, not just a benefit.
7
Say what's out for day one, close on the number
Say it like this
"So: a rough, one-week prototype, tested on real speech, built to catch what a spec can't, not to prove it can handle scale. I wasn't testing concurrent sessions or store noise that week, that's a separate test for later. The number I'd point to: nine of the fourteen customers who tried it used a condition or a correction in their very first sentence, something the spec never once wrote down."
Why this works
Interviewers remember the last line most, and this one hands them something they can check, not just a mood.

Let's learn

What happens the first time a real person talks to a feature that's only ever been read on paper?

Kettlebridge Grocers built QuickSay, a way to fill your pickup cart just by talking to your phone.

A hand-sketch of a shopper at her kitchen counter, typing grocery items into her phone one at a time while stirring a pot, a paper list taped to the cupboard above her with a couple of items crossed off in red pencil
Before QuickSay, one item typed at a time

Before QuickSay, the spec took three weeks to write. Three people signed off on it in a review: the engineering lead, the design lead, and Nkechi, QuickSay's PM. On paper it covered a lot: say an item, it gets added, say a quantity, it adjusts the count, say something unclear, it asks you to repeat. Engineering estimated six weeks to build it.

Knowledge spark: what makes something a prototype, and not a demo? A demo is scripted ahead of time, so nothing goes wrong. A prototype is rough, but a real person can actually use it and say anything to it. A demo protects itself from surprises. A prototype invites them, on purpose, before anyone else has to find them the hard way.

Nkechi asked for one week first. Not to build QuickSay properly, just to wire a rough version, real speech recognition, a general model, running live at the store's pickup counter. Fourteen real customers tried it while they waited for their order. Nine of them, in their very first sentence, said something the spec never listed at all.

Here is the important part. Those nine sentences were not full of new words the model didn't know. Every item in them was something it heard correctly. Bananas. Oat milk. Coffee filters. The problem was what people said around those items: "only if," "not the," "the other one," "unless." Small words. The spec never had a box for them, so nobody had ever tested what the model did with them.

Fourteen customers, first sentence at the QuickSay counter
14 customers tested 9 5 used a condition the spec never wrote a case for matched exactly what the spec expected
used a condition the spec never listed matched what the spec expected
One week, one pickup counter, fourteen real customers. Nearly two out of three said something in their first sentence that a three-week, three-person spec review never once wrote a case for.
The spec was never wrong. It was just never spoken to.

At its worst, this doesn't just cost one odd grocery order. Say QuickSay had shipped straight from the signed-off spec. The first time it happened in production wouldn't be a test with an engineer standing nearby. It would be a shopper picking up a bag of bananas she specifically said she didn't want, with nobody there to explain why the app sounded so sure of itself. Fixing that after launch means pulling the feature back mid-rollout, which costs more than the one week Nkechi asked for ever would have.

The decision that mattered Build a rough, working prototype and test it on real speech with real customers, watching for what the model drops and how confident it sounds while it drops it, before the six-week build starts.

The choice I would take back. When the team set up how QuickSay's specs get approved, they folded two separate checks into one meeting: is the spec well written, and is the feature actually ready to build. That worked fine back when every feature in the app was a button doing one obvious thing. It stopped working the moment the feature meant a model listening to a full sentence, because a full sentence is exactly the thing a meeting room can't test.

What I would leave alone. QuickSay also shipped a "reorder your last basket" button around the same time, one tap, same items as last time, no listening involved. Nobody needed a live prototype for that. There's nothing a real customer could say to it that the spec didn't already cover, because there's nothing to say.

The lesson. A spec describes what a feature should do. It can't describe what a real person will actually say to it, because nobody writes a spec by imagining every way a sentence gets interrupted, half-finished, or qualified halfway through. That gap doesn't show up on paper. It only shows up the first time someone talks to the thing, so somebody has to make them talk to it before the six weeks start, not after.

Now here is the same thing as a story

Read this one when you've got a few minutes. The short version is above. This is for when you want to feel why it mattered.

The pickup counter at Kettlebridge Grocers' Elm Street store gets a rush right around 5:30, when everyone's grabbing dinner on the way home.

Nkechi Adeyanju ran product for ordering features at Kettlebridge Grocers. She'd been the PM three review meetings running where nobody found a real hole in her spec, because she wrote down every case she could think of before anyone else got to it.

QuickSay's spec took three weeks. It covered the item, the quantity, the unclear-audio case, the empty-cart case. In the review, the engineering lead read it twice, asked two questions, and signed it. The design lead signed it the same afternoon. Six weeks of build time got booked on the calendar.

Nkechi asked for one week first anyway. Not because she doubted the spec. Because she'd learned, on an earlier feature, that a page full of cases can still miss the one nobody thought to write down, and the cheapest time to find that is before the calendar fills up, not after.

So instead of six weeks of engineering, one person spent four days wiring a rough version: off-the-shelf speech recognition, a general model, no polish, running live at the Elm Street store's pickup counter for a week.

For the first two days it looked almost boring. Shoppers said an item, the app added it, they moved on. Nkechi stood near the counter with a clipboard, checking each order against what the shopper actually said.

Then, on the third afternoon, a shopper picking up her order leaned toward the phone the associate was holding and said, "Can you get me bananas, but only if they're not too ripe, and my usual oat milk, not the vanilla one."

Close hand-sketch of a phone at the pickup counter with two speech bubbles: a green one showing what the shopper actually said, including the ripeness condition and the not-vanilla correction, and a red one showing what the app added to the cart with both conditions missing
The anchor: watch what it drops, not just what it hears

The prototype answered in about two seconds. "Got it. One bunch of bananas, one oat milk."

It had heard both nouns clean. It dropped both conditions. It said "got it" the exact same way it said "got it" to every simple order that afternoon, with nothing in its voice or its text to say it had left anything out.

The spec was never wrong. It was just never spoken to.

Nkechi checked her clipboard against the rest of that week's fourteen customers. Nine of them, she found, had done the same thing in their very first sentence: a "but only if," a "not the," a "the other one." Small words, tacked onto the end of an order, that the three-week spec never once wrote a case for, because nobody sitting in a review room thinks to write down the ways people actually talk.

So here is the decision Nkechi took back.

Months earlier, when Kettlebridge set up how features got approved, the team folded two separate questions into one meeting: is the spec well written, and is the feature ready to build. That made sense back when most features were one tap doing one obvious thing, a button, a toggle, a reorder. There was nothing a real customer could say to a button that the spec hadn't already covered.

QuickSay wasn't a button. It was a model listening to a full sentence, and a full sentence is exactly the thing a meeting room can't test on paper.

Nkechi split the gate in two. A spec could pass review clean and still not be ready to build, not until a rough version of it had been spoken to by real customers, out loud, for at least a few days. Not months. Not a full production trial. Just enough people, saying real sentences, to find the nine-out-of-fourteen problem before the six-week build started instead of after it shipped.

Two panels: left labeled old, adds it anyway, a card reading one bunch bananas with the ripeness condition dropped and never asked about; right labeled fixed, reads it back, the same card now reading it back and checking ripeness before confirming
The day it's wrong, checked against the fix

Run the same week forward with the split gate in place. The one-week prototype catches the qualifier problem on day three. The team spends two extra days teaching the model to read a dropped condition back before confirming it: "I heard bananas, but I didn't catch the part about ripeness, want me to ask?" The six-week build starts a week later than planned, with the actual problem already fixed, instead of finishing on schedule and discovering it live.

One process trusts a clean spec review as proof a feature is ready. The other checks, on purpose, for the one thing a review room can never produce: someone talking to it who doesn't know what it expects to hear.

And the thing I'd want to tell myself, back when we drew up that first single-meeting gate: we built a check for whether the spec made sense to the people who wrote it. We never built one for whether it made sense to the person it would actually be talking to, and that person was always going to find the hole first.

SPARK, before the build starts

This question asks what a prototype teaches you that a spec can't, not asking you to design the feature itself, but SPARK still fits: it's one concrete decision about the exact moment a spec's confidence and a real sentence come apart. A question asking how you'd measure QuickSay's overall accuracy across every store would reach for LEAD instead.

S, situation. Before a working prototype exists, whoever reviews a voice feature's spec checks it against cases written from memory: item, quantity, unclear audio. Nobody in the room is actually talking to it.
P, payoff. Not "a smoother voice feature." The habit worth building: catch what a real customer says that the spec never wrote down, before six weeks of engineering commit to it, not after it ships.
A, anchor. Build a rough version and put it in front of real customers for a few days. Test specifically for what happens when someone drops a condition into the sentence, and whether the model still sounds sure of itself once it's dropped that condition.
R, risk. Lean on the prototype too hard, and one good week at one counter gets mistaken for proof it's ready for every store and every busy Saturday. Lean on it too little, and the spec sails through review and the build starts before anyone has actually spoken to it.
K, keep out. No attempt yet to prove the prototype can handle real traffic, every accent, or a packed Saturday. That's a production-scale question, for a production-scale test, not a one-week counter test.
Why the anchor survives the risk Check it against the near miss. Does testing real speech still catch the dropped-condition problem? Yes, because the test is built to watch for confident silence around a condition, not just whether the nouns were heard right. Does it avoid mistaking a good week for proof of scale? Yes, because K keeps production reliability explicitly off this test's job.

And if you want to be sure it really works, try it somewhere else

A city's after-hours road report line runs on a completely different desk, but the same gap between what got heard and what got understood shows up in a phoned-in pothole report.

S. Corwin Delahaye dispatches for the City of Windhollow's after-hours line, where residents call in to report potholes and downed streetlights. Today, without a live-tested prototype, whoever reviews the phone-report spec checks it against cases written from a whiteboard: street name, problem type, cross street. Nobody on the review calls it to actually report anything.
P. The habit worth building: catch what a resident actually says on a real call, before six weeks of engineering are booked to build it.
A. Same shape, different sentence. A resident calls in and says, "There's a pothole on Cedar, right past the school, but only on the side going toward downtown." The spec has boxes for street and problem type. It has no box for "only on the side going toward downtown," so the rough prototype logs a work order for the wrong side of the street, and reads it back with total confidence.
R. Test the prototype for one quiet afternoon and it looks ready when the real trouble only shows up during a storm, when fifty calls come in about the same three streets and callers talk fast and over each other. Wait for a whole storm season before ever spec'ing the feature, and nothing ever gets built, because the city can't spec a fix and sit on it for a year to find out if it's needed.
K. No attempt yet to handle a flood of calls during a storm, or a caller speaking a language the model doesn't know well. That's a scale and coverage question for a later test.

A hand-sketch of a dispatcher at a desk wearing a headset, beside a work order card with three rows: street checked green, problem type checked green, and which side of the street marked with a red X, not asked
Same anchor, a different desk, streets and sides instead of groceries and conditions

It took a second crew getting sent to the wrong side of Cedar Avenue, twice in one week, before anyone noticed that "street name" and "which side" were never the same box.

Swap the trigger and it still runs

  • Speed: even if QuickSay answered in half a second instead of two, a fast wrong cart is still wrong. Speed never fixes what the prototype is testing.
  • Cost: if the one-week prototype cost nothing and needed no engineer's time at all, that still wouldn't tell you which nine customers add a condition. It still needs real people talking to it, not zero.
  • The model gets better: if QuickSay's speech recognition became perfect, catching every word without fail, it could still drop the condition around those words with total confidence, because hearing every word and understanding what a condition changes are two different jobs.

Where people run it wrong

  • Testing the prototype with typed-out sample sentences instead of real speech, so nobody ever hits an interruption or a half-finished condition.
  • Treating one clean week at one counter as proof the feature is ready for every store, instead of what it actually is: proof the interaction pattern works.
  • Trying to prove the prototype can handle a packed Saturday and every accent in the first week, so the useful, narrow test never finishes because it's carrying a job it was never built for.

How to use it live

If you're asked this cold, ask what the spec assumes a person will say, then ask what happens the moment they say something else. That question, asked out loud, usually finds the missing case faster than trying to list every possible sentence from a blank page.

Flashcards (click a card to flip it)

1 · THE SITUATION
What's the situation, before this prototype existed?
Tap to flip
ANSWER
A spec for QuickSay covered the item, the quantity, and unclear audio, and passed a clean review by three people who never once spoke to a working version of it.
2 · THE PAYOFF
What's the real habit this prototype is trying to build?
Tap to flip
ANSWER
Catching a real interaction problem before six weeks of engineering get booked to build it, not after it ships.
3 · THE ANCHOR
What did the one-week prototype actually test?
Tap to flip
ANSWER
Whether the model stayed confident-sounding after dropping a spoken condition, not just whether it heard the right words.
4 · THE RISK
What breaks if a team leans on a prototype too hard, or not at all?
Tap to flip
ANSWER
Too hard, and one good week at one counter gets mistaken for proof it's ready for every store. Not at all, and the spec sails through review with nobody ever having spoken to it.
5 · THE PROOF
What almost got missed at the pickup counter?
Tap to flip
ANSWER
A shopper asked for bananas "only if they're not too ripe" and different oat milk "not the vanilla one." The prototype heard both nouns, dropped both conditions, and said "got it" anyway.
6 · THE NUMBER
___ of ___ customers said something in their first sentence that the spec never listed.
Tap to flip
ANSWER
9 of 14. Nearly two out of three customers who tried QuickSay used a condition or a correction the written spec had no case for.
7 · THE REPLAY
Same week, split gate in place. What changes?
Tap to flip
ANSWER
The prototype catches the condition problem on day three. Two extra days teach the model to read a dropped condition back before confirming it. The six-week build starts a week late, with the real problem already fixed.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor test?
Tap to flip
ANSWER
City of Windhollow's after-hours pothole report line. Its anchor tests the same shape of gap: whether "only on the side toward downtown" survives being logged, not just whether the street name gets heard right.

Check yourself Score: 0 / 0

True or false
1. True or false: The QuickSay spec failed its review because two of the three reviewers found missing cases in it.
  • True
  • False
Show hint
Look at what the story says happened in the review meeting itself, not what happened later at the counter.
Show answer
False. It passed clean. All three reviewers signed off. The missing case wasn't something anyone in that room could have found, because none of them were actually speaking to it.
Fill in the blank
2. Before the split gate, ___ of the ___ customers who tried the QuickSay prototype used a condition or correction in their very first sentence that the spec never listed.
Show hint
This is the number the whole argument leans on. It shows up twice, once in the story, once in the chart.
Show answer
9; 14. Nearly two out of three customers who tried QuickSay used a condition the written spec had no case for.
Multiple choice
3. Which design matches the anchor this answer argues for?
  • A. A longer written spec, with more example sentences added to it before the build starts.
  • B. A rough, working prototype tested on real speech for a few days before the build starts, watching for confident but incomplete answers.
  • C. A full production-scale trial across every store, run before writing any spec at all.
  • D. A checklist the engineering lead reads aloud during the spec review meeting.
Show hint
The anchor needs someone actually talking to a working version, not reading about it.
Show answer
B. A adds more paper, not a real listener. C mixes up the day-one job with the keep-out, production scale is a separate later test. D is still spoken in a review room, not to a live model.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Think about why one meeting covering "is the spec good" and "is it ready to build" sounded like enough, back when features were mostly buttons.
Show answer
Model answer: Kettlebridge had folded two checks into one meeting: is the spec well written, and is the feature ready to build. That made sense when most features were one tap doing one obvious thing, with nothing a customer could say that the spec hadn't already covered. It stopped making sense once a feature meant a model listening to a full sentence.
Short answer, apply it yourself
5. Pick something at your own work that gets approved from a written plan alone, a proposal, a script, a form. What's the one thing a real person interacting with it might do that the plan never wrote down?
Show hint
Look for whether your plan assumes people follow a straight line through it, in order, without ever jumping ahead or answering out of turn.
Show answer
Model answer: "I write onboarding email sequences from a plan that assumes people read every email in order. A one-week test with five real signups showed three of them replied to email two with a question before email three ever went out, something the plan never had a case for."
Multiple choice
6. Based on this answer's own numbers, if only 2 of 14 customers had used a condition instead of 9, would the one-week prototype still have been worth building?
  • A. Yes, because even a small pattern found in one week is cheaper than the same fix found after a six-week build ships.
  • B. No, because with only 2 of 14, the written spec would already have covered it.
  • C. No, a prototype only earns its place once more than half of customers hit the same problem.
  • D. Yes, but only if those two customers happened to complain loudly about it.
Show hint
Compare the cost of catching this on day three of a one-week test against catching it after a six-week build has already shipped.
Show answer
A. The size of the pattern changes how urgent the fix is, not whether the prototype was worth running. Catching even a small, real gap before the build starts is still cheaper than catching it in production.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more