ConceptIntermediateShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #6

Describe the difference between a demo prototype and a learning prototype.

The direct answer
A demo prototype is built to make the room feel the idea is exciting. A learning prototype is built to answer one real question you don't know the answer to yet, even if it looks rough doing it. If you can't name that one question before you start building, you've built a demo, not a learning tool, no matter how finished it looks.
Do this, in order
  1. Name one real, unanswered question before you build anything.Why: this is the whole switch the answer turns on. Skip it and a "prototype" is just a rehearsal with extra steps.
  2. Pull your test cases from real, messy data, not from a set that always works.Why: a prototype that only ever meets cases you already know it can handle never gets a real chance to fail.
  3. Keep the pre-build dry run and the real test as two separate passes.Why: merging them is the decision that quietly turns every prototype into a demo, because whatever survives rehearsal is, by definition, whatever gets shown.
  4. Let a learning prototype end in a number, even an ugly one, not applause.Why: a number you can act on beats a reaction you can't do anything with.
  5. Let some prototypes stay demos, on purpose.Why: an early pitch built to win budget doesn't need to answer a hard question yet. That isn't the job it's doing.

How to answer this, stage by stage

Six moves. Most of the weight sits in stage five, the four sentences where a real number finally shows up. Every stage has the actual words to say.

1
Ground it in one concrete prototype and person
Say it like this
"Let me make this concrete. Say a fintech startup called Ledgerfox is building a categorizer, a tool that reads one line on an expense report and guesses whether it's travel, meals, software, or client entertainment. Simone Aldridge owns the prototype that's supposed to prove it works."
Why this works
Nobody can judge "demo versus learning" without a real product and a real person deciding what counts as proof.
2
Say your structure in one breath
Say it like this
"Five things, fast. Who's building it. What she stopped letting into the room. The switch with no middle setting. The call I'd take back. And the same demo cycle, replayed around one real question instead of a good reaction."
Why this works
A named route up front tells the interviewer you have a plan, not a story you're inventing live.
3
Reframe what the question is actually testing
Say it like this
"This isn't really asking me to define two kinds of prototype. It's asking what a prototype is even for, because a demo prototype and a learning prototype can look completely identical from the outside, and only one of them tells you anything true."
Why this works
That's the gap between a real answer and a dictionary definition, which is the trap most people fall into on this question.
4
Give the one decision
Say it like this
"Concretely: before I build a prototype, I write down the one thing I genuinely don't know the answer to yet. A demo prototype is built to make people feel the idea works. A learning prototype is built to answer that one question, even if the rest of it looks unfinished."
Why this works
There's a real mechanism in that sentence, a named question, not just "test it more."
5
Prove it with the compressed failure
Say it like this
"Say Simone's team ran three demo cycles of the categorizer, each one built from the same fifteen clean receipts that always worked. Great reactions every time. Then a new engineer asked her, after the third one, what they'd actually learned. Her honest answer was 'people like it.' When they finally tested it on forty real, messy receipts, it only got twenty-six right."
Why this works
Four sentences, and it still lands on the exact question that exposed the gap.
6
Say what you'd leave alone, then close on the one line
Say it like this
"Not every prototype needs to be a learning prototype. The very first pitch that gets you three more months of budget is allowed to just be exciting. But once you're building to find out if something works, name the question first. A demo prototype proves the idea is exciting. A learning prototype answers the one thing you don't know, even if it looks rough doing it."
Why this works
Shows judgment instead of a blanket rule, and ends on the sentence an interviewer can repeat back to their own team.
If you remember one thing A demo prototype is judged by the reaction in the room. A learning prototype is judged by whether it answered the one question you built it to answer.

Let's learn

Here's what happens when a prototype gets great at one thing, being watched, and never gets tested at the other thing it's actually for.

Ledgerfox is a fintech startup. It's building a feature called the categorizer: a tool that reads one line on an expense report, a receipt, a vendor name, an amount, and guesses which spending category it belongs to. Travel. Meals. Software. Client entertainment. Guess right, and an employee never has to touch the dropdown.

Knowledge spark: what's a grey-zone receipt? One where the right category isn't obvious from the receipt alone. An Uber ride home from the office and an Uber ride to a client dinner can look almost the same: same app, similar fare, similar time of night.

For one quarter, the categorizer went through three demo cycles. Every single one ran on the same fifteen receipts, picked because they always worked. Every single one got a great reaction. Nobody in the room ever saw it get anything wrong.

Categorizer accuracy, curated demo receipts versus real, ambiguous ones
0% 25% 50% 75% 100% 100% Curated demo set, 15 receipts 65% Real grey-zone set, 40 receipts
Same tool, same model, run the same week. The gap wasn't the categorizer getting worse. It was the difference between a folder built to look good and a set of receipts built to be honest.

Then a new engineer, three weeks into the job, asked Simone a plain question after the third demo. "So what did we actually learn building this?" Simone opened her mouth to answer and found she had exactly one sentence. People like it.

Here is the turn. The fifteen perfect receipts were never the real problem. The real problem was what Simone's team had quietly stopped doing every single time they got ready for a room full of people: trying to find the receipt that would break it.

Three demos in, and the only thing we'd learned was that people like a good demo.

At its worst, this costs exactly what the categorizer was built to save. Ship it without ever finding the hard cases, and the first time it guesses wrong on a real client dinner, someone in Finance has to notice, unwind it, and recode it by hand anyway, which is slower than if the categorizer had never existed at all.

The decision that mattered Stop letting the pre-demo dry run double as the real test. Before building the next version, name one specific question nobody in the room can answer yet, and build only to answer that.

What I would leave alone. The very first demo Simone ever ran, the one that got Ledgerfox's leadership to approve three more months of budget for the categorizer, that one is fine being built purely to excite. Nobody was going to ship code off the back of it. They were deciding whether to keep paying for it.

The lesson. A prototype that only ever meets the cases you already know it can handle isn't rough, it's decorated. Rough is fine. Untested is not the same thing as rough, and it's easy to mix the two up when the room keeps clapping.

Now here is the same thing as a story

Pull this one out when there's more time and you want the room to feel it, not just note it down.

Simone Aldridge spent four years coding expense reports by hand before Ledgerfox hired her onto the categorizer team. She can tell a personal Uber ride from a client dinner ride most days, just from the time stamp and the fare. She built the first version of the categorizer's demo herself, on a Sunday, from fifteen receipts pulled out of her own old expense history, because she already knew exactly how each one should be coded.

The first demo, in January, was small. Four people in a conference room, fifteen receipts, fifteen correct guesses. Somebody said it felt like magic. Simone went home early that day for the first time in months.

By the second demo, in February, the room was bigger, two board observers this time, and Simone added one flashy touch: a receipt in Japanese yen, converted and categorized correctly on screen. She'd tested that one receipt six separate times that week to be sure it would land. It landed. A board observer asked when they could pilot it with real employees.

By the third demo, in March, they let the room submit a live receipt on the spot. It felt spontaneous. It wasn't, quite. An engineer had quietly scanned every submitted photo an hour early and swapped out two that looked risky for backups pulled from a folder of twenty receipts they already knew worked.

Nobody set out to build a folder of receipts that couldn't fail. It just kept being the fastest way to have a good Thursday.

After the third demo, the new engineer, three weeks into the job, caught Simone in the hallway. "That was great," she said. "What did we actually learn today?"

Simone started to answer and stopped. She had reactions. Nods, a couple of laughs at the yen receipt, one board member asking about pricing. She did not have one single fact about the categorizer that she hadn't already known back in January.

"People like it," she said, and heard how small it sounded the second it left her mouth.

Three demos in, and the only thing we'd learned was that people like a good demo.

Here's the decision I'd take back. Back in December, when the team was setting up how demos would work, somebody suggested running a dry run the night before each one, just to make sure nothing embarrassing happened live. That made sense. Nobody wants a bad moment in front of the board. But the dry run and the actual test of the categorizer became the same meeting. Whatever survived the dry run was, by definition, whatever got shown. There was never a separate pass whose only job was trying to make the thing fail.

Left, a dial with a needle among many fine marks, labelled many settings, how good did it feel in the room. Right, a plain square switch, labelled two settings, answered the real question or didn't, sitting on the same desk.
A prototype is a switch, not a dial

I'd take that back, and split it in two. Keep the dry run, for the room's sake. Add a second pass, a week earlier, whose only goal is finding the ugliest real case in the pile.

Run the fourth demo cycle again with that fixed. This time Simone doesn't start from her own fifteen easy receipts. She starts from one real question: can the categorizer tell a personal Uber ride from a client dinner Uber ride, when the receipt alone doesn't say which one it is? She pulls forty real receipts, already sitting in Ledgerfox's own expense system, that a human had to sit and puzzle over before coding.

Cold, before any fix, the categorizer gets twenty-six of the forty right. Not fifteen of fifteen. Twenty-six of forty.

That's the ugly number nobody wanted in a demo. It's also the first true thing the team has learned about the categorizer since January.

The fix turns out to be small: when the model isn't sure, don't guess silently, ask the employee one tap, personal or client? Run the same forty receipts again. Thirty-seven right, and the three still wrong are flagged, not hidden.

The fourth demo ends at nine at night with a number on the screen instead of applause. Thirty-seven out of forty, three known misses, a plan for each one. Nobody claps. Simone doesn't mind. For the first time since January, she knows something about the categorizer that she didn't already know before she built it.

If I'm honest, picking fifteen easy receipts for that first demo wasn't the mistake. Anybody would have, in January, with four people in a room and nothing built yet. The mistake was never asking, three demos later, whether easy was still the point.

The five letters, applied to a demo that never got tested

The letters matter less than which one breaks first. Here's the same five steps, mapped onto Simone's hallway conversation.

Five stacked rows, F L I P S, each a small icon in a coloured circle, a step name, and a short question. The I row's icon is a question mark on a card.
FLIPS, five rows
FFind the person
Who actually gets asked what the demo proved?
Not "the team" in the abstract. Whoever has to answer, out loud, for what got learned.
In this answer: Simone Aldridge, product manager at Ledgerfox, four years coding expense reports by hand before this, who built the categorizer's first demo herself.
LLocate the habit
What did she stop letting into the room once the demos kept landing?
Look for the pass that quietly disappeared, not her overall effort. A reaction that's always good earns the trust that later breaks something.
In this answer: She stopped letting an untested receipt anywhere near a demo. Three cycles in, nothing shown had ever been allowed to fail in front of anyone.
IIdentify the flip
What two setting switch snaps, with no middle?
"She got less curious about the hard cases" describes a feeling, not an action. Name the exact two states with nothing between them.
In this answer: Builds the next prototype to make the room feel the idea works, or builds it to answer one specific question nobody can answer yet, even if it looks rough doing it. Once the new engineer asked what got learned, there was no version of "a bit of both" left standing.
PPinpoint the old decision
Which call only made sense before anyone needed a real answer?
Look for a narrow, defensible call from the early days. "Add more rehearsal" doesn't count, that's a bigger dial.
In this answer: Back in December, the team decided the pre-demo dry run and the real test of the categorizer would be the same meeting, so nothing that survived rehearsal ever got a separate chance to fail.
SShow the replay
Same fourth demo, one real question. Better ending?
Run the identical trigger through the fixed design and see where it stops. A count, not an adjective.
In this answer: Built around one question, tested on 40 real receipts. Twenty-six right cold, thirty-seven right after a one-tap fix, three flagged instead of hidden. Nine at night, a real number on the screen instead of applause.
Two panels. Left, the demo folder, the same fifteen easy receipts, cycle after cycle, drifting further from the real mix. Right, what Simone does, trusts the reaction in the room, or asks one real question, nothing between.
The folder barely changed. What she did about it snapped in one move.

"She got less curious about the hard cases" is a diagnosis anyone can offer after the fact. The harder part is naming the exact thing that had to stop first, asking a question nobody had answered yet, and showing there was no smaller version of it left once three demos in a row had already gone well.

And if you want to be sure it really works, try it somewhere else

Suncrest is a residential solar startup, nowhere near fintech. Its prototype is a roof screener: a tool that looks at one photo of a roof and predicts whether it's a good candidate for panels. Same question, a different flip this time. Nobody stops checking because the news is good. The eval photos themselves quietly get groomed.

F. Anouk Delvaux, ops lead at Suncrest, six years estimating roofing jobs before this, who can guess a roof's pitch from the sidewalk.
L. Every quarterly review reused the exact same folder of thirty roof photos, all from Suncrest's own flat, sunny, tree-free service area, because collecting new ones ate a whole afternoon nobody had.
I. A different flip from Simone's. Nobody rides good news into skipping a check. The eval set itself quietly narrows. Tests the roof screener on the same friendly thirty photos every time, or tests it on the specific roofs the team is quietest about, the shaded ones, the steep ones, the ones with an awkward chimney, with nothing in between once a sales rep needs a real answer standing in front of a customer.
P. When the eval folder was first built, an engineer picked thirty clean, representative photos to keep quarterly reviews quick. The folder just never got refreshed after that, because building a new one wasn't anyone's job.
S. Pull 25 real, hard roof photos that field reps had flagged over the past year. Cold, the screener is only right on 9 of those 25. Add a required manual check for anything shaded or steep instead of a silent guess. Re-run: 22 of 25 called correctly. The next investor review runs on the hard 25, not the easy 30, and ends with a real, slightly imperfect number instead of a clean sweep.

Suncrest's roof screener, correct calls on the hard 25, before and after
Before, tested only on the friendly 30-photo folder
Before
9 / 25 correct
After, hard 25 added, shaded and steep roofs get a manual check
After
22 / 25 correct
Before: right on 9 of the 25 hardest roofs, the rest were confident guesses on shade and pitch it had never actually been tested against. After: 22 of 25, once shaded and steep roofs got a required manual check instead of a silent guess.
A second decision worth taking back Reusing the same easy eval folder every quarter is itself a decision, not a fact about how photo testing works. A folder that gets refreshed with the roofs everyone's nervous about is what makes a review worth running.

Swap the trigger and it still runs

  • Speed: if Ledgerfox only ran one demo a year instead of one a month, the untested folder would sit even longer before anyone had a reason to open it.
  • Cost: if pulling forty real receipts took a data engineer's whole week instead of an afternoon, Simone would ration which questions she ever bothered to ask, saving the real test for the biggest launches only.
  • The model got better: this is close to what happened at both companies. Neither the categorizer nor the roof screener got worse. Both got shown more, and "it's worked every time so far" is exactly what makes an untested folder invisible.

Where people run it wrong

  • Blaming the presenter for a bad demo, when the folder of easy cases was the actual decision.
  • Building a fancier demo instead of asking whether anyone has tried to break this yet.
  • Waiting for a customer complaint to reveal the gap, instead of naming the one hard question before the next round of building even starts.

How to use it live

Buy yourself a few seconds by naming the reframe before the fix: "The real question isn't whether we have a working prototype, we clearly do. It's whether it's ever been asked something we don't already know the answer to." Say that, and the rest of the answer is just naming the question.

Flashcards (click a card to flip it)

Eight fixed slots, pulled straight from the answer above.

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
The applause flip. Builds the next prototype to earn a good reaction in the room, or builds it to answer one open question, with nothing in between once someone actually asks what got learned.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Simone Aldridge, product manager at Ledgerfox, four years coding expense reports by hand before this. She can tell a personal Uber ride from a client one most days, just from the fare and the time.
3 · THE HABIT
What did they stop doing because it worked?
Tap to flip
ANSWER
She stopped letting an untested receipt anywhere near a demo. Three cycles in, nothing shown had ever been allowed to fail in front of anyone.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
A demo prototype, built to make the room feel the idea works. Or a learning prototype, built to answer one specific open question, even if it looks rough doing it. No setting between the two.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
In December, the team decided the pre-demo dry run and the real test of the categorizer would be the same meeting, so nothing that survived rehearsal ever got a separate chance to fail.
6 · THE NUMBER
On the curated demo receipts, the categorizer got 15 of 15 right. On 40 real, ambiguous receipts, cold, it got only ___ right.
Tap to flip
ANSWER
26 of 40, 65 percent. That gap is the whole reason "people liked it" was never a real answer to what the team had learned.
7 · THE REPLAY
Same fourth demo, one real question, what changes?
Tap to flip
ANSWER
Tested on 40 real receipts, not 15 easy ones. 26 right cold, 37 right after a one-tap fix for uncertain cases, the last 3 flagged instead of hidden. Nine at night, a real number instead of applause.
8 · CROSS-PRODUCT
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Suncrest's roof screener, a solar site prototype. It uses a curation flip: the team keeps reusing the same easy eval photos instead of the news-driven applause flip from Simone's story.

Check yourself Score: 0 / 0

True or false
1. True or false: every demo Ledgerfox ever ran for the categorizer needed to be a learning prototype.
  • True
  • False
Show hint
Think about the very first demo Simone ran, the one that only had to win three more months of budget.
Show answer
False. The first demo was built to win funding, not to answer a hard question. Nobody was going to ship code off it. That one is fine staying a demo. The mistake was letting every later demo, the ones meant to prove the tool actually worked, stay a demo too.
Multiple choice
2. What was the flip in Simone's story, and what were its two settings?
  • A. She becomes less confident in the categorizer after the hallway conversation.
  • B. She builds the next prototype to make the room feel the idea works, or builds it to answer one specific open question, with nothing in between once someone actually asks what got learned.
  • C. The categorizer's underlying model gets worse at guessing categories over the three demo cycles.
  • D. She asks an engineer to sit in on every future demo and take notes.
Show hint
Look for something Simone does with her own hands and her own choice of test cases, not a feeling or a change in the model.
Show answer
B. C never happens, the model never gets worse, it just never gets tested on the hard cases. A describes a feeling, and the story never shows Simone losing confidence, only realizing she had nothing real to report. D is a reasonable fix, but it isn't the switch that actually flips in the story.
Fill in the blank
3. The decision this answer takes back is that in December, the team let the pre-demo ______ and the real ______ of the categorizer become the same ______, so nothing that survived it ever got a separate chance to ______.
Show hint
Think about why merging those two things made sense in December, before anyone needed a real answer.
Show answer
Dry run, test, meeting, fail. Nobody was wrong to want a smooth demo in front of the board. The mistake was never giving the categorizer a separate pass whose only job was trying to break it.
Multiple choice
4. Why couldn't Simone have just "kept demoing and see how it goes" instead of naming one specific question for the fourth cycle?
  • A. Because Ledgerfox's leadership had already stopped approving budget for more demos.
  • B. Because more demos built from the same easy folder would keep producing the same good reaction and teach the team nothing new, the way it had already happened three times in a row.
  • C. Because the categorizer's model needed retraining before another demo could run at all.
  • D. Because the new engineer refused to attend a fourth demo.
Show hint
This is the flip versus dial mistake. "Keep doing the same thing, but more" is a bigger dial, not a different setting.
Show answer
B. A, C, and D all invent constraints the story never states. The real reason is simpler: repeating the same easy test, no matter how many times, was never going to produce a new fact, because the folder itself was the problem.
Short answer, apply it yourself
5. Think of a product you use yourself that started as somebody's prototype: a feature, an app, a tool. What's one question about it that a demo would never have answered, but a real test on messy, everyday use would?
Show hint
Picture the polished first version you'd have shown to get it approved, then ask what it would have hit the first ordinary week it was actually used.
Show answer
Model answer: "A grocery app's photo-based receipt scanner would demo perfectly on a flat, well-lit store receipt. The real question a demo never answers is whether it can read a crumpled receipt pulled out of a coat pocket a week later, with the ink half faded. That's the one you'd only find by testing real receipts, not the three clean ones picked for the pitch." Any honest example counts, as long as it names a specific real-world condition the polished demo case would never have shown.
Fill in the blank, do the math
6. Of the 40 grey-zone receipts, the categorizer got 26 right cold and 37 right after the one-tap fix. How many additional receipts did that one small fix correctly catch?
Show hint
Subtract the cold score from the fixed score.
Show answer
11 more receipts. A single small design change, asking one tap instead of guessing silently, closed most of the gap between the demo's fifteen easy cases and the real forty. The three still wrong got flagged instead of hidden, which is its own kind of progress.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more