How do you keep a prototype from setting unrealistic expectations?
- Run one live, unscripted input during every prototype demo, picked by someone in the room right then.Why: it's the only thing that actually resets the promise the audience walks away holding, instead of the promise the curated examples happened to make.
- Say out loud any step where a person checks or edits the output before it reaches anyone.Why: hidden latency from a person in the loop reads as a broken product the moment real turnaround shows up, if nobody was told that step existed.
- Recut early usage numbers by input type before reporting one blended score.Why: a number that still looks mostly fine can be one input type running almost unusable underneath it.
- Show the room one deliberate failure case before they ever see the tool live.Why: without ever seeing "wrong," people have no way to recognize it when it actually happens.
- Trace a stakeholder's complaint back to the literal words used at demo time before deciding they're asking for too much.Why: the complaint is almost always an echo of a promise someone actually made out loud.
How to answer this, stage by stage
Eight moves. The trap in this question is answering it with a general rule about honesty, so most of these stages exist to prove that a demo is a promise, not a preview, and to find exactly where that promise breaks.
Let's learn
Marginwise is a tool that reads a student's essay draft and hands back inline comments, the way a tutor marks a paper with a red pen, except it does it in seconds.
Before Marginwise, a Hollowmere student wanting feedback booked a twenty-minute slot with a tutor, and the whole writing center could only see about 60 students a week. Most students never got a look at their draft until the night before it was due, if they got an appointment at all.
At the demo, Winona ran Marginwise on five essays in front of twelve tutors and the writing center's director, Otis Vandermeer. Every comment came back in under five seconds. Tutors kept 47 of its 49 comments. Otis signed off on a semester pilot two weeks later: three hundred students, three courses.
Here is the turn. That slow slide from 88 to 68 percent was never the real problem, and it never was. The real problem showed up the moment somebody split that number apart by course instead of reading it as one number for the whole writing center.
At its worst, this costs Hollowmere a chemistry TA's fifteen students, all handed the same wrong comment on lab reports overnight instead of the instant feedback they were promised, and it nearly gets the whole pilot shut down two weeks before finals, taking Marginwise away from the 86 percent of students it was actually working for.
The choice I would take back. Winona's team built the entire demo from Comp 101 essays, because that was the class the pilot started in, and never wrote down that this was a curated sample rather than a fair test. I would take that back. I would put one line in the launch plan: this demo used five essays we picked ourselves; a randomly chosen input, live, is the actual bar, before any stakeholder signs off on a rollout.
What I would leave alone. Comp 101, the class Marginwise was actually tested on, doesn't need any of this. Its usefulness rate held at 86 percent the whole five weeks. Slowing that down to rebuild trust for a problem that isn't happening there would just cost the students it's already working for.
The lesson. A demo that never once fails and never once waits isn't proof a tool is ready. It's proof nobody has shown it anything hard yet, and the room has no way to know the difference.
The six weeks between the demo and the complaint
Read the short version above if you're short on time. This is the long version, for the part where you feel exactly what nearly went wrong.
Winona Castellijo spent three days picking those five essays. Not cherry-picking, exactly, she'd tell you. Just choosing the ones that looked like what most Comp 101 drafts looked like: a clear thesis in the second paragraph, three body points, a conclusion that circled back to the opening line.
The demo ran on a Monday in late February, in a conference room with a bad projector and twelve tutors on folding chairs. Winona fed in the first essay. Four seconds later, Marginwise had six comments, each one specific: "This sentence buries your strongest point, move it up." "You never define 'accessible' before using it three times." Tutors leaned forward. By the third essay, they were finishing Marginwise's comments before it did.
Otis Vandermeer, the writing center's director, signed the pilot paperwork before the meeting even ended. Three hundred students, three courses: Comp 101, and two lab-report classes, Bio 110 and Chem 101, because the writing center had always served the whole campus and nobody thought to ask whether a lab report looks anything like an essay.
For the first two weeks, the numbers looked fine. Usefulness, the share of comments tutors sampled and kept, sat at 88 percent, then 84. Nobody was watching closely. Midterms were coming, and the writing center staff had their own students to see.
By week four it was 74. By week five, 68. Still, on paper, "mostly fine." Nobody split it by course, because nobody had a reason to look.
Then, in the sixth week, a Chem 101 TA emailed Otis directly. Fifteen students had submitted their lab reports to Marginwise the night before they were due. Every single one came back with some version of the same comment: "Add a clearer thesis statement in your opening paragraph." Lab reports don't have a thesis statement. They have a hypothesis. And the comments hadn't come back in seconds, the way Winona had shown the room in February. They'd come back the next morning.
Otis forwarded the email to Winona with one line: "This isn't what you showed us."
Her first instinct was to wonder if he was being unfair. Marginwise hadn't changed. So she pulled the recording of the February demo. Fourteen minutes in, she heard her own voice: "You can throw literally anything at this and it'll come back in seconds with something useful." She'd said it. She'd meant it about Comp 101 essays. Nobody in that room had any way to know that.
She checked the model next, because that's the easy thing to blame. Same frozen version, running everywhere, nothing had changed. So she looked at what the lab reports actually were, next to what she'd demoed.
She took one real Chem 101 report, still in the queue, and ran it live in front of her own team. It scored below Marginwise's confidence line and queued for a human tutor to check, the same queue that had never once fired on her five demo essays because none of them had ever been that unsure. Eleven hours later, it came back: "Add a clearer thesis statement."
So here is the decision I would take back. When Winona's team picked the demo essays, they picked the ones that looked like the class the pilot would start in. That made sense in February. What nobody did was write down that this was a curated sample, and run even one input nobody had chosen, live, before three hundred students and a chemistry TA's trust were riding on it.
And the part I'd want to tell myself, if I could go back: we tested Marginwise against the shape of writing we already had. We never once tested it against the shape of writing someone else would actually turn in.
What splitting it by course actually showed
Before trusting the course-level gap, Winona's team checked whether Marginwise's grading was even right. Two tutors hand-checked 20 of the lab-report comments against what an actual instructor would have flagged. They agreed with Marginwise's own confidence flag on 18 of 20. The grading wasn't the problem. That left the input.
Three reasons a demo over-promises, and the one that was true
Not because anyone was careless. Each of these, on its own, looks like a normal decision at demo time. Together, they're why a prototype that never once stumbled can still make a promise the real product can't keep.
Marginwise was tuned and shown against five-paragraph drafts with an explicit thesis sentence in the second paragraph. Chem 101 and Bio 110 lab reports don't have a thesis statement at all. They have a hypothesis and a results section, a genuinely different shape of writing.
Any comment scoring under Marginwise's confidence line gets queued for a human tutor to check before a student ever sees it, usually 10 to 14 hours overnight. None of the five demo essays ever dropped below that line, so the queue never fired, and nobody in the room was ever told it existed.
Across all five demo essays, every comment was specific enough that nobody in the room ever saw Marginwise hedge, miss, or give a generic answer. Tutors had no picture of what "wrong" looks like, so they had no way to recognize it when it actually arrived.
TRACE, run against five weeks of Marginwise
This reads like a question that wants a rule of thumb about honest demos, but the real job is diagnosis: work out why a prototype that never once stumbled can still set an expectation the real product can't keep, and prove exactly where that expectation broke.
Same blind spot, a rabbit nobody demoed
Brindlewood Veterinary Group, a chain of five clinics, is piloting a triage prototype called Pawcheck: a vet types in symptoms and it suggests how urgent the case is. Basil Renfro leads product for it. It launched with a live demo on six clean dog and cat cases, then rolled out clinic-wide five weeks later.
T. Basil demoed Pawcheck on six dog and cat symptom cases in month one; vets agreed with its urgency call 34 of 36 times. Clinic-wide rollout followed in month two. The blended agreement rate, Pawcheck's urgency call matching what the vet actually decided, crept from 92 to 78 percent over the next six weeks, and nobody split it apart until a clinic manager escalated.
R. Recut by species. Dog and cat cases: 91 percent, steady the whole time. Exotic species, birds, reptiles, rabbits: 30 percent.
A. Same frozen model scored both groups. A manual review of 15 flagged exotic-species cases agreed with Pawcheck's own confidence flag on 13. The grading held up. Basil pulled the demo recording: he'd told the room "bring it literally any patient." The clinics weren't asking for more than that.
C. Three candidates, the same shape as before: only dog and cat cases were ever demoed, never an exotic species; the on-call vet review queue for low-confidence calls, invisible in the demo because none of the six cases ever triggered it; and no failure case ever shown, so staff had no picture of Pawcheck being unsure.
E. One real rabbit case, reduced appetite, queued for the on-call vet and came back six hours later, flagged routine instead of urgent, the exact kind of miss the clinic manager reported. Re-run against a case built from a dog with the same symptom pattern: correctly flagged urgent in under a second. Same model, only the species changed, which pointed straight at the input Pawcheck had never once been shown, not at a model that can't reason about rabbits.
Swap the trigger and it still runs
- Speed: Winona could have rushed Marginwise campus-wide in one week to hit a semester deadline instead of rolling out over four months. TRACE still starts by asking what shipped at demo time and when the real friction reached someone, not by how fast the rollout happened.
- Cost: leadership could have skipped the live unscripted check to save a day before the pitch. The check still has to happen eventually, just after a TA's complaint instead of before one.
- The model really did get better: say Marginwise's next version genuinely got sharper on every essay type, better comments company-wide, in the very same stretch a new course's writing shape still caught it flat-footed. TRACE still finds the gap, because the recut isolates one course even while the overall trend looks like good news.
Where people run it wrong
- Trusting a blended number that's still technically "mostly fine," without ever cutting it apart by course or input type.
- Treating one bad comment as proof the model needs more training, before checking whether the input even matched what got demoed.
- Fixing the visible symptom, retraining on more lab-report examples, instead of the actual gap: what got shown at demo time, and what got quietly hidden.
How to use it live
Buy yourself ten seconds by naming the split out loud. "So there's the demo everyone remembers, and there's whatever it never got tested against. Let me say how I'd check whether that gap is already showing up." That's not stalling. That's where the real diagnosis starts.
Flashcards (click a card to flip it)
This is a case question about keeping a prototype honest, worked as a diagnosis, so these eight test the TRACE moves and the real numbers behind them.
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Prototyping with LLMs and rapid POCs
- #1 What can you learn from a prototype that you cannot learn from a spec?
- #2 Describe how you would build a working prototype of an AI feature in a day.
- #3 What are the risks of a PM prototyping without engineering involvement?
- #4 Explain when a Wizard of Oz prototype beats a real model.
- #6 Describe the difference between a demo prototype and a learning prototype.
- #7 What should you test with a prototype before writing the PRD?