CaseIntermediateShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #5

How do you keep a prototype from setting unrealistic expectations?

The direct answer
Before any prototype leaves the room, run it once, live, on an input someone in the room picks that minute, and say out loud any step where a person quietly checks the output before it ships. A demo that never fails and never waits is not proof a tool is ready. It is proof nobody has shown it anything hard yet.
Do this, in order
  1. Run one live, unscripted input during every prototype demo, picked by someone in the room right then.Why: it's the only thing that actually resets the promise the audience walks away holding, instead of the promise the curated examples happened to make.
  2. Say out loud any step where a person checks or edits the output before it reaches anyone.Why: hidden latency from a person in the loop reads as a broken product the moment real turnaround shows up, if nobody was told that step existed.
  3. Recut early usage numbers by input type before reporting one blended score.Why: a number that still looks mostly fine can be one input type running almost unusable underneath it.
  4. Show the room one deliberate failure case before they ever see the tool live.Why: without ever seeing "wrong," people have no way to recognize it when it actually happens.
  5. Trace a stakeholder's complaint back to the literal words used at demo time before deciding they're asking for too much.Why: the complaint is almost always an echo of a promise someone actually made out loud.

How to answer this, stage by stage

Eight moves. The trap in this question is answering it with a general rule about honesty, so most of these stages exist to prove that a demo is a promise, not a preview, and to find exactly where that promise breaks.

1
Anchor it in one product and one number
Say it like this
"Let me put a number on this. Say a university, Hollowmere, has a writing center. A product manager named Winona Castellijo builds a tool called Marginwise: a student pastes in a draft, it hands back inline comments, like a tutor with a red pen, in seconds. At the demo, she ran it on five essays. Tutors kept forty-seven of its forty-nine comments."
Why this works
A real product, a real person, and a real number turn "how do you keep a demo honest" into something you can actually trace back to a decision.
2
Say your structure out loud
Say it like this
"I want to run this as a diagnosis. Timeline, recut, assume nothing, cause candidates, evidence test. TRACE. The question sounds like it wants a rule of thumb. What it's really asking is why a demo that goes great can quietly promise something the real product never agreed to."
Why this works
Naming the plan up front tells the interviewer this isn't going to be five minutes of guessing out loud.
3
Say what the question is actually testing
Say it like this
"Here's the real test in this question. A demo isn't a preview of the product. It's a promise, made out loud, whether anyone means it that way or not. The gap between what got shown and what ships isn't a bug. It's the actual expectation you set, minus the thing you actually built."
Why this works
This sentence is the whole answer compressed. Skip it and everything after sounds like a checklist instead of a reason.
4
Give the direct decision, straight
Say it like this
"So here's what I'd actually do. Before any prototype leaves the room, run it once, live, on an input somebody in the room picks right then. And say out loud, plainly, anywhere a person quietly checks or edits the output before it goes out."
Why this works
Naming the concrete action before the story proves it means nobody has to wait for the ending to know the answer.
5
Walk the timeline, then recut it
Say it like this
"Winona demoed Marginwise on February 24th. The pilot launched March 10th, to Comp 101 and two lab-report courses. For five weeks the number everyone watched just crept down, 88 to 84 to 79 to 74 to 68 percent useful. Looked like a normal, slow slide. So I recut it by course. Comp 101 held at 86 percent the whole time. The lab-report courses were at 22."
Why this works
A blended number hiding a segment that's basically broken is the single strongest move in a diagnosis, and it's the one most answers skip.
6
Rule out the audience before blaming the product
Say it like this
"Before I say the writing center staff were just expecting too much, I check. I pull the actual demo recording. Fourteen minutes in, Winona says, word for word, 'you can throw literally anything at this and it'll come back in seconds with something useful.' The pitch deck sent to the STEM faculty used the exact same five screenshots. So no, they weren't being unreasonable. The demo made that exact promise."
Why this works
Ruling out "the audience is just difficult" before blaming the design is what separates a real diagnosis from an excuse.
7
Name the causes, then run the one test that proves it
Say it like this
"Three reasons a demo like this over-promises. One, only clean essays got demoed, never a lab report. Two, the demo was instant because low-confidence comments get quietly queued for a person to check, and none of the five demo essays ever triggered that. Three, nobody in the room ever once saw Marginwise get something wrong. So here's the test. Take one real Chem 101 lab report, run it live, right now. It queues. Eleven hours later it comes back with 'add a clearer thesis statement,' the exact complaint. One test, and it's the whole story."
Why this works
Naming three candidates and then narrowing to the one the evidence actually confirms is the hardest, strongest move in the whole method.
8
Close on the one line that matters
Say it like this
"So that's the answer. Marginwise never once broke on stage. It just never had the chance to, because nobody gave it anything it hadn't already been shown. Run one unscripted test before you ever demo again, and say out loud wherever a person is quietly checking the output. That's the whole fix."
Why this works
Ends on the decision, not a recap, which is the line an interviewer actually remembers.

Let's learn

Marginwise is a tool that reads a student's essay draft and hands back inline comments, the way a tutor marks a paper with a red pen, except it does it in seconds.

Knowledge spark: what's a low-confidence queue? Some of Marginwise's comments come back less sure than others. Anything under a set bar gets set aside for a real tutor to check before a student ever sees it. That check takes hours, not seconds.

Before Marginwise, a Hollowmere student wanting feedback booked a twenty-minute slot with a tutor, and the whole writing center could only see about 60 students a week. Most students never got a look at their draft until the night before it was due, if they got an appointment at all.

At the demo, Winona ran Marginwise on five essays in front of twelve tutors and the writing center's director, Otis Vandermeer. Every comment came back in under five seconds. Tutors kept 47 of its 49 comments. Otis signed off on a semester pilot two weeks later: three hundred students, three courses.

The blended usefulness rate, week by week
Share of comments tutors rated worth keeping, all courses combined
68%, and nobody had split it by course
wk1, 88%wk2, 84%wk3, 79%wk4, 74%wk5, TA complaint, 68%

Here is the turn. That slow slide from 88 to 68 percent was never the real problem, and it never was. The real problem showed up the moment somebody split that number apart by course instead of reading it as one number for the whole writing center.

Marginwise never once broke in front of anyone. It just never once met anything it hadn't already been shown.

At its worst, this costs Hollowmere a chemistry TA's fifteen students, all handed the same wrong comment on lab reports overnight instead of the instant feedback they were promised, and it nearly gets the whole pilot shut down two weeks before finals, taking Marginwise away from the 86 percent of students it was actually working for.

A hand sketch horizontal timeline with four marks: Winona demoing Marginwise on five clean essays, the pilot launching to 200 students across three courses, usefulness quietly slipping week by week with nobody watching, and a TA complaint marked in red-orange where lab reports were told to add a thesis statement, with a bracket under the gap between the first and last marks labeled nobody looked here.
The gap between when the promise got made and when anyone actually checked it

The choice I would take back. Winona's team built the entire demo from Comp 101 essays, because that was the class the pilot started in, and never wrote down that this was a curated sample rather than a fair test. I would take that back. I would put one line in the launch plan: this demo used five essays we picked ourselves; a randomly chosen input, live, is the actual bar, before any stakeholder signs off on a rollout.

The decision that mattered Run one unscripted input, live, before any prototype leaves the room, and say out loud any step where a person is quietly checking the output. Both of those, together, before rollout, not after a TA has to escalate.

What I would leave alone. Comp 101, the class Marginwise was actually tested on, doesn't need any of this. Its usefulness rate held at 86 percent the whole five weeks. Slowing that down to rebuild trust for a problem that isn't happening there would just cost the students it's already working for.

The lesson. A demo that never once fails and never once waits isn't proof a tool is ready. It's proof nobody has shown it anything hard yet, and the room has no way to know the difference.

The six weeks between the demo and the complaint

Read the short version above if you're short on time. This is the long version, for the part where you feel exactly what nearly went wrong.

Winona Castellijo spent three days picking those five essays. Not cherry-picking, exactly, she'd tell you. Just choosing the ones that looked like what most Comp 101 drafts looked like: a clear thesis in the second paragraph, three body points, a conclusion that circled back to the opening line.

The demo ran on a Monday in late February, in a conference room with a bad projector and twelve tutors on folding chairs. Winona fed in the first essay. Four seconds later, Marginwise had six comments, each one specific: "This sentence buries your strongest point, move it up." "You never define 'accessible' before using it three times." Tutors leaned forward. By the third essay, they were finishing Marginwise's comments before it did.

Otis Vandermeer, the writing center's director, signed the pilot paperwork before the meeting even ended. Three hundred students, three courses: Comp 101, and two lab-report classes, Bio 110 and Chem 101, because the writing center had always served the whole campus and nobody thought to ask whether a lab report looks anything like an essay.

For the first two weeks, the numbers looked fine. Usefulness, the share of comments tutors sampled and kept, sat at 88 percent, then 84. Nobody was watching closely. Midterms were coming, and the writing center staff had their own students to see.

By week four it was 74. By week five, 68. Still, on paper, "mostly fine." Nobody split it by course, because nobody had a reason to look.

We didn't build a worse tool for the lab reports. We built the same tool and quietly assumed the whole campus wrote like Comp 101.

Then, in the sixth week, a Chem 101 TA emailed Otis directly. Fifteen students had submitted their lab reports to Marginwise the night before they were due. Every single one came back with some version of the same comment: "Add a clearer thesis statement in your opening paragraph." Lab reports don't have a thesis statement. They have a hypothesis. And the comments hadn't come back in seconds, the way Winona had shown the room in February. They'd come back the next morning.

Otis forwarded the email to Winona with one line: "This isn't what you showed us."

Her first instinct was to wonder if he was being unfair. Marginwise hadn't changed. So she pulled the recording of the February demo. Fourteen minutes in, she heard her own voice: "You can throw literally anything at this and it'll come back in seconds with something useful." She'd said it. She'd meant it about Comp 101 essays. Nobody in that room had any way to know that.

She checked the model next, because that's the easy thing to blame. Same frozen version, running everywhere, nothing had changed. So she looked at what the lab reports actually were, next to what she'd demoed.

She took one real Chem 101 report, still in the queue, and ran it live in front of her own team. It scored below Marginwise's confidence line and queued for a human tutor to check, the same queue that had never once fired on her five demo essays because none of them had ever been that unsure. Eleven hours later, it came back: "Add a clearer thesis statement."

So here is the decision I would take back. When Winona's team picked the demo essays, they picked the ones that looked like the class the pilot would start in. That made sense in February. What nobody did was write down that this was a curated sample, and run even one input nobody had chosen, live, before three hundred students and a chemistry TA's trust were riding on it.

And the part I'd want to tell myself, if I could go back: we tested Marginwise against the shape of writing we already had. We never once tested it against the shape of writing someone else would actually turn in.

What splitting it by course actually showed

Before trusting the course-level gap, Winona's team checked whether Marginwise's grading was even right. Two tutors hand-checked 20 of the lab-report comments against what an actual instructor would have flagged. They agreed with Marginwise's own confidence flag on 18 of 20. The grading wasn't the problem. That left the input.

Same five weeks, cut by course instead of blended
86%
22%
Comp 101
the shape Marginwise was demoed on
Bio 110 and Chem 101 lab reports
never once shown at the demo
Courses using the tested input shape
The courses that never used it
The blended number read 68 percent because lab reports were still under a third of all submissions in week five, midterm lab crunch aside. Their own number never showed up until someone cut it out on its own.
Lab report, as submitted
1 real Chem 101 report, week six, live test
11 hrsqueued, then a generic thesis comment
Same content, rewritten as an essay
Same report, restructured with an explicit thesis line, same model
4 secspecific, on-topic comment, no queue

Three reasons a demo over-promises, and the one that was true

Not because anyone was careless. Each of these, on its own, looks like a normal decision at demo time. Together, they're why a prototype that never once stumbled can still make a promise the real product can't keep.

Three hand-sketched labelled boxes: a stack of papers, circled in red-orange, for the assumption that only clean essays were ever demoed, confirmed as the true cause; a clock icon for the hidden human-review queue that made the demo instant but real answers take hours; and a question-mark icon for the assumption that Marginwise never once had to show it getting a comment wrong.
Three separate, checkable causes, only one of them confirmed by the live test
Cause 1
Only clean Comp 101 essays, ever demoed.

Marginwise was tuned and shown against five-paragraph drafts with an explicit thesis sentence in the second paragraph. Chem 101 and Bio 110 lab reports don't have a thesis statement at all. They have a hypothesis and a results section, a genuinely different shape of writing.

How you'd check it: take a sample of real lab reports, rewrite them into essay form with an explicit thesis line, and rerun the same frozen model. If the comments turn specific and on-topic, this is the assumption that broke.
Cause 2
Instant on stage, hours for real.

Any comment scoring under Marginwise's confidence line gets queued for a human tutor to check before a student ever sees it, usually 10 to 14 hours overnight. None of the five demo essays ever dropped below that line, so the queue never fired, and nobody in the room was ever told it existed.

How you'd check it: check what share of each course's submissions actually route to the human queue. If lab reports queue far more often than Comp 101, the hidden wait was always going to land there first.
Cause 3
Never once shown getting it wrong.

Across all five demo essays, every comment was specific enough that nobody in the room ever saw Marginwise hedge, miss, or give a generic answer. Tutors had no picture of what "wrong" looks like, so they had no way to recognize it when it actually arrived.

How you'd check it: ask whether the demo ever included one clearly imperfect comment. If the honest answer is no, nobody had any way to spot the failure the first time it showed up for real.

TRACE, run against five weeks of Marginwise

This reads like a question that wants a rule of thumb about honest demos, but the real job is diagnosis: work out why a prototype that never once stumbled can still set an expectation the real product can't keep, and prove exactly where that expectation broke.

T, timeline. Winona demoed Marginwise on February 24th, five essays, all Comp 101. The pilot launched March 10th, to Comp 101 and two lab-report courses. Nobody flagged the difference in writing shape at the time; it looked like the writing center serving its usual campus, not a product decision.
R, recut. The same five weeks, split by course instead of blended. Comp 101: steady at 86 percent the whole time. Bio 110 and Chem 101 lab reports: 22 percent. The blended number of 68 percent hid a segment running roughly four times worse than the rest.
A, assume nothing. Before blaming the writing center for expecting too much, rule that out. Pull the actual demo recording: Winona's own line, fourteen minutes in, promises the tool works on "literally anything." The pitch deck sent to STEM faculty reused the same five screenshots. The demo, not an unreasonable audience, set the expectation.
C, cause candidates. Three, named and separate: only clean essays were ever demoed, never a lab report; the human-review queue's hidden latency, invisible because none of the demo essays ever triggered it; and no failure case ever shown, so nobody had a picture of what "wrong" looks like.
E, evidence test. Take one real Chem 101 lab report, run it live with the same frozen model. It queues, and comes back eleven hours later with the same generic thesis comment the TA complained about. Rewrite the same content as an essay, with an explicit thesis line, and rerun: four seconds, a specific comment, no queue. Same model, only the shape of the writing changed, which is what proves it's the input-shape cause, not a model that's forgotten how to grade.
Why the live test is the hard step Anyone can suspect the lab reports were the problem. The live test turns that suspicion into two results off the same model, the raw report and the rewritten one, and shows exactly how much of the gap the writing's shape explains, instead of a hunch dressed up as a finding.

Same blind spot, a rabbit nobody demoed

Brindlewood Veterinary Group, a chain of five clinics, is piloting a triage prototype called Pawcheck: a vet types in symptoms and it suggests how urgent the case is. Basil Renfro leads product for it. It launched with a live demo on six clean dog and cat cases, then rolled out clinic-wide five weeks later.

T. Basil demoed Pawcheck on six dog and cat symptom cases in month one; vets agreed with its urgency call 34 of 36 times. Clinic-wide rollout followed in month two. The blended agreement rate, Pawcheck's urgency call matching what the vet actually decided, crept from 92 to 78 percent over the next six weeks, and nobody split it apart until a clinic manager escalated.
R. Recut by species. Dog and cat cases: 91 percent, steady the whole time. Exotic species, birds, reptiles, rabbits: 30 percent.
A. Same frozen model scored both groups. A manual review of 15 flagged exotic-species cases agreed with Pawcheck's own confidence flag on 13. The grading held up. Basil pulled the demo recording: he'd told the room "bring it literally any patient." The clinics weren't asking for more than that.
C. Three candidates, the same shape as before: only dog and cat cases were ever demoed, never an exotic species; the on-call vet review queue for low-confidence calls, invisible in the demo because none of the six cases ever triggered it; and no failure case ever shown, so staff had no picture of Pawcheck being unsure.
E. One real rabbit case, reduced appetite, queued for the on-call vet and came back six hours later, flagged routine instead of urgent, the exact kind of miss the clinic manager reported. Re-run against a case built from a dog with the same symptom pattern: correctly flagged urgent in under a second. Same model, only the species changed, which pointed straight at the input Pawcheck had never once been shown, not at a model that can't reason about rabbits.

Swap the trigger and it still runs

  • Speed: Winona could have rushed Marginwise campus-wide in one week to hit a semester deadline instead of rolling out over four months. TRACE still starts by asking what shipped at demo time and when the real friction reached someone, not by how fast the rollout happened.
  • Cost: leadership could have skipped the live unscripted check to save a day before the pitch. The check still has to happen eventually, just after a TA's complaint instead of before one.
  • The model really did get better: say Marginwise's next version genuinely got sharper on every essay type, better comments company-wide, in the very same stretch a new course's writing shape still caught it flat-footed. TRACE still finds the gap, because the recut isolates one course even while the overall trend looks like good news.

Where people run it wrong

  • Trusting a blended number that's still technically "mostly fine," without ever cutting it apart by course or input type.
  • Treating one bad comment as proof the model needs more training, before checking whether the input even matched what got demoed.
  • Fixing the visible symptom, retraining on more lab-report examples, instead of the actual gap: what got shown at demo time, and what got quietly hidden.

How to use it live

Buy yourself ten seconds by naming the split out loud. "So there's the demo everyone remembers, and there's whatever it never got tested against. Let me say how I'd check whether that gap is already showing up." That's not stalling. That's where the real diagnosis starts.

Flashcards (click a card to flip it)

This is a case question about keeping a prototype honest, worked as a diagnosis, so these eight test the TRACE moves and the real numbers behind them.

1 · THE FRAMEWORK
Which framework fits "how do you keep a prototype from setting unrealistic expectations," and why?
Tap to flip
ANSWER
TRACE. It sounds like a request for a rule of thumb, but the real job is diagnosis: working out why a demo that goes great can quietly promise something the real product never agreed to, then finding exactly where that promise breaks.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Winona Castellijo, product manager for Marginwise, an AI writing-feedback tool. She picked and ran the February 24th demo herself, for Hollowmere University's writing center.
3 · THE HABIT
What did nobody on Winona's team do before the semester pilot launched?
Tap to flip
ANSWER
Run Marginwise once on an input nobody had picked in advance. Every demo essay was chosen and checked by Winona herself, all in the same shape as Comp 101's usual drafts.
4 · THE THREE CAUSES
Name the three reasons a demo like this over-promises.
Tap to flip
ANSWER
Only clean Comp 101 essays were ever demoed, the human-review queue's latency was hidden because it never fired on the demo essays, and Marginwise was never once shown getting a comment wrong.
5 · THE NUMBER
Comp 101's usefulness rate held at 86 percent while the lab-report courses fell to ______ percent.
Tap to flip
ANSWER
22 percent. The blended, campus-wide number only read 68 percent, because lab reports were still under a third of all submissions in week five.
6 · THE CHECK
Name the one test that proved it was the input, not the model.
Tap to flip
ANSWER
Running one real Chem 101 lab report live: it queued and came back 11 hours later with the generic thesis comment. Rewriting the same content as an essay with an explicit thesis line came back in 4 seconds, specific, no queue.
7 · THE FIX
What should have happened before Marginwise ever left the demo room?
Tap to flip
ANSWER
One live, unscripted input chosen by someone in the room, plus saying out loud that a person quietly reviews any low-confidence comment before a student ever sees it.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what's the number?
Tap to flip
ANSWER
Pawcheck, a vet triage prototype for Brindlewood Veterinary Group. Dog and cat agreement held at 91 percent while exotic-species cases, never demoed, fell to 30 percent.

Check yourself Score: 0 / 0

Fill in the blank
1. Marginwise's usefulness rate on the lab-report courses fell to ______ percent by week five, while the blended, all-course number only read ______ percent.
Show hint
Look at the recut chart, the two bars split by course, next to the blended line chart above it.
Show answer
22 and 68. The gap only showed up once someone cut the number by course instead of reading the campus-wide average.
Multiple choice
2. Why couldn't Winona's team have just added a line like "results may vary" to the pitch deck instead of building the live, unscripted check?
  • A. Because a vague disclaimer never shows anyone what "wrong" or "slow" actually looks like, so it never resets the real expectation the demo created.
  • B. Because disclaimers aren't allowed in university procurement contracts.
  • C. Because it would have made Marginwise look worse than competing tools.
  • D. Because tutors never read anything attached to a pitch deck.
Show hint
Ask what a vague disclaimer actually shows the room, versus what a live, unscripted test shows them.
Show answer
A. A disclaimer is words about uncertainty. A live, unscripted test is uncertainty the room actually watches happen, which is the only thing that resets a promise someone believed.
True or false
3. True or false: since the blended usefulness rate only slid from 88 to 68 percent over five weeks, that proves Marginwise's model was getting worse at giving feedback.
  • True
  • False
Show hint
Look at which single thing stayed frozen across the whole five weeks.
Show answer
False. The same frozen model ran the whole time. The recut and the live test both point to the lab reports' writing shape, never demoed, not the model getting worse.
Short answer
4. Name a place in Hollowmere's use of Marginwise where this same fix would NOT matter, and say why.
Show hint
Think about the course whose essays already match what got shown at the demo.
Show answer
Model answer: "Leave Comp 101 alone. Its essays match exactly what got demoed, and its usefulness rate held at 86 percent across all five weeks. Rebuilding anything there spends effort on a gap that isn't happening."
Short answer, apply it yourself
5. Think of a demo or pitch you've seen for a real tool. What's one thing about it that made the promise feel bigger than what actually shipped?
Show hint
Look for a case where the demo never showed the tool waiting, hedging, or getting something wrong.
Show answer
Model answer: "A scheduling assistant demo only ever booked meetings between two calm, empty calendars. It never showed what happens with a packed week and three conflicting invites, which turned out to be exactly when I actually needed it." Any honest answer works if it names a real gap between what the demo showed and what you actually needed the tool to handle.
Fill in the blank
6. If lab reports had made up 60 percent of submissions in week five instead of about a quarter, the blended number would have read about ______ percent instead of 68.
Show hint
Weight 86 percent and 22 percent by 40 percent Comp 101 and 60 percent lab reports.
Show answer
About 48. 0.6 × 22 plus 0.4 × 86 comes out to roughly 48 percent, low enough that nobody could have called it "mostly fine." The blend only stayed reassuring because lab reports were still a minority of the traffic.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more