Critique a prototype that only shows the happy path.
- Run one live, unscripted input before any go/no-go review, picked by someone who wasn't building the demo.Why: it's the only thing that resets the promise the room walks away holding, instead of the promise the curated sample happened to make.
- Recut early accuracy by input type before reporting one blended number to leadership.Why: a number that still looks "mostly fine" can be one input type running almost useless underneath it.
- Name out loud which input shapes never got tested, before anyone asks for budget.Why: a gap you named on purpose is a known risk; a gap nobody named is a surprise waiting for a controller to find.
- Show the deciding stakeholders one deliberate miss before they ever see the tool run clean.Why: without ever seeing "wrong," a room has no way to recognize it once it happens for real.
- Trace the first real failure back to which input shape it actually was, before assuming the model needs retraining.Why: retraining fixes nothing if the real gap was which inputs got tested, not how the model scores on them.
How to answer this, stage by stage
Seven moves. The trap in this question is answering it with a general rule about honest demos, so most of these stages exist to prove the demo made a promise, out loud, and to find exactly where that promise broke.
Let's learn
Codeframe is a tool that reads a vendor invoice, line by line, and decides which GL code, which account bucket, each charge belongs to.
Before Codeframe, an AP clerk at Thornbury coded about 220 invoice lines a day by hand, checking each one against the vendor's catalog and the right plant's cost center. Getting through a week of invoices took the whole team, plus two extra evenings before month-end close.
At the demo, Idun ran Codeframe on fifteen invoices from three vendors, in front of the CFO, Casper Rundgren, and the finance controller, Bastian Vintner. Every invoice came back coded in under a second. Forty-six of its forty-eight lines were right. Casper signed off on the full build that same afternoon: every vendor, every invoice, live by the end of the month.
Here is the turn. That slide from 88 to 60 percent was never the real problem, and it never was. The real problem showed up the moment somebody split that number apart by vendor instead of reading it as one number for the whole AP queue.
At its worst, this costs Thornbury a plant controller catching a $38,400 freight charge coded to the wrong cost center during month-end close, and it nearly gets the whole build shelved two weeks before the fiscal year rolls over, taking Codeframe away from the ninety percent of invoices it was actually coding right.
The choice I would take back. Idun's team built the entire demo from the three highest-volume vendors, because those were the invoices IT already had a clean feed for, and never wrote down that this was a curated sample rather than a fair test. I would take that back. I would put one line in the funding request: this demo used fifteen invoices we picked ourselves; one invoice nobody chose, run live, is the actual bar, before anyone signs the budget.
What I would leave alone. The three vendors Codeframe was actually tested on don't need any of this. Their straight-through rate held at 90 percent the whole five weeks. Slowing that down to rebuild trust for a problem that isn't happening there would just cost the invoices it's already coding right.
The lesson. A demo that never once gets an invoice wrong isn't proof a tool is ready. It's proof nobody has shown it an invoice it doesn't already recognize, and the room has no way to tell the difference.
The five weeks between the sign-off and the reversal
Read the short version above if you're short on time. This is the long version, for the part where you feel exactly what nearly went wrong.
Idun Falkenberg spent four days picking those fifteen invoices. Not cherry-picking, exactly, she'd tell you. Just choosing the ones that looked like what most of Thornbury's invoices looked like: one page, one vendor logo she recognized, one cost center, a freight line and a materials line and nothing stranger than that.
The demo ran on a Thursday afternoon in April, in a glass-walled conference room with the good coffee machine. Idun fed in the first invoice. Under a second later, Codeframe had it coded: freight to 6100, packaging to 5400, each line labeled with the plant it belonged to. Casper leaned back in his chair. By the fifth invoice, Bastian Vintner, the controller, had stopped watching the screen and started watching Casper.
Casper signed the funding request before the meeting even ended. Full build, all forty vendors, live invoices flowing straight into the ledger by the end of the month. Nobody in the room asked what the other thirty-seven vendors' invoices actually looked like, because the writing center, in this case the AP team, had always coded every vendor the same way, and nobody thought to ask whether a freight bill from a trucking company looks anything like an invoice from an office-supply vendor.
For the first two weeks, the numbers looked fine. Straight-through, the share of lines Codeframe coded with nobody correcting it, sat at 88 percent, then 82. Nobody was watching closely. Month-end close was coming, and the AP team had their own invoices to chase down.
By week four it was 68. By week five, 60. Still, on paper, "mostly fine." Nobody split it by vendor, because nobody had a reason to look.
Then, in the fifth week, Bastian Vintner found it during month-end close. A freight invoice from Harborlight Freight Lines, a vendor never once in the demo, had come in split across three plants on one page: a shared trailer run, one line item per plant. Codeframe had coded the whole $38,400 charge to Plant 2. It belonged to Plant 3. The mistake hadn't shown up in seconds, the way Idun had shown the room in April. It had come back overnight, queued behind a low-confidence flag nobody in that conference room had ever seen fire.
Bastian forwarded the posting to Idun with one line: "This isn't what you showed us."
Her first instinct was to wonder if he was being unfair. Codeframe hadn't changed. So she checked the model version first, because that's the easy thing to blame. Same frozen build, running everywhere, nothing had changed since April. So she looked at what the Harborlight invoice actually was, next to what she'd demoed.
She pulled the recording of the April demo. Fourteen minutes in, she heard her own voice: "Drop any invoice in the folder into this and it'll code every line by the time you blink." She'd meant it about the three vendors she'd tested. Nobody in that room had any way to know that.
She took one more real invoice from a vendor never demoed, still sitting in the queue, and ran it live in front of her own team. It scored below Codeframe's confidence line and queued for a human coder to check, the same queue that had never once fired on her fifteen demo invoices because none of them had ever been that unsure. Nine hours later, it came back: coded to the wrong plant, same mistake shape as the Harborlight charge.
So here is the decision I would take back. When Idun's team picked the demo invoices, they picked the ones that looked like the vendors IT already had a clean feed for. That made sense in April. What nobody did was write down that this was a curated sample, and run even one invoice nobody had chosen, live, before a year of engineering budget and a plant controller's trust were riding on it.
And the part I'd want to tell myself, if I could go back: we tested Codeframe against the shape of invoices we already had a clean feed for. We never once tested it against the shape of invoices the rest of the company actually received.
What splitting it by vendor actually showed
Before trusting the vendor-level gap, Idun's team checked whether Codeframe's grading was even right. Two AP clerks hand-checked 20 of the miscoded lines against what a controller would have flagged. They agreed with Codeframe's own confidence flag on 18 of 20. The model wasn't the problem. That left the input.
Three reasons a demo over-promises, and the one that was true
Not because anyone was careless. Each of these, on its own, looks like a normal decision at demo time. Together, they're why a prototype that never once stumbled can still make a promise the real product can't keep.
The fifteen demo invoices all came from the three highest-volume vendors, the ones with a clean, standard, single-page template and one cost center per invoice. Freight invoices split across plants, credit memos, and multi-page vendor statements never made it into the sample.
The team had nine working days between kickoff and demo day. Every hour of it went into tuning Codeframe's accuracy against the fifteen-invoice sample. Nobody was ever assigned the separate job of trying to find an invoice that would break it.
The demo audience was the CFO, deciding whether to fund a year of engineering headcount. Idun believed that showing one miscoded line, on purpose, would read as "this isn't ready," not as "here's the exact edge we're going to close before rollout."
TRACE, run against Codeframe's first five weeks
This reads like a question that wants a general rule about honest demos, but the real job is diagnosis: work out why a prototype that never once stumbled could still walk a finance team into funding the wrong build, and prove exactly where that promise broke.
Same blind spot, a bolt of fabric nobody demoed
Thistlebrook Loom Textiles, a garment manufacturer, is piloting a defect classifier called SeamCheck: a QA camera photographs a finished garment and flags stitching or fabric defects before it ships. Ingunn Saether leads QA tooling for it. It launched with a live demo on twelve clean cotton-shirt photos, then rolled out floor-wide four weeks later.
T. Ingunn demoed SeamCheck on twelve cotton-shirt photos in month one; floor supervisors agreed with its call 35 of 36 times. Floor-wide rollout followed three weeks later. The blended agreement rate, SeamCheck's flag matching what a supervisor actually decided, crept from 94 to 71 percent over the next five weeks, and nobody split it apart until a line supervisor escalated.
R. Recut by fabric type. Cotton shirts: 93 percent, steady the whole time. Stretch knits and dark denim, never once demoed: 27 percent.
A. Same frozen model scored both groups. A manual review of 15 flagged stretch-knit photos agreed with SeamCheck's own confidence flag on 13. The grading held up. Ingunn pulled the demo recording: she'd told the room "point it at literally anything coming off the line." The floor wasn't asking for more than that.
C. Three candidates, the same shape as before: only cotton-shirt photos were ever demoed, chosen because they photographed cleanly under the line's lighting; no time was set aside to go hunt a garment that would trip the camera; and the demo audience was the plant director, deciding whether to fund cameras for every line, and a visible miss would have read as "this isn't ready to leave the lab."
E. One real stretch-knit photo, queued for a supervisor and came back four hours later, passed as clean when it had a dropped stitch, the exact kind of miss the line supervisor reported. Re-shot against a cotton shirt with the same dropped-stitch pattern stitched in for the test: correctly flagged in under a second. Same model, only the fabric changed, which pointed straight at the input SeamCheck had never once been shown, not at a model that can't see a dropped stitch.
Swap the trigger and it still runs
- Speed: Casper could have pushed Codeframe to all forty vendors in one week to hit a quarter-close deadline instead of rolling out over two weeks. TRACE still starts by asking what shipped at demo time and when the real friction reached someone, not by how fast the rollout happened.
- Cost: the team could have skipped the live unscripted check to save a day before the pitch to Casper. The check still has to happen eventually, just after a $38,400 posting gets reversed instead of before one.
- The model really did get better: say Codeframe's next version genuinely got sharper on every vendor type, better coding company-wide, in the very same stretch a new freight vendor's invoice shape still caught it flat-footed. TRACE still finds the gap, because the recut isolates one vendor group even while the overall trend looks like good news.
Where people run it wrong
- Trusting a blended number that's still technically "mostly fine," without ever cutting it apart by vendor or invoice type.
- Treating one bad posting as proof the model needs more training, before checking whether the input even matched what got demoed.
- Fixing the visible symptom, retraining on more freight examples, instead of the actual gap: what got shown at demo time, and what got quietly left out.
How to use it live
Buy yourself ten seconds by naming the split out loud. "So there's the demo everyone remembers, and there's whatever it never got tested against. Let me say how I'd check whether that gap is already showing up." That's not stalling. That's where the real critique starts.
Flashcards (click a card to flip it)
This is a critique question about a happy-path prototype, worked as a diagnosis, so these eight test the TRACE moves and the real numbers behind them.
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Prototyping with LLMs and rapid POCs
- #1 What can you learn from a prototype that you cannot learn from a spec?
- #2 Describe how you would build a working prototype of an AI feature in a day.
- #3 What are the risks of a PM prototyping without engineering involvement?
- #4 Explain when a Wizard of Oz prototype beats a real model.
- #5 How do you keep a prototype from setting unrealistic expectations?
- #6 Describe the difference between a demo prototype and a learning prototype.