Artifact critiqueIntermediateShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #16

Critique a prototype that only shows the happy path.

The direct answer
Before anyone signs off on building it out, run the prototype once, live, on one input picked by someone who did not build the demo. And say out loud, in the same room, which shapes of real input never once got tested. A demo that never gets anything wrong is not proof the tool is ready. It's proof nobody has shown it anything hard yet.
Do this, in order
  1. Run one live, unscripted input before any go/no-go review, picked by someone who wasn't building the demo.Why: it's the only thing that resets the promise the room walks away holding, instead of the promise the curated sample happened to make.
  2. Recut early accuracy by input type before reporting one blended number to leadership.Why: a number that still looks "mostly fine" can be one input type running almost useless underneath it.
  3. Name out loud which input shapes never got tested, before anyone asks for budget.Why: a gap you named on purpose is a known risk; a gap nobody named is a surprise waiting for a controller to find.
  4. Show the deciding stakeholders one deliberate miss before they ever see the tool run clean.Why: without ever seeing "wrong," a room has no way to recognize it once it happens for real.
  5. Trace the first real failure back to which input shape it actually was, before assuming the model needs retraining.Why: retraining fixes nothing if the real gap was which inputs got tested, not how the model scores on them.

How to answer this, stage by stage

Seven moves. The trap in this question is answering it with a general rule about honest demos, so most of these stages exist to prove the demo made a promise, out loud, and to find exactly where that promise broke.

1
Anchor it in one product and one number
Say it like this
"Let me put a number on this. Say a wholesale distributor, Thornbury Supply, has an AP team drowning in invoices. A PM, Idun Falkenberg, builds a tool called Codeframe: it reads a vendor invoice line by line and picks the GL code, the account bucket, for every charge. At the demo, it ran on fifteen invoices from three vendors. Forty-six of forty-eight lines came back coded right."
Why this works
A real product, a real person, and a real number turn "critique a happy-path prototype" into something you can actually trace back to a decision.
2
Say your structure, and reframe the question
Say it like this
"I want to run this as a diagnosis, T-R-A-C-E: timeline, recut, assume nothing, cause candidates, evidence test. Because the real question underneath 'critique a happy-path prototype' isn't 'was the demo dishonest.' It's why a demo that never once broke could still walk a whole finance team into building the wrong thing."
Why this works
Naming the plan and the reframe in one breath tells the interviewer this won't be five minutes of general advice about honesty.
3
Give the direct decision, straight
Say it like this
"So here's what I'd actually do. Before anyone signs a budget off the back of a prototype, run it once, live, on one input nobody on the build team picked. And say out loud, in the same room, which input shapes never got tested."
Why this works
Naming the concrete action before the story means nobody has to wait for the ending to know the answer.
4
Walk the timeline, then recut it
Say it like this
"Idun demoed Codeframe on April 9th, fifteen invoices, all from three vendors. Casper Rundgren, the CFO, signed off on the full build that same afternoon. Rollout to all forty vendors started April 23rd. For five weeks the straight-through rate just crept down, 88 to 82 to 75 to 68 to 60 percent. Looked like normal settling-in. So I recut it by vendor. The three demoed vendors held at 90 percent the whole time. The other thirty-seven sat at 20."
Why this works
A blended number hiding a segment that's basically broken is the single strongest move in a diagnosis, and it's the one most answers skip.
5
Rule out the model before blaming the input
Say it like this
"Before I say the other thirty-seven vendors' invoices were just messier, I check. Same frozen model, the whole five weeks, nothing changed. I pull the demo recording. Idun says, word for word, 'drop any invoice in the folder into this and it'll code every line by the time you blink.' So the model didn't drift. The promise was made about every invoice, and only three vendors were ever asked to keep it."
Why this works
Ruling out the model before blaming the input is what separates a real diagnosis from a guess that happens to sound right.
6
Name the causes, then run the one test that proves it
Say it like this
"Three reasons this happens. One, the fifteen demo invoices all came from the three highest-volume, cleanest vendors, cherry-picked because they were single-page, one cost center each. Two, the team had nine days between kickoff and demo and spent every hour tuning against that sample, nobody was ever assigned to go find an invoice that would break it. Three, the room was the CFO deciding whether to fund a year of engineering, and a visible miscode would have read as 'not ready,' not 'here's the edge we're closing.' Here's the test. Pull one real freight invoice from a vendor never in the demo set, split across three plants, run it live. It comes back miscoded, thirty-eight thousand four hundred dollars charged to the wrong plant. Rewrite the same invoice as one cost center, one page, rerun it: right, in under a second. One test, and it's the whole story."
Why this works
Naming three candidates and then narrowing to the one the evidence actually confirms is the hardest, strongest move in the whole method.
7
Close on the one line that matters
Say it like this
"So that's the answer. Codeframe never once got an invoice wrong in that conference room. It just never once met one it hadn't already been shown. Run one input nobody picked before you ever ask for the budget, and name out loud which shapes never got tested. That's the whole fix."
Why this works
Ends on the decision, not a recap, which is the line an interviewer actually remembers.

Let's learn

Codeframe is a tool that reads a vendor invoice, line by line, and decides which GL code, which account bucket, each charge belongs to.

Knowledge spark: what's a GL code? A short number that tells the accounting system which bucket a cost belongs in. Freight goes in one bucket, raw materials in another, so the books add up right at month end.

Before Codeframe, an AP clerk at Thornbury coded about 220 invoice lines a day by hand, checking each one against the vendor's catalog and the right plant's cost center. Getting through a week of invoices took the whole team, plus two extra evenings before month-end close.

At the demo, Idun ran Codeframe on fifteen invoices from three vendors, in front of the CFO, Casper Rundgren, and the finance controller, Bastian Vintner. Every invoice came back coded in under a second. Forty-six of its forty-eight lines were right. Casper signed off on the full build that same afternoon: every vendor, every invoice, live by the end of the month.

The blended straight-through rate, week by week
Share of invoice lines Codeframe coded with no human correction, all vendors combined
60%, and nobody had split it by vendor
wk1, 88%wk2, 82%wk3, 75%wk4, 68%wk5, controller finds it, 60%

Here is the turn. That slide from 88 to 60 percent was never the real problem, and it never was. The real problem showed up the moment somebody split that number apart by vendor instead of reading it as one number for the whole AP queue.

We didn't build a product that got invoices right. We built a demo that only ever met invoices it already knew how to get right.

At its worst, this costs Thornbury a plant controller catching a $38,400 freight charge coded to the wrong cost center during month-end close, and it nearly gets the whole build shelved two weeks before the fiscal year rolls over, taking Codeframe away from the ninety percent of invoices it was actually coding right.

A hand sketch horizontal timeline with four marks: Codeframe demoed on three clean vendors, the full build approved the same afternoon, rollout going live to all forty vendors, and a controller catching a thirty-eight thousand four hundred dollar miscode marked in red, with the gap between the approval and the miscode labeled as the stretch nobody watched.
The gap between when the build got funded and when anyone actually checked it

The choice I would take back. Idun's team built the entire demo from the three highest-volume vendors, because those were the invoices IT already had a clean feed for, and never wrote down that this was a curated sample rather than a fair test. I would take that back. I would put one line in the funding request: this demo used fifteen invoices we picked ourselves; one invoice nobody chose, run live, is the actual bar, before anyone signs the budget.

The decision that mattered Run one unscripted invoice, live, before any funding ask, and name out loud which invoice shapes never got tested. Both of those, together, before the budget is signed, not after a controller has to reverse a posting.

What I would leave alone. The three vendors Codeframe was actually tested on don't need any of this. Their straight-through rate held at 90 percent the whole five weeks. Slowing that down to rebuild trust for a problem that isn't happening there would just cost the invoices it's already coding right.

The lesson. A demo that never once gets an invoice wrong isn't proof a tool is ready. It's proof nobody has shown it an invoice it doesn't already recognize, and the room has no way to tell the difference.

The five weeks between the sign-off and the reversal

Read the short version above if you're short on time. This is the long version, for the part where you feel exactly what nearly went wrong.

Idun Falkenberg spent four days picking those fifteen invoices. Not cherry-picking, exactly, she'd tell you. Just choosing the ones that looked like what most of Thornbury's invoices looked like: one page, one vendor logo she recognized, one cost center, a freight line and a materials line and nothing stranger than that.

The demo ran on a Thursday afternoon in April, in a glass-walled conference room with the good coffee machine. Idun fed in the first invoice. Under a second later, Codeframe had it coded: freight to 6100, packaging to 5400, each line labeled with the plant it belonged to. Casper leaned back in his chair. By the fifth invoice, Bastian Vintner, the controller, had stopped watching the screen and started watching Casper.

Casper signed the funding request before the meeting even ended. Full build, all forty vendors, live invoices flowing straight into the ledger by the end of the month. Nobody in the room asked what the other thirty-seven vendors' invoices actually looked like, because the writing center, in this case the AP team, had always coded every vendor the same way, and nobody thought to ask whether a freight bill from a trucking company looks anything like an invoice from an office-supply vendor.

For the first two weeks, the numbers looked fine. Straight-through, the share of lines Codeframe coded with nobody correcting it, sat at 88 percent, then 82. Nobody was watching closely. Month-end close was coming, and the AP team had their own invoices to chase down.

By week four it was 68. By week five, 60. Still, on paper, "mostly fine." Nobody split it by vendor, because nobody had a reason to look.

We didn't build a worse tool for the other thirty-seven vendors. We built the same tool and quietly assumed every vendor invoiced like the three we picked.

Then, in the fifth week, Bastian Vintner found it during month-end close. A freight invoice from Harborlight Freight Lines, a vendor never once in the demo, had come in split across three plants on one page: a shared trailer run, one line item per plant. Codeframe had coded the whole $38,400 charge to Plant 2. It belonged to Plant 3. The mistake hadn't shown up in seconds, the way Idun had shown the room in April. It had come back overnight, queued behind a low-confidence flag nobody in that conference room had ever seen fire.

Bastian forwarded the posting to Idun with one line: "This isn't what you showed us."

Her first instinct was to wonder if he was being unfair. Codeframe hadn't changed. So she checked the model version first, because that's the easy thing to blame. Same frozen build, running everywhere, nothing had changed since April. So she looked at what the Harborlight invoice actually was, next to what she'd demoed.

She pulled the recording of the April demo. Fourteen minutes in, she heard her own voice: "Drop any invoice in the folder into this and it'll code every line by the time you blink." She'd meant it about the three vendors she'd tested. Nobody in that room had any way to know that.

She took one more real invoice from a vendor never demoed, still sitting in the queue, and ran it live in front of her own team. It scored below Codeframe's confidence line and queued for a human coder to check, the same queue that had never once fired on her fifteen demo invoices because none of them had ever been that unsure. Nine hours later, it came back: coded to the wrong plant, same mistake shape as the Harborlight charge.

So here is the decision I would take back. When Idun's team picked the demo invoices, they picked the ones that looked like the vendors IT already had a clean feed for. That made sense in April. What nobody did was write down that this was a curated sample, and run even one invoice nobody had chosen, live, before a year of engineering budget and a plant controller's trust were riding on it.

And the part I'd want to tell myself, if I could go back: we tested Codeframe against the shape of invoices we already had a clean feed for. We never once tested it against the shape of invoices the rest of the company actually received.

What splitting it by vendor actually showed

Before trusting the vendor-level gap, Idun's team checked whether Codeframe's grading was even right. Two AP clerks hand-checked 20 of the miscoded lines against what a controller would have flagged. They agreed with Codeframe's own confidence flag on 18 of 20. The model wasn't the problem. That left the input.

Same five weeks, cut by vendor instead of blended
90%
20%
The 3 demoed vendors
the shape Codeframe was demoed on
The other 37 vendors
never once shown at the demo
Vendors using the tested invoice shape
The vendors that never used it
The blended number read 60 percent because the other 37 vendors were still under half of total invoice volume in week five. Their own number never showed up until someone cut it out on its own.
Freight invoice, as it actually arrived
1 real Harborlight invoice, week five, live test
9 hrsqueued, then coded to the wrong plant
Same charges, reshaped to match the demo set
Same invoice, restructured to one cost center, same model
0.6 seccorrect plant code, no queue

Three reasons a demo over-promises, and the one that was true

Not because anyone was careless. Each of these, on its own, looks like a normal decision at demo time. Together, they're why a prototype that never once stumbled can still make a promise the real product can't keep.

Three hand-sketched labelled panels: a stack of papers for the cherry-picked, clean vendors chosen to look good, confirmed as the true cause; a small dial for the nine days spent tuning with no time budgeted to hunt for a failure case; and a person icon for the CFO audience that would have read a shown failure as weakness.
Three separate, checkable causes, only one of them confirmed by the live test
Cause 1
Cherry-picked inputs, chosen to look good.

The fifteen demo invoices all came from the three highest-volume vendors, the ones with a clean, standard, single-page template and one cost center per invoice. Freight invoices split across plants, credit memos, and multi-page vendor statements never made it into the sample.

How you'd check it: take a sample of real invoices from vendors never demoed, and check how many share the same shape, one page, one cost center, as the demo set. If almost none do, the sample was chosen to look good, not to represent the real queue.
Cause 2
No time budgeted to hunt for a failure case.

The team had nine working days between kickoff and demo day. Every hour of it went into tuning Codeframe's accuracy against the fifteen-invoice sample. Nobody was ever assigned the separate job of trying to find an invoice that would break it.

How you'd check it: ask whether one hour of the build schedule, anywhere, was explicitly set aside to go find a failure case. If the honest answer is no, the team only ever practiced succeeding.
Cause 3
A stakeholder audience that would have read a failure as weakness.

The demo audience was the CFO, deciding whether to fund a year of engineering headcount. Idun believed that showing one miscoded line, on purpose, would read as "this isn't ready," not as "here's the exact edge we're going to close before rollout."

How you'd check it: ask what a deliberately shown miss would have actually cost that day, against what an undiscovered one cost five weeks later. If the second number is bigger, the fear of looking unready was the wrong fear.

TRACE, run against Codeframe's first five weeks

This reads like a question that wants a general rule about honest demos, but the real job is diagnosis: work out why a prototype that never once stumbled could still walk a finance team into funding the wrong build, and prove exactly where that promise broke.

T, timeline. Codeframe was demoed and the full build was approved on the same afternoon, April 9th. Rollout to all forty vendors began April 23rd. The first real failure, the $38,400 miscode, surfaced during month-end close in week five, roughly seven weeks after the approval. Nobody flagged the gap between the invoice shapes tested and the invoice shapes rolling out; it looked like the AP team serving its usual vendor list, not a product decision.
R, recut. The same five weeks, split by vendor instead of blended. The three demoed vendors: steady at 90 percent the whole time. The other thirty-seven vendors, never once shown at the demo: 20 percent. The blended number of 60 percent hid a group running roughly four and a half times worse than the rest.
A, assume nothing. Before blaming the other thirty-seven vendors' invoices for being messier, rule the model out. Same frozen build ran the whole five weeks. Pull the actual demo recording: Idun's own line, fourteen minutes in, promises the tool works on "any invoice in the folder." The model didn't drift. The demo set the expectation, and only three vendors were ever asked to meet it.
C, cause candidates. Three, named and separate: the demo invoices were cherry-picked from the cleanest vendors to look good; no time was ever budgeted to hunt for a failure case in the nine days before demo day; and the CFO audience deciding on a year of budget would have read a shown miss as "not ready" rather than "here's the edge we're closing."
E, evidence test. Take one real Harborlight Freight invoice, run it live with the same frozen model. It queues, and comes back nine hours later coded to the wrong plant, the same mistake shape Bastian found. Restructure the same charges as one cost center, one page, and rerun: 0.6 seconds, correct plant, no queue. Same model, only the shape of the invoice changed, which is what proves it's the input-shape cause, not a model that's forgotten how to code freight.
Why the live test is the hard step Anyone can suspect the other thirty-seven vendors were the problem. The live test turns that suspicion into two results off the same model, the raw invoice and the reshaped one, and shows exactly how much of the gap the invoice's shape explains, instead of a hunch dressed up as a finding.

Same blind spot, a bolt of fabric nobody demoed

Thistlebrook Loom Textiles, a garment manufacturer, is piloting a defect classifier called SeamCheck: a QA camera photographs a finished garment and flags stitching or fabric defects before it ships. Ingunn Saether leads QA tooling for it. It launched with a live demo on twelve clean cotton-shirt photos, then rolled out floor-wide four weeks later.

T. Ingunn demoed SeamCheck on twelve cotton-shirt photos in month one; floor supervisors agreed with its call 35 of 36 times. Floor-wide rollout followed three weeks later. The blended agreement rate, SeamCheck's flag matching what a supervisor actually decided, crept from 94 to 71 percent over the next five weeks, and nobody split it apart until a line supervisor escalated.
R. Recut by fabric type. Cotton shirts: 93 percent, steady the whole time. Stretch knits and dark denim, never once demoed: 27 percent.
A. Same frozen model scored both groups. A manual review of 15 flagged stretch-knit photos agreed with SeamCheck's own confidence flag on 13. The grading held up. Ingunn pulled the demo recording: she'd told the room "point it at literally anything coming off the line." The floor wasn't asking for more than that.
C. Three candidates, the same shape as before: only cotton-shirt photos were ever demoed, chosen because they photographed cleanly under the line's lighting; no time was set aside to go hunt a garment that would trip the camera; and the demo audience was the plant director, deciding whether to fund cameras for every line, and a visible miss would have read as "this isn't ready to leave the lab."
E. One real stretch-knit photo, queued for a supervisor and came back four hours later, passed as clean when it had a dropped stitch, the exact kind of miss the line supervisor reported. Re-shot against a cotton shirt with the same dropped-stitch pattern stitched in for the test: correctly flagged in under a second. Same model, only the fabric changed, which pointed straight at the input SeamCheck had never once been shown, not at a model that can't see a dropped stitch.

Swap the trigger and it still runs

  • Speed: Casper could have pushed Codeframe to all forty vendors in one week to hit a quarter-close deadline instead of rolling out over two weeks. TRACE still starts by asking what shipped at demo time and when the real friction reached someone, not by how fast the rollout happened.
  • Cost: the team could have skipped the live unscripted check to save a day before the pitch to Casper. The check still has to happen eventually, just after a $38,400 posting gets reversed instead of before one.
  • The model really did get better: say Codeframe's next version genuinely got sharper on every vendor type, better coding company-wide, in the very same stretch a new freight vendor's invoice shape still caught it flat-footed. TRACE still finds the gap, because the recut isolates one vendor group even while the overall trend looks like good news.

Where people run it wrong

  • Trusting a blended number that's still technically "mostly fine," without ever cutting it apart by vendor or invoice type.
  • Treating one bad posting as proof the model needs more training, before checking whether the input even matched what got demoed.
  • Fixing the visible symptom, retraining on more freight examples, instead of the actual gap: what got shown at demo time, and what got quietly left out.

How to use it live

Buy yourself ten seconds by naming the split out loud. "So there's the demo everyone remembers, and there's whatever it never got tested against. Let me say how I'd check whether that gap is already showing up." That's not stalling. That's where the real critique starts.

Flashcards (click a card to flip it)

This is a critique question about a happy-path prototype, worked as a diagnosis, so these eight test the TRACE moves and the real numbers behind them.

1 · THE FRAMEWORK
Which framework fits "critique a prototype that only shows the happy path," and why?
Tap to flip
ANSWER
TRACE. It sounds like it wants a general rule about honest demos, but the real job is diagnosis: working out why a demo that never once broke could still walk a team into funding the wrong build, then finding exactly where that promise broke.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Idun Falkenberg, product manager for Codeframe, an AI invoice line-item coder. She picked and ran the April 9th demo herself, for Thornbury Supply's accounts-payable team.
3 · THE HABIT
What did nobody on Idun's team do before the full build got funded?
Tap to flip
ANSWER
Run Codeframe once on an invoice nobody had picked in advance. Every demo invoice was chosen and checked by Idun herself, all from the three vendors IT already had a clean feed for.
4 · THE THREE CAUSES
Name the three reasons a demo like this over-promises.
Tap to flip
ANSWER
The demo invoices were cherry-picked from the cleanest vendors to look good, no time was budgeted to hunt for a failure case, and the CFO audience deciding on the budget would have read a shown miss as weakness.
5 · THE NUMBER
The three demoed vendors held at 90 percent straight-through while the other 37 fell to ______ percent.
Tap to flip
ANSWER
20 percent. The blended, company-wide number only read 60 percent, because the other 37 vendors were still under half of total invoice volume in week five.
6 · THE CHECK
Name the one test that proved it was the input, not the model.
Tap to flip
ANSWER
Running one real Harborlight freight invoice live: it queued and came back 9 hours later coded to the wrong plant. Restructuring the same charges as one cost center, one page, came back in 0.6 seconds, correct plant, no queue.
7 · THE FIX
What should have happened before Codeframe ever got its full-build budget?
Tap to flip
ANSWER
One live, unscripted invoice chosen by someone who didn't build the demo, plus naming out loud which invoice shapes never got tested.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what's the number?
Tap to flip
ANSWER
SeamCheck, a garment defect classifier for Thistlebrook Loom Textiles. Cotton-shirt agreement held at 93 percent while stretch knits and dark denim, never demoed, fell to 27 percent.

Check yourself Score: 0 / 0

Fill in the blank
1. Codeframe's straight-through rate on the other 37 vendors fell to ______ percent by week five, while the blended, all-vendor number only read ______ percent.
Show hint
Look at the recut chart, the two bars split by vendor, next to the blended line chart above it.
Show answer
20 and 60. The gap only showed up once someone cut the number by vendor instead of reading the company-wide average.
True or false
2. True or false: since the blended straight-through rate only slid from 88 to 60 percent over five weeks, that proves Codeframe's model was getting worse at coding invoices.
  • True
  • False
Show hint
Look at which single thing stayed frozen across the whole five weeks.
Show answer
False. The same frozen model ran the whole time. The recut and the live test both point to the shape of the never-demoed vendors' invoices, not the model getting worse.
Multiple choice
3. Why couldn't Idun's team have just added a line like "results may vary by vendor" to the funding request instead of running the live, unscripted check?
  • A. Because a vague disclaimer never shows anyone what "wrong" or "slow" actually looks like, so it never resets the real expectation the demo created.
  • B. Because disclaimers aren't allowed in corporate budget requests.
  • C. Because it would have made Codeframe look worse than a competing tool.
  • D. Because Casper never reads anything attached to a funding request.
Show hint
Ask what a vague disclaimer actually shows the room, versus what a live, unscripted test shows them.
Show answer
A. A disclaimer is words about uncertainty. A live, unscripted test is uncertainty the room actually watches happen, which is the only thing that resets a promise someone believed.
Short answer
4. Name a place in Thornbury's use of Codeframe where this same fix would NOT matter, and say why.
Show hint
Think about the vendors whose invoices already match what got shown at the demo.
Show answer
Model answer: "Leave the three demoed vendors alone. Their invoices match exactly what got shown, and their straight-through rate held at 90 percent across all five weeks. Rebuilding anything there spends effort on a gap that isn't happening."
Short answer, apply it yourself
5. Think of a demo or pitch you've seen for a real tool. What's one thing about it that made the promise feel bigger than what actually shipped?
Show hint
Look for a case where the demo never showed the tool waiting, hedging, or getting something wrong.
Show answer
Model answer: "An expense-report scanner demo only ever read printed restaurant receipts. It never showed a crumpled gas-station receipt with faded ink, which turned out to be exactly what I actually hand it most weeks." Any honest answer works if it names a real gap between what the demo showed and what you actually needed the tool to handle.
Fill in the blank
6. If the other 37 vendors had made up 65 percent of the invoice queue in week five instead of about 43 percent, the blended number would have read about ______ percent instead of 60.
Show hint
Weight 90 percent and 20 percent by 35 percent demoed vendors and 65 percent the rest.
Show answer
About 45. 0.35 × 90 plus 0.65 × 20 comes out to roughly 45 percent, low enough that nobody could have called it "mostly fine." The blend only stayed reassuring because the other 37 vendors were still under half of the queue.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more