ConceptIntermediateShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #14
Explain how prototyping shortens the feasibility debate.
The direct answer
When a feasibility debate stalls, stop arguing and run the model on your worst real examples for one afternoon. A pass or fail on the hardest cases you actually have settles whether something is possible faster and more honestly than any amount of discussion. It only proves the thing can be done at all, not that it is ready, cheap, or reliable enough to trust yet.
Do this, in order
Stop debating and run the model on real, hard examples for one afternoon.Why: a real attempt at the task is the only signal that isn't just someone's opinion.
Pick the messiest, worst cases you actually have, on purpose.Why: a pass on the easy cases proves nothing about whether it is possible at all.
Treat a pass as answering "can it be done," never "should we build it."Why: cost and reliability at real scale are a separate question the quick test cannot touch.
Time-box the test to hours, not another week of meetings.Why: a fast, rough answer that actually gets run beats a careful one that never does.
Bring the room a result, not a feeling.Why: a shown pass or fail is much harder to argue with than a verbal opinion.
Skip this for debates nobody is actually having.Why: don't spend an afternoon proving something everyone already agrees on.
How to answer this, stage by stage
Seven moves. Name what a fast test can and cannot settle before the story, or the answer sounds like advice instead of a decision.
1
Scope it to one product and one stuck meeting
Say it like this
"Let's make this concrete. Say a legal-ops team at Marrowgate Power, a regional utility, wants a tool that reads a stack of vendor contracts and flags any clause that drifts from the company's standard language. Torvald Amsel runs product for that team, and right now the group is stuck arguing about whether that's even possible."
Why this works
A question about prototyping stays abstract until it's tied to one real argument that's actually going nowhere.
2
Say the plan out loud
Say it like this
"I'll walk this through LEAD. L is the real outcome a prototype protects here. E is the early signal it gives you. A is how the debate gets gamed without one. D is what that quick read still can't tell you."
Why this works
Naming the four letters up front tells the interviewer you're about to make a call, not share a feeling about prototypes.
3
Say what the question is actually checking
Say it like this
"This sounds like a question about whether prototypes are useful. It's really asking whether a decision about capability should come from a real attempt at the task, or from whoever in the room sounds most sure of themselves."
Why this works
This moves the answer from "prototypes are good practice" to naming exactly what a prototype replaces: an opinion contest.
4
Give the decision straight
Say it like this
"Here's the answer. Stop arguing and hand the model your worst real cases for one afternoon. A real pass or fail on the hardest version of the task ends the feasibility question honestly, in hours instead of weeks. It only tells you the thing can be done at all. It does not tell you if it's reliable or cheap enough to actually ship."
Why this works
This names exactly what the quick test settles and what it leaves open, instead of a vague "prototyping helps."
5
Prove it with the number that moved
Say it like this
"Here's why that matters. Marrowgate's legal-ops group argued about this for three weeks across six meetings and never reached a decision. Torvald pulled the ten worst vendor contracts he could find, handwritten riders, faxed scans, the ones a paralegal groans at, and tested the model on them in one afternoon. By the fourth hour it had correctly flagged eight of the ten. The debate that had gone nowhere in three weeks ended in a twelve-minute meeting the next day."
Why this works
A real result that ends a three-week argument in twelve minutes does more work than a paragraph about the value of testing early.
6
Name the gaming path
Say it like this
"Nobody was acting in bad faith. In week three, the most senior lawyer in the room said flatly, 'there's no way a model can tell a real deviation from boilerplate,' and nobody could out-argue that with anything but another opinion. That's how a feasibility debate gets gamed without anyone cheating: whoever sounds most certain wins the room, whether or not they've actually opened a contract."
Why this works
Naming the exact mechanism, confidence beating evidence, is what separates this from a generic "test things early" answer.
7
Name the hard limit, close on one line
Say it like this
"And here's what that afternoon still couldn't tell Torvald. Whether eighty percent on ten contracts is good enough to actually trust. What it costs to get from eighty to something you'd stake a real workflow on. How it holds up across all four hundred contracts instead of ten. That's a separate question, for a separate week. So the one thing I'd actually do: when a feasibility debate stalls, stop talking and run the model on your worst real examples. Let a result decide it, not whoever's loudest."
Why this works
Closing on the same decision from stage four means it's the last thing the interviewer hears, with the boundary named clearly.
If you only get through two stages
Stages 4 and 7 are the answer. Say the decision, run the model on your worst real cases for an afternoon instead of debating, then name the boundary: it proves feasible-at-all, not feasible-at-acceptable-cost-and-reliability. Everything else here is how you defend that under pushback.
Let's learn
What happens when two reasonable people argue for three weeks about whether something is possible, and neither one has actually tried it?
That's what was happening at Marrowgate Power, a regional electric utility with a legal-ops team that manages several hundred vendor contracts, tree-trimming crews, transformer suppliers, linemen contractors, each with its own version of the standard liability and indemnification language. Someone proposed a tool that would read a stack of those contracts and flag any clause that had quietly drifted from Marrowgate's approved wording.
Knowledge spark: what's an indemnification clause?
The part of a contract that says who pays if something goes wrong on the job. Vendors often try to soften it in their favor. Catching a softened one before signing is the whole point of the tool.
Before anyone tested anything, the debate ran on opinions. One camp said obviously a model could read contract language, it reads everything else. The other camp said legal nuance is exactly the kind of thing a model fakes convincingly and gets wrong. Six meetings across three weeks produced sharper arguments on both sides and zero decisions.
Here's the turn. The problem was never really about the model. It was about the fact that nobody in that room had actually pointed it at a real contract.
We were never debating whether the model could do it. We were debating who sounded more sure in the room.
At its worst, that kind of stall doesn't just waste three weeks. A confident "no way this works," said plainly enough and repeated often enough, starts to feel like a finding instead of a guess. A workable tool gets quietly shelved because the loudest person in the room never opened a single contract, and nobody else had anything but another opinion to set against theirs.
Pass rate on the ten hardest contracts, tracked hour by hour through one afternoon
Small sample, looked strong
More of the worst cases added, settling near 80%
By the fourth hour, the model had correctly flagged eight of the ten worst contracts Torvald could find. That single afternoon produced more real evidence than the three weeks of meetings before it.
The debate itself had a cost that never showed up in a spreadsheet. Every extra meeting made the confident "no" sound a little more settled, and a little harder for anyone to walk back.
Time to an actual decision, debate alone against one afternoon of testing
3 weeks
Six meetings, arguing only, no decision reached
4 hours
One afternoon of testing, decision reached the next day
Three weeks of arguing settled nothing. Four hours of actually running the model settled the question the meetings never touched.
The confident answer won the room. The contract that could have answered it stayed shut.
The choice I'd take back
We tried to settle whether the tool was possible with more meetings, reading case studies from other companies and comparing opinions, instead of pulling Marrowgate's own worst contracts and running the model on them. That felt like ordinary diligence in week one. It stopped being fine once six meetings had passed and the room had stronger opinions but still no real answer.
What I'd leave alone. A feasibility question nobody is actually arguing about doesn't need this. Whether the same tool could reformat a contract into a plain summary table was never in dispute. Nobody needed an afternoon of testing to settle a question the room already agreed on.
The lesson. A debate about whether something is possible is not evidence about whether something is possible. It's evidence about who's in the room that day. The fastest honest way to end it is to remove the argument and hand the model the worst examples you actually have.
Now here is the same thing as a story
Use this version when you've got a few minutes. The short version is above. This is for when the twelve-minute meeting needs to actually land.
Torvald Amsel can spot a padded timeline before anyone's finished the first slide. Three years running product for Marrowgate Power's legal-ops tooling group taught him that: if someone can't tell him what a proposal does to one real file, it isn't real yet.
So when the clause-comparison idea came up, the first two meetings were genuinely good. People built real arguments, brought real examples of contract language that had drifted, disagreed in ways that moved the conversation somewhere. Torvald liked those meetings. They felt like the team doing its job.
By the fifth meeting, in the third week, nothing new was being said. Everyone was restating a position they'd already stated twice before. Then Marrowgate's most senior in-house counsel, arms crossed at the head of the table, said it plainly: "There's no way a model can tell a real deviation from boilerplate. I've read enough of these to know." Nobody in the room had a contract in front of them. Nobody could counter an opinion with anything but another opinion. The sixth meeting, the following Tuesday, opened exactly where the fifth one had ended.
Torvald didn't wait for a seventh. Wednesday morning, before the tooling group's stand-up, he pulled the ten ugliest vendor contracts he could find in Marrowgate's files himself, not the clean ones anyone would have handed him as "representative." Handwritten margin edits. A faxed scan so blurry the footer was barely legible. A liability clause split across two non-standard paragraphs. He ran the model on them, one at a time, and timed himself.
Hour one, two contracts, both flagged correctly. Hour two, five contracts, four right. Hour three, eight contracts, six right. By four o'clock, ten contracts, eight right. Eighty percent, on the worst cases in the building, gathered by one person in an afternoon.
The cost was never the two contracts the model got wrong. It was the three weeks Marrowgate nearly let a confident opinion stand in for evidence.
He printed the ten contracts with the model's flags marked in the margin and brought them to what would have been meeting seven. It lasted twelve minutes. Nobody argued the model was flawless, it had missed two out of ten. But two mistakes on paper, sitting in front of the room, did something six meetings of argument never managed: it turned "is this even possible" into "how good does it need to be before we trust it," which was a completely different, much better question.
Two months earlier, in the kickoff meeting for this idea, someone had asked whether they should just test it on a real contract right away. Torvald remembers agreeing that made sense, but suggesting they first read a few vendor case studies from other companies "to get a feel for it," so nobody would feel put on the spot testing something half-built. Nobody ever circled back to say the case studies had quietly become the whole plan.
I would take that back. I'd pull the worst real examples on day one of the debate, not three weeks and six meetings into it.
Here's the replay. Same idea, same skeptical senior counsel, but Torvald runs the afternoon test the week the idea is first raised, before a single meeting about whether it's "even possible." Meeting one becomes the twelve-minute meeting instead of meeting seven. Three weeks of arguing never happens. The team spends that time deciding what accuracy they'd actually need to trust the tool, which is the real question, six weeks earlier than it got asked the first time.
One path hands the room an opinion. The other hands it a result.
And the thing I'd tell myself, back in that kickoff meeting: the fastest way to end an argument about whether something's possible is to stop having the argument.
LEAD, for a debate that has no number yet
This sounds like a question about whether prototypes are useful. Underneath, it's still asking whether a team can act on a real result instead of an opinion. That's LEAD, run on an argument instead of a dashboard.
L, link. What a prototype actually protects here. Not whether the model looks capable in conversation. Whether the decision to build, or not build, gets made because someone watched it try the real task, not because someone sounded certain. → Here, that's whether Marrowgate greenlit the clause tool because it read ten real contracts, not because a senior voice sounded sure.
E, early signal. A real pass or fail on the hardest version of the task, gathered in hours, not another week of argument. → Torvald's afternoon: eight of ten correctly flagged on Marrowgate's worst contracts, by four o'clock the same day.
A, abuse. How a feasibility debate gets gamed without a prototype. Whoever argues most confidently wins the room, regardless of whether they've tried it. → The senior counsel's flat "no way" carried two extra meetings, and nobody could out-argue an opinion with another opinion.
D, decision. What a quick prototype's read genuinely cannot settle. Whether it's feasible at acceptable cost and reliability, only whether it's feasible at all. → Eighty percent on ten contracts in an afternoon proved it could be done. It said nothing about whether eighty is good enough to trust, or what it costs to close that gap.
The check that proves the read is real
Pull the cases the room is actually arguing about, not the clean ones. If the model still gets most of them right, the possibility question is answered. Everything argued after that is a cost and reliability question, not a "can it be done" question.
And if you want to be sure it really works, try it somewhere else
A hospital radiology department had been debating for two months of grand rounds whether an AI second-read tool could reliably flag missed findings in mammograms before anyone would fund a real pilot. Ottokar Fenwright leads product for imaging tools at Quennell General Hospital, and watched the argument circle the same three points every session.
L. Whether Quennell trusts the tool enough to fund a real pilot because it read real disputed cases, not because a senior voice in grand rounds sounded certain.
E. A pass or fail on the fifteen hardest case files, the ones two radiologists had already disagreed about, gathered over one on-call weekend.
A. Seniority wins grand rounds without anyone opening a case file. The most senior voice in the room doesn't need evidence when nobody present outranks the confidence in their tone.
D. It can't say whether the tool is reliable enough to rely on across thousands of scans, or what a missed finding costs at that reliability. Only that catching a disputed case at all is possible.
Knowledge spark: what's a missed finding?
Something in the scan a radiologist should have flagged and didn't. Rare, and exactly the kind of case a demo built on easy scans would never surface.
Ottokar pulled fifteen case files from the hospital's own near-miss log, cases where two radiologists had genuinely disagreed, and ran the tool over one weekend instead of waiting for grand rounds to reach a verdict on its own. It correctly flagged eleven of the fifteen. The following week's session, expected to be argument number nine, lasted six minutes.
The debate's clock barely moved in two months. The test's clock reached a real answer in a weekend.
Swap the trigger and it still runs
The model answers faster. Doesn't help on its own. A faster wrong opinion is still an opinion, you still need the worst real cases run for real.
The review budget shrinks. Doesn't help either. Fewer people in the room just means the loudest opinion meets less pushback, not that it's more likely to be right.
The model genuinely gets better at the task. Run the hard-case test anyway. That's the one situation where the quick read should actually improve, and only running it tells you it did.
Where people run it wrong
Testing the easy, clean cases to "prove" it works, instead of the ones the room is actually arguing about.
Treating a pass on the hard cases as proof it's ready to ship, skipping the cost and reliability question entirely.
Letting a stalled debate run for weeks before anyone thinks to just try it.
How to use it live
Say the split first, out loud. "Before we argue this any further, I want to separate two questions: can it do this at all, and is it good enough to trust. Let's answer the first one this week, with our worst examples, and only then argue about the second." That's not stalling. It names which question is actually stuck, and it buys you room to give a real boundary instead of a flat yes or no.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what does each letter stand for here?
Tap to flip
ANSWER
LEAD. L is the real outcome, a decision made from a real attempt, not an opinion. E is the early signal, a pass or fail on the hardest cases, gathered in hours. A is how it gets gamed, whoever argues most confidently wins. D is what it still can't tell you, whether it's feasible at acceptable cost and reliability.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Torvald Amsel, who runs product for the legal-ops tooling group at Marrowgate Power, and who ended a three-week feasibility debate with one afternoon of testing.
3 · THE HABIT
What kept the debate stuck for three weeks?
Tap to flip
ANSWER
The team kept trying to settle the question with more meetings and outside case studies instead of testing the model on Marrowgate's own contracts, because arguing about it felt like progress.
4 · THE EARLY SIGNAL
What real signal ended the debate, and how fast did it arrive?
Tap to flip
ANSWER
A pass or fail read on the ten worst contracts Torvald could find, tested in one afternoon. It landed at eight of ten correct by the fourth hour.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Trying to settle the feasibility question with more meetings and outside case studies instead of pulling Marrowgate's own worst contracts and testing the model on them from day one.
6 · THE NUMBER
The debate ran ______ across ______ meetings with no decision. The prototype settled it in ______.
Tap to flip
ANSWER
Three weeks. Six meetings. Four hours, with eight of ten contracts correctly flagged.
7 · THE REPLAY
Same debate, prototype run on day one instead of week three, what changes?
Tap to flip
ANSWER
Meeting one becomes the twelve-minute meeting instead of meeting seven. Three weeks of arguing never happens, and the team spends that time on the real question, what accuracy they'd need to trust it, six weeks earlier than before.
8 · THE TRANSFER
Section four runs this same question again for a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
Quennell General Hospital's radiology second-read tool. Its early signal is a pass or fail on the fifteen hardest, previously-disputed case files, gathered over one on-call weekend.
Check yourself Score: 0 / 0
Multiple choice
1. Which of these is the strongest read of what Torvald's afternoon test actually proved?
A. The model is ready to run this task across all of Marrowgate's contracts.
B. The model can flag deviations correctly on contracts shaped like the ten worst ones Torvald tested.
C. The senior counsel was wrong about everything, and the debate had no value.
D. The two contracts it missed prove the whole idea should be dropped.
Show hint
Look for the option that names exactly what got tested, not a claim about production readiness.
Show answer
B. It proved feasibility on the hardest cases Torvald had. It didn't prove the tool was ready to ship, and it didn't make the earlier debate pointless, it just gave the debate something real to argue about.
True or false
2. True or false: once Torvald's prototype passed eight of ten hard contracts, Marrowgate had enough to know the tool was reliable enough to use across all four hundred vendor contracts.
True
False
Show hint
Separate "can it be done" from "is it good enough to trust at scale."
Show answer
False. Ten contracts in an afternoon settled whether the idea was possible at all. It said nothing about reliability across four hundred contracts, or what it would cost to get there. That's the D step, the part a quick prototype genuinely can't answer.
Short answer
3. Name a feasibility question in this same legal-ops pipeline where testing this rigorously wouldn't matter.
Show hint
Look for a claim nobody in the room was actually disputing.
Show answer
Model answer: "Whether the tool could reformat a contract into a plain summary table. Nobody was arguing that was impossible, so there was no stalled debate to end and no reason to spend an afternoon proving it."
Fill in the blank
4. The feasibility debate ran for ______ weeks across ______ meetings with no decision. Torvald's afternoon test settled it, correctly flagging ______ of the ten worst contracts.
Show hint
The numbers behind both charts in Section 1.
Show answer
Three. Six. Eight. Six meetings of pure argument produced no decision. One afternoon of real testing produced a result the room could act on the very next day.
Short answer, apply it yourself
5. Think of a time you saw a team argue about whether something was "even possible" without anyone actually trying it. What's the fastest real test that could have ended that argument?
Show hint
Look for the hardest, messiest version of the task sitting closest to hand.
Show answer
Model answer: "A team argued for a week about whether a chatbot could handle refund requests, without anyone running one through it. The fastest test would have been pulling the five angriest, most confusing real refund emails on file and seeing what it did with them that same day."
Multiple choice
6. If Torvald had run his afternoon test the week the idea was first raised instead of after three weeks of debate, what would most likely have changed?
A. Nothing, the model's accuracy on the ten contracts would have been the same either way.
B. The debate would have taken even longer, because testing slows a team down.
C. The team would have skipped the three weeks of arguing entirely and spent that time on the real question of what accuracy they needed to trust it.
D. It depends only on who was in the room, not on when the test happened.
Show hint
Think about what running the test earlier actually buys: a faster real question, not a different model result.
Show answer
C. The model's accuracy doesn't change based on when you test it, but running the test early replaces three weeks of opinion with one afternoon of evidence. The replay in Section 2 shows exactly this: meeting one becomes the twelve-minute meeting instead of meeting seven.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.