How do you decide when to stop prototyping and commit?
- Write down the exact questions before round one, and commit the moment they are answered.Why: this is the whole switch the answer turns on. Skip it and there is no way to know when you are done, so you never are.
- Tie every question to one real decision it will actually change.Why: a question nobody is waiting on is not worth writing down. It just adds a box to check.
- Treat "one more round" as a real cost, not a free option.Why: extra rounds only feel cheap because nobody is counting the calendar time or the launch window they are eating into.
- Check whether the last two rounds changed your recommendation at all.Why: if they did not, you are not learning anymore. You are stalling and calling it care.
- Let anything with no live decision behind it stay loose, on purpose.Why: a market you are not launching this quarter does not need a deadline, because nothing is actually waiting on it.
- Name who pays the second cost, not just the hours.Why: the person reviewing round after round with no visible end stops paying attention long before anyone notices.
How to answer this, stage by stage
Seven moves. Most of the weight sits in stage five, the four sentences where the flat line finally gets named out loud. Every stage has the actual words to say.
Let's learn
Here is what happens when a team keeps running the exact same test long after it stopped teaching them anything new.
Say a global e-commerce company builds a tool that takes one English social caption and writes a localized version of it for nine other markets, matching each market's own slang and tone. Before the tool, four regional social managers did this by hand for every big sale, about three hours of translating and rewriting, spread across people who had other jobs that week. The company built a prototype that did it in under two minutes, and the first version got close: about six in ten captions still needed a heavy rewrite from a native speaker before they could post, but the other four in ten already nailed the tone on the first try.
By round four, in week five, the tool's rewrite rate had dropped to nine percent, under the ten percent bar the team cared about. It stayed between seven and nine percent for the next seven rounds. Nothing that happened after round four ever moved the number again.
Here is the turn. Those extra seven rounds were never the real problem. The real problem was what the team did instead of naming what "done" meant: they kept running the same test, waiting for the number to feel finished on its own, which it can never do, because a feeling has no line to cross.
At its worst, this does not end in a slightly better caption. It ends in a caption tool that missed the one thing it was built for. The company's year-end sale needed captions localized by week ten. The team was still testing in week ten, so the four regional managers translated everything by hand again, the same three hours per event as before the tool ever existed. The tool's first real launch happened after the sale, for ordinary week to week posts, where the win mattered far less.
What I would leave alone. The tone experiments Faelan's team was running for two markets with no launch date scheduled this year, those are fine staying loose and open ended. Nobody is waiting on those to ship anything, so there is no clock running yet, and no reason to force a stopping rule onto exploring for its own sake.
The lesson. A prototype does not stop because it finally feels done. It stops because the questions you built it to answer have real answers. If nobody writes those questions down, there is no way to know you are finished, so you never are.
Now here is the same thing as a story
Pull this one out when there is more time, and you want the interviewer to feel it, not just note it down.
Faelan Sandvik has run regional social campaigns for six years. Hand her a caption in any of the nine markets Threndle Commerce sells in, and she can tell in one read whether it lands as a joke or an insult, just from the word choice and the emoji. She built the first test set for the caption tool herself, out of forty real captions her own team had posted the year before, because she already knew exactly how each one should sound in every market.
Round one, in early spring, was rough and everyone expected it to be. Sixty one percent of the captions needed a full rewrite from a native reviewer before they could post. Round two dropped that to forty four percent. Round three, twenty seven. By round four, in week five, the number was nine percent, under the ten percent bar Faelan had privately been aiming for since the project started.
Nobody said the word "done." There had never been a meeting where anyone wrote down what done meant, so there was nothing on paper for round four to match against. It just felt like a good number in a project that was still going, and the team moved on to round five.
Round five was reasonable. They wanted to see the nine percent hold for a second round before trusting it. It held, at eight percent.
Round seven and eight turned into something else. Someone suggested handling emoji localization too, since they were already in there. That felt like diligence. It was really just a new task wearing the same badge as the old one.
By round ten and eleven, nobody in the weekly sync could say, if you had asked them directly, what specific thing they still did not know about the caption tool. "One more round" had become the answer to a question nobody was asking anymore.
Preeti Nair reviewed every Hindi-market caption in every one of those rounds. Around round eight, in week nine, she pulled Faelan aside after a sync. "This is the fourth week in a row I am reviewing basically the same caption," she said. "What am I actually supposed to be finding at this point?" Faelan did not have a real answer for her. Preeti kept reviewing, but she stopped reading closely. She started approving on a kind of autopilot, because nothing about the job had changed since week five, and nobody had told her it ever would.
The moment that finally broke it had no drama in it at all. In week fourteen, prepping the deck for round twelve's steering committee update, Faelan copied over the slide template from the last review and typed the recommendation line: "Launch human reviewed, revisit full automation in Q2." She stopped. That sentence looked familiar. She scrolled back to round six's deck, eight weeks and five rounds earlier. Same sentence, word for word. She checked rounds seven through ten. Same line, every single time.
I want to say the problem was that the caption tool was not good enough yet. It was already good enough, five rounds and nine weeks earlier. That is not really the story. Faelan never had a number that told her when to stop. She had a feeling, and a feeling only has two settings: this could still teach us something, or this is just making us feel better about a decision we already made. Once she saw the identical slide, there was no version of "test a little less" left standing.
So here is the decision I would take back.
Back at kickoff, when the caption tool project first got greenlit, the brief said "prototype the localizer, make it good." Nobody wrote down the specific things it had to prove. That felt right at the time. They were exploring something new, and pinning down exact questions before anyone had touched the tool once felt too early to be useful.
I would put those questions in the kickoff doc. Two of them: can it catch idioms that do not translate word for word, and will a native reviewer approve it with only light edits, not a full rewrite. Both tied to one real decision: launch human reviewed now, or wait and go fully automated.
Run the project again with that fixed. Round four still hits nine percent. Both questions still get answered, in week five. This time, because the answer is written down and not just felt, the team commits. Preeti reviews four rounds, not eleven, and she knows round four is the last one before ship, not the start of an open-ended series, so she stays sharp through all of it instead of going numb by round eight.
The tool launches in week five, five weeks ahead of the year-end sale window instead of four weeks behind it. Nine weeks and seven rounds get handed back. Nobody in the sync ever has to reuse an old slide, because there is no round twelve to prep a slide for.
If I am honest, not writing the questions down at kickoff was not the mistake. Anybody would have skipped it, in the excitement of a new project with nothing built yet. The mistake was never going back, five rounds and nine weeks later, and asking whether "make it good" still meant anything specific at all.
FLIPS, worked backward from a slide that did not change
The letters matter less than which one breaks first. Here is the same five steps, mapped onto Faelan's steering committee slide.
"She was too cautious to commit" is a diagnosis anyone can offer after the fact. The harder part is naming the exact thing that was missing from day one, a written question tied to a real decision, and showing there was no smaller fix once eight rounds in a row had already produced the same slide.
And if you want to be sure it really works, try it somewhere else
Baytree Agricultural Co is nowhere near e-commerce or marketing. Its prototype is a handheld crop-disease scanner: a tool field agronomists point at a leaf to check for early blight before it spreads. Same question, a different flip this time. Nobody argues about whether to commit. The pilot users just quietly stop showing up.
F. Sunita Gurung, reliability engineer at Baytree Agricultural Co, eight years diagnosing crop disease by eye in the field before this, who can name a blight from ten feet away.
L. For the first six weeks, all forty pilot agronomists opened the scanner every morning, photographing any leaf that worried them before they walked the row themselves.
I. A different flip from Faelan's. Nobody stops trusting the scanner outright. They just quietly stop opening it, one at a time, with no complaint filed, because the scoring model changed almost every Friday and nobody could learn what a given score meant from one week to the next.
P. The team never froze a version for the pilot. Shipping a fresh model every Friday felt safe, keep improving as fast as possible, nobody gets stuck defending a bad version. It also meant no agronomist ever tested the same tool for more than seven days straight.
S. Freeze one version for a fixed four week window, with two questions written down first: does it catch what a trained eye would miss, and are agronomists still opening it in week four. Both answered by week four, thirty four of forty still opening it daily. The team ships that frozen version as v1 in week five, instead of the real seven months of endless Friday updates that quietly thinned the pilot down to six users.
Swap the trigger and it still runs
- Speed: if Threndle only reviewed the caption tool once a quarter instead of every week, the habit would take longer to form, but the missing stopping questions would still cost them a launch window eventually, just a slower one.
- Cost: if each extra round of the caption tool cost real money, an outside agency billing per round instead of two people's spare days, Faelan's team would have stopped much sooner. It was not cheap testing that caused the drift. It was testing that only felt cheap.
- The model got better: if the caption tool had kept climbing past round four instead of flattening out, the same habit shows up even faster, because "it is still getting better" is an easier reason to keep going than "it might not be ready yet."
Where people run it wrong
- Blaming Faelan for being too cautious, when the real gap was that nobody wrote down what "done" meant.
- Adding a hard deadline with no questions attached, which just moves the guessing to a new date instead of removing it.
- Waiting for a missed launch to force the decision, instead of naming the questions before round one starts.
How to use it live
Buy yourself a few seconds by naming the reframe before the fix: "The real question is not whether the tool is good enough. We could argue about that forever. It is whether anything we are still testing could change what we would decide." Say that, and the rest of the answer is just naming the two questions.
Flashcards (click a card to flip it)
Eight fixed slots, pulled straight from the answer above.
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Prototyping with LLMs and rapid POCs
- #1 What can you learn from a prototype that you cannot learn from a spec?
- #2 Describe how you would build a working prototype of an AI feature in a day.
- #3 What are the risks of a PM prototyping without engineering involvement?
- #4 Explain when a Wizard of Oz prototype beats a real model.
- #5 How do you keep a prototype from setting unrealistic expectations?
- #6 Describe the difference between a demo prototype and a learning prototype.