Describe a prototype test you would run with five real users.
- Pick five testers who cover five different kinds of work, not five people who happen to be free.Why: variety in who tests it is what turns five sessions into five different kinds of proof, not one kind repeated five times.
- Give each tester a real, unfinished piece of their own work, not a demo script.Why: a script only ever breaks the way you wrote it to break; real work breaks the way real work actually breaks.
- Time-box each session and its write-up so five sessions fit inside a day you can also make sense of.Why: a session you can't summarize the same day gets summarized wrong, from memory, days later.
- State the range before promising a number: how many real problems five users could plausibly find, low and high.Why: five users on one kind of work turns up far fewer real problems than five spanning five kinds, and a flat number hides that swing.
- Sanity-check the day against how many sessions you can actually run and write up, not how many you'd like to run.Why: a sixth or seventh session sounds like more proof, but past about five in a day the notes get thinner, not richer.
- Know that who you pick moves the count more than how many you run.Why: a sixth tester in a category already covered buys one problem; swapping one tester into a category nobody's tested yet can buy several.
How to answer this, stage by stage
Nobody is grading whether you land on exactly twelve problems found. They're grading whether you can defend who's in the five chairs, whether the hour math holds up, and whether you close with a number the room can check. Seven moves get you there.
Let's learn
The product is a tool inside a nonprofit's grant system that reads a funder's application questions and drafts real answers, pulled from the org's own past reports and program numbers, instead of a grant writer starting from a blank page every time.
Before a prototype like this gets tested, a team told to "go find five users" usually just asks whoever's in the room. No plan for who those five people are, just five warm bodies and a Friday afternoon. That habit costs almost nothing when the product only ever does one kind of task. It costs the whole test when the tool's job changes shape depending on who's using it, because five people doing the same task mostly find the same three or four things wrong, no matter how carefully each one works.
Here's the turn. Running five sessions is not the hard part. Choosing five people who will actually break the tool in five different ways is. A test with five sessions and one real problem repeated five times teaches you almost nothing you didn't already suspect.
At its worst, that gap shows up after the tool ships to the whole team, not during the test. Everyone hears "we tested it with five real users" and assumes it's ready for anyone, when it was only ever ready for the one kind of grant those five people happened to be writing that week.
The choice I would take back. On the first pass, I picked five testers by convenience: three people in my own office plus two others who happened to be around that week. All five, it turned out, were mid-way through the same annual renewal grant. That felt efficient. It wasn't a real five-user test. It was a one-user test with four extra chairs.
What I would leave alone. If an org genuinely only wrote one kind of grant, five people all testing that one kind would be exactly right, no need to go hunting for variety that doesn't exist. The fix isn't "always spread testers across categories." It's "spread them across however many real categories the work actually has."
The lesson. Five is not a sample size that protects you by itself. It only works if the five people cover five different ways the tool can be asked to do its job. Same headcount, wildly different test, depending only on who's in the five chairs.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why picking the five testers is the whole decision, not a scheduling detail before the real work starts.
Anwen Pryce runs the development team at Milltown Family Services, a nonprofit that runs job training, food assistance, and housing referrals for a mid-size town. Three staff writers report to her, plus two contract writers who come on during grant season. Between the five of them, they carry federal renewals, foundation letters, corporate sponsorship asks, the annual progress report, and the handful of major-donor appeals the board leans on every fall. Anwen has been doing this for eleven years. Ask her which funder wants the budget narrative in a table instead of prose and she'll tell you before you finish the question.
Her executive director asked for a prototype: an assistant that reads a funder's application questions and drafts real answers from the org's own past reports, so nobody starts from a blank page at ten at night. One week, and then show it to the whole team.
Anwen needed five real users to test it properly. She picked the two writers at the next desk, one who was in the break room, and two contract writers who happened to be in that week finishing the same thing she was: the county's annual federal block-grant renewal, due in nine days. It was the fastest way to get five people in front of the tool by Thursday.
The test went well. All five ran a real renewal-grant question through the assistant. It handled the page limits fine, kept the required certification language intact, and only tripped on one small formatting quirk twice. Four problems, total, across five sessions. Anwen wrote them up, fixed the formatting quirk that afternoon, and told the team Monday's meeting would open with a live demo.
Monday came. Anwen opened the demo with the renewal-grant example, and it went perfectly, the way it had all week. Then a colleague who hadn't been part of the test, behind on a foundation letter of inquiry due Wednesday, asked to try it live on her own draft, in front of everyone. The assistant flattened her organization's actual story into three generic sentences about "empowering communities," and repeated the same origin-story paragraph twice, once in the answer to a question that hadn't asked for it.
Nobody blamed the tool out loud. But the room went quiet in the particular way a room goes quiet when something that was supposed to be finished clearly isn't, and the meeting moved on to other business fast.
Anwen ran the test again, properly this time. Five testers, one real category each: the federal renewal, a foundation letter, a corporate sponsorship ask, last year's progress report, and a major-donor appeal. Each one worked through their own unfinished draft for forty-five minutes. She wrote up what broke right after each session, twenty minutes apiece, while it was still fresh.
Federal turned up three problems, the same formatting quirk plus two new ones under real conditions. Foundation turned up two, including the flattened-voice issue from Monday. Corporate turned up three, the sharpest one being that the assistant buried the actual sponsorship ask in paragraph three instead of leading with it. The progress report turned up two, numbers that didn't match what the org had actually reported the year before. The donor letter turned up two, reading like a form letter and dropping the one personal detail the writer had fed it. Twelve distinct problems, from five sessions, in five hours and twenty-five minutes, with two hours left in the day to sort them into what to fix first.
She fixed all twelve over the following week. The next time a colleague pulled up a real, unscripted draft in front of the whole team, foundation, this time, it worked. Not because the tool had gotten smarter on its own. Because twelve real problems had already been found and closed before anyone else's draft could find them live.
The thing I'd tell myself, the week I picked those first five testers: five people in a room feels like proof. It isn't, until you know what one-fifth of the real work they actually represent.
BOUND, counted out before you send five invites
This is a sizing question about how many real problems five sessions can plausibly surface, and what that count actually depends on, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. Five sessions is not the number that matters. What matters is how many different kinds of real input those five sessions carry between them, and how many genuinely different ways the tool can fail inside each kind. A federal grant fails on page limits and required language. A foundation letter fails on voice. A sponsorship ask fails on burying the number. A renewal fails on matching last year's own report. A donor letter fails on sounding like a form. Five testers, five categories, and each category tends to surface two or three failures none of the other categories would ever produce.
O, own the numbers. Five sessions, one tester each, one real category apiece: a federal renewal, a foundation letter of inquiry, a corporate sponsorship ask, last year's progress report, and a major-donor appeal. Each tester brings whatever real, unfinished draft they're already behind on this month, works through it with the tool for forty-five minutes, then spends twenty minutes right after naming everything that felt wrong.
U, use a range. Before running anything, the honest range is wide. If the org's grant work barely varies, five sessions might turn up only four real problems, all versions of the same issue. If the five testers genuinely span five different kinds of grant, the same five sessions can turn up twelve to fourteen. Same headcount, a three-times difference in what you learn, depending only on who's in the five chairs.
N, nail the sanity check. Forty-five minutes plus twenty to write it up is sixty-five minutes a session. Five sessions is five hours twenty-five, inside a seven-and-a-half-hour day, leaving roughly two hours to sit with all twelve problems together and sort them into what to fix first, before anyone forgets which failure came from which draft. Three more sessions chasing categories the org doesn't actually have would cost another three hours, past the day, for almost nothing new.
D, direction. Two assumptions could change this number a lot, and two barely move it. Whether the five testers actually span five categories swings the count by eight. Whether each one gets a real, unfinished draft instead of a script swings it by about five. Adding a sixth session, or three more sessions in categories already covered, moves it by only one or two. Under time pressure, protect who the five people are and what they bring with them, not how many extra sessions you can squeeze in.
And if you want to be sure it really works, try it somewhere else
A public library system prototypes a tool that reads a newly catalogued item and drafts subject headings and a short description for it, instead of a cataloger typing every field from scratch. Five real users, five different branches: children's picture books, academic reference, local special collections, government documents, and general fiction.
B, break it down. Same shape, different work. A children's book fails on missing the reading-level tag. Academic reference fails on the wrong subject-heading standard. Special collections fails on losing a provenance note or misdating a rare item. Government documents fails on the wrong agency-of-origin field. General fiction rarely fails at all, it's the closest thing this system has to an easy category.
O, own the numbers. Five catalogers, one collection type each, each working a real new item off their own shelf for thirty minutes, then ten minutes to name what the drafted heading got wrong.
U, use a range. Five catalogers who all happen to work general fiction might find two or three problems, most of them minor. Five spanning all five collection types found eleven: two in children's books, three in academic reference, three in special collections, two in government documents, one in fiction.
N, nail the sanity check. Thirty minutes plus ten to write up is forty minutes a session. Five sessions is three hours twenty, comfortably inside a working day, with time left over even for a sixth session if a new collection type turns up.
D, direction. Same lever as the grant assistant, a different domain. Which collection types get tested swings the count by roughly nine. Adding extra sessions inside fiction, the easy category, barely moves it, because fiction was already close to solved before the test began.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: pick five testers covering five real categories of the org's actual work, not five convenient people.
Cost: instead of a full day, there's only an afternoon, three hours. Cut the number of categories covered to the three most common kinds of work, not the number of testers, since two testers each doing a category nobody's tested is still worth more than five testers repeating one.
The model got better: a newer model rarely misses federal page limits or certification language anymore. That doesn't remove the need for five testers, it moves which categories are worth double-checking, foundation voice and sponsorship framing become the ones still worth watching closely.
Where people run it wrong.
They recruit five people by convenience instead of by category.
They hand testers a demo script instead of their own real, unfinished draft.
They count sessions run as proof, without checking how many of the findings were actually different from each other.
How to use it live. Say the equation before naming any numbers: "the count a five-user test finds isn't about five, it's about how many different kinds of work those five people represent, so before I give you a number, I'd ask what the five categories actually are." That buys the room to count instead of guessing a figure that sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Prototyping with LLMs and rapid POCs
- #1 What can you learn from a prototype that you cannot learn from a spec?
- #2 Describe how you would build a working prototype of an AI feature in a day.
- #3 What are the risks of a PM prototyping without engineering involvement?
- #4 Explain when a Wizard of Oz prototype beats a real model.
- #5 How do you keep a prototype from setting unrealistic expectations?
- #6 Describe the difference between a demo prototype and a learning prototype.