CaseIntermediateShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #11

Describe a prototype test you would run with five real users.

The direct answer
Run five sessions, but pick the five testers to cover five different kinds of work the product actually has to handle, not five people who were simply free that week. Hand each one a real, unfinished piece of their own work, not a demo script, and watch where the tool trips. Five testers spanning five real categories found twelve distinct problems in early testing; five testers doing the same kind of task found four, because the other eight were the same mistake said five times.
Do this, in order
  1. Pick five testers who cover five different kinds of work, not five people who happen to be free.Why: variety in who tests it is what turns five sessions into five different kinds of proof, not one kind repeated five times.
  2. Give each tester a real, unfinished piece of their own work, not a demo script.Why: a script only ever breaks the way you wrote it to break; real work breaks the way real work actually breaks.
  3. Time-box each session and its write-up so five sessions fit inside a day you can also make sense of.Why: a session you can't summarize the same day gets summarized wrong, from memory, days later.
  4. State the range before promising a number: how many real problems five users could plausibly find, low and high.Why: five users on one kind of work turns up far fewer real problems than five spanning five kinds, and a flat number hides that swing.
  5. Sanity-check the day against how many sessions you can actually run and write up, not how many you'd like to run.Why: a sixth or seventh session sounds like more proof, but past about five in a day the notes get thinner, not richer.
  6. Know that who you pick moves the count more than how many you run.Why: a sixth tester in a category already covered buys one problem; swapping one tester into a category nobody's tested yet can buy several.

How to answer this, stage by stage

Nobody is grading whether you land on exactly twelve problems found. They're grading whether you can defend who's in the five chairs, whether the hour math holds up, and whether you close with a number the room can check. Seven moves get you there.

1
Scope it to a real product and a real definition of "user"
Say it like this
"Let's ground this in one thing. Say we're building a tool inside a nonprofit's grant system that reads a funder's application questions and drafts real answers from the org's past reports. Five real users means five people who actually write grants there, not five people who happened to answer my email first."
Why this works
Stops the answer from staying abstract before a single session gets booked.
2
Say the structure out loud before naming any numbers
Say it like this
"The number that matters here isn't five sessions. It's how many different kinds of work those five people bring with them, and how many genuinely different ways the tool can fail inside each kind. Five people doing the same kind of task mostly find the same mistake, five times over."
Why this works
Shows the equation before the arithmetic, so the numbers that follow read as a plan, not a guess.
3
Reframe what a five-user test is actually testing
Say it like this
"This isn't really 'did I run five user tests.' It's 'did those five people, between them, break the tool in five different ways.' A test that only ever proves the tool works on one kind of grant hasn't been tested with five real users. It's been tested with one, five times."
Why this works
Separates a real five-user test from five people doing the same thing in different chairs.
4
Own the real test structure
Say it like this
"Here's the structure I'd actually run: five sessions, one tester each, one real category apiece: a federal renewal, a foundation letter of inquiry, a corporate sponsorship ask, last year's progress report, and a major-donor appeal. Each one works through their own real, unfinished draft for forty-five minutes, then spends twenty minutes with me right after naming everything that felt wrong."
Why this works
A real proposed structure, not "get five users in a room," is what the O step of BOUND actually asks for.
5
Give the range, not one fixed number
Say it like this
"If all five testers happen to be mid-way through the same kind of grant, five sessions might only turn up four real problems, all the same one said differently. If the five genuinely span five different kinds of grant, the same five sessions can turn up twelve to fourteen. Same five people on paper, three times the finding, depending only on who's actually in the room."
Why this works
A single number here would claim confidence the answer doesn't have.
6
Sanity check the day against real sessions and real write-up time
Say it like this
"Forty-five minutes a session plus twenty to write up what broke, right after, while it's fresh, is sixty-five minutes a session. Five sessions is five hours twenty-five, inside a seven-and-a-half-hour day, which leaves about two hours to actually sort twelve problems into what gets fixed first. Add three more sessions chasing categories the org doesn't have and that's another three hours, past the day, for almost nothing new."
Why this works
This is the step most five-user plans skip, and it's what turns "I ran five sessions" into "I know what to fix Monday morning."
7
Name what moves it most, then close on the one line
Say it like this
"If I had to bet on what changes this number most, it's not the session count. It's whether the five people actually cover five different kinds of work. That alone swings the count by eight. So here's the line I'd actually say: run five sessions, but spend the time before any of them choosing five testers who cover five real categories of the work, hand each one their own unfinished draft, and count how many genuinely different problems come back, not how many sessions you ran."
Why this works
Closes on the literal ask, a test someone could actually check, not a vibe about "getting user feedback."
If you remember one thing Pick the five testers to cover five real categories of the work, and give each one their own unfinished draft. That's what turns five sessions into five different kinds of proof instead of one kind found five times.

Let's learn

The product is a tool inside a nonprofit's grant system that reads a funder's application questions and drafts real answers, pulled from the org's own past reports and program numbers, instead of a grant writer starting from a blank page every time.

Before a prototype like this gets tested, a team told to "go find five users" usually just asks whoever's in the room. No plan for who those five people are, just five warm bodies and a Friday afternoon. That habit costs almost nothing when the product only ever does one kind of task. It costs the whole test when the tool's job changes shape depending on who's using it, because five people doing the same task mostly find the same three or four things wrong, no matter how carefully each one works.

Knowledge spark: what's a distinct failure pattern? A specific way the tool gets something wrong that's different from the ways it's already gotten something wrong. Not the same mistake showing up twice with a different name on it.

Here's the turn. Running five sessions is not the hard part. Choosing five people who will actually break the tool in five different ways is. A test with five sessions and one real problem repeated five times teaches you almost nothing you didn't already suspect.

Five sessions on one kind of grant found four problems. Five sessions on five kinds of grant found twelve.

At its worst, that gap shows up after the tool ships to the whole team, not during the test. Everyone hears "we tested it with five real users" and assumes it's ready for anyone, when it was only ever ready for the one kind of grant those five people happened to be writing that week.

The decision that mattered Pick the five testers to cover the five real categories of grant work the team actually does, not the five people who were free. Variety in who tests it, not headcount, is what a five-user test is actually buying you.

The choice I would take back. On the first pass, I picked five testers by convenience: three people in my own office plus two others who happened to be around that week. All five, it turned out, were mid-way through the same annual renewal grant. That felt efficient. It wasn't a real five-user test. It was a one-user test with four extra chairs.

What I would leave alone. If an org genuinely only wrote one kind of grant, five people all testing that one kind would be exactly right, no need to go hunting for variety that doesn't exist. The fix isn't "always spread testers across categories." It's "spread them across however many real categories the work actually has."

The lesson. Five is not a sample size that protects you by itself. It only works if the five people cover five different ways the tool can be asked to do its job. Same headcount, wildly different test, depending only on who's in the five chairs.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why picking the five testers is the whole decision, not a scheduling detail before the real work starts.

Anwen Pryce runs the development team at Milltown Family Services, a nonprofit that runs job training, food assistance, and housing referrals for a mid-size town. Three staff writers report to her, plus two contract writers who come on during grant season. Between the five of them, they carry federal renewals, foundation letters, corporate sponsorship asks, the annual progress report, and the handful of major-donor appeals the board leans on every fall. Anwen has been doing this for eleven years. Ask her which funder wants the budget narrative in a table instead of prose and she'll tell you before you finish the question.

Her executive director asked for a prototype: an assistant that reads a funder's application questions and drafts real answers from the org's own past reports, so nobody starts from a blank page at ten at night. One week, and then show it to the whole team.

Anwen needed five real users to test it properly. She picked the two writers at the next desk, one who was in the break room, and two contract writers who happened to be in that week finishing the same thing she was: the county's annual federal block-grant renewal, due in nine days. It was the fastest way to get five people in front of the tool by Thursday.

The test went well. All five ran a real renewal-grant question through the assistant. It handled the page limits fine, kept the required certification language intact, and only tripped on one small formatting quirk twice. Four problems, total, across five sessions. Anwen wrote them up, fixed the formatting quirk that afternoon, and told the team Monday's meeting would open with a live demo.

Hand-sketched comparison. Left panel, five grey identical folder icons each labeled Federal renewal, captioned 4 problems found. Right panel, five green folder icons labeled Federal, Foundation, Corporate, Renewal, and Donor letter, captioned 12 problems found.
Same five chairs, two different tests. One buys proof for a fifth of the work. The other buys proof for all of it.

Monday came. Anwen opened the demo with the renewal-grant example, and it went perfectly, the way it had all week. Then a colleague who hadn't been part of the test, behind on a foundation letter of inquiry due Wednesday, asked to try it live on her own draft, in front of everyone. The assistant flattened her organization's actual story into three generic sentences about "empowering communities," and repeated the same origin-story paragraph twice, once in the answer to a question that hadn't asked for it.

We didn't build a bad assistant. We tested it on a fifth of the work and called it done.

Nobody blamed the tool out loud. But the room went quiet in the particular way a room goes quiet when something that was supposed to be finished clearly isn't, and the meeting moved on to other business fast.

Anwen ran the test again, properly this time. Five testers, one real category each: the federal renewal, a foundation letter, a corporate sponsorship ask, last year's progress report, and a major-donor appeal. Each one worked through their own unfinished draft for forty-five minutes. She wrote up what broke right after each session, twenty minutes apiece, while it was still fresh.

Federal turned up three problems, the same formatting quirk plus two new ones under real conditions. Foundation turned up two, including the flattened-voice issue from Monday. Corporate turned up three, the sharpest one being that the assistant buried the actual sponsorship ask in paragraph three instead of leading with it. The progress report turned up two, numbers that didn't match what the org had actually reported the year before. The donor letter turned up two, reading like a form letter and dropping the one personal detail the writer had fed it. Twelve distinct problems, from five sessions, in five hours and twenty-five minutes, with two hours left in the day to sort them into what to fix first.

She fixed all twelve over the following week. The next time a colleague pulled up a real, unscripted draft in front of the whole team, foundation, this time, it worked. Not because the tool had gotten smarter on its own. Because twelve real problems had already been found and closed before anyone else's draft could find them live.

The thing I'd tell myself, the week I picked those first five testers: five people in a room feels like proof. It isn't, until you know what one-fifth of the real work they actually represent.

BOUND, counted out before you send five invites

This is a sizing question about how many real problems five sessions can plausibly surface, and what that count actually depends on, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. Five sessions is not the number that matters. What matters is how many different kinds of real input those five sessions carry between them, and how many genuinely different ways the tool can fail inside each kind. A federal grant fails on page limits and required language. A foundation letter fails on voice. A sponsorship ask fails on burying the number. A renewal fails on matching last year's own report. A donor letter fails on sounding like a form. Five testers, five categories, and each category tends to surface two or three failures none of the other categories would ever produce.
O, own the numbers. Five sessions, one tester each, one real category apiece: a federal renewal, a foundation letter of inquiry, a corporate sponsorship ask, last year's progress report, and a major-donor appeal. Each tester brings whatever real, unfinished draft they're already behind on this month, works through it with the tool for forty-five minutes, then spends twenty minutes right after naming everything that felt wrong.
U, use a range. Before running anything, the honest range is wide. If the org's grant work barely varies, five sessions might turn up only four real problems, all versions of the same issue. If the five testers genuinely span five different kinds of grant, the same five sessions can turn up twelve to fourteen. Same headcount, a three-times difference in what you learn, depending only on who's in the five chairs.
N, nail the sanity check. Forty-five minutes plus twenty to write it up is sixty-five minutes a session. Five sessions is five hours twenty-five, inside a seven-and-a-half-hour day, leaving roughly two hours to sit with all twelve problems together and sort them into what to fix first, before anyone forgets which failure came from which draft. Three more sessions chasing categories the org doesn't actually have would cost another three hours, past the day, for almost nothing new.
D, direction. Two assumptions could change this number a lot, and two barely move it. Whether the five testers actually span five categories swings the count by eight. Whether each one gets a real, unfinished draft instead of a script swings it by about five. Adding a sixth session, or three more sessions in categories already covered, moves it by only one or two. Under time pressure, protect who the five people are and what they bring with them, not how many extra sessions you can squeeze in.

The build-up: five testers, five categories, twelve problems
Federal renewal3
+ Foundation letter5
+ Corporate sponsorship ask8
+ Progress report renewal10
+ Major-donor appeal12
Each of the five real categories adds two or three problems none of the others would have found. That's where the twelve comes from, not from running five sessions as such.
What moves the total most (swing in distinct problems found)
All five testers share one category instead of five−8
Each tester gets a demo script instead of a real draft−5
A sixth tester added in a brand-new category+2
Three more sessions, same five categories repeated+1
Who the five testers are and what they bring swings the count by five to eight. Adding more sessions on top of the same five categories barely moves it at all.

And if you want to be sure it really works, try it somewhere else

A public library system prototypes a tool that reads a newly catalogued item and drafts subject headings and a short description for it, instead of a cataloger typing every field from scratch. Five real users, five different branches: children's picture books, academic reference, local special collections, government documents, and general fiction.

B, break it down. Same shape, different work. A children's book fails on missing the reading-level tag. Academic reference fails on the wrong subject-heading standard. Special collections fails on losing a provenance note or misdating a rare item. Government documents fails on the wrong agency-of-origin field. General fiction rarely fails at all, it's the closest thing this system has to an easy category.
O, own the numbers. Five catalogers, one collection type each, each working a real new item off their own shelf for thirty minutes, then ten minutes to name what the drafted heading got wrong.
U, use a range. Five catalogers who all happen to work general fiction might find two or three problems, most of them minor. Five spanning all five collection types found eleven: two in children's books, three in academic reference, three in special collections, two in government documents, one in fiction.
N, nail the sanity check. Thirty minutes plus ten to write up is forty minutes a session. Five sessions is three hours twenty, comfortably inside a working day, with time left over even for a sixth session if a new collection type turns up.
D, direction. Same lever as the grant assistant, a different domain. Which collection types get tested swings the count by roughly nine. Adding extra sessions inside fiction, the easy category, barely moves it, because fiction was already close to solved before the test began.

Same shape, different lever At Milltown, the thing that decided the test was whether the five testers spanned five kinds of grant. At the library, the collection type barely changes the shape of the argument at all: the count still lives or dies on whether the work being tested is actually spread across what the system has to handle, not on how many sessions get booked.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: pick five testers covering five real categories of the org's actual work, not five convenient people.
Cost: instead of a full day, there's only an afternoon, three hours. Cut the number of categories covered to the three most common kinds of work, not the number of testers, since two testers each doing a category nobody's tested is still worth more than five testers repeating one.
The model got better: a newer model rarely misses federal page limits or certification language anymore. That doesn't remove the need for five testers, it moves which categories are worth double-checking, foundation voice and sponsorship framing become the ones still worth watching closely.

Where people run it wrong.
They recruit five people by convenience instead of by category.
They hand testers a demo script instead of their own real, unfinished draft.
They count sessions run as proof, without checking how many of the findings were actually different from each other.

How to use it live. Say the equation before naming any numbers: "the count a five-user test finds isn't about five, it's about how many different kinds of work those five people represent, so before I give you a number, I'd ask what the five categories actually are." That buys the room to count instead of guessing a figure that sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a "test with five real users" question, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question, how many real problems five sessions can plausibly surface, and what that number depends on, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Anwen Pryce, who runs the development team at Milltown Family Services, a nonprofit running job training, food assistance, and housing referrals. Eleven years writing grants there.
3 · THE HABIT THAT COST HER
What did Anwen do on her first user test that limited what it found?
Tap to flip
ANSWER
She picked five testers by convenience, three colleagues nearby plus two contract writers around that week, and all five happened to be working the same federal renewal grant, so the test only found problems that one category has.
4 · THE BUILD-UP, IN THIS STORY
What's the five-session structure this answer turns on?
Tap to flip
ANSWER
Five testers, one real category each: federal renewal, foundation letter, corporate sponsorship ask, progress report, major-donor appeal. Forty-five minutes working plus twenty minutes write-up, per session.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Choosing five testers by who was nearby and available instead of by which category of work they'd bring. It felt efficient, and it was the fastest way to get five people in front of the tool by Thursday.
6 · THE NUMBER
Fill in the blank: five testers covering five real categories found ___ distinct problems; five testers all doing the same category found ___.
Tap to flip
ANSWER
Twelve distinct problems across five categories. Four when all five testers shared the same category.
7 · THE REPLAY
Same live demo, a colleague pastes in a real draft, second attempt. What changes?
Tap to flip
ANSWER
The twelve real problems, including the flattened-voice issue, were already found and fixed the week before. The colleague's live foundation-letter draft works, not because the model improved on its own, but because it had already been tested against that category.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the lever that swings its count?
Tap to flip
ANSWER
A public library's subject-heading drafting tool, tested across five branch collection types. The lever is the same: which collection types get tested, not how many sessions get booked.

Check yourself Score: 0 / 0

Fill in the blank
1. Anwen's first test used five testers who were all mid-way through the same ___ grant, which is why it only surfaced ___ distinct problems instead of the twelve the second test found.
Show hint
Look at what category all five testers in the first attempt shared, and check the number right before the redo.
Show answer
Federal renewal; 4. All five testers in the first pass were working the same annual federal block-grant renewal, so the test only ever found that category's problems.
Multiple choice
2. Why does which five categories get tested move the total problem count more than how many sessions get run?
  • A. Sessions take too long to schedule, so extra ones rarely happen anyway.
  • B. A new session in a category already covered mostly repeats problems already found, while a session in a brand-new category tends to surface genuinely new ones.
  • C. Five is a fixed rule for user testing that can never be exceeded.
  • D. Problems found in a session get double-counted automatically, so extra sessions inflate the number for free.
Show hint
Compare the size of the swing in the sensitivity chart's top two rows against the bottom two.
Show answer
B. Category variety swings the count by five to eight. Adding sessions on top of categories already covered only moves it by one or two.
True or false
3. True or false: the assistant broke in front of the whole team because the model itself got a fact wrong.
  • True
  • False
Show hint
Check what category the five original testers covered, versus what the colleague's live draft was.
Show answer
False. It broke because it had never been tested against a foundation letter, a category none of the first five testers touched. The model wasn't so much wrong as untested on that kind of work.
Short answer, the number question
4. If Milltown's five testers had covered six categories instead of five, say adding event-sponsorship letters worth about two more distinct problems, roughly how many total problems might the test have found, and would that change the D-step conclusion?
Show hint
Add the extra category's problems to the twelve already found, then check whether it changes which assumption matters most.
Show answer
Roughly 14 distinct problems. Twelve plus about two more from the new category, still inside the predicted four-to-fourteen range. The D-step conclusion wouldn't change: which categories get covered still swings the count more than adding another session would.
Short answer, apply it yourself
5. Think of a product you'd want to test with five real users. What five different kinds of task or five different kinds of user would you want those five people to represent, instead of five people who are simply easy to reach?
Show hint
Name five real categories of the work, not five names, and check whether the people you'd normally grab actually cover them.
Show answer
Model answer: "A small clinic's intake-form assistant. I'd want five testers covering five real intake types: a first-time patient, a returning patient with an update, a non-English speaker using the translated form, a parent filling it in for a child, and someone entering it from a phone in the waiting room. Five front-desk staff who all handle the same walk-in type would miss four of those five failure modes entirely."
Short answer
6. Anwen's five sessions took about five hours twenty-five minutes total, inside a workday of roughly seven and a half hours. Why does that leftover two hours matter, and what happens if a manager pushes her to add three more sessions "to be thorough"?
Show hint
Think about what the leftover time actually gets used for, and check the sensitivity chart's row for extra sessions in categories already covered.
Show answer
Model answer: The two hours left is what lets her sort the twelve problems into what to fix first, the same day, while she still remembers which failure came from which draft. Three more sessions repeating categories she's already covered would cost close to three more hours, blow past the day, and by the numbers above buy only about one more distinct problem, trading real synthesis time for almost nothing new found.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more