ConceptIntermediateShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #13

What is the right fidelity for an AI prototype at the discovery stage?

The direct answer
Build the low fidelity version first: a rough, mostly faked calendar with captions a person wrote by hand to look right. But also run the real model on a handful of real examples for two days, before anyone signs anything real. Skip that quick check, and a client can fall in love with a demo the model can never actually make.
How to pick a prototype's fidelity at discovery, in order
  1. Build the low-fidelity mockup first, faked captions a person wrote, not the real model.Why: it tests whether anyone wants the idea, cheaply, before a single line of real integration gets built.
  2. Run a two-day scripted test of the real model on real past posts before treating the mockup's yes as proof of anything.Why: only the real model, on real examples, can show it can hit a client's actual voice.
  3. Name who feels each kind of miss before you pick.Why: a wrong layout gets caught in the room; a wrong voice gets caught only after a client signs.
  4. Set the kill line before you build anything real.Why: without a number, "prove the model works" quietly turns into "build it and find out."
  5. Save the full integration for after both checks pass.Why: five weeks of engineering time is the expensive way to learn something a two-day test would have told you for free.
  6. Leave the interface fidelity low for as long as it still teaches you something.Why: the calendar layout and the click flow don't need to be real to learn if people want them.

How to answer this, stage by stage

This is a yes-or-no about which question to test first, not a rulebook for every prototype, so PICK does the work here.

1
Scope it to one real decision
Say it like this
"Let's make this real. Say Hollis Creative, a marketing agency, wants to build Calendra, a tool that reads a client's brand guide and past posts and writes a month of captions. Before anyone builds anything real, I need to decide what to show a client first."
Why this works
Stops the answer floating at a dictionary definition of "fidelity" and gives the interviewer one real product to push on.
2
Say your structure out loud
Say it like this
"I'll pick a fidelity first, then say who feels it if I'm wrong on each side, then say which kind of miss is worse, then say what would change my mind. That's PICK, and I'll go in that order."
Why this works
Signals a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Reframe what the question is really testing
Say it like this
"This isn't really 'how polished should the first version look.' It's 'which of two different things am I actually trying to learn here, whether people want this, or whether the model can do it,' because those two questions need very different amounts of real."
Why this works
Shows the interviewer you see past the surface ask to the real judgment being tested.
4
State the position, with the real mechanism in it
Say it like this
"My pick: build the shell low fidelity, faked captions a copywriter wrote by hand, four days of work. But also run the real model, no interface at all, against ten of the client's actual past posts, two more days. Skip that second part, and you've only tested the idea, never the model."
Why this works
PICK rewards a real mechanism and a real number, not a vague promise to "test more."
5
Name who feels each kind of miss, with a real case behind it
Say it like this
"Here's the split. Say the mockup's calendar layout is wrong. The design lead sees it in the room, same afternoon, and redraws it. Now say we skip the model check instead. A client signs an add-on off the mockup. Five weeks later the real captions come back sounding like nobody at that company, and it takes a full billing cycle before anyone official even flags it."
Why this works
A real case turns "asymmetry" from a word into something the interviewer can picture happening to an actual client.
6
Say what you'd leave alone
Say it like this
"I wouldn't fake the model on that ten-post test. No copywriter pretending to be Calendra can tell you if the real model can do this job. But I'd happily fake the whole calendar interface, the drag and drop, the approval screen. None of that needs to be real to learn if a client wants it."
Why this works
Shows judgment instead of treating every part of a prototype as needing the same amount of real.
7
Name the kill criteria and close on one sentence
Say it like this
"I'd flip to the full five-week build once the real model hits at least eighty percent of sample captions passing a brand-voice check without a heavy rewrite, checked twice on two different batches of real posts. Below that line, five weeks of integration answers a question nobody's confirmed the model can pass yet."
Why this works
Ends on the line the interviewer remembers, and shows the pick isn't permanent, it's a decision that can be revisited with evidence.

A last note before the walkthrough ends: this pick is about order, not a ban on ever going real. Most candidates hear "low fidelity or high fidelity" and answer like it's permanent. Name the one thing only a real model can prove, and you've shown judgment instead of reciting a definition.

Let's learn

Calendra is a tool a marketing agency wants to build. It would read a client's brand guide and their last few months of posts, then write a full month of social captions and image briefs onto a calendar.

Knowledge spark: what does fidelity mean for a prototype? How close a test version is to the real thing. Low fidelity means most of it is faked, sometimes by a person quietly doing the work behind the screen. High fidelity means the real model, and the real hookups to real systems, are actually running.

Before Calendra, a strategist at Hollis Creative built one client's month of captions and image briefs by hand. It took about 14 hours, and every caption read exactly like that client, because a person who knew the brand wrote every line.

The agency tested Calendra a different way first. A copywriter hand wrote twenty example captions, in four days, and a designer dropped them into a calendar mockup. No model touched any of it. Three client teams saw the mockup. All three loved it. One, Cortland Coffee Roasters, signed a new retainer add-on the same week, off the mockup alone.

Here is the turn. That signed retainer is not proof the tool works. It only proves people want the idea. Nobody had checked yet whether the real model could actually write a caption that sounded like Cortland, because there still wasn't a real model anywhere near the demo.

Days before the team knew if Calendra could really do the job
0 days 55 days Mockup miss, caught in the room Model miss, caught when the client cancels
The green bar is small on purpose: a wrong layout in a faked mockup shows up the moment someone in the room reacts to it, and it costs an afternoon to fix. The red bar is what skipping the model check actually cost: 25 days to build the real integration, then 30 more days, one full billing cycle, before Cortland canceled the retainer the mockup had just sold.
Cortland didn't sign because Calendra worked. Cortland signed because a person wrote captions that sounded like Cortland.

At its worst, the team skips the model check entirely, builds the real integration straight off the mockup's good reaction, and finds out only after a client has already signed, already gone live, and already started noticing the captions sound like nobody they know.

The choice I would take back Hollis Creative skipped a two-day step: run the real model, no interface at all, against ten of a client's real past posts, before promising anything. I would put that step back in. Two days, against five weeks of building the wrong thing first.

What I would leave alone. The calendar interface itself, the drag and drop, the approval screen. None of that needs to be real for a client to say yes to the idea. Fake it freely, it costs nothing and teaches plenty.

The lesson. A prototype answers one question at a time. A mockup answers "does anyone want this." Only the real model, on real examples, answers "can it actually do this." Fake the first and skip the second, and you can end up selling something you never checked exists.

Two panels side by side. Left, a laptop screen showing a calendar grid, labeled what Cortland saw, Calendra's calendar. Right, a hand-drawn person at a desk with a sheet of paper, labeled what was really there, a copywriter, not the model.
What the client saw, and what was actually behind it

The mockup that sold itself

You don't need this to answer the question. Read it if you want to feel why the split has to happen on a real test, not on how good a meeting went.

Zinnia Rooksby can read a caption for four seconds and tell you whether it sounds like the client or like nobody at all. Five years running growth accounts at Hollis Creative taught her that trick, long before Calendra ever existed.

Cortland Coffee Roasters has been one of her accounts for three of those years, a small wholesale coffee brand with a voice so specific the team keeps rules for it. No exclamation marks, ever. Always call the owner Dev, never "the founder." A joke about the roaster's burnt first batch, worked into a caption at least once a month.

Hollis Creative wanted to sell a new kind of retainer: Calendra, a tool that would read a client's brand guide and their last ninety days of posts, then draft a full month of captions and image briefs onto a calendar. Every client account at once, not just Cortland. If it worked, it was the biggest thing the agency had shipped in years.

Before any of that got built, Zinnia had to decide what to actually show a client first.

She picked the cheap route. A copywriter sat down and hand wrote twenty captions for Cortland, the kind Calendra was supposed to make on its own, and a designer dropped them into a calendar mockup, four days of work. No model touched any of it. It just had to look like one had.

Three client teams saw the mockup that week, Cortland among them. All three loved it. Cortland's marketing lead signed the retainer add-on before the meeting even ended.

Leadership heard that and did the obvious thing. Skip more testing, they said, the clients already said yes, just build it. Engineering scoped the real integration, reading a live brand kit, pulling ninety days of real posts, writing drafts straight into the client's scheduling tool, at five weeks before a single real caption came out the other end.

Five weeks later, it did.

The team ran the real model against Cortland's actual brand guide and its last ninety posts for the first time. Ten sample captions came back. Nine needed a full rewrite. One used an exclamation mark. Three called Dev "the founder." None mentioned the burnt first batch, not once.

Nobody in the room had checked, before those five weeks, whether the model could even do this job. They had checked whether a client wanted it. Those turned out to be two completely different questions, and only one of them had ever gotten an answer.

The retainer went live anyway, patched captions and all. Cortland's marketing lead flagged the tone within the first week. One full billing cycle later, thirty days, Cortland canceled, not just the new add-on, the whole retainer, three years, gone.

We didn't lose nine bad captions. We lost three years of Cortland trusting us to sound like them.
Two boxes of unequal size. A small plain box labeled mockup miss, wrong layout, caught same day, felt by the design lead in the room. A much larger jagged red-orange shape labeled model miss, found on day 55, client cancels, felt by the whole account weeks later.
Same product, two very different sizes of wrong

I want to say the mistake was believing the mockup. It wasn't. The mockup did exactly its job: it proved Cortland wanted a tool like this. The mistake was letting that proof stand in for a different proof it was never built to give.

So here is the choice I'd take back.

When we scoped the pilot, we skipped a two-day step: run the real model, no interface, against ten of Cortland's actual past posts, and have a strategist grade the captions before anyone signs anything. Two days. We had the time. We just didn't think we needed it, because the mockup had already gone so well.

I'd put that step back in. Fake the calendar, fake the drag and drop, fake the whole screen, none of it needs to be real for a client to say yes to the idea. But the ten captions that tell you whether the model can actually do the job, those have to come from the model, not from someone hiding behind it.

And the thing I'd tell myself, back in that scoping meeting: a client loving the idea and a model being able to do the idea are two different yeses. We only ever asked for one of them.

PICK, spelled out for a coffee client

This is a yes-or-no about which question to test first, not a rule for every prototype Hollis Creative ever builds, so PICK carries the weight here.

P, position. Fake the interface first, a hand-written calendar mockup. But run the real model, no interface, against ten of the client's real past posts, before anyone signs a retainer built on the idea.
I, impact. A wrong layout in the mockup is felt by the design lead, caught in the room, fixed that afternoon. A model that can't hit a client's brand voice is felt by the whole account, caught only after a client signs, goes live, and starts noticing, weeks later.
C, cost asymmetry. The mockup miss is cheap and visible. It shows up the moment someone in the room reacts to it, and it costs an afternoon to fix. The model miss is hidden and expensive. The model sounds just as sure writing a caption that misses the brand as one that nails it, so nothing flags it until a signed client cancels.
K, kill criteria. Flip to the full build once the real model hits at least 80 percent of sample captions passing a brand-voice check without a heavy rewrite, checked on two separate batches of at least 30 real posts each. Below that line, five weeks of integration answers a question nobody's confirmed the model can pass.
Knowledge spark: why not just skip the mockup and ask the client if they'd want this? Because a client can't picture something that doesn't exist yet from a question in a meeting. Seeing captions that sound like their own brand, even faked ones, tells you far more than a survey ever could. The mockup earns its four days.
Sample captions passing Cortland's brand-voice check, by real posts tested
Share passing without a full rewrite
Kill line: 80%, held on a second batch
0% 50% 100% 80% kill line 18 32 44 58 68 76 81 (held) 83 85 5 posts 15 posts 25 posts 35 posts 45 posts
The pass rate climbs steadily from 18 percent on 5 real posts to 81 percent on 35 posts, crossing the 80 percent kill line for the first time, then holds above it at 40 and 45 posts on a second batch. That two-batch hold, not a date on the project plan, is the real signal the model was ready for the full build.

Try the same four letters on a newsroom pitch

Bertrand Kowalczyk edits the Mossbank Weekly Gazette, a small paper that covers one town's council meetings, school board, and little else. He wants to build Lineup, a tool that reads the week's wire feed and the council's meeting minutes, then drafts three story pitches for the paper's freelance stringers to choose from.

Before wiring Lineup to a live feed, Bertrand tests it the cheap way. He hand-picks three real past weeks of council minutes and writes, himself, the three pitches he thinks a good tool would have found. He hands the fake pitch list to two stringers and watches which one they pick.

P. Fake the pitch list first, real minutes, pitches written by a person. Keep the real model out until a separate, narrow test proves it can find a real angle in real minutes without inventing one.
I. A weak fake pitch is felt by a stringer, who shrugs and picks a different one off the list, no cost. A live model's confident wrong pitch, telling a stringer a routine budget line is a scandal, is felt by the stringer's whole week chasing a story that was never there.
C. The fake-pitch miss is cheap and visible, a stringer notices in five minutes and moves on. The live-model miss is hidden and expensive, a stringer can spend a full week reporting a pitch out before an editor catches, at copy desk, that the story doesn't exist.
K. Flip to the real feed once the model finds a real, checkable angle in at least 8 of 10 past weeks of minutes, tested blind, with an editor grading each one against what actually happened.

What I would leave alone, at the Gazette The layout of the pitch list itself, whether it's an email, a shared doc, or a one-pager, stays fake as long as needed. That's a question about how stringers like to receive work, and it has nothing to do with whether the model can read a council agenda correctly.

Swap the trigger and it still runs

  • Speed: if Hollis Creative needed Calendra live in two weeks instead of a full quarter, the pick doesn't move, a two-day model test fits inside two weeks and a five-week integration never does.
  • Cost: if the real model turned out to cost more per caption at scale than expected, the pick still doesn't move, the two-day test was never about running cost, it was about whether the model could do the job at all.
  • The model got better: if a model already reliably matched a specific brand's voice with no extra tuning, proven on ten real posts, that's exactly the evidence that flips the pick straight to the full build.

Where people run it wrong

  • Skipping the model check once the mockup gets a yes, so nobody learns if the model can do the job until it's already been sold to a client.
  • Making the mockup itself too real, wiring in a live model early, so nobody can tell if a client is reacting to the idea or to one lucky caption.
  • Running the model check once and calling it proven, with no second batch, so a lucky sample gets mistaken for a real pattern.

Buy yourself two seconds, out loud

Say the reframe before you answer with a definition. "Give me a second, I want to separate what we're actually testing here, whether anyone wants this, or whether the model can do it." That's true, it's already stage three of the walkthrough, and it buys you the time to find the real split instead of reciting a textbook line.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's the hardest step to nail?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a hidden brand-voice miss costs more than a mockup miss caught in the room.
2 · THE PERSON
Who decides what to build first at Hollis Creative, and what has she done for five years?
Tap to flip
ANSWER
Zinnia Rooksby, the senior product manager who has run growth accounts at Hollis Creative for five years.
3 · THE HABIT
What step did the team skip once the mockup landed so well with three clients?
Tap to flip
ANSWER
The two-day scripted test of the real model on ten real posts. Leadership went straight to the five-week integration build instead, since the mockup had already gotten a yes.
4 · THE ASYMMETRY
Name the two kinds of miss here and what each one costs.
Tap to flip
ANSWER
A mockup miss: a wrong layout, caught by the design lead, fixed the same afternoon. A model miss: captions off Cortland's brand voice, caught only 55 days later, when the client canceled a signed retainer.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Fake the interface first, but run the real model on real posts for two days before anyone signs anything real off the idea.
6 · THE NUMBER
Of the first ten sample captions the real model wrote for Cortland, ______ needed a full rewrite.
Tap to flip
ANSWER
9. Only one of the model's first ten captions was usable as written, the exact test the team skipped before promising Cortland anything.
7 · THE KILL CRITERIA
What evidence would flip this pick to the full five-week build?
Tap to flip
ANSWER
The real model hitting at least 80 percent of sample captions passing a brand-voice check, checked on two separate batches of at least 30 real posts each. Below that, the model isn't proven yet.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the position land there?
Tap to flip
ANSWER
Lineup, a story-pitch tool at the Mossbank Weekly Gazette. The fake pitch list stays until the real model finds a checkable angle in at least 8 of 10 tested weeks of real minutes.

Check yourself Score: 0 / 0

True or false
1. True or false: this position means Hollis Creative should never build the real Calendra integration.
  • True
  • False
Show hint
Think about what the kill line is actually for.
Show answer
False. The position is about order and evidence: fake the interface first, then run the real model on real examples, then build the full integration once both checks pass, not never build one at all.
Multiple choice
2. Which of these is the actual mechanism behind this answer's pick?
  • A. Skip the mockup and build the real integration straight away, since it's needed eventually anyway.
  • B. Fake the interface with hand-written captions, but also run the real model on ten real posts for two days before anyone signs anything.
  • C. Build the mockup and the real integration side by side and compare them every week.
  • D. Ask clients in a survey whether they'd trust an AI captions tool, instead of building anything.
Show hint
Three of these either skip the model question entirely or never produce evidence anyone could act on.
Show answer
B. A never checks if the model can do the job. C spends the five weeks anyway, on top of the mockup. D never touches a real caption. Only B tests both questions cheaply, in order.
Fill in the blank
3. Fill in the blank: before Calendra, a strategist built one client's month of captions and image briefs by hand, and it took about ______ hours.
Show hint
It's the number from the old, fully manual process, before any mockup or model existed.
Show answer
14. Fourteen hours by hand was the baseline the whole pitch for Calendra was trying to beat, and it's the number that made the mockup's four days look so appealing on its own.
Multiple choice
4. Why not just skip the two-day model check and build the full integration, since the mockup already got a yes from three clients?
  • A. Because engineers are more expensive than copywriters, so cost isn't really the reason.
  • B. Because a mockup only proves someone wants the idea; it says nothing about whether the real model can hit a specific client's brand voice.
  • C. Because Cortland's contract legally required a two-day test before launch.
  • D. Because the real model technically cannot write captions until the interface is finished.
Show hint
Think about who actually learns something during the five weeks a real integration takes to build, and what they learn.
Show answer
B. Building the integration first spends real engineering weeks answering a technical question before anyone's checked whether the model itself, not a person faking it, can write on-brand captions at all.
Short answer
5. If Cortland's cancellation had landed 10 days after go-live instead of a full 30-day billing cycle later, would the same position still hold? Walk through it.
Show hint
Think about whether the fix protects against the exact number of days lost, or against the fact that nothing in a live model catches its own confident, wrong caption before a client notices.
Show answer
Yes, still worth it. The 55 days is evidence of how expensive the gap is, not the reason it exists. Even at 10 days, the real problem stands: nothing in the model's own design would catch a wrong caption before a client acted on it. The day count changes how loud the alarm should ring. It doesn't decide whether the two-day check needed doing.
Short answer, apply it yourself
6. Pick a product you use yourself. Name one question about it that a low-fidelity fake could safely test first, and one that only the real model could.
Show hint
Look for the question that's really about whether you'd want the thing at all, versus the question that's really about whether a model can technically do the specific task.
Show answer
Model answer: "A grocery app's new meal-planning feature. Whether shoppers actually want a week of meals suggested to them at all is something a few hand-picked, human-written sample plans could test in a day. Whether the model can build a plan that actually fits what's in stock at their specific store is a different question, and that one needs the real model checked against real store data, not a person guessing on its behalf." Any answer works if it names the question whose true answer only exists once the real model is doing the actual work.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more