CaseAdvancedShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #17

What prototype would you build to test a pricing hypothesis?

The direct answer
Build a real, working feature and put a real, enforced price behind it, even if you send the bill by hand, then watch what people actually do when the charge is real. A "yes" to a hypothetical price costs the person nothing to give you. A real payment is the only signal that tells you what you'll actually make, and what it takes to earn it.
Do this, in order
  1. Build a working feature gated behind a real, enforced price, even a manually invoiced one, and test it on real people, not a survey.Why: a stated yes and a paid yes measure two different things, and only one of them predicts revenue.
  2. Watch what someone does and says right at the moment they hesitate at the price, not just whether they paid.Why: the hesitation usually points at what the price is really being weighed against, which a plain conversion number hides.
  3. Keep the test narrow: one feature, one price, real money, small enough to run in weeks.Why: a broad test builds months of billing infrastructure before anyone's confirmed a single person will pay for it.
  4. Don't let a strong survey number book engineering time before a real price has been tested.Why: a spoken yes costs the responder nothing. A real yes costs them their own money, right now.
  5. Leave the final packaging, subscription versus per-session, tier names, bundling, alone on day one.Why: that's a business-model decision the small test can't answer, and locking it in early bakes in a guess.
  6. Feed what the paywall drop-off shows back into the product before anyone designs a tier.Why: fixing why people hesitate is cheaper before the billing system gets built around the wrong offer.

How to answer this, stage by stage

Seven moves. Anchor it to one real price test, not a hypothetical one, since that's what an interviewer can actually push on.

1
Scope it to one concrete price test
Say it like this
"So I'm the PM for RoomDraft at Copperfield & Vine, a home-goods retailer. I'm not going to talk about pricing tests in general. I'll walk through the actual $25 per-session hypothesis we had, and the one real prototype that told us something the survey never could."
Why this works
A scoped example gives the interviewer something to picture and push back on.
2
Say the structure out loud
Say it like this
"Here's how I'll go: what we thought we knew from the survey, why a real price prototype tells you something a survey can't, the one decision I'd actually make, what happens if that test runs too narrow or too broad, and what I'd leave for a later decision, not day one."
Why this works
Two seconds of structure tells the interviewer you have a plan, so they follow instead of guessing where you're going.
3
Reframe what a pricing hypothesis actually needs
Say it like this
"Most people think testing a pricing hypothesis means asking people what they'd pay. That's not really the test. Asking what someone would pay costs them nothing to answer. Asking them to actually pay costs them real money, right now, for a real thing. Those are different questions, and only one of them predicts what you'll actually make."
Why this works
This is the actual insight being tested. Skip it and you're just saying "test the price," which nobody disagrees with and nobody learns from.
4
Give the anchor: the real paywall
Say it like this
"So the anchor is this: before we booked a single week of billing engineering, I put a real Stripe link behind RoomDraft's full plan. Free preview, blurred layout, a few product picks visible. Then one button: 'Unlock your full plan for $25,' a real card form, a real charge. I wasn't asking if $25 sounded fair. I was watching what a real shopper did the second the charge was real."
Why this works
Naming the specific mechanism, out loud, is what separates a real design decision from a vague "let's test pricing" gesture.
5
Prove it with what the real test found
Say it like this
"Here's what the survey missed. 610 of 900 loyalty members told us by email they'd pay $25 for this. But when 340 real shoppers hit the actual paywall, only 31 paid. And when I read what the other 309 typed right before they left, three out of four of them asked some version of the same thing: can I actually buy these exact items, or is this just a pretty picture? They weren't saying no to $25. They were saying no to a plan they couldn't tell was real."
Why this works
A real number, with the actual words people typed, does more work than "conversion was lower than expected" ever will.
6
Name the risk, in both directions
Say it like this
"Run this test too narrow, and you try one price on one small feature for two weeks and decide the whole idea is dead, when really the plan just needed to prove it was buyable. Run it too broad, and you spend three months building subscription tiers, saved designs, and referral credits around a number from an email survey, before a single real dollar has changed hands."
Why this works
Naming both failure directions shows you understand the test has a limit, not just a benefit.
7
Say what's out for day one, close on the number
Say it like this
"So: a real $25 paywall, real money, watching what real shoppers do and say the moment they hesitate. I wasn't building tiers or a subscription plan that month, that's a packaging decision for later, once we know people will pay something at all. The number I'd point to: 31 of 340 real shoppers paid, against 610 of 900 who said they would on paper, and three in four of the ones who didn't pay were really asking one specific question."
Why this works
Interviewers remember the last line most, and this one hands them something they can check, not just a mood.

Let's learn

What happens the first time a real price stands between a shopper and something they only said, on paper, they'd pay for?

Copperfield & Vine, a home-goods retailer, built RoomDraft: upload a photo of your room, answer a few questions, and it draws you a furniture layout with real, in-stock products from their catalog.

A hand-sketch of a shopper sitting on the floor with a hand-drawn floor plan on graph paper, two crooked pencil rectangles standing in for furniture, a tape measure beside her, and paint swatches taped to a wall in the background, no screen anywhere in the scene
Before RoomDraft, a room planned by hand

Before the prototype, the team had one thing to go on: an email survey to 900 loyalty members, asking if they'd pay $25 for a full, guided design session. 610 said yes. That's 68 percent. Someone put the number on a slide, everyone nodded, and engineering started scoping months of work: full account billing, subscription tiers, saved designs, referral credits.

Knowledge spark: what is a real price test? Asking someone if they'd pay costs them nothing. Putting a real charge in front of them and watching what they do costs them something real. Only the second one shows you what they'd actually do with their own money.

Fionnuala Brandvold, RoomDraft's PM, asked for one thing before the roadmap got locked: a rough version, wired to a real $25 Stripe link, shown to real shoppers on the actual site, not the survey list.

Here is what it found. 340 real shoppers reached the paywall over the pilot. Only 31 paid. Not 68 percent. About 9.

What people said they'd pay versus what they actually paid
100% 0% 50% 68% survey said yes 610 of 900 9% actually paid $25 31 of 340
said yes to a hypothetical price paid a real price
Nearly seven in ten shoppers said yes on paper. Fewer than one in ten paid for real. The gap is the whole reason a survey can't stand in for a prototype.

Here is the important part. Most of those 309 who didn't pay weren't objecting to the amount. Fionnuala had the chat log capture the last thing each of them typed before they left the paywall. Three out of four had asked some version of one question: can I actually buy these exact items, or is this just a picture?

610 people said yes to a price. 31 people paid for a plan.

At its worst, this doesn't just cost one small pilot. Say the team had trusted the 68 percent and spent the quarter building subscription tiers, saved designs, and referral credits around it. They'd have shipped a billing system nobody needed yet, while the actual problem, that shoppers couldn't tell if the plan was real, sat there untouched, quietly capping every price they ever tried.

The decision that mattered Build a rough, working feature and gate it behind a real, enforced price, even a manually sent invoice, on real people, watching for what they say and do right at the moment they hesitate, before a single week of billing engineering gets booked.

The choice I would take back. Months earlier, when the team scoped the RoomDraft pilot, they treated the 610-of-900 survey number as decision-grade evidence on its own, good enough to start booking engineering time for full billing and tiers. That made sense at the time. The survey was cheap, fast, and every smaller feature at Copperfield & Vine had shipped off less evidence than that. It stopped making sense the moment real money, not a stated opinion, became the thing being tested.

What I would leave alone. RoomDraft's intake also asks how many rooms a shopper is furnishing, one versus the whole apartment, to route them to the right flow. Nobody needs a real-price prototype for that. No money changes hands over the answer, so a quick round of five-minute interviews tells you everything a real test would, for a fraction of the cost.

The lesson. A spoken yes and a paid yes are different signals, and only the second one comes with a reason attached for what almost stopped it. If a price is involved, the test has to be too. Anything short of that is just a nicer-sounding guess.

Now here is the same thing as a story

Read this one when you've got a few minutes. The short version is above. This is for when you want to feel why it mattered.

Every Monday morning, Fionnuala pulled the loyalty-survey results up on the big screen in the product review, and for four months running, the same slide got the same nod.

She'd run product for RoomDraft at Copperfield & Vine since the idea was three sketches on a whiteboard: upload a photo of your room, answer a few questions, get a furniture layout pulled from the real catalog, real prices, real stock. Everyone liked the idea. The open question was whether anyone would pay for the full version instead of the free, blurry taste of it.

So the team ran a survey. 900 loyalty members, one email: if Copperfield & Vine offered a full AI design session, twenty minutes, personalized layout, exact products with prices, would you pay $25 for it? 610 said yes. Sixty-eight percent. The engineering lead read the number twice in review and started sketching a roadmap: account billing, three subscription tiers, saved designs across sessions, referral credits. Three months of work, booked on the calendar.

Fionnuala asked for one week first anyway. Not because she doubted the number. Because a stated yes had never once told her what it costs someone to change their mind.

So instead of the three-month build, one engineer spent four days wiring a rough version: the same intake questions, the same AI layout, but this time the full plan sat behind a real button. "Unlock your full plan for $25." A real Stripe checkout. A real charge to a real card.

Close hand-sketch of a phone screen showing a blurred layout with wavy lines standing in for blur, three grey product thumbnails below it, and a solid red button reading twenty five dollars unlock your full plan, with a caption noting it is a real Stripe link and not a survey question
The anchor: a real charge, not a hypothetical one

For the first two days it looked almost boring. A shopper would upload a photo, answer the questions, look at the blurred preview, and either tap the button or close the tab. Nothing dramatic either way.

Then Fionnuala started reading the chat logs. Not just who paid. What people typed right before they left.

The same question kept showing up. "Can I actually buy these exact items, or is this just inspiration?" Sometimes politely. Sometimes bluntly. Over and over, in different words, from shoppers who never clicked the $25 button at all.

610 people said yes to a price. 31 people paid for a plan.

By the end of the pilot, 340 real shoppers had reached the paywall. Only 31 paid. Nine percent, against the 68 the survey had promised. But three of every four who didn't pay had asked that same question first. They weren't rejecting twenty-five dollars. They were rejecting a plan they had no way to tell was real.

So here is the decision Fionnuala took back.

Months earlier, when the pilot got scoped, the team had treated the survey number as enough on its own to start booking engineering time. That made sense back then. The survey was cheap and fast, and every small feature they'd shipped before RoomDraft had gone out on less evidence than a 610-of-900 result. It stopped making sense the moment real money became the thing actually being tested. A stated opinion and a paid decision are not the same fact wearing different clothes. They're two different experiments, and only one of them has a real dollar in it.

Two panels side by side. Left, labeled if we had trusted the survey, shows five months booked solid on a roadmap grid under the number sixty eight percent said yes. Right, labeled real price real answer, shows the same paywall screen still working, with a speech bubble reading can I actually buy these exact items pinned beside it
The day it's wrong, and the anchor still tells you something

Fionnuala didn't kill RoomDraft's price. She fixed what the price was actually being judged against. The team added one line to the paywall screen, right above the button: "Every item shown is exactly what you can buy today, in stock, at this price." Nothing else changed. Same $25. Same catalog. Same layout.

The next 60 shoppers who reached the paywall: 11 paid. About 18 percent, close to double the earlier rate, from one sentence that answered the actual question people had been asking all along.

And the thing I'd want to tell myself, back when that first slide got its nod in review: a number people say costs them nothing to say. We built three months of a roadmap on a number nobody had to spend a cent to give us.

SPARK, run against a real price tag

This question sounds like it wants a metric, how would you measure demand, or an estimate, how many would pay. It's really asking for one concrete decision about how you'd learn the truth before you commit to building the wrong thing, so SPARK fits. A question asking how you'd measure RoomDraft's ongoing revenue per session once it's live would reach for LEAD instead.

S, situation. Before a real-price prototype exists, a pricing question gets answered by asking people what they'd pay, in a survey or an interview, where saying yes costs nothing.
P, payoff. Not "know the right price." The habit worth building: learn what people will really pay, with real money, before any billing infrastructure gets built around a guess.
A, anchor. A working feature gated behind a real, enforced price point, even a manually invoiced one, tested on real people, watching what they do and say at the moment the charge is real.
R, risk. Too narrow, and one flat no at one price kills an idea that just needed proof the plan behind it was real. Too broad, and months get spent on subscription tiers before a single real dollar changes hands.
K, keep out. The final packaging: subscription versus per-session, tier names, bundling with a purchase. That's a decision for once you know people will pay something at all, not day one.
Two side by side boxes. Left, solid, shows the working twenty five dollar paywall labeled solid one real price. Right, a dotted greyed out box labeled tiers, subscription, bundle, not day one
What we left for later, kept visibly separate from day one
Why the anchor survives the risk Check it against the near miss. Does a real paywall still teach you something even at a low conversion rate? Yes, because it's built to capture what people say the moment they hesitate, not just whether they clicked. Does it avoid the broad-commitment trap? Yes, because K keeps the tier and packaging decision explicitly off this test's job.

And if you want to be sure it really works, try it somewhere else

A boutique wedding-planning company runs on a completely different calendar, but the same gap between a stated yes and a paid yes shows up in a couple's very first real invoice.

S. Desmond Kilbride runs product at Amberlea Weddings, testing DayFlow, an AI tool that builds a minute-by-minute wedding-day timeline synced with every vendor. Today, without a real-price prototype, whoever decides the price checks it against friendly phone calls with past clients, where saying yes costs nothing.
P. The habit worth building: learn what an engaged couple will really pay, before Amberlea books months of engineering to bundle DayFlow into every planning package.
A. Same shape, a different desk. Desmond offers DayFlow to couples currently mid-planning, gated behind a real $400 invoice their coordinator sends by hand. No subscription system, just a real bill for a real thing.
R. Test it on five weddings over one slow month, and a quiet yes rate looks like proof it's ready for every couple, when the real test only ran during the calm season. Wait for a full year of weddings before ever testing a real price, and Amberlea never learns anything, because nobody can sit on an idea that long.
K. No attempt yet to decide whether DayFlow sells standalone or gets folded into the premium planning package. That's a packaging call for later, once real invoices show people will pay at all.

Close hand-sketch of a checkout screen reading four hundred dollars reserve your timeline, with a pinned sticky note beside it reading seventy one percent said yes and thirty eight of three hundred ten paid, captioned a real invoice sent by her wedding coordinator
Same anchor, a different desk, an invoice instead of a checkout link

Desmond had asked 40 past clients in casual follow-up calls if they'd have paid $400 for this. 28 said yes, 71 percent. When 310 real couples were offered the real invoice, 38 paid. Most who declined weren't balking at the price. They asked some version of the same question: will an actual person double-check this the week before the wedding, or is it just the app? The doubt wasn't about money. It was about whether a human still stood behind the plan, the same shape of question RoomDraft's shoppers had asked about their furniture.

Swap the trigger and it still runs

  • Speed: even if the checkout finished in one tap and two seconds, a fast real charge tests the exact same thing a slow one does. Speed doesn't fix a shopper's doubt about whether the plan is real.
  • Cost: if the Stripe link cost nothing to set up, that still wouldn't tell you which nine of your fourteen customers actually trust the plan enough to pay. You still need real money changing hands, not zero.
  • The model gets better: if RoomDraft's layout suggestions became perfect, that still wouldn't fix a shopper's doubt about whether the products shown are buyable, because accuracy and believability are two different jobs.

Where people run it wrong

  • Testing the price with a fake "would you pay this" button that never actually charges anyone, so nobody ever has to reach for their wallet.
  • Treating one strong survey number as proof enough to book months of billing engineering before a single real payment happens.
  • Trying to nail the final packaging, tiers, bundles, subscription versus per-session, inside the same two-week pricing test, so the narrow question, will anyone pay at all, never gets a clean answer.

How to use it live

If you're asked this cold, ask what happens to the answer if the person has to actually spend their own money to give it. That question, asked out loud, usually finds the real test faster than trying to write the perfect survey question.

Flashcards (click a card to flip it)

1 · THE SITUATION
What's the situation, before this pricing prototype existed?
Tap to flip
ANSWER
Copperfield & Vine had only asked shoppers what they'd pay for RoomDraft in a survey. 610 of 900 loyalty members said yes to $25, and the team was ready to book months of engineering on that number alone.
2 · THE PAYOFF
What's the real habit this pricing prototype is trying to build?
Tap to flip
ANSWER
Learning what people will really pay, with real money, before any billing infrastructure gets built around a guess.
3 · THE ANCHOR
What did the real $25 paywall actually test?
Tap to flip
ANSWER
Not whether $25 sounded fair, but what a real shopper did the moment the charge was real, and what they said right before they walked away from it.
4 · THE RISK
What breaks if the pricing test runs too narrow, or too broad a commitment?
Tap to flip
ANSWER
Too narrow, and one flat no at $25 kills an idea that just needed proof it was buyable. Too broad, and months get spent on subscription tiers before a single real dollar changes hands.
5 · THE PROOF
What did the real shoppers do at the paywall that the survey never showed?
Tap to flip
ANSWER
Only 31 of 340 paid, but three of every four who didn't pay asked some version of one question first: can I actually buy these exact items, or is this just a picture?
6 · THE NUMBER
___ of ___ real shoppers paid the $25, against ___ of ___ who said they would on the survey.
Tap to flip
ANSWER
31 of 340 paid for real. 610 of 900 said yes on paper.
7 · THE REPLAY
Same paywall, one line added. What changes?
Tap to flip
ANSWER
A line reassuring shoppers every item shown is real and in stock at that price, right above the $25 button. The next 60 shoppers who reached it: 11 paid, close to double the earlier rate.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor test?
Tap to flip
ANSWER
Amberlea Weddings' DayFlow timeline tool, at a real $400 per-wedding invoice. Its anchor tests the same shape of doubt: whether a human coordinator still backs up the plan, not just whether the price feels fair.

Check yourself Score: 0 / 0

True or false
1. True or false: Most of the 309 shoppers who didn't pay were reacting to the $25 price being too high.
  • True
  • False
Show hint
Look at what the chat logs actually captured people typing right before they left the paywall.
Show answer
False. Three of every four who didn't pay asked some version of "can I actually buy these exact items," not a complaint about the amount. They doubted the plan was real, not that $25 was fair.
Fill in the blank
2. ___ of ___ real shoppers paid the $25 price, against ___ of ___ who said they'd pay it in the survey.
Show hint
This is the number the whole argument leans on. It shows up twice, once in the story, once in the chart.
Show answer
31 of 340; 610 of 900. Nearly two-thirds said yes on paper. Fewer than one in ten paid for real.
Multiple choice
3. Which design matches the anchor this answer argues for?
  • A. A longer survey sent to more loyalty members before deciding.
  • B. A real, working feature gated behind a real, enforced price, even if invoiced by hand, tested on real people.
  • C. A full subscription billing system with three tiers, built and launched immediately.
  • D. A focus group shown mockups of the $25 paywall screen.
Show hint
The anchor needs someone actually spending real money, not talking about hypothetical money.
Show answer
B. A is still a survey, just a bigger one. C mixes up the day-one job with the keep-out, packaging is a separate later decision. D is still a picture of a charge, not a real one.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Think about why treating the survey as decision-grade evidence sounded fine, back before real money was the thing being tested.
Show answer
Model answer: The team treated the 610-of-900 survey number as enough on its own to start booking engineering time for billing and tiers. That made sense when the survey was cheap and fast, and every smaller feature had shipped off less evidence than that. It stopped making sense once real money, not a stated opinion, became the thing being tested.
Short answer, apply it yourself
5. Pick something at your own work that gets priced or greenlit from a stated opinion alone, a survey, a poll, a "would you use this" question in an interview. What would change if you had to test it against a real, if small, payment or commitment instead?
Show hint
Look for whether the stated opinion cost the person anything at all to give you.
Show answer
Model answer: "I run a workshop that people say they'd pay for when I ask in a survey. A ten-dollar deposit link, sent to the next twenty interested people, would show me who actually shows up with money down, not just who's polite in an email."
Multiple choice
6. Based on this answer's own numbers, if only 6 of 340 real shoppers had paid instead of 31, would the decision to hold off on tiers still make sense?
  • A. Yes, because the real question is whether the reason people don't pay is fixable, not whether the raw rate cleared some number, and packaging still isn't the thing to decide first either way.
  • B. No, at 6 of 340 the team should ship the full subscription tiers exactly as originally planned.
  • C. No, a lower number means the $25 price was too high and should simply be raised.
  • D. Yes, but only if the 6 who paid also complained that the price was too low.
Show hint
Compare what a lower number changes about the drop-off reason against what it changes about whether packaging is ready to decide.
Show answer
A. A lower rate makes the believability problem more urgent to fix, not less real. Either way, tiers and packaging were never the day-one decision, so a worse number doesn't suddenly make them one.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more