ConceptIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #19

What is the risk of demoing with hand-picked examples, and what would you demo instead?

GUARD · what Vytautas's three favorite products never had to prove

Winnow is Driftglass Analytics' tool. It reads every review on a product listing and writes shoppers a short summary card instead of six hundred reviews to scroll through. Vytautas Solomou is the sales engineer who runs Winnow's live demos, always on the same three easy products. Karolina Nkomo signed a contract for her company, Nettlewood Provisions, after watching one of those calls. Danae Ogilvie, the product manager who owns Winnow, got pulled in five weeks later, once Nettlewood's best-selling hot sauce started selling less, for no reason anyone at Nettlewood could see.

The direct answer
A demo built from a handful of hand-picked examples isn't a sample of what the product does. It's a sample of what the demo-builder already knows will look good, and that gap stays invisible to whoever's deciding based on it, until their own product lands in the part of the range the demo quietly avoided. Pull the demo example live, at random or stratified, from the real evaluation set every time, and include at least one known hard case, flagged honestly, instead of a curated highlight reel. A demo that can't survive showing its own hard case isn't ready to be shown at all.
Do this, in order
  1. Pull every demo example live, at random or stratified, from the real evaluation set.Why: this is the direct answer. A fixed "reliable three" is a selection bias, not a fair sample.
  2. Include at least one known hard case in every demo, flagged honestly with its real number.Why: a demo with zero failures teaches the room nothing about where the product actually struggles.
  3. Before locking a demo script, find where the real distribution actually gets hard, and check whether the script ever touches it.Why: Nettlewood's whole catalog lived exactly where "the reliable three" was chosen to avoid.
  4. Track the gap between the accuracy a demo reports and what an account's 90-day production data shows.Why: a demo culture that's drifted looks perfectly healthy until someone measures this gap on purpose.
  5. Watch whether demo calls keep pulling a fresh example or slide back to the same fixed three.Why: a script that never changes again is the surest sign the habit crept back in.
  6. Leave the polished screenshots in a blog post or a conference talk alone.Why: matching the bar to the stakes is the point. Nobody signs a contract off a slide deck.

How to answer this, stage by stage

Nobody is grading whether you believe in honesty. They're grading whether you can name the actual mechanism, a biased sample, and turn it into a specific, checkable demo policy.

1
Pin the demo to one product and one audience
Say it like this
"Let me make this concrete. Say a company sells a tool that reads every review on a listing and writes shoppers a short summary. A sales engineer runs the same three easy products in every demo. A prospect signs based on what she saw. That's the scenario I'll run, because 'hand-picked demo' means nothing until there's an actual three products and an actual signature."
Why this works
Keeps the answer from turning into a lecture about honesty in general.
2
Say your structure out loud
Say it like this
"I'll run this as GUARD. Groups, who's actually deciding based on the demo. Unequal, whose real catalog lands where the demo never went. Ability to contest, can the audience even tell what they're not being shown. Reduce, what the demo should pull from instead. Detect, how you'd know the habit crept back in."
Why this works
Two seconds of structure signals a method, not a hot take about honesty.
3
Name the bias for what it is
Say it like this
"This sounds like a question about honesty. It's really about sampling. Picking the three products that always work isn't a neutral peek at the product, it's a biased sample, the same way polling only your happiest customers isn't market research. The bias is built into the picking, nobody has to lie for it to mislead someone."
Why this works
This is the reframe. It's what separates a real answer from a platitude about always being honest.
4
Give the one decision
Say it like this
"Here's what I'd actually build: every demo pulls its example live, at random or stratified, from the real evaluation set, and it always includes at least one listing I know the model handles badly, flagged honestly on screen. Not a fixed script. A live draw, every time."
Why this works
This matches the direct answer word for word. If it doesn't, the interviewer notices before you do.
5
Prove it with the compressed failure
Say it like this
"Here's what happens without it. The model nails the three demo products ninety-seven times out of a hundred. On the real catalog it's eighty-one, and on listings with mixed or sarcastic reviews, it's thirty-eight. The customer never saw that split. She turned it on for her whole catalog, including her best-selling product, and its sales dropped for five weeks before anyone traced it."
Why this works
This is the story below, cut to four sentences. The long version proves it happened for real.
6
Say what you'd still hand-pick, on purpose
Say it like this
"I wouldn't hold every piece of content to this bar. A screenshot in a blog post or a conference slide isn't what anyone signs a contract on. The line is simple: anything a real decision gets made against needs the honest sample. Anything else can stay polished."
Why this works
Shows judgment about scope instead of blanket paranoia about every screenshot everywhere.
7
Close on something checkable
Say it like this
"You'll know it's working when the gap between what a demo promises and what an account sees ninety days in shrinks close to zero. You'll know it's slipping if the same three products keep showing up on every call, because that's not a coincidence, that's a script nobody's touched."
Why this works
Ends on a test the interviewer could actually run, not a promise that the process is fine.

Let's learn

Picture a product listing page with six hundred reviews under it. Nobody reads all six hundred. Winnow reads them for you and writes three sentences.

Hand sketched labeled parts diagram titled What's on a Winnow summary card. A central document icon labeled Winnow card, with four callouts: Sentiment badge, Two line gist, Tag call-outs, and No confidence flag, missing.
A sentiment badge, a two-line gist, and a few tag call-outs. Nowhere on the card does it say how sure Winnow actually is.

Winnow is Driftglass Analytics' tool for turning a pile of customer reviews into something a shopper can read in ten seconds. In sales calls, Vytautas Solomou always runs it live on the same three products: a water bottle, a phone case, a desk lamp. Every reviewer on those three basically agrees. Either it works or it doesn't. Nobody's being sarcastic about a phone case. Checked by a human rater against what the reviews actually say, Winnow's summary lands right ninety-seven times out of a hundred on those three.

Knowledge spark: what's a stratified eval set? A test batch built on purpose to cover the hard parts of the job, not just the easy ones. Driftglass keeps two hundred real seller listings for this, picked so the mix includes plenty of the messy, contradictory, sarcastic reviews that a random scoop of listings might barely touch. It exists so a number like "ninety-seven percent accurate" can't hide behind only ever being tested on the easy stuff.

Across that full two-hundred-listing batch, Winnow is right eighty-one times out of a hundred. On the forty-two listings where the reviews genuinely contradict each other or read as jokes, exaggeration, or backhanded praise, it drops to thirty-eight.

Hand sketched quadrant diagram titled Which listings demo well, and which are the real test. X axis, how mixed or sarcastic the reviews are, from clean one-sided to contradictory sarcastic. Y axis, how well Winnow holds up, from breaks down to solid. Water bottle, desk lamp, and phone case demo products sit in the upper left, clean and solid. Reaper's Regret hot sauce and gag gift novelty mugs sit in the lower right, contradictory and breaking down.
The three products Vytautas always demos sit in one corner. The products that actually test the model sit in the opposite one, and nobody watching a sales call ever sees that second corner.

Here's the turn. Thirty-eight isn't really the problem on its own. The problem is that nobody watching a demo ever sees that number, so whoever's deciding based on the demo has no way to know which corner their own catalog falls into. Karolina Nkomo watched the three clean products, signed for Nettlewood Provisions, and turned Winnow on for her whole catalog, the same way the demo taught her to.

She didn't take a chance. She did the sensible thing with the only three products anyone had ever shown her.

What that costs at its worst: Nettlewood's best-selling product is a hot sauce called Reaper's Regret, six hundred forty reviews, most of them five-star raves written like complaints, "this ruined my life," "worst decision I ever made, ordering three more." Winnow read the words, not the joke, and its summary told shoppers the product mostly disappointed people. Sales on that one listing fell for five weeks before anyone traced it back to the card sitting on the page.

The choice I would take back During Winnow's first quarter, Driftglass's sales team found that leading every demo with the same three clean products gave the smoothest, most reliable calls, so they locked it into the standard script. That made sense with one small demo library and every new rep needing a fast, safe default. It stopped making sense once the locked script became the only version of Winnow any prospect, or any executive watching the win rate, ever actually saw.

What I would leave alone: a screenshot of those same three products in a conference talk or a blog post doesn't need this bar. Nobody signs a contract off a blog post. Gating marketing collateral the same way a sales demo gets gated just slows down harmless content for no real protection.

The lesson: a demo that always wins isn't proof the product is ready. It's proof nobody has checked yet whether it can survive its own hardest case in front of someone about to pay for it.

Now here is the same thing as a story

Stage five above compresses this into four sentences. Here's the eight months underneath it, for the part that doesn't fit in a stand-up answer.

Every sales call Vytautas runs starts the same way: pull up three products, hit demo, watch the prospect's shoulders drop an inch. He's been doing technical demos for Driftglass for four years, and he's good at reading a room, good enough to know within ninety seconds whether a prospect is sold.

He never planned to reuse the same three products forever. In Winnow's first weeks he rotated between five or six items depending on what a prospect actually sold. Then he noticed the water bottle, the phone case, and the desk lamp closed at a noticeably higher rate than anything else he tried, so those three became his opener. Then, a few months later, his only three. By the time new hires joined, the training deck listed them by name, "the reliable three," and nobody rotated anymore.

Hand sketched comparison diagram titled Two people, one selector. Left panel, a person labeled Vytautas, caption picks which three products get demoed. Right panel, a person labeled Karolina, caption only ever sees what he picked.
Vytautas holds the dial that decides what a prospect ever sees Winnow do. Karolina, on the other end of the call, has empty hands.

Karolina Nkomo founded Nettlewood Provisions eleven years ago, a small e-commerce shop selling spice blends and hot sauces to people who like their food to hurt a little. She watched Vytautas's demo on a Thursday afternoon, the water bottle, the phone case, the lamp, each summary landing exactly right, and signed the following week. She turned Winnow on across her whole catalog the same day the contract cleared, forty-one listings at once, because why would you turn on ninety percent of a good thing and leave the rest for later.

A new hire joined Driftglass's sales team that same month. In her second week, sitting in on one of Vytautas's calls, she leaned over afterward and asked, half curious, half joking, "do we ever demo on something like theirs," nodding at a prospect's spicy-snack catalog on the shared screen. Vytautas laughed and said no, not really, the reliable three always closes better. Neither of them thought about it again for five more weeks.

Hand sketched timeline diagram titled Eight months from a locked script to a traced drop. Five milestones: Demo script locked, caption bottle case lamp. Karolina signs, caption Nettlewood Provisions. Winnow goes live, caption full catalog. Conversion drifts down, caption weeks 2 to 4. Karolina traces it, caption week 5, to the summary, this last milestone emphasized in gold.
Two of these five moments had already happened before Karolina ever saw a demo. Nobody connected the dots until the fifth one.

Reaper's Regret started sliding in week two. Karolina's team tried the obvious explanations first: a seasonal dip, a competitor's price cut, a shipping delay somewhere upstream. None of it held up. It took until week five, when someone finally opened the listing page itself and actually read the AI summary sitting under the star rating, to find the real cause: "Reviewers report largely negative experiences with this product." Six hundred forty reviews, most of them glowing, and Winnow had turned the loudest, most sarcastic ones into a warning label.

It was never really about the thirty-eight percent. It was about who got to see that number, and who never did.

Karolina called Driftglass ready to cancel the whole contract. Danae Ogilvie, the product manager who owns Winnow, took the call herself. She pulled the account's real numbers that afternoon and found exactly what the stratified eval set had already been quietly showing for months: this wasn't a bug in Karolina's account, it was the known, measured gap between "the three products in the demo" and "the catalog Nettlewood actually sells."

Hand sketched flow diagram titled Where a demo should get checked, and doesn't. Five connected boxes reading Demo runs, Contract signed, Goes live, No check ran, Hard cases surface, with the fourth box emphasized in red.
Four of these five steps happened exactly as designed. The missing one, a check on whether the demo was ever representative, was never built.

The decision Danae would take back sits eight months earlier, in the meeting where "the reliable three" got written into the new-hire training deck as company policy, not just Vytautas's personal habit. Nobody in that meeting asked what should happen when a prospect's real catalog looked nothing like the three products doing all the demoing. At the time, nobody had a reason to ask.

So Danae rebuilt the demo script. Every call now pulls its example live from the two-hundred-listing eval set, at random, and it's rigged so at least one pull every time is a listing from the hard, contradictory-review bucket, shown with its real confidence number on screen, not hidden.

Hand sketched icon list diagram titled What a live, honest demo has to include now. Four rows: a document icon for pulled fresh from the real eval set not a fixed script, a gauge icon for at least one known hard case shown on purpose, a question mark icon for a flag on any summary under the confidence floor, and a scale icon for the real accuracy number stated out loud, not just the demo's.
Four things a call has to include now. A demo missing any one of them is running on the old script, whether anyone admits it or not.

Run the same eight months again, with one change. Vytautas's demo for Nettlewood pulls up Reaper's Regret's own listing type live, a spice product with mixed, joking reviews, right there in the call, flagged as one Winnow gets right about four times in ten. Karolina signs anyway, now expecting exactly what she gets. When the early numbers on her hot sauce come in soft in week two, Danae's account team already has that listing tagged for extra review, and the confidence threshold gets tightened for her account within three days, not five weeks.

What I'd tell myself, back in the meeting where "the reliable three" became policy: the three products were never the actual problem. The fact that they were the only three anyone outside Driftglass ever saw, that was the whole thing.

GUARD, run against Vytautas's opening three

This was never really about whether Vytautas meant to mislead anyone. GUARD is for naming who's actually deciding on a biased sample, and turning "show real examples" into a policy specific enough to check.

GGroups. Who's actually deciding, and on what.
Karolina Nkomo, who signs a contract and later stakes her best-selling listing's sales on what she believed Winnow could do across her whole catalog. Driftglass's own sales leadership, who scoped their entire outbound motion around the win rate "the reliable three" produced, so the bias stopped being one rep's habit and became the company's growth plan. And, further out, Nettlewood's own shoppers, who read a Winnow card and never learn it was machine-written or that the model struggles with sarcasm at all. They just decide not to buy.
Name the people making real decisions off the demo, not just the person who built it. Most answers only name the second.
UUnequal. Where the gap actually lands.
A seller whose whole catalog is plain home goods never notices any of this, because their reviews were never going to read as sarcastic in the first place. The gap concentrates on sellers like Nettlewood, whose products, hot sauces, novelty items, anything built to be extreme or funny, draw exactly the mixed, tongue-in-cheek reviews "the reliable three" was chosen to avoid.
The sellers most likely to get hurt by this are the ones whose real catalog was never anywhere near the demo to begin with.
AAbility to contest. Who can tell what they're not being shown.
Karolina watching a polished call had no way to know the three products were picked for being easy. The demo format itself carries no label saying "these were chosen." She couldn't ask "show me on something like my hot sauce," because she had no reason to think there was anything else to ask for. Driftglass's own leadership couldn't ask either, since the win-rate number reaching them was built from the same hand-picked calls.
This is GUARD's sharpest question here: not "did anyone lie," but "could the audience have known what they weren't seeing."
RReduce. The actual fix, not a slogan.
Every demo pulls its example live, at random or stratified, from the real two-hundred-listing eval set, and always includes at least one listing from the hard, contradictory-review bucket, shown with its real number on screen. The bar isn't "the demo has to be perfect." It's "the honest number, on the real distribution, clears an agreed floor, and everyone in the room, including the prospect, hears that number, not the shiny one."
The alternative worth naming and rejecting: hold every demo until Winnow's handling of sarcasm improves past ninety percent. That sounds responsible, but it stalls the roughly eighty percent of the real distribution the product already handles fine, in exchange for hiding a real limitation instead of disclosing it. A disclosed weakness a customer can plan around beats a hidden one they discover in production.
Winnow summary accuracy, by which listings get demoed and which get the real test
100% 50% 0% 97% 3 hero demo products 81% Full eval set 200 listings 38% Mixed or sarcastic 42 of the 200
What the demo showsThe honest averageThe part nobody demos
The three colors are three different truths about the same model. A prospect only ever hears the green one.
DDetect. How you'd know it's happening, before an account almost cancels.
Track two numbers every month, not one. First, the gap between the accuracy a demo reports and what an account's own 90-day production data shows. Before the fix, that gap could run wide and stay invisible: 97 percent reported against 38 percent actually experienced by an account whose catalog leaned hard, a 59-point gap nobody was measuring. Second, whether demo calls are actually pulling a fresh example each month, or quietly settling back onto the same three. Before the fix, that number was zero percent live pulls, one hundred percent fixed script, for eight straight months.
The failure worth naming plainly: a rising demo win rate looks like good news and can mean the bias just got worse, since "the reliable three, unchanged" is exactly what a rising win rate would look like too.
Reaper's Regret, weekly conversion rate, before and after Winnow's summary card went live
5% 2.5% 0 summary goes live Wk -1 Wk 1 Wk 2 Wk 3 Wk 4 Wk 5, traced Wk 6, fixed
Weekly conversion rate on Reaper's Regret
Nothing on Driftglass's own dashboard flagged this drop. It only shows up when you go look at one seller's one listing, which is exactly what a demo culture built on three easy products never trains anyone to do.
The trade-off this accepts on purpose Switching to live, honest demos dropped Driftglass's demo-to-signed-contract rate from 68 percent to 54 percent. That's real revenue given up on purpose. In exchange, the share of new customers filing a complaint inside their first 90 days about "this isn't what we saw in the demo" fell from 22 percent to 6 percent. Fewer contracts close. The ones that do, hold.

And if you want to be sure it really works, try it somewhere else

Same five letters, a farm cooperative instead of a spice shop, and this time the hidden case is a blurry leaf photo instead of a sarcastic review.

Leafmark, built by Cindermoss Agritech, reads a photo of a crop leaf and tells a farmer which disease it likely has and how urgent treatment is. Nazariy Yeneva, Cindermoss's field rep, demos it on three textbook-clear photos every time: sharp focus, one obvious lesion, good light.

Tobiasz Rindale runs a forty-farm cooperative and signed a season-long contract after watching that same clean demo. Real field photos from his farmers came back blurry, shot in bad light, and a few showed two diseases overlapping on one leaf, exactly the case the demo photos never included. Leafmark's confidence dropped hard on those, and on several early cases it named only one of the two diseases present, delaying treatment for the second one on plants that were already stressed.

Hand sketched decision tree titled Which leaf photo does a crop disease demo show. Root: picking a photo for the sales demo. Three branches: a clear textbook shot leads to looks perfect every time, a real blurry farmer photo leads to confidence drops might mislabel, and two overlapping diseases on one leaf leads to model names one misses the other.
Three different photos a demo could show. Cindermoss only ever reached for the top branch.

Same rank, mapped onto Leafmark: pull the demo photo live from Cindermoss's real intake queue every time, guaranteed to include one blurry or overlapping case shown honestly with its lower confidence stated, instead of the fixed three textbook shots.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: pull the demo live, include a known hard case, state the honest number out loud.
Cost: no budget this quarter to build a live-pull system. Rotate three fixed examples weekly from a short list that always includes one hard case, cheaper than true randomness, and it still breaks the illusion of one permanently safe set.
The model got better, for real: say Winnow's handling of sarcasm genuinely climbs to seventy percent. That's a reason to update which listing counts as "the hard one" to show, not a reason to stop showing one, since there will always be some harder slice once today's hardest slice is fixed.

Where people run it wrong.
They treat "the demo worked" as proof the product works, instead of proof the three chosen examples work.
They let the person who benefits from a high win rate also be the person who decides what gets shown.
They fix this once, after a near-cancellation, instead of tracking the demo-to-production gap every month for free.

How to use it live. Before answering, ask yourself out loud: "if I demoed this right now on a random pull from the real eval set, would I be comfortable with what came up?" If the honest answer is no, that's not a demo problem. That's a product gap the demo is currently hiding.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
Which framework fits "what's the risk of demoing with hand-picked examples, and what would you demo instead"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real test isn't whether anyone lied, it's who's making a real decision off a biased sample, and how you'd fix the sampling itself.
2 · THE PEOPLE
Who are the three people this answer names?
Tap to flip
ANSWER
Vytautas Solomou, the sales engineer who always demos the same three products. Karolina Nkomo, founder of Nettlewood Provisions, who signed based on that demo. Danae Ogilvie, the product manager who owns Winnow and rebuilt the demo process.
3 · THE REAL SPLIT
What are the three real accuracy numbers behind one "the demo works" claim?
Tap to flip
ANSWER
97 percent on the three hero demo products. 81 percent across the full 200-listing eval set. 38 percent on the 42 listings with mixed or sarcastic reviews, the bucket the demo never touched.
4 · THE UNEQUAL COST
Which sellers absorb this gap, and which never notice it at all?
Tap to flip
ANSWER
A seller with a plain, one-sided catalog never notices, because their reviews were never going to be contradictory. Sellers like Nettlewood, whose products draw mixed, sarcastic reviews, absorb the whole gap, invisibly, until sales drop.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Locking "the reliable three" into Driftglass's standard training deck as official policy, not just one rep's habit, so the bias became the company's growth plan instead of something anyone could easily undo.
6 · THE NUMBER
Fill in the blank: Reaper's Regret's weekly sales fell from ___ percent to ___ percent over ___ weeks before anyone traced it to the summary card.
Tap to flip
ANSWER
From 4.1 percent to 2.6 percent, over 5 weeks. Nothing on Driftglass's own dashboard flagged it, because the drop only shows up on one seller's one listing.
7 · THE REPLAY
Same eight months, new demo process, what changes?
Tap to flip
ANSWER
The demo pulls a live example from the hard bucket, flagged at roughly four-in-ten accuracy. Karolina signs knowing that. When her hot sauce dips in week two, it's already flagged, and the fix ships in 3 days instead of 5 weeks.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which one, and who plays the equivalent roles?
Tap to flip
ANSWER
Leafmark, Cindermoss Agritech's crop-disease tool. Nazariy Yeneva plays Vytautas's role, demoing clean textbook leaf photos. Tobiasz Rindale plays Karolina's role, signing a co-op contract before his own farmers' blurry photos exposed the gap.

Check yourself Score: 0 / 0

Fill in the blank
1. On the three hero demo products, Winnow is accurate ___ percent of the time. Across the full 200-listing eval set, it's ___ percent. On the 42 listings with mixed or sarcastic reviews, it drops to ___ percent.
Show hint
Check the bar chart in the GUARD recap, and the "Let's learn" section.
Show answer
97, 81, and 38 percent. Three true numbers about the same model. A prospect only ever hears the first one.
Multiple choice
2. Why doesn't a high demo win rate prove the "reliable three" script was safe to keep using?
  • A. Because win rate isn't a real business metric.
  • B. Because Vytautas was intentionally lying to prospects on every call.
  • C. Because a rising win rate built on the same easy examples can mean the sampling bias got worse, not that the product improved.
  • D. Because Winnow's accuracy on the three hero products was actually fake.
Show hint
Check the Detect step in the GUARD recap.
Show answer
C. The 97 percent on the hero products was real, that rules out D. Nobody lied, that rules out B. The problem is what a healthy-looking win rate can quietly be hiding.
True or false
3. True or false: once the fix is in place, no product should ever get demoed on a curated, hand-picked example again.
  • True
  • False
Show hint
Check "what I would leave alone" in Let's learn, and stage 6 of the walkthrough.
Show answer
False. A blog post or a conference slide isn't what anyone signs a contract on. The fix targets anything a real decision gets made against, not every piece of content that mentions the product.
Short answer, name the rejected alternative
4. Besides pulling demos live from the real eval set, what alternative fix does this answer name and reject, and why does it lose?
Show hint
Look at the Reduce step in the GUARD recap.
Show answer
Model answer: Holding every demo until Winnow's sarcasm handling improves past ninety percent. It loses because that stalls the roughly eighty percent of the real distribution the product already handles fine, in exchange for hiding a limitation instead of disclosing it.
Short answer, apply it yourself
5. Think of a product demo you've watched, for anything, not just AI. What's one hard or messy case you'd bet was never shown, and how would you find out?
Show hint
Look for a demo that always used the same one or two examples every time you saw it.
Show answer
Model answer: A resume-screening tool always demoed on tidy, single-page resumes. The messy case, someone with a career gap or a non-linear path, never appears. You'd find out by asking the vendor to run it live on a resume like that, on the spot.
Fill in the blank, work the number
6. If Nettlewood's catalog of 41 listings has the same 21 percent hard-case rate as the 42-of-200 eval set, about how many of Karolina's own listings sit in Winnow's hard, 38-percent-accurate bucket?
Show hint
42 out of 200 is 21 percent. Apply that rate to 41 listings.
Show answer
About 9 listings. Nine listings quietly running at roughly 38 percent accuracy is not a rare fluke inside one account, it's close to a quarter of her whole catalog.
Before you close the answer
Why this works
Tests whether you can name selection bias as the actual mechanism, not just say "be honest," and turn that into a specific, measurable demo policy instead of a mood. Most candidates say "show real examples" without saying which ones, how they'd be chosen, or how you'd catch the habit creeping back.
Follow-up traps
"Isn't showing a known failing case just going to scare the customer off?" Response: a prospect who sees the hard case with an honest caveat signs knowing what they're buying. A prospect who never sees it finds it in production instead, and that costs more, in trust and eventually in the contract.

"Doesn't a lower demo win rate hurt the business?" Response: yes, and that's the trade-off, accepted on purpose: win rate fell from 68 to 54 percent, while the 90-day new-customer complaint rate fell from 22 to 6 percent. Fewer contracts close. The ones that do, hold.
If pressed
Winnow's sentiment step scores each review on a confidence scale before writing the summary. Reviews it scores near its own uncertainty floor are disproportionately the sarcastic or contradictory ones, the exact same ones the "reliable three" was picked to avoid. Driftglass already had that number sitting in the pipeline, unused, before this fix ever exposed it to a sales call or a customer.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more