Explain why 'it worked in the demo' is a systematically misleading signal for AI features.
Brieflane drafts a tailored resume and cover letter for a job seeker in under a minute. Rasmine Kolisch runs product there. For eight months, Corwyn Odutola ran the same twelve resumes for investors and at hiring fairs, and Brieflane never once got a fact wrong. Then it shipped to the public, and Adaugo Ferriday sent a cover letter that invented two years of a job title she never held.
- Treat a demo's success as evidence about only the inputs it was shown, never about the model in general.Why: this is the whole reversal the rest of the answer works out.
- Before general release, compare the demo's input diversity against the eval set's real coverage.Why: this one check would have shown an eight-month streak sitting inside a sliver of the real shapes.
- Recut any success number by variety and how hard the input pushes back, not by how many times you ran it.Why: ninety reruns of twelve resumes is still twelve shapes, not a sample.
- Watch for a demo-runner's own unconscious habit of avoiding inputs they know the model struggles with.Why: Corwyn's instinct to reach for the "safe" resumes hid the exact failure production later found.
- Deliberately throw adversarial and messy inputs at any demo used to justify a launch decision.Why: nobody ever tried a gapped or self-employed resume, so the ceiling stayed hidden until real users found it.
- Leave deterministic post-processing steps alone.Why: Brieflane's grammar and spelling pass doesn't depend on career-history shape, so this scrutiny doesn't apply to it.
How to answer this, stage by stage
Nobody is grading whether you can say "demos are misleading." They're grading whether you can name the actual gap between what got tested and what real users send, and the one check that would have caught it before launch.
Let's learn
What happens when a tool passes every test anyone ever gave it, and the tests were never the point.
Brieflane reads a job seeker's work history and a job posting, and drafts a tailored resume and cover letter in under a minute. Before it existed, most people spent two to three hours per application, longer if they were rewriting a cover letter from scratch each time. With Brieflane, that dropped to about ten minutes: paste in the history, review the draft, send it.
Here's the turn. The extra mistakes that showed up in production were never really the problem. The real problem was who they landed on, and why. Brieflane's investor demo ran the same twelve resumes, over and over, for eight months, and never once invented a fact. That streak felt like proof. It wasn't proof of anything except that twelve specific resumes, all shaped the same way, are safe to show a room full of people who are deciding whether to fund you.
Real job seekers don't arrive in one shape. Some have a single steady job. Many don't: a caregiving gap, a career change, two part-time roles held at once, a business they ran for a few years before it closed. Brieflane's demo set never included any of that, not because anyone decided to exclude it, but because it never came up.
Someone in this room could point out that ninety run-throughs of those twelve resumes, over eight months, sounds like a real track record. It isn't. It's the same twelve items, rerun ninety times. That number measures how well Corwyn remembered what worked. It says nothing about the input distribution real users would actually send, because it never drew from it.
What it cost at its worst: Adaugo Ferriday spent six years as a warehouse supervisor at Larkbrook Distribution, then took a fourteen-month gap to care for a parent, then came back working two part-time roles at once, a logistics coordinator role and a weekend forklift-safety instructor gig. Brieflane's cover letter smoothed her real history into a tidier one: it invented a promotion to "Operations Supervisor" at Larkbrook and quietly extended her employment dates through the gap, so the timeline would read clean. She skimmed the draft under a deadline and sent it. The hiring manager checked her history against Larkbrook's own records, found the mismatch, and rejected her, citing inaccurate application materials.
What I would leave alone: Brieflane's grammar and spelling pass doesn't need any of this scrutiny. It's a separate, deterministic step, and it behaves the same whether someone's history is one job or five. Auditing it by resume shape would be time spent on a place this problem never touches.
The lesson: a demo that never fails hasn't proven the model is ready. It's proven the demo never asked a hard question. Those are very different claims, and only one of them is safe to build a launch decision on.
Now here is the same thing as a story
The short version above is what you say out loud. Read this one when you want to feel exactly what a "clean" cover letter cost Adaugo.
Every Thursday evening, Corwyn ran the same twelve resumes through Brieflane before Friday's investor call, the way you'd run a soundcheck before a show. He'd built the set himself, back when Brieflane was three people and a laptop: a barista turned UX designer, a teacher turned data analyst, ten more like them, each with one steady job leading cleanly into the next. He picked them because they showed the product at its best. Nobody ever told him not to add a messier one. He just never got around to it, and the twelve kept working, so there was never a reason to.
For eight months, that soundcheck never once hit a wrong note. Investors watched a cover letter build itself in nine seconds, accurate down to the dates. At a hiring fair in April, a recruiter fed in her own resume live, on the spot, and Brieflane got every line right. Corwyn started opening pitches with that story.
It hadn't always been twelve. Early on, Corwyn tried closer to forty, pulled from a folder of real applicant resumes a friend had donated. A handful of those forty produced something strange: a job title that didn't quite match, a date that shifted by a year. He dropped them from the rotation, the way you'd cut a shaky song from a set list, not because he decided real-world resumes were unsafe to demo, but because a demo has one job, and it isn't finding the ceiling.
Brieflane opened to the public in June. The first few weeks looked like more of the same. Then, slowly, they didn't. A support ticket here, tagged "wrong dates." Another, "made-up job title." Nothing alarming on its own. By week seven, someone on support noticed all of them shared something: none of the people filing them had a simple, single-job history.
Adaugo Ferriday's cover letter went out in week eight. Six years at Larkbrook Distribution, a fourteen-month gap to care for her father, then two part-time roles at once while she looked for something full time. She opened Brieflane's draft the night before the deadline, skimmed the top half, liked the tone, and sent it. She had no reason to check the middle paragraph line by line. The product had a reputation, even if she'd never seen the investor demos that built it.
The hiring manager who read her letter did what any careful hiring manager does before an offer: cross-checked her history against Larkbrook's own employment records. Larkbrook confirmed six years, warehouse supervisor, ending fourteen months before Adaugo said it did. Brieflane's letter said "Operations Supervisor," continuous, no gap. She got a rejection email citing inaccurate application materials, with no room to explain that the mistake wasn't hers.
Rasmine pulled the numbers the following week. Flagged for a fabricated detail: two percent of letters built from a single, unbroken job history. Thirty-four percent of letters built from a resume with a gap or a career change. Forty-one percent from resumes with self-employment or two concurrent jobs. The demo's twelve resumes, run ninety times, had tested exactly the two percent slice.
Corwyn's twelve resumes were never wrong. That was the whole trap. A wrong demo gets fixed. A demo that's simply too narrow gets trusted, for exactly as long as nobody asks what it never tried.
What I'd tell myself, the week we picked those twelve: a demo isn't there to prove the product works. It's there to show a product working. Those look identical from the audience seats. They are not the same claim, and only one of them was true.
TRACE: the five checks that separate a good demo from real evidence
Not a way to spot a bad demo. TRACE is what you run when a demo looks perfect, because a perfect demo is exactly the one that hides this.
The recap, one line per letter: eight months of demos, run against production's actual mix. Twelve shapes tested, out of a mix that's mostly something else. Ninety reruns is not a sample. Three habits kept the gap invisible. One coverage check would have shown it before launch.
And if you want to be sure it really works, try it somewhere else
Same five checks, a farm instead of a job board, and the missing shape is a blurry photo instead of a messy resume.
Sporewatch reads a photo of a crop leaf and tells a farmer which disease it has, if any. Zophia Marsboom demoed it at agricultural trade shows for most of a growing season, using the same twenty reference photos every time: one leaf, centered, in daylight, one disease clearly visible. Sporewatch called every one of them correctly, every time, for five straight months.
Mapped straight onto TRACE: the timeline is a launch that looked clean for five months before the first blurry, low-light photo exposed the gap. The recut is demo photos (clear, single-leaf, daylight) against real farm photos (blurry, shadowed, sometimes two diseases at once). Assume nothing rules out "twenty photos shown at every trade show all season" as a sample, since it's the same twenty photos, not a random draw from real fields. The cause candidates are the same three habits in new clothes: the twenty photos stayed because they were reliably sharp, Zophia reached for good lighting without deciding to, and nobody ever demoed a genuinely bad phone photo on purpose. The evidence test is identical in shape: check the demo photos against the eval set's own lighting and multi-disease coverage tags before trusting the season's flawless streak.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: compare demo-input diversity against real traffic before trusting any demo streak, full stop.
Cost: no budget to build a bigger eval set this quarter. Tag the eval set you already have by input shape first, that's cheap, then compare demo coverage against it before spending on new data.
The model got better, for real: say Brieflane's fabrication rate on gapped resumes drops to five percent. Keep the coverage check anyway. A smaller gap is still a gap, and it still deserves to be measured instead of assumed away.
Where people run it wrong.
They count demo run-throughs and call the count a sample.
They let the same person who built the demo also decide whether it's representative.
They treat "it demoed fine for months" as evidence about the model, when it's only evidence about the twelve or twenty things it was shown.
How to use it live. When an interviewer says "the demo went great, what's the risk," ask one question back before answering: "how many genuinely different shapes did the demo actually try?" That question alone usually tells you whether the risk is real or already covered.
One alternative worth naming and rejecting: Brieflane could have added a mandatory human review step on every letter before sending. That was on the table. It got rejected, because it breaks the entire reason job seekers use the product, a draft in under a minute, and it doesn't scale past a small user base. The coverage gate catches the same root problem earlier and cheaper, without adding friction to every single letter. The AI-specific failure mode here is hallucination filling a narrative gap the model has barely seen the shape of, and the guardrail is the coverage gate itself: compare input diversity before launch, then tag production complaints by resume shape so a concentrated pattern surfaces in weeks, not months. None of this is free. Building a properly shape-tagged eval set and running the comparison before every release costs real time before a launch date, and keeping it current costs ongoing labeling work as the user base shifts. That cost is worth it, because the alternative is a job seeker losing a real offer to a fact the product invented and never flagged. And the bar itself has to be calibrated, not absolute: ship once the flagged rate on a stratified sample, weighted toward non-linear histories, holds under an agreed line, not when it hits zero, because zero was never real, it was just untested.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't this really just about needing a bigger eval set, nothing to do with demos specifically?" Response: A bigger eval set alone doesn't help if nobody ever compares the demo's coverage against it. The gate has to explicitly check the demo against the eval set, or a narrow demo still gets treated as launch-ready.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on What changes when the product is probabilistic
- #1 Name three product decisions that change when a feature's output is probabilistic rather than deterministic.
- #2 A traditional feature either works or has a bug. Explain why that framing breaks for an LLM feature.
- #3 What does 'correct' mean for a summarization feature? Give a definition your engineering team could test against.
- #4 QA files a bug that reads: the model gave a wrong answer once. How do you triage it?
- #5 Explain the difference between a defect and an acceptable error rate to a non-technical executive.
- #6 Why can you not write an acceptance criterion like 'the output must be accurate' for a generative feature?