ConceptFoundationalModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #3

Describe the difference between a demo and a product, using a concrete example.

PICK · Veilure's virtual makeup try-on matched every shade in the launch demo, then missed nearly a third of them once real faces showed up

Veilure is a phone app. Point your camera at your own face, pick a lipstick or a foundation shade, and it renders that shade onto your skin, live, before you buy anything. Winsome Hallowell owns the try-on model surface. Domhnall Onwuachi leads the small team that built the shade-matching model underneath it. When a rival's launch teaser forced Veilure's own launch decision up by two days, Winsome had to decide what actually counted as proof the product was ready, and the evidence she reached for first was not the evidence that would have told her the truth.

The direct answer
A demo only has to work on a small set of inputs you picked yourself, once, in conditions you controlled. A product has to keep working on inputs you did not pick, over and over, for people you will never watch use it. Before calling anything a product, run it against a large, representative sample of real inputs, including the messy and unflattering ones, not the curated set that made the demo look good.
Do this, in order
  1. Test on a large, representative sample of real inputs before calling it done.Why: this is the whole decision. Everything below only exists to protect it.
  2. Build the sample to match the real spread of skin tones, lighting, and camera angles, not the easy end of it.Why: a hand-picked set of inputs can't reveal a failure it was never built to contain.
  3. Treat a hidden mismatch as the expensive mistake, and pay to find it before launch, not after.Why: that's the actual cost asymmetry, and it only points one way.
  4. Set the pass bar on the worst-performing group, not the average.Why: an average hides exactly the group a curated demo never included in the first place.
  5. Never let a room's applause double as a launch decision.Why: applause measures whether forty photos looked good. It doesn't measure the product.
  6. Where an input genuinely can't vary, skip the extra testing.Why: not every surface carries the same risk, and testing the ones that don't wastes the time you need for the ones that do.

How to answer this, stage by stage

Nobody is grading whether you can define "demo" and "product" like a dictionary. They're grading whether you know that a room clapping is not the same kind of evidence as a real sample, and that mixing the two up is a modeling mistake, not a marketing one.

1
Ground it in one real product and one real decision
Say it like this
"Let's ground this. Say there's a company called Veilure. Its app lets you point your phone at your own face and see what a lipstick or foundation shade would actually look like on you, live, before you buy it. The PM, Winsome, has to decide what actually proves the shade-matching model is ready for everyone, not just for the room she's about to present it to."
Why this works
Keeps the interviewer from grading a textbook definition instead of a real, specific decision.
2
Say your structure out loud
Say it like this
"I'll use PICK. Position, the real distinction between a demo and a product. Impact, what breaks if you treat one like the other. Cost asymmetry, which mistake is cheap and which one is expensive. Kill criteria, the one test that tells you which of the two you're actually holding."
Why this works
Two seconds of structure signals a method, not a mood, before the story does any persuading.
3
State the real distinction, plainly, before any example
Say it like this
"My position: a demo is built to look good on a small set of inputs you picked yourself, once, in conditions you controlled. A product has to look good on inputs you didn't pick, over and over, for strangers, including the ones with bad lighting, dark skin, and a phone held at an odd angle. Forty perfect photos prove the demo works. They don't prove the product does."
Why this works
This is the direct answer, said out loud, before the story has a chance to blur it into "test more."
4
Prove the distinction with a real number
Say it like this
"Winsome's launch demo used forty photos, shot in even studio light, on a narrow band of skin tones. Every shade matched. Ten days after the real launch, across twelve hundred real sessions, the shade came out visibly wrong in thirty one percent of them, and forty eight percent for the darkest skin tones under warm indoor light, the exact combination the demo never once tried."
Why this works
A real before-and-after number turns "demo versus product" from a slogan into something you could check yourself.
5
Name what breaks on each side
Say it like this
"Treat a demo as proof, and you launch something that fails in public, on exactly the inputs you never tested, and you find out from a one-star review instead of a test run. Go the other way and never trust any demo at all, and you never ship anything, because every real product started as a demo once. The actual skill is knowing which one you're looking at."
Why this works
Naming both failure directions stops the answer from collapsing into "just test more," which isn't a decision.
6
Name which mistake is cheap and which is expensive
Say it like this
"Running a slightly less flashy demo, one that's honest about what it hasn't tested yet, costs you some excitement in the room. That's cheap, and everyone forgets it by next week. Shipping something that looked demo-ready but wasn't costs real people a product that made them look bad on their own camera, plus refunds, one-star reviews, and a lot more work to earn the trust back."
Why this works
This is the center of PICK: the two mistakes don't cost the same, and the asymmetry is what tells you which one to risk.
7
Give the one test, out loud
Say it like this
"Before I call anything a product instead of a demo, I ask one question: does it hold up on a large, representative sample of real inputs, including the messy and adversarial ones, not just the curated set that made the demo look good. If the answer's no, it's still a demo, no matter how loud the room was."
Why this works
A kill test with no real check behind it is just a preference wearing a framework's clothes.
8
Close on the one line
Say it like this
"So the whole answer: a demo proves an idea can work once, for the people who built it, under conditions they picked. A product proves it works for the people who didn't build it, on their own bad lighting, on an ordinary Tuesday. Until you've tested the second one, you don't have a product yet. You have a really good demo."
Why this works
Restates the position and hands the interviewer the exact line that would survive a follow-up.

Let's learn

Hand sketched labeled parts diagram titled What Veilure actually does. A rounded square icon in the center labeled Veilure app, with five labels radiating out: camera reads your face, model picks a shade match, render blends it onto skin, ships in under a second, and no person checks it first.
A face goes in, a shade comes back rendered onto it, and it ships with nobody checking first. That last part is why the gap between the demo and the real thing mattered.

Veilure turns your own camera feed into a makeup mirror. Point it at your face, pick a shade, and it shows you what that exact lipstick or foundation would look like on you, rendered live, before you spend a rupee on it.

For the launch review, Winsome built a demo from forty photos. She and her team picked them by hand: bright, even studio light, faces looking straight at the camera, a narrow band of skin tones. Every single shade matched. The whole board watched a foundation render onto a face in under half a second and look exactly right. They approved a full launch, same afternoon.

Hand sketched comparison diagram titled What the demo tested vs who actually shows up. Left panel, a document icon labeled The demo, forty photos, studio light, picked by the team. Right panel, a person icon labeled The real world, 1200 sessions, kitchens, cars, strangers.
The demo answered one question. The real launch asked a much bigger one, on a Tuesday, in a kitchen, under a warm bulb nobody chose.

Ten days after launch, Veilure pulled a sample of twelve hundred real sessions: actual users, actual phones, actual bathrooms and kitchens and car mirrors. The shade came out visibly wrong, too orange, too pale, chalky, ashy, in thirty one percent of them. For the darkest skin tones under warm indoor light, the exact pairing missing from every one of the forty demo photos, it was forty eight percent.

Knowledge spark: why would a model look so sure and still be wrong? The shade-matching model learned mostly from well-lit, front-facing, lighter-toned faces, because that's what most of its training photos looked like. Warm indoor light on darker skin was a corner of the real world it barely saw. It doesn't know that corner exists. It renders with the same confidence everywhere, because nothing in its training ever told it to hesitate there.

Here's the turn. Those extra wrong renders were never really the problem. The problem is that a demo built from forty easy photos got treated as proof the product was ready for twelve hundred hard ones, and nobody had actually checked.

We didn't test whether it worked. We tested whether it worked on the easiest version of the problem, and called that the same thing.
The choice I would take back I would take back treating the forty-photo demo as the launch evidence. Domhnall's team had already built a much bigger, representative eval set, twelve hundred sessions across the real range of skin tones, lighting, and camera angles. It just wasn't finished running when the board meeting got moved up. Winsome shipped on the number that existed and looked good, not the one that would have actually told her something true.

What I would leave alone: the flat little color swatch shown under each shade's thumbnail in the shop doesn't need any of this. It's not a model prediction, it's the manufacturer's own listed color, shown exactly as given. Nothing about a person's face, lighting, or camera angle can make that wrong, so testing it against a thousand real sessions would test nothing real.

The lesson: a demo proves an idea can work once, for the people who built it, under conditions they picked. A product proves it works for people who didn't build it, on their own bad lighting, on an ordinary Tuesday. If the only evidence you have is the first kind, you haven't actually tested the thing that matters yet.

Now here is the same thing as a story

The short version above is what you actually say out loud. Read this one for what it cost Veilure to learn it the slow way.

Winsome Hallowell has run product for Veilure's try-on surface for two and a half years. She can spot a bad render in about half a second, mostly by watching the jawline, where a wrong foundation shade shows up first as a hard, wrong-colored line against real skin.

Domhnall Onwuachi leads the small team underneath her that built the model doing the actual matching: reading a face from camera pixels and deciding which of Veilure's four hundred shades belongs there.

The Tuesday before the board meeting was supposed to be a quiet one. Winsome spent the morning doing the thing she was good at: going through a folder of hundreds of internal test photos and pulling the best forty, the ones where the render landed clean every time. By lunch she had a demo that made a lipstick look correct on every single face in it. She was proud of it. She had earned being proud of it.

Then, that same morning, a rival app dropped a teaser for a near-identical live try-on feature, slicker marketing, a countdown clock on the landing page. Leadership panicked a little. The board meeting that was supposed to happen Friday got moved to that same afternoon, with one line attached: if the demo's ready, we go the same day.

Hand sketched horizontal timeline titled Winsome's launch week. Four milestones: rival's teaser drops, that morning. Board meeting moved up, same afternoon. Forty photos, every shade right, board says go, this milestone emphasized. 1200 real sessions, day 10, 31 percent wrong.
Two hours between the teaser and the meeting. Two things sitting on Winsome's desk. She only carried one of them into the room.

Two things were sitting on Winsome's desk that afternoon. The forty-photo demo, finished and beautiful. And Domhnall's twelve-hundred-session eval, the representative one, maybe sixty percent run, no clean number yet, definitely not ready to present. She didn't have time to wait for the second one. She walked into the meeting with the first.

The board watched every one of the forty photos land correctly. Winsome herself got swept up in how good it looked, the way you do when a thing you built goes exactly right in front of people whose opinion matters. She recommended launching that day. The room agreed in about four minutes.

By the following Thursday, eleven days after launch, the app store reviews had started using the words grey and ashy. A beauty writer with a decent following posted a screenshot of her own render next to Veilure's own marketing image, side by side, captioned "guess which one is the ad." Refund requests came in faster than new sign-ups for most of that week.

The forty photos never lied. They just answered a question nobody in that room was actually asking.

I want to say the problem is that the model got things wrong. It did get things wrong. But that's not really the story. Nobody at Veilure had a number in their head for "ready." They had a feeling, and the feeling only had two settings: the room looked convinced, or it didn't. Forty good photos flipped it to convinced. Nothing in that room was built to notice that the forty photos and the twelve hundred real ones were answering two completely different questions.

Months earlier, when Domhnall's team first proposed the representative eval set as a standing gate before every launch decision, it got scoped down. Building it properly meant real photography across nine skin-tone bands, four lighting setups, and three head angles, real people, real time, and nobody wanted to hold up a fast-moving roadmap for it. So it became a thing the team ran when there was time, not something a launch decision actually had to clear first.

Hand sketched quadrant diagram titled Where the go-live decision actually sat. X axis, how real the sample was, from hand-picked to the real distribution. Y axis, how sure the room felt, from unsure to certain. Four items placed: the 40-photo demo, high on certainty, low on realism, the danger corner. Twelve good demos in a row, also high certainty, low realism. One flawless render, low on both. The 1200-session eval, unfinished, high on realism, middling on certainty.
Everything that made the room feel certain sat in the same dangerous corner: hand-picked, and convincing because of it.

Here's the replay. Say the board meeting again, but this time Winsome walks in with both numbers instead of one. The forty-photo demo, spotless. And the real eval, seventy four percent complete, current read: nineteen percent mismatch overall, climbing to forty four percent in the hardest cell. She recommends a two-week hold instead of a launch.

Domhnall's team spends those two weeks doing exactly what the real numbers point at, mostly recalibrating how the model reads warm indoor light against darker skin. They don't touch the render engine itself, which never had a problem. Mismatch rate on the full representative set: nineteen percent at week zero, eleven at week one, four at week two, under the five percent line the team had agreed on beforehand. Veilure launches in week three instead of that afternoon, to a market that never sees a bad render make the news.

One version of this story ships on the strength of forty photos and spends its first two weeks fighting reviews. The other spends two weeks fixing a known number before a single stranger sees it. Same model, same team, same two weeks either way. The only thing that changed was which number got to be the evidence.

What I'd tell myself, watching Winsome carry those forty photos into that meeting: a demo that makes a room go quiet is real evidence of something. It's just not evidence of the thing everyone in the room decided it was evidence of.

PICK: what forty perfect photos can't prove

Not a rule about being cautious. PICK only earns its keep here if it makes you name, out loud, which of the two mistakes actually costs more, and it isn't the one that feels riskier in the room.

Hand sketched comparison diagram titled The asymmetry, drawn. Left panel, a gauge icon labeled Honest, less flashy demo, costs some applause, forgotten by next week. Right panel, a question mark icon labeled Shipped on a demo that wasn't ready, costs refunds, 1-star reviews, a viral screenshot.
One mistake shows up the same week and costs a little pride. The other hides behind a good-looking number until strangers find it for you.
PPosition. The real distinction.
A demo is optimized to look good on a small, often hand-picked set of inputs, once, in a controlled setting. A product has to work reliably, repeatedly, across the full messy real-world distribution of inputs, including edge cases, adversarial users, and scale.
This isn't a rule about being modest. It's about what the evidence you're holding actually proves, and forty curated photos and twelve hundred real sessions never prove the same thing, no matter how similar they look on a slide.
Say the position before any story. A position built backward from what already broke looks like it was reverse engineered from the ending.
IImpact. What breaks each way.
Mistake a demo for a product, and you greenlight a launch on evidence that only tested the easy inputs. It fails visibly and repeatedly on the real distribution the demo never had to face, and the failure damages trust far more than a slower, honest launch ever would have.
Go the other way and never trust a demo at all, and nothing ships, since every real product started life as something that first worked once, for its own builders, under conditions they picked.
Naming both losses keeps this from reading as "test more," which sounds like advice but isn't a decision.
CCost asymmetry. The heart of it.
Running a slightly underwhelming but honest demo costs some excitement in the room. Cheap, and it's forgotten within a week regardless of who was disappointed. Shipping something that looked demo-ready but wasn't costs real user harm, in Veilure's case a chalky, ashy render that made real people look worse than they do, plus refunds, reviews, and a trust recovery that takes far longer than the launch delay would have. Start by risking the cheap, visible mistake. Only risk the expensive one once the representative evidence actually says it's safe to.
KKill criteria. The one test.
Does it hold up on a large, representative, not hand-picked sample of real inputs, including the messy and adversarial ones, not just the curated set that made the demo look good. Winsome's team considered one other fix after the launch: quietly route the hardest skin-tone and lighting combinations to a static, pre-approved swatch instead of the live render, until the model caught up. It tested fine internally and it got rejected, because it would have quietly given the worst experience to exactly the users the live try-on most needed to work for, and nobody outside the team would ever have known it was happening.
Hand sketched decision tree diagram titled Does it ship, or is it still a demo. Root box, a demo just worked, branching to four labeled conditions: only tested on a picked set, leading to still a demo, test more. Tested broad, skipped the hardest group, leading to not done yet. Tested on the real spread, worst group included, leading to ready to call it a product. No sample, just the room's reaction, leading to not evidence, don't ship.
Four branches, one question asked before anything ships: did the evidence cover the real spread, and did it check the worst group, not just the average.
Cost, by the numbers: catching it before launch versus cleaning up after
260 hrs 130 hrs 0 90 hrs Finish the eval, before deciding 260 hrs Cleanup after, the real launch
Cheap, paid up frontExpensive, paid in public
Finishing the representative eval before the board meeting would have cost about ninety hours of work Domhnall's team had already started. Shipping on the demo alone cost about two hundred sixty hours of hotfixes, support replies, and refund processing in the first two weeks.
The kill line, charted: worst-cell mismatch rate during the held-launch retest
45% 22% 0% kill line: 5% 44% hardest cell 11% 4%, cleared Week 0 Week 1 Week 2
Above the kill lineCleared the kill line
The team tracked the hardest cell, not the average, because the average was already sitting near six percent and would have looked fine on a dashboard the whole time.

The trade worth saying out loud: building and running the representative eval before a launch decision costs real calendar time, in this case about two to three weeks of real photography and labeling across skin tone, lighting, and angle, time a forty-photo demo simply doesn't need. That's slower and more expensive up front. It's worth paying, because the alternative cost, a public launch that fails on exactly the users it needed to work for, is bigger every single time.

And if you want to be sure it really works, try it somewhere else

Same four letters, an HVAC diagnostics tool instead of a makeup mirror, and the fix a field team almost trusted would have been wrong for just as specific a reason.

Coilwise runs a phone app for HVAC field technicians. Point it at a rooftop unit, and it tells you which part is likely failing, a bad capacitor, a clogged coil, a dying motor, before the technician opens the panel. Its demo, built from photos of a handful of clean, well-lit units on a training rooftop, diagnosed every fault correctly. Real units live in dim mechanical rooms, coated in dust, lit by a phone's own flash, across a dozen brands the demo never touched.

Hand sketched metaphor scene titled The same shape, a different shop. Left panel, a document icon labeled The demo, bright rooftop photos, one clean unit. Right panel, a person icon labeled The real job, dim plant rooms, every brand, dust and glare.
Coilwise's PM had the same two options Winsome did. One tests the model against the job it will actually be asked to do. The other tests it against the job that was easiest to photograph.

Position: a demo unit, clean and well lit, proves the model can read a fault when nothing is fighting it. A product has to read the fault through dust, glare, and a brand the model has barely seen. Impact: launching on the demo alone risks technicians trusting a wrong diagnosis and swapping the wrong part, which costs a customer a second visit and Coilwise a technician's trust in the tool. Cost asymmetry: a slower rollout, tested first against a real sample of dim, dirty, mixed-brand units, costs a few weeks of field data collection. A wrong diagnosis in the field costs a wasted service call and a technician who stops trusting the app on every job after, not just the ones it got wrong. Kill criteria: does the model hold up against a large, representative sample of real field photos across brand, lighting, and unit condition, not just the demo units it was shown on a clean rooftop.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the kill test before anything else gets discussed.
Cost: no budget for a full representative sample before the deadline. Fine, but test the highest-risk cells first, the darkest skin tones or the dirtiest mechanical rooms, not a bigger pile of the same easy inputs.
The model got better, for real: say a new checkpoint scores near-perfect even on the hard cells. Still don't skip the representative retest. "Better on average" can still hide one bad cell nobody looked at.

Where people run it wrong.
They mistake a room going quiet for the product actually working, when it only proves the forty photos were chosen well.
They test more of the same easy inputs to make the number look better, instead of harder new ones to learn something true.
They patch visible failures one ticket at a time as users report them, instead of asking whether a representative test was ever run at all.

How to use it live. If you're ever handed a glowing demo mid-interview and asked whether you'd ship it, buy yourself a second by asking one thing out loud: "how was that test set actually chosen." That question is the whole method, said as a question instead of a rule.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question asking you to tell a demo apart from a product?
Tap to flip
ANSWER
PICK: state the real distinction as a position, name what breaks on each side, find which mistake is cheap versus expensive, then give the one test that tells you which one you're actually looking at.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Winsome Hallowell, who owns the try-on model surface at Veilure, an app that renders makeup shades onto a user's own face through their phone camera before they buy.
3 · THE REAL DISTINCTION
What actually separates a demo from a product?
Tap to flip
ANSWER
A demo has to work on a small, hand-picked set of inputs, once, in a controlled setting. A product has to work reliably, repeatedly, on the full messy real-world distribution of inputs, including the edge cases and the adversarial ones.
4 · THE SPLIT
What breaks in each direction if you mix a demo up with a product?
Tap to flip
ANSWER
Treat a demo as proof, and you launch something that fails visibly and repeatedly on the real inputs it was never tested on. Never trust any demo at all, and nothing ships, since every product starts as a demo once.
5 · THE ASYMMETRY
Which mistake is cheap here, and which one is expensive?
Tap to flip
ANSWER
An honest, less flashy demo costs some excitement in the room, cheap and forgotten within a week. Shipping on a demo that wasn't ready cost real user harm, refunds, and a much harder trust recovery, harder to undo than the excitement was to lose.
6 · THE NUMBER
Fill in the blank: the launch demo used ___ photos with ___ percent of shades matching. The real launch, sampled across ___ real sessions, missed ___ percent overall and ___ percent for the hardest skin-tone and lighting combination.
Tap to flip
ANSWER
40 photos, 100 percent matched. 1,200 real sessions, 31 percent wrong overall, 48 percent for the darkest skin tones under warm indoor light, the exact combination missing from the demo.
7 · THE KILL TEST
What's the one test that tells you whether you're looking at a demo or a product?
Tap to flip
ANSWER
Does it hold up on a large, representative, not hand-picked sample of real inputs, including the messy and adversarial ones, not just the curated set that made the demo look good.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of the forty curated photos there?
Tap to flip
ANSWER
Coilwise's HVAC fault-diagnosis app. The role goes to the handful of clean, well-lit demo units on a training rooftop, standing in for the dim, dusty, mixed-brand units its technicians actually photograph in the field.

Check yourself Score: 0 / 0

Fill in the blank
1. Veilure's launch demo used ___ photos and matched ___ percent of shades. The real launch, checked across ___ real sessions, missed the shade in ___ percent overall.
Show hint
Look at the numbers under the direct answer and in the story's opening.
Show answer
40 photos, 100 percent matched. 1,200 real sessions, 31 percent wrong. The gap between those two numbers is the whole reason a demo can't stand in for a product.
Multiple choice
2. Why did Veilure's real launch fail on exactly the darkest skin tones under warm indoor light?
  • A. The render engine can't handle warm lighting of any kind.
  • B. That combination was missing from the forty-photo demo set, so nobody had checked whether the model handled it before launch.
  • C. Veilure's camera hardware doesn't work well indoors.
  • D. Darker skin tones were deliberately excluded from the product on purpose.
Show hint
Check what the demo's forty photos actually covered, and what they didn't.
Show answer
B. The demo's studio-lit, narrow-skin-tone photos never included that combination, so the launch decision had no real evidence either way about how the model handled it.
True or false
3. True or false: once Winsome's forty-photo demo landed every shade correctly in front of the board, that was strong evidence the model was ready to launch to everyone.
  • True
  • False
Show hint
Check what forty hand-picked photos can and can't prove about the real distribution of users.
Show answer
False. A small, hand-picked set proves the demo works. It says nothing about the far larger, messier range of real skin tones, lighting, and angles the product would actually face.
Short answer, name the rejected alternative
4. After the failed launch, what fix did Winsome's team consider and reject, and why did it lose?
Show hint
Look at the K, kill criteria, section of the framework recap.
Show answer
Model answer: Quietly route the hardest skin-tone and lighting combinations to a static, pre-approved swatch instead of the live render, until the model caught up. It tested fine internally and got rejected because it would have quietly given the worst experience to exactly the users the live try-on most needed to prove it worked for, with nobody outside the team ever knowing.
Short answer, apply it yourself
5. Think of an AI feature you've used that impressed you the first time. What would you actually need to see, beyond that one good moment, before you'd call it reliable?
Show hint
Think about what conditions your one good moment happened under, and what it left untested.
Show answer
Model answer: A photo-cleanup app removed a stranger from my beach photo perfectly on the first try. Before I'd trust it, I'd want to see it handle a cluttered background, a person overlapping the subject, and low light, not just one clean shot with a plain sky behind the person removed.
Short answer, work the number
6. If the demo's forty photos had included even five real examples of dark skin under warm indoor light, and all five had rendered visibly wrong, would the board likely have still approved a same-day launch? Why or why not?
Show hint
Think about what a small but honest sample would have shown, versus a small curated one.
Show answer
Probably not. Five wrong renders out of forty, visible in the room, would have been a different kind of evidence entirely. It doesn't take twelve hundred sessions to catch a real problem, it takes a sample that actually includes the hard cases instead of quietly leaving them out.
Before you close the answer
Why this works
Tests whether you understand that "demo versus product" is really a question about what a small, hand-picked sample can and can't prove about a model's behavior on the real distribution. Most candidates answer with generic MVP-versus-final-product advice. The real judgment is knowing that a model can look confidently right on the easy inputs and confidently wrong on the ones it never saw.
Follow-up traps
"Isn't it unrealistic to test every possible input before any launch?" Response: nobody's asking for every input, just a representative sample large enough to include the hardest real cells, and a kill line set on the worst cell, not the average. That's a bounded, doable amount of work, not an infinite one.

"What if the board had insisted on launching anyway, evidence or not?" Response: then the honest move is to say plainly what the demo does and doesn't prove, in the room, before the decision gets made, not after the reviews come in. Silence at that moment is what turned a demo into an accidental launch decision.
If pressed
The real fix split into two parts: the render engine itself never had a problem, so it was left untouched. The actual gap was in how the shade-matching step read warm, low color-temperature light against higher melanin skin, which needed more training examples in that specific lighting band, not a bigger model or a different architecture.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more