CaseIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #1

Your CEO saw a demo on social media and wants that feature in six weeks. Structure your response.

The direct answer
Don't say yes and don't say no. Run the demo's own approach against real, messy input first, so you know exactly where it already works and where it doesn't. Then bring back a scoped pilot that fits six weeks, built only from the part that already works, plus an honest number for the rest.
Do this, in order
  1. Test the demo's approach against real, messy input before you answer at all.Why: a demo proves one thing works once. It says nothing about the range of things a real product has to handle every day.
  2. Check your own gut number before you blame the demo for being unfair.Why: if six weeks really is wrong, you want to know that because you did the math, not because it felt too fast.
  3. Name the specific reasons the demo looked further along than it was.Why: "the demo was cherry-picked" is a feeling. Three named, checkable reasons is something a CEO can act on.
  4. Bring back a scope that fits six weeks, not a flat no.Why: structuring your response means giving the CEO something to say yes to, not just a reason the ask was wrong.
  5. Keep a person in the loop on every result until the real numbers earn the right to remove them.Why: the part of the demo that skipped this is exactly the part that would hurt someone if it shipped as shown.
  6. Give the CEO something shipping this week, not just a longer date.Why: a pushback with no near-term win reads as stalling, even when the math behind it is right.

How to answer this, stage by stage

Seven moves. This is a diagnosis wearing a deadline's clothes, so most of the work is finding out what's actually true before you push back on anything.

1
Scope it to one product and one message
Say it like this
"Let me put a real shape on this. Say we run a callback line for patients who couldn't get through, and our nurses turn each voicemail into a note by hand. Our CEO forwards a video of an AI tool doing that in four seconds and says, 'get this live in six weeks.'"
Why this works
A vague "structure your response" question invites a vague answer. Naming the product and the exact ask gives every later step something real to check against.
2
Say what the demo actually proved, not what it implied
Say it like this
"Before I plan anything, I'd separate two claims. The demo proved the model can turn one clean voicemail into a tidy note. It did not prove it can do that for the range of voicemails we actually get. Those are different facts, and the CEO is reacting to the second one on the strength of the first."
Why this works
Naming this gap out loud, calmly, is the whole skill in one sentence. It stops the conversation from turning into "you don't believe in AI" versus "you're moving too fast."
3
Rule out your own bad guess before you doubt the demo
Say it like this
"And before I go to the CEO with 'that's not realistic,' I'd check whether I'm the one who's wrong. So I'd spend two hours with an engineer actually breaking six weeks into parts, because 'that feels too fast' is not something I want to say out loud without doing the arithmetic first."
Why this works
This is the move most people skip. Pushing back on a deadline without checking your own number first sounds exactly like someone protecting their own comfort, whether or not it is.
4
Name three specific reasons the demo looked done
Say it like this
"Three things, and they're each checkable. One, the demo used one voicemail, recorded for the video, in a quiet room, about one thing. Two, it never showed what happens when the model gets a detail wrong. Three, it ran once, by hand, never touching our real phone system or our real call volume."
Why this works
"The demo was cherry-picked" is an opinion. Three named conditions turn it into something the CEO can go check themselves, which is what actually earns trust back.
5
Run the one test that turns a guess into a number
Say it like this
"Here's the test. Same kind of AI, no new engineering yet, run against fifty real voicemails we already have on file. Count how many come out as a note a nurse could use without listening to the call again. That's the number I bring back, not a feeling."
Why this works
This is the strongest single move in the whole answer. It replaces "I don't think this is realistic" with a number nobody in the room can argue with, because it came from real calls.
6
Bring back a scope, not a no
Say it like this
"Here's what I'd actually propose. In six weeks, we ship this for the calls it's already good at: one concern, said clearly, and a nurse still checks the note before it's saved. The harder calls, the ones the demo never showed, need about nine weeks, because that's where the real engineering is."
Why this works
This is what "structure your response" is really asking for. Refusing the deadline is easy. Handing back a version of it that's actually true is the hard, valuable part.
7
Close on the line that turns a deadline into a decision
Say it like this
"So six weeks isn't wrong, it's just pointed at the wrong slice. We can hit six weeks on the calls that look like the demo. We need three more for the calls that don't, because that's exactly where getting it wrong would hurt someone."
Why this works
It gives the CEO a yes to say. Most candidates end on a warning. Ending on a decision the CEO can actually approve is what separates a good answer from a defensive one.
If you remember one thing Stages 3 and 5 are what the interviewer is grading. Check your own number before you doubt the demo's. Then run one test that turns "this feels too fast" into a number from real calls. The scope you bring back is only as good as that test.

Let's learn

The tool listens to a voicemail a patient leaves when they can't get through, and turns it into a short note a nurse can read in seconds instead of playing the whole message.

Petra Lindqvist has been the product manager for the nurse tools at Coverpath Health for fourteen months. Right now, the callback line gets about forty-five voicemails a day. Four nurses take turns listening and writing a note by hand, about three minutes each, call and note together. That's a little over two hours of listening spread across the team, every single morning.

Three weeks ago, a rival health-tech company posted an eighteen-second video. One voicemail, a patient asking to move an appointment, goes in. A clean, three-line note comes out in about four seconds. It got shared everywhere. On Tuesday, the CEO forwarded it in Slack: "Can we have this in six weeks?"

Petra didn't have a number yet, so she went and got one. She took the same kind of AI the demo used, no new engineering, and ran it against fifty real voicemails already sitting in Coverpath's files.

Timeline: demo posted with one clean call, then the CEO's ask to ship in six weeks, then fifty real calls tested the same week, then a scoped pilot proposed at six weeks with the rest at nine.
The demo happened weeks before anyone tested it on a real call

Here is the turn. The model isn't bad, and the demo didn't lie about what it showed. The problem is that the demo picked the easiest half of the job and let everyone assume it was the whole job. Twenty-seven of the fifty real voicemails, about half, came out as a note a nurse could use without listening again. The other half needed real, careful review, and some needed the nurse to replay the message from the start.

The demo didn't need a better model. It needed the twenty-three calls it never played.

What made those twenty-three hard was never about the model getting the words wrong. Patients calling a health line don't talk like the demo's actor did. They mention a medication and a symptom and a reschedule in one rambling message. A toddler is crying in the background. Someone pauses mid-sentence to find their pharmacy's name. None of that showed up in one clean, scripted call.

Knowledge spark: what "usable without a rewrite" means here Petra's team read every AI note next to the real voicemail. A note counted as usable if a nurse could act on it straight away, no re-listening, no missing detail. A note that read fine but left out a medication, or got a symptom slightly wrong, did not count, even though it looked done on the screen.

The choice I would take back. Fourteen months ago, when the callback line first launched, Coverpath decided to delete the audio recording thirty days after a note was filed. It was the sensible call at the time: less patient data sitting around is safer, and storage costs money. Nobody was thinking about an AI needing to be double-checked against the original call someday.

What I would put back Keep the recording linked to every AI-written note for as long as the note stays in the chart, not just thirty days. If a nurse ever needs to check what the AI heard against what the patient actually said, the only source of truth should not have quietly expired.

What I would leave alone. The twenty-two voicemails that were one clear concern in a quiet room worked fine, seventeen of twenty-two usable straight away. There's no reason to slow that part down and make it wait for the hard half to catch up. Ship the easy slice fast. Holding it back until the whole feature is perfect helps nobody, and it wastes the one part that's already working.

The lesson. A demo is a promise about one moment, not about a Tuesday morning with forty-five real calls in the queue. My old self would have argued about whether six weeks was fair. My new self runs the fifty calls first, and lets the number do the arguing.

Three reasons a demo gets ahead of itself

Not because anyone lied. Because a demo is built to show the best case, and nobody in the room asked what got left out.

Three panels: one clean call recorded in a quiet room for one concern, no wrong note ever shown, and the demo run once offline with no real call load.
Three separate, checkable reasons, not one reason wearing three hats
Reason 1
One clean call. The demo's input was made for the video.

Someone recorded that voicemail to be shown, not to be typical. Quiet room, one concern, said clearly, no restarts. A real patient's voicemail is none of those things most of the time, and the demo never had to prove it could handle the other kind.

How you'd check it: run the exact same pipeline against real voicemails and see how the result changes when the room gets noisy or the message covers more than one thing.
Reason 2
No wrong note, ever shown. The demo never had a bad day.

An eighteen-second video has no room for the model getting a dose or a symptom wrong. So nobody watching it ever saw what a mistake looks like, or what happens next when one shows up. A feature that's ready to ship needs an answer for that moment. The demo simply never had one.

How you'd check it: look for notes that read as complete and clean but leave out or misstate a real detail from the call. Those are the ones a nurse would sign off on by mistake.
Reason 3
One offline run. The demo never touched the real system.

The demo's clip was handed to a model once, by a person, off to the side. It never went through a real phone system, never had to keep up with forty-five calls landing in a morning, and never had to meet the record-keeping rules a health company has to follow. None of that shows up on video, and none of it is optional in production.

How you'd check it: ask what the demo's pipeline would need to plug into a live phone line and store notes the way the rest of the record system already does. If the answer is "nothing built yet," that's real work, not polish.

The test that turned a guess into a number

Reason one is the strongest, and it's the one you can actually measure without building anything new. Petra's team split the fifty real voicemails by what kind of call they were, then checked how many came out usable in each group.

Usable notes, out of fifty real voicemails, by call type
17
One concern, quiet room
17 of 22 usable
6
Two or more concerns
6 of 16 usable
4
Noisy or rambling
4 of 13 usable
Overall, 27 of 50 real calls came out usable, 54 percent, next to the demo's one call at 100 percent. The average number hides that the tool is already good at one thing and not yet good at two others.

That's the number that changes the conversation. It isn't "the demo was fake." It's "the demo showed us the 22-call segment, and we needed to know about the other 28."

What six weeks actually buys, and what it doesn't

The by-segment number tells Petra where the tool already works. It doesn't tell the CEO what a real build costs. So she and an engineer broke the honest build into parts, in weeks, and lined it up against the six-week ask.

Weeks needed for a full build, against the six-week ask
2 wk
2 wk
3 wk
1.5 wk
6 weeks
Tune the model on real calls
Nurse review screen for flagged notes
Phone system + record-keeping
Test on real calls, one site first
The honest full build is about 8.5 weeks. The six-week mark lands inside the phone system and record-keeping work, before it's even finished, let alone tested.

So the counter isn't "six weeks is impossible." It's "six weeks buys you the first two pieces, on the calls that are already working." That's exactly what Petra brings back.

The six-week pilot
One concern calls only
Nurse confirms every note
17of 22 already usable
The full build
All call types
Phone system + records done
8.5weeks, honestly

The CEO gets something real in six weeks. The nurses get a tool that never quietly hides a symptom, because every note still passes a human before it's saved. And the hard 28 calls get the three extra weeks the demo never told anyone they'd need.

TRACE, in the order Petra actually used it

This is a diagnosis question dressed up as a deadline, so TRACE is the framework, not a story about a habit changing. A "what if the error rate doubled" question would reach for FLIPS instead.

T, timeline. The demo posted three weeks before the CEO's message. Petra's team had real voicemails on file the whole time, just never run through an AI. The gap between "a demo exists" and "someone tested it" is the whole diagnosis.
R, recut. Split the calls by what kind of call they are. One clean number, 54 percent usable, hides a 77 percent segment and a 31 percent segment sitting on either side of it.
A, assume nothing. Before blaming the demo, Petra checked whether six weeks was simply her own wrong guess. It wasn't, quite: a narrower six-week version turned out to be true, once she stopped assuming the ask meant the whole feature.
C, cause candidates. Three, named and separate: one clean call, no wrong note ever shown, one offline run. Not a list of every possible complaint about demos.
E, evidence test. The same kind of AI, run against fifty real voicemails, split by call type. The strongest move in the whole framework, and the one the CEO can't argue with, because it came from real calls, not an opinion.
Why E is the hard step Anyone can say a demo was cherry-picked. A test earns its place by using the demo's own approach, unchanged, against your own real data. If you change the model and the data at the same time, you can't tell which one moved the number.

Run TRACE on a different deadline

An HVAC dispatch startup's CEO sees a viral photo demo: point a phone at a broken air conditioner, and an AI names the exact failed part in five seconds. The trade show is in four weeks, and the CEO wants it in the app by then.

T. The photo in the demo was posted by an enthusiast five weeks earlier, well-lit, the broken part fully exposed and already labeled online.
R. Real technician photos split into three kinds: a clear shot of an open panel, a blurry phone photo taken fast in a cramped attic, and a photo of the outside case before anything's been opened at all.
A. The dispatch PM checks their own instinct that four weeks is impossible, and finds that a narrower version, first-pass suggestion only, technician still decides, might actually fit.
C. Three causes: the demo's photo was staged and labeled, the demo never showed a wrong guess, and the demo used one photo picked by its own maker instead of a real batch.
E. Run the same model against forty real photos technicians already uploaded last month for parts orders, and count how many land a correct guess in the top three, split by photo type.

Swap the trigger and it still runs

  • Speed: the CEO wants it by a specific event date instead of a fixed number of weeks. TRACE still starts with the demo's own conditions, not the calendar.
  • Cost: the CEO wants it built cheaply, using an off-the-shelf model, because the demo used one too. The build-up chart still has to show what a cheap model skips, not just what it costs.
  • The model really is good: the case on this page. The tool isn't fake or broken. It's just been tested on a fifth of the real job, and everyone assumed the rest for free.

Where people run it wrong

  • Answering "yes" or "no" to the deadline before running any test. Both answers are guesses dressed up as a decision.
  • Arguing that the demo was unfair instead of naming what, specifically, it skipped. A feeling convinces nobody; three checkable reasons do.
  • Bringing back a scope with no numbers attached. "It'll take longer" without the by-segment split or the week build-up is just a longer guess.

If you are asked this cold

Buy yourself ten seconds by naming the two claims out loud. "So there's what the demo proved, and what the CEO is hoping it proved, and those aren't quite the same thing yet. Let me say what I'd actually check first." That's not stalling. That's stage one, and in nearly every version of this question, that's where the real answer starts.

Flashcards (click a card to flip it)

This is a diagnosis question, not a flip story, so these eight test the TRACE moves and the real numbers instead of a flip family.

1 · THE FRAMEWORK
Which framework fits "the CEO wants the demo feature in six weeks," and why?
Tap to flip
ANSWER
TRACE. The real question hiding under the deadline is a diagnosis: why does the demo look done when it isn't yet. FLIPS is for "what if something changed," and nothing has changed here yet.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Petra Lindqvist, product manager for the nurse tools at Coverpath Health, fourteen months in. She owns the callback line the CEO wants rebuilt around AI.
3 · RULING HERSELF OUT
What does Petra check about her own judgment before she doubts the demo?
Tap to flip
ANSWER
Whether her own gut feeling that six weeks is too fast is actually correct. She times out a real estimate with an engineer instead of just trusting the feeling.
4 · THE THREE REASONS
Name the three reasons the demo looked further along than it was.
Tap to flip
ANSWER
One clean call made for the video. No wrong note ever shown. One offline run that never touched the real phone system or call volume.
5 · THE NUMBER
Of the fifty real voicemails tested, how many came out usable without a rewrite?
Tap to flip
ANSWER
27 of 50, 54 percent. That splits into 17 of 22 for single-concern calls, 6 of 16 for multi-concern calls, and 4 of 13 for noisy or rambling ones.
6 · THE CHECK
Name the one test that turned a guess into a number, and what it holds still.
Tap to flip
ANSWER
The exact same kind of AI the demo used, unchanged, run against 50 real voicemails already on file. Same approach, real data, so the result can't be blamed on a different model.
7 · THE COUNTER-SCOPE
What does Petra actually bring back to the CEO?
Tap to flip
ANSWER
A six-week pilot for single-concern calls only, always confirmed by a nurse before saving, on one site. And an honest 8.5-week timeline for the harder calls the demo never showed.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE on a different product with a different deadline. Which one, and what's the twist?
Tap to flip
ANSWER
An HVAC dispatch app whose CEO wants a photo-based part-diagnosis feature for a trade show in four weeks. The demo used one staged, well-lit photo of an already-labeled part.

Check yourself Score: 0 / 0

True or false
1. True or false: since the demo already proved the model works, the fastest path is to start building the full feature right away.
  • True
  • False
Show hint
Think about what the demo actually tested, versus what it looked like it tested.
Show answer
False. The demo proved the model can handle one clean call. It says nothing about the range of real calls a shipped feature has to survive. Building straight from the demo skips the test that tells you where it's actually safe to ship first.
Multiple choice
2. Petra runs the same kind of AI the demo used against fifty real voicemails instead of just judging the demo on sight. Why?
  • A. To collect more training data for the model.
  • B. To find out how the model does on the real range of calls the demo never showed.
  • C. To prove the CEO wrong about wanting AI at all.
  • D. To make the pilot's launch date look better on paper.
Show hint
One of the three reasons the demo looked done gets tested directly by this move.
Show answer
B. This test turns reason one, that the demo used one curated call, into a number: 27 of 50 real calls came out usable. That number is what makes the counter-scope believable instead of defensive.
Fill in the blank
3. Of the fifty real voicemails Petra tested, ______ produced a note a nurse could use without a rewrite.
Show hint
17 of 22 for single-concern calls, 6 of 16 for multi-concern calls, 4 of 13 for noisy ones.
Show answer
27, or 54 percent. That overall number hides a real split: 77 percent usable on the easy segment, down to 31 percent on the noisy one. The average is the least useful number in the whole test.
Short answer
4. Name a segment of real voicemails where the tool is already good enough to ship fast, and say why holding it back would be a mistake.
Show hint
Look at the by-segment chart. One bar is already close to the demo's own number.
Show answer
Model answer: "Single-concern, quiet-room calls. 17 of 22 already come out usable. Making that segment wait for the hard 28 to catch up wastes the one part of the feature that's already proven, and it gives the CEO nothing to see for six weeks instead of something real."
Multiple choice
5. A colleague offers four explanations for why the demo looked more finished than it was. Which one is a genuinely separate, checkable reason, not just a restatement of "they're better than us"?
  • A. The demo's AI model is simply more advanced than what we could build.
  • B. Their engineering team is more talented than ours.
  • C. The demo never showed the tool handling more than one concern in a single message.
  • D. Their marketing budget made the video look more polished.
Show hint
Three of these end with "so they're just better than us." One names a specific thing you can go test.
Show answer
C. A, B and D are all versions of "they're better," which nobody can act on. C names an actual gap you can go check against real calls, which is exactly what turned into the by-segment chart.
Short answer, apply it yourself
6. Think of a demo you've seen for something you use yourself, a phone feature, an app, anything that looked like magic in the video. What's one thing about your own daily use that the demo never had to handle?
Show hint
Think about noise, multiple things happening at once, or a case the demo never showed going wrong.
Show answer
Model answer: "Voice-to-text on my phone. The demo I saw was one person, one sentence, in a quiet studio. My actual use is in the car, with the radio on, saying two things at once because I'm in a hurry. The demo never had to survive the radio." Any honest answer works if it names a real condition the demo skipped, not just "it doesn't work as well for me."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more