CaseAdvancedModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #15

What do you do when leadership wants an AI feature primarily for the press release?

FLIPS · scoping Nowcast, Corvenna's press-ready trend-prediction feature for marketing teams

Corvenna reads a brand's social accounts and tells its marketing team which content angles are about to take off. Nowcast is the feature that predicts trends before they peak. Palmira Vosberg owns Nowcast's roadmap. Three weeks before Corvenna's Series B announcement, her CEO wants Nowcast as the centerpiece slide, and the number it was actually built on turns out to be three accounts, not fourteen hundred.

The direct answer
Let the press date stand, but never let it be the whole spec. The moment a feature gets scoped for an announcement, pair it with a second bar: a real, measured user-value target, checked against a genuinely representative sample of accounts, not the cherry-picked ones in the demo. Nothing goes to everyone until both bars clear, even when the demo already looks like a standing ovation.
Do this, in order
  1. Set a real, measured user-value bar the moment the press date gets set, not after.Why: this is the actual reversal; skip it and "it'll demo well" quietly becomes the entire spec.
  2. Treat the announcement as one legitimate constraint, not something to refuse.Why: blocking it outright reads as obstruction and burns trust you'll need for the next hard ask; the job is to add a bar, not remove one.
  3. Validate against a representative, stratified sample of accounts, never the flagship demo accounts alone.Why: a hit rate on three cherry-picked accounts says nothing about the other 1,397.
  4. Split the launch gate: a demo can ship as a labeled preview before the eval bar clears, but general availability cannot.Why: this is where a merged launch checklist gets un-merged; a demo passing and an eval passing are two different questions.
  5. Give thin-data accounts an honest "still gathering signal" state instead of a confident wrong call.Why: an honest gap costs nothing; a wrong, confident call costs a customer's actual ad budget.
  6. Skip the two-bar ritual for genuinely low-stakes AI polish features.Why: the overhead is worth paying when a wrong call costs money or trust, not on a caption-suggestion tweak nobody's betting a press cycle on.

How to answer this, stage by stage

Nobody is grading whether you can describe a polite way to slow leadership down. They are grading whether you'll notice that "make it announceable" quietly ate the whole spec, and catch it before a demo that worked on three accounts ships confidently wrong to the other fourteen hundred.

1
Scope it to one product, one asker, one moment
Say it like this
"Let's ground this. Corvenna does social analytics for marketing teams; Nowcast is the AI feature that predicts trends before they peak. Palmira Vosberg owns its roadmap. Three weeks out from a Series B announcement, her CEO wants Nowcast as the centerpiece, and the number behind it is three accounts, not the fourteen hundred that'll actually use it."
Why this works
Naming the product, the feature, and the person stops the answer from staying a general statement about handling a demanding stakeholder.
2
Say your structure out loud
Say it like this
"I'll run this as FLIPS. Find the person whose scoping habit changes. Locate what she used to optimize for without saying so. Identify the flip, the exact moment 'announceable' stops being the whole spec. Pinpoint the old launch decision that only made sense before. Show the replay with the new gate in place."
Why this works
Two sentences of structure signal a plan before any story starts.
3
Reframe what's actually being tested
Say it like this
"This isn't really asking whether I can push back on an exec. It's asking whether I'll notice that a demo proves a feature can be right, never how often, and stop those two questions from getting collapsed into one before the feature reaches everyone."
Why this works
Compresses the whole answer into one breath before a single detail can bury it.
4
Give the one decision
Say it like this
"Here's what I'd actually do. Announceable stays a real requirement, but it becomes one of two. The second is a hit rate on a stratified sample of real accounts, not the flagship ones, with a real number attached. A demo can ship early as a clearly labeled preview. Nothing rolls out to every account until both bars clear."
Why this works
This is the direct answer, said plainly, before the story arrives to earn it.
5
Prove it with a compressed failure
Say it like this
"Here's what happens without it. Nowcast calls nine of ten trend windows right on three flagship demo accounts. Corvenna ships it to all 1,400 accounts the same week. A stratified sample of 180 accounts comes back at 38 percent. A skincare brand with 2,800 followers bets a full week's $16,000 ad budget on one of Nowcast's confident calls, and it flops, because their account never had enough history for the model to have real signal."
Why this works
Four sentences carry a whole incident that a full retelling would take a page to earn.
6
Name the AI-specific detail you'd hold onto
Say it like this
"The real fix wasn't a smarter model. It was admitting three cherry-picked accounts are not an eval set. We built a stratified sample across follower bands, held it to a real pass bar, and added a guardrail: accounts under 4,000 followers or under 90 days of history get 'still gathering signal' instead of a confident call. That's not a promise Nowcast is always right. It's a threshold, on a real sample, that it has to clear before it gets to sound confident."
Why this works
Shows the judgment is about eval-set representativeness, not model quality in the abstract.
7
Close on the one line
Say it like this
"So: 'leadership wants it for the press release' is never the whole scope, it's the first bar of two. Somebody has to name the second one, in writing, before a single account outside the demo sees the feature, or the team just gets better and better at shipping things that were only ever tested on the three accounts built to make them look good."
Why this works
Leaves the interviewer with the actual decision, not just a well-told story about one demo.

Let's learn

Corvenna is a dashboard that watches a brand's social accounts and tells its marketing team what to post next. Nowcast is the feature inside it: three predicted trend windows a day, tap to add one straight to the content calendar. Before Nowcast, a marketing team found trends by scrolling their own feed mid-morning and guessing. After Nowcast, it's a card with a confidence-sounding badge on it.

Hand sketched two panel comparison titled The I step, in one picture. Left panel a gauge icon labeled The small move, caption a press-tied AI ask lands, Palmira scopes it to whatever will demo best, same as always. Right panel a question mark box icon in rust orange labeled The big snap, caption now announceable is one bar, not the whole spec, a real measured user bar has to clear too, every time.
Nothing about the ask itself changed. What Palmira required before treating it as scoped changed completely.

In the internal demo, built on Corvenna's three best-instrumented flagship accounts, Nowcast called nine of the next ten trend windows correctly. Leadership loved it enough to put it on the Series B deck. Palmira shipped it to all 1,400 accounts the same week the demo passed, because Corvenna's launch checklist had one gate: the demo works, ship it everywhere.

Nowcast hit rate: three flagship demo accounts vs. a representative 180-account sample
100% 50% 0 90% Demo, 3 accounts 38% Representative, 180 accounts
Flagship demo accountsStratified real sample
A demo that looked like a slam dunk was measuring a population of three. The other 1,397 accounts were a different question entirely.

Here's the turn. Those misses weren't spread evenly. When Corvenna finally stratified that 180-account sample by follower count, accounts under 4,000 followers came back at 21 percent. Accounts over 50,000 followers, the same size as the demo accounts, came back at 61 percent. Nowcast wasn't a little worse everywhere. It was excellent exactly where the demo happened to look, and thin everywhere else, because smaller accounts simply don't have enough posting history for the model to have real signal to work from.

At its worst, Orla Duchamps, the marketing manager at Petalcombe, a skincare brand with about 2,800 followers, bets the brand's full week of paid-boost spend, $16,000, on a Nowcast call recommending a specific content angle. It flops. Engagement lands well under Petalcombe's own baseline. Orla wasn't careless. The card looked exactly as confident as it did for the flagship accounts in the demo, because it was built from the same badge, the same UI, no visible difference at all.

We didn't ship Petalcombe a trend prediction. We shipped them the confidence of an account eight times their size, wearing their badge.
The choice I would take back Corvenna's launch checklist for press-tied AI features had one gate: the demo passes, it ships to everyone. That was fine for two years, because everything before Nowcast was caption suggestions and best-time-to-post nudges, small utility features where being wrong cost someone a re-type, not a week's ad budget. Nobody had ever needed a second gate, so nobody built one. The same shortcut that kept features shipping fast is what let a three-account demo stand in for 1,400 real ones.
Knowledge spark: what makes an eval set "representative"? It means the accounts you test on look like the accounts that'll actually use the feature: a real spread of follower counts, industries, and history lengths, not the three best-documented ones you happened to have on hand. A model that passes on a narrow slice tells you it can pass a narrow slice, nothing more.

What I would leave alone: a caption-suggestion tweak nobody's betting a press cycle on doesn't need this ritual. If it gets the tone wrong 1 time in 10, someone edits it, costs a few seconds. Being wrong there is cheap. The two-bar gate earns its keep on anything a customer might act on with real money or real trust behind it, not on everything with "AI" in its name.

The lesson: a demo proves a feature can be right. It never proves how often. Those are two different questions, and for two years Corvenna only ever asked the first one, because asking the second one used to be free: nothing had ever been expensive enough to need it.

Now here is the same thing as a story

The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what $16,000 and one bad week costs, not just hear the number.

Palmira Vosberg has run Nowcast's roadmap at Corvenna for two years, and before that she shipped the caption-suggestion tool and the best-time-to-post nudge, the two features that made her the person leadership handed vague ideas to. She was good at it: turn a hallway comment into a build, get a demo in front of the room fast, ship it when the room nodded. For two years, nodding was enough. A good demo and a good feature had, in her experience, always been the same thing.

Then Godfrey Winterbach, Corvenna's CEO, asked for something with a bigger name attached: Nowcast, an AI that predicts trends before they peak, as the centerpiece of the Series B announcement, three weeks out. Palmira did what always worked. She scoped it to whatever would demo best, fast: three flagship accounts with the richest posting history Corvenna had, the ones every internal tool already ran cleanest on. In the room, Nowcast called a specific meme format two days before it peaked, live, on screen. People actually stood up. Nobody, including Palmira, asked whether it did that for the other 1,397 accounts, because in two years, a demo that good had never once turned out to be hollow.

Hand sketched horizontal timeline titled Palmira's habit, narrowing into demo-only scoping. Four milestones: The early features, caption caption suggestions, low stakes. The habit sets, caption a good demo means ship it. Nowcast, three accounts, caption 9 of 10, standing ovation, this milestone emphasized. Petalcombe's week, caption $16,000, one bad call.
Nobody decided to stop asking. It stopped being asked over two good years, the way most habits actually do.

Corvenna's launch checklist had exactly one gate for a press-tied AI feature: the demo passes, engineering flips it on for every account. That gate had never once been wrong before, so Nowcast rolled out to all 1,400 accounts the same week the room stood up.

For the flagship-sized accounts, the good months held. Then, five weeks in, a routine account-health call surfaced something odd: Petalcombe, a skincare brand with 2,800 followers, had spent its entire week's paid-boost budget, $16,000, chasing a Nowcast call that never landed. Orla Duchamps, who runs their marketing, hadn't filed a complaint. She'd just quietly asked her account manager, on a call about something else entirely, whether Nowcast's calls were "actually tested on brands our size, or just the big ones."

Nobody at Corvenna could answer that question with a number. That was the whole problem in one sentence.

Hand sketched full page metaphor titled What Corvenna assumed, and what was true. Left panel a gauge icon labeled DIAL, caption we assumed a great demo means a great feature, by degrees. Right panel a box icon in rust orange labeled SWITCH, caption either both bars clear before GA, or nothing ships, nothing between.
The whole answer, in one picture. Nobody designed a dial. Everybody got a switch, and for two years nobody had to notice.

Palmira built the real eval set that week: 180 accounts, stratified across follower bands and industries, the population Nowcast actually served instead of the three that had impressed a room. It came back at 38 percent overall. Broken out by follower band, accounts under 4,000 sat at 21 percent. Accounts over 50,000, the size of every demo account, sat at 61 percent. Nowcast hadn't gotten a little worse for smaller accounts. It had never really been tested on them at all.

We hadn't shipped a feature that was sometimes wrong. We'd shipped a feature that was only ever right for the three accounts we happened to already trust.

The easy fix would have been a disclaimer: add a line under the badge saying results may vary by account size. Palmira actually drafted one. She killed it within a day. Orla hadn't read fine print before betting a week's budget. She'd read a badge that looked exactly as confident as the one Corvenna's CEO had watched call a meme format two days early. A caveat doesn't change what the badge looks like at 9am on a Monday. It just moves the blame after the money's already spent.

Here's the decision I'd take back instead, and it isn't really Palmira's. Back when Corvenna's launch checklist was written, every AI feature it covered was small: caption tone, posting time, things where a wrong call cost a person a re-type. One gate, demo passes, ship it, was a sensible rule for that population of features. Nobody ever split "the demo is ready" from "the feature is ready for everyone," because nothing before Nowcast had been expensive enough to make the difference visible. The same shortcut that kept two years of small features shipping fast is what let a three-account demo stand in for fourteen hundred real ones.

Run the same kind of ask again, eight weeks later. Godfrey wants Nowcast Live, real-time competitor benchmarking, unveiled at a marketing conference keynote, six weeks out. This time Palmira splits the gate in two before a single line of the brief gets written. An announce track: the same three flagship accounts, honestly labeled a curated preview, ready in time for the keynote. An eval track, running in parallel: a fresh stratified sample of 220 accounts, with a real pass bar, 65 percent overall and 50 percent within the smallest follower band, before anything reaches every customer. Two weeks before the keynote, the eval track flags the same weak spot: accounts under 4,000 followers and under 90 days of history. This time it's caught before a customer ever sees it. The team ships a guardrail, those accounts get "still gathering signal" instead of a confident call, and the keynote goes ahead on schedule with the honest preview. General availability follows three weeks later, once the stratified hit rate clears 71 percent overall and 58 percent in the smallest band. Zero accounts get a wrong, confident call during the gap.

What I'd tell the version of myself who wrote that first launch checklist, back when every AI feature Corvenna shipped was small enough that being wrong cost an afternoon: the checklist that gets you through two good years doesn't announce when it's stopped fitting the feature you're about to ship through it. Somebody has to go looking for the moment it stopped fitting, on purpose, because a demo that works will never once raise its own hand and say it isn't enough.

FLIPS, or the two bars a press date almost skipped

Not a trick to sound structured. It's the difference between a feature that survives contact with 1,400 real accounts and one that only ever had to survive a room of people who wanted to believe it.

Hand sketched numbered list titled FLIPS, five questions before a press date gets a build. Five rows: F, find the person, whose scoping habit is this. L, locate the habit, what did she skip checking. I, identify the flip, what verb snaps. P, pinpoint the old decision, which gate got merged. S, show the replay, same ask, two gates now.
Four setup and payoff letters, and one question that only had to be asked once it actually mattered.
FFind the person. Whose habit is this?
Palmira Vosberg, the PM who owns Nowcast's roadmap at Corvenna, the person whose first response shapes every press-tied AI ask before it becomes a build.
The flip belongs to whoever decides what "ready" means before anyone ships anything, not whoever happens to be closest to the model.
LLocate the habit. What did she stop doing?
She scoped press-tied AI asks purely around what would demo best, because a good demo had, for two years, always turned out to be a good feature. She never separately defined what "works for a real customer" would mean.
That habit cost nothing while the features were small. It stopped being free the moment one carried a customer's ad budget.
IIdentify the flip. What verb snaps?
Old setting: "this needs to be announceable" is treated as the whole scope, and a strong demo on a handful of accounts is accepted as proof, no further check required. New setting: announceable becomes one constraint among several. A real, measured user-value bar, checked against a representative sample, has to clear too, every single time, before the feature reaches anyone outside the demo. Nothing in between: there's no version where a great demo alone is enough, because the second bar exists independently of how good the first one looked.
This is the answer to the question in one line. "Leadership wants it for the press release" only becomes a real scope once it has a second bar standing next to it.
PPinpoint the old decision. Which choice made sense before?
Corvenna's launch checklist merged two steps into one: "the demo is press-ready" and "the feature ships to every account" were a single gate, because for two years every AI feature it covered was small enough that the merge never showed.
"Add a disclaimer" would be a new dial bolted onto the same merged gate. Splitting the two steps apart is the reversal actually taken back.
SShow the replay. Same ask, better ending?
A similar ask returns: Nowcast Live, wanted for a keynote six weeks out. This time the announce track and the eval track run in parallel from day one, and the eval track catches the same low-follower weakness two weeks before anyone outside Corvenna sees it.
The replay ends in a count: the keynote ships on schedule with an honest preview, and general availability follows three weeks later once the stratified hit rate clears 71 percent overall and 58 percent in the smallest band, with zero wrong confident calls in the gap.
Hand sketched left to right flow diagram titled The two-gate design, after. Five boxes connected by arrows: Press-tied ask lands. Announce track, 3 curated accounts. Eval track, 220-account stratified check, this step emphasized. Guardrail, thin accounts flagged. GA once both bars clear.
The announcement never got smaller. It just stopped being the only thing standing between a demo and 1,400 real accounts.
Nowcast's stratified hit rate, week by week, as the guardrail rolled in
100% 50% 0 38% guardrail ships 71%, GA Wk 1 Wk 3 Wk 6
Stratified hit rate, that weekGuardrail shipsGA
A three-account demo would have called this finished in week one, at 38 percent. The eval track kept it open for six.

The AI-specific failure worth naming plainly is eval-set representativeness: Nowcast wasn't wrong so much as untested outside the population that happened to impress a room, so a model that looked calibrated on three rich-history accounts turned out to have almost no real signal for accounts without that history. The guardrail is the follower-count and history-length threshold: thin accounts get an honest "still gathering signal" state instead of a confident call, checked against a real stratified sample before every general release, not promised to be right, held to a target rate on purpose. There's a real trade-off, accepted on purpose: the two-track gate costs two to three weeks of delay and real eval-labeling effort before general availability, in exchange for never shipping a confidently wrong call to an account that never had enough data to earn one. For a caption-tone tweak nobody's betting a press cycle on, that overhead costs more than it's worth, which is exactly why it's a gate for expensive, announceable asks, not a rule for every feature with a model in it. One alternative got a real look and got rejected: a disclaimer under the badge, "results may vary by account size." It never shipped, because it doesn't change what the badge looks like at 9am on a Monday, it just moves the blame after the money's already spent.

And if you want to be sure it really works, try it somewhere else

Same five letters, a completely different flip family this time. Nobody scopes to a good-looking demo here. A check just quietly stops seeing the cases that would embarrass it.

Sedgemoor sells farm-advisory software, and Bloom Forecast is its AI feature: it predicts the best harvest window for each of Sedgemoor's roughly 600 partner farms. Eamon Cuthbertson owns Bloom Forecast's roadmap. When leadership wanted it as the headline feature for a regional ag-tech conference, Eamon didn't scope around a flashy demo. The trap here was quieter: his data team had, over months, learned which three or four farms had the cleanest sensor telemetry, no gaps, no missed readings, and had started pulling only those farms into every pre-release check, without ever deciding to.

Hand sketched labeled parts diagram titled How a pre-edited check hides its own gap. Central gauge icon labeled Bloom Forecast check, with four radiating labels: Only clean-sensor farms pulled in. Messy farms filtered out first. The number looks stable. The real gap stays invisible.
Same five questions, a completely different way the flip hides. This time the check itself quietly stopped seeing the mess.

It wasn't a decision anyone made in a meeting. A messy farm's data dragged the pre-release number down and slowed the demo prep, so whichever engineer was running the check that week would quietly swap it for a cleaner one. Within four months, the same three farms had been in every single check, and the messy 90 percent of Sedgemoor's real farms had been in none.

Hand sketched two panel comparison titled Same five letters, a different I both times. Left panel a person icon labeled Palmira, caption Corvenna, over-trust flip, a great looking demo is treated as proof until it isn't. Right panel a document icon labeled Eamon, caption Sedgemoor, pre-editing flip, the check quietly stops seeing the messy farms at all.
Both stories run F through S. Only one letter, the I, tells you which way the habit actually broke.

F · Eamon Cuthbertson, who owns Bloom Forecast's roadmap at Sedgemoor.
L · The data team stopped pulling whichever farms were due for a check and started pre-selecting the three or four with the cleanest telemetry, because a messy farm always dragged the number down.
I · The pre-editing flip, a different shape from Palmira's. Old setting: any farm due for a check gets pulled in, including the messy ones, and the check logs what actually breaks. New setting: the population gets filtered before the check ever runs, so a "check" always meant a check of the easy cases, and the number stopped measuring anything real. Nothing in between: either the check sees the real population, or it sees a hand-picked one.
P · Sedgemoor's validation pipeline never had a fixed, mandatory sample. Any engineer could pick which farms to include. That was fine while Bloom Forecast was a 40-farm internal beta and all 40 looked fairly similar. Nobody ever built a fixed roster, because the population itself used to be uniform.
S · Six weeks before a regional keynote, Eamon locks a fixed, stratified 85-farm sample, drawn once and reused for every release, spanning every sensor-quality band Sedgemoor actually has. Nineteen of those 85 farms have sensor readings that arrive more than six hours late at least once a week, and the fixed sample catches something the cherry-picked one never could: late readings were being silently treated as "no frost risk" instead of flagged unknown. The team ships a freshness guardrail nine days before the conference. Bloom Forecast's new feature, Frost Alert, launches on the fixed sample's real number, not a hand-picked one, and it holds in the field afterward.

What finally surfaced it A routine data audit, unrelated to the conference, asked why the same three farm IDs showed up in every pre-release report for four straight months. Nobody had an answer that wasn't "they're the reliable ones."

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: an AI ask scoped purely to look good in a demo will always find three accounts willing to make it look good, no matter how good the model actually is.
Cost: no budget to build a full stratified eval pipeline from scratch. Reuse the customer population itself; sampling 180 existing accounts costs almost nothing extra, it's refusing to skip the sampling that costs something.
The model got better, for real: say Nowcast's overall accuracy climbs on its own next release. Doesn't matter, maybe matters more. A model getting quietly better is exactly when nobody thinks to check whether the improvement reached the accounts that needed it most.

Where people run it wrong.
They treat "we have an eval set" as proof it's representative, instead of checking whether it actually mirrors who'll use the feature.
They let a genuinely impressive demo substitute for the second bar, instead of treating the demo and the eval as two separate questions.
They wait for a customer's bad week to force the second bar into existence, when the whole point of setting it early is that you don't need one to show up first.

How to use it live. If an interviewer asks how you'd handle a press-driven AI ask, ask yourself one thing before answering: if this shipped to every account tomorrow morning, could I already say what number would prove it was working? If the honest answer is no, that's the whole question, answered.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
An over-trust flip. Palmira's habit was to stop scrutinizing a press-tied AI ask once the demo looked genuinely strong, because a great demo had, for two years, always turned out to be a great feature. The trigger here is good news, not bad, which is what made it easy to miss.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Palmira Vosberg, the PM who owns Nowcast's roadmap at Corvenna, a social-media analytics platform for marketing teams.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
She scoped press-tied AI asks purely around what would demo best, and never separately defined what "works for a real customer" would mean, because that shortcut had always been enough before.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Old: announceable is treated as the whole scope; a strong demo on a few accounts counts as proof. New: announceable is one bar of two; a real, measured user-value bar on a representative sample has to clear too, every time, before wider rollout. No version ships on the demo bar alone.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Corvenna's launch checklist merged "the demo is press-ready" and "the feature ships to every account" into one gate. That was fine while every AI feature it covered was small enough that being wrong cost a re-type, not a customer's ad budget.
6 · THE NUMBER
Fill in the blank: Nowcast's demo hit rate on three flagship accounts was ___%. On a representative 180-account sample, it came back at ___%.
Tap to flip
ANSWER
90%, then 38%. A demo measuring three accounts said almost nothing about the other 1,397.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
A similar ask (Nowcast Live, for a conference keynote) returns. The announce track and eval track run in parallel from day one. The eval track catches the same low-follower weakness two weeks early, a guardrail ships before the keynote, and GA follows three weeks later at 71% overall and 58% in the thinnest band, with zero wrong confident calls in the gap.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs FLIPS again on a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Bloom Forecast, Sedgemoor's harvest-window AI. The pre-editing flip: the validation check quietly stopped seeing any farm except the three or four with the cleanest sensor data, so the number it produced stopped measuring the real population.

Check yourself Score: 0 / 0

Fill in the blank
1. Nowcast's hit rate on the three flagship demo accounts was ___%. On a representative, stratified sample of 180 real accounts, it came back at ___%.
Show hint
Check the bar chart in "Let's learn."
Show answer
90%; 38%. A demo measuring three accounts told the team almost nothing about the other 1,397, and by follower band the gap was even sharper: 21% under 4,000 followers versus 61% over 50,000.
True or false
2. True or false: Palmira's flip was that she started double-checking every AI feature more carefully before any demo, regardless of how it performed.
  • True
  • False
Show hint
Look at the I step, and which family this flip belongs to.
Show answer
False. This is an over-trust flip, not a verification flip. Before the incident, a strong demo was itself what made her stop scrutinizing further. The fix isn't "check more always," it's requiring a second, separate bar every time, regardless of how good the first one looks.
Multiple choice
3. Why couldn't Palmira have just added a disclaimer under Nowcast's badge instead of building a real eval-set gate?
  • A. Corvenna's legal team blocked disclaimers by policy.
  • B. A disclaimer doesn't change what the badge looks like at the moment someone acts on it; it just moves the blame after the money's already spent, without changing the model's actual accuracy for that account.
  • C. Nowcast's interface had no space to render a disclaimer.
  • D. The engineering team refused to ship any UI copy changes that quarter.
Show hint
Look at the paragraph in the story where Palmira drafts, then kills, the disclaimer.
Show answer
B. Orla never read fine print before spending Petalcombe's budget; she read a confident-looking badge. A caveat doesn't fix the underlying gap, it just relocates who's blamed for it.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Corvenna's launch checklist merged "the demo is press-ready" and "the feature ships to every account" into a single gate. It made sense while every AI feature it covered was small and low-stakes, caption tone, posting time, where being wrong cost a re-type, not a customer's ad budget.
Short answer, apply it yourself
5. Think of a request you or your team gets where something impressive in a demo or pitch gets treated as proof it's ready for everyone. What's the one number you'd ask for before believing it?
Show hint
Look for the population the demo was actually run on, not the population it'll be judged by.
Show answer
Model answer: For a "this pitch deck slide killed in the room" moment: how many of the actual customers we're targeting saw a version of this message, and did any of them act on it, versus how many people in that specific room happened to already agree with us going in.
Multiple choice
6. Which of these Nowcast-adjacent asks would you NOT put through the full two-bar gate?
  • A. "We want Nowcast Live, real-time competitor benchmarking, for the keynote."
  • B. "Let's have the AI suggest three caption tone variants a marketer can pick from before posting."
  • C. "We want Nowcast to auto-recommend which accounts should boost spend this week."
  • D. "Let's ship the trend-window predictor to every account by default, no opt-in."
Show hint
Look at "What I would leave alone" in Let's learn.
Show answer
B. Getting a caption tone wrong costs a marketer a few seconds to edit. The two-bar gate earns its keep where a wrong call costs real money or real trust, not on a low-stakes suggestion nobody's betting a press cycle on.
Before you close the answer
Why this works
Tests whether you'll treat a press-driven AI ask as an instruction to demo well, or as a scope with a missing half. Most candidates describe how they'd manage the stakeholder relationship. Fewer notice that a demo answers "can it," never "how often," and that those are genuinely different questions for anything with a model in it.
Follow-up traps
"Isn't requiring a second bar just going to blow the press date?" Response: only if the eval track starts after the announce track. Run them in parallel from day one and a labeled preview can still ship on schedule; what changes is that general availability waits for its own bar, not the demo's.

"What if leadership just overrules you and ships to everyone off the demo alone?" Response: then I'd want that decision made explicitly, in writing, by someone with the authority to accept the risk, not made by default because nobody asked the second question. The point isn't to win every argument, it's to make sure the gap gets named out loud before it becomes a customer's bad week.
If pressed
The 220-account eval sample for Nowcast Live wasn't drawn once and reused forever. It was redrawn stratified by follower band and industry each release, because Corvenna's own account mix shifted quarter to quarter, and a sample built for last quarter's population would have quietly gone stale the same way the three-account demo did.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more