InterviewAdvancedModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #16

Explain how to say no to an AI initiative without being labelled the person who blocks innovation.

PICK · Draymark's robots go free-roam in wide aisles with zero collisions in 40,000 merges. Cadwell's aisles are not wide, and nobody had checked the one number that mattered there

Draymark Robotics builds the pathfinding and collision-avoidance model that drives a warehouse's fleet of small delivery robots between racks. Zubin Kalantari owns that model. Renate Odukoya runs field deployment and wants to switch Cadwell Grocery Distribution's whole fleet to fully autonomous "free-roam" routing in twelve days, the same week a rival announced they'd done it first. Zubin has to say no to Cadwell without sounding like the one person in the building who's afraid of shipping something new.

The direct answer
Don't say no because free-roam feels risky. Say no by naming the exact gap in the evidence and the exact number that would close it. At Cadwell, that's: the collision model's confidence has never been checked against a blind-corner merge, only a wide-open one, and a shadow log sitting on the shelf already shows it's wrong more often than it thinks it is there. Name that gap, name the number that flips it to yes, and set a date to go check.
Do this, in order
  1. Name the exact evidence gap behind the no, and the number that would close it.Why: this is the whole decision. A no with no attached number is a feeling, and feelings are what get you labelled the blocker.
  2. Say yes to what already cleared the bar, loudly, before saying no to what hasn't.Why: it proves the no isn't reflexive. You're not against the initiative, you're against shipping the one piece nobody's tested yet.
  3. Pull real numbers from whatever log already exists on the exact situation in question.Why: a felt sense of danger convinces nobody in a room with a deadline. A number pulled from real merges does.
  4. Write the flip condition down before the deadline pressure hits the room.Why: a condition set under pressure gets renegotiated under the same pressure. Put it in writing while it's still calm.
  5. Offer a cheap, bounded way to get the missing evidence fast, instead of a flat no.Why: "not yet, here's how we get to yes in eight weeks" reads as a plan. A flat no reads as a wall.
  6. Anchor the objection to one specific, checkable risk, never to general caution.Why: general caution invites "you always say that." A specific number invites "let's go look."

How to answer this, stage by stage

Nobody in the room is grading whether you sound confident saying no. They're grading whether your no is checkable, or whether it's just a feeling wearing a decision's clothes.

1
Scope it to one real product and one real decision
Say it like this
"Let's ground this in something real. Draymark builds the routing model that drives a fleet of small delivery robots through a warehouse. Zubin owns that model. His field-deployment lead, Renate, wants to switch one client's whole fleet to fully autonomous routing in twelve days. He has to say no to part of that plan without being the guy who kills the launch."
Why this works
Keeps the interviewer grading a real decision instead of a general theory about office politics.
2
Say your structure out loud
Say it like this
"I'll use PICK. Position, what actually separates a calibrated no from a reflexive one. Impact, what's lost on each side if you get that wrong. Cost asymmetry, which mistake is cheap and which one is expensive. Kill criteria, the one number that would flip a no into a yes."
Why this works
Two seconds of structure tells the interviewer they're about to hear a method, not a mood.
3
State the real distinction, plainly, before any story
Say it like this
"My position: a no that names a specific, checkable reason, and a specific thing that would change your mind, reads as judgment. A no with no reason attached, or a reason nobody can check, reads as obstruction. And honestly, it usually is obstruction, even when it doesn't feel that way from the inside."
Why this works
This is the direct answer, said before the story gets a chance to blur it into "be careful with new tech."
4
Prove the distinction with a real number
Say it like this
"Free-roam ran clean at the pilot site: zero collisions across 40,000 merges, in aisles wide enough to see a robot coming six seconds out. Cadwell's aisles are 1.4 meters wide with blind corners. Draymark had been quietly logging what the model would have decided at Cadwell's blind merges for ten weeks, 3,200 of them, without anyone acting on it. When the model said it was over 90% sure, it was actually right 79% of the time there, against 96% at the pilot site."
Why this works
A real before-and-after number turns "I have concerns" into something the room can check for itself.
5
Name what's lost on each side
Say it like this
"Say no with no number attached, and you burn the credibility to raise a real concern next time, because it reads as nerves. Say yes just to avoid that label, and you ship a robot that thinks it's 91% sure at a blind corner where it's actually right about eight times out of ten. One costs you a meeting. The other costs someone their footing in a warehouse aisle."
Why this works
Naming both losses stops the answer from collapsing into "just be cautious," which isn't a real decision.
6
Name which mistake is cheap and which is expensive
Say it like this
"Holding Cadwell for eight weeks while the model gets recalibrated on real blind-corner data costs about 120 hours of work Draymark's already scheduled. That's cheap, and it's forgotten by the next release. Shipping free-roam blind cost a sister site, Sedgewick, about 640 hours: an incident report, a paused fleet, and a client meeting nobody wanted to be in. That's the asymmetry. It only points one way."
Why this works
This is the center of PICK: the two mistakes don't cost the same, and naming which one is expensive is what makes the no defensible.
7
Give the one test, and offer the path to yes
Say it like this
"Here's my one test: does the model's confidence at blind-corner merges match its real accuracy, within about 5 points, on Cadwell's own aisles, not a proxy from somewhere else. Right now the gap is 11 points and climbing at the tightest corners. Recalibrate on real blind-merge data, close that gap, and I'll sign off on a phased pilot in one aisle with the old marshal still watching for the first two weeks. That's not a no. That's a not-yet with a date on it."
Why this works
A kill test with a path back to yes is what separates a calibrated no from someone just digging in.

Let's learn

Hand sketched flow diagram titled What happens at a blind-corner merge. Five steps left to right: robot A rounds the corner, robot B rounds the same corner, cameras catch a partial view, model reports 91 percent sure (highlighted), they merge with no marshal watching.
Five steps, and the whole argument sits inside the fourth one: the number the model reports, and whether that number is actually true where it's being reported.

What happens the first time a robot is sure about something it shouldn't be sure about?

Draymark's robots move totes through a warehouse. Right now most clients run them in "lane mode": geofenced paths, and a central system that sequences who goes first at any merge, especially the blind ones. It's safe and it's a little slow, because every robot waits its turn even when nobody's actually coming.

Three months ago, Draymark piloted "free-roam" at a produce site with wide, open aisles: no lanes, no central sequencer, just the model's own read on who has the right of way, decided robot to robot in real time. It worked. 40,000 merges, zero collisions, and robots got where they were going 18% faster because nobody was waiting on a turn that didn't need waiting for.

Hand sketched comparison diagram titled Two aisle geometries, side by side. Left panel, a gauge icon labeled The pilot site, 2.4 meter aisles, about 6 seconds of clear sightline before any merge. Right panel, a scale icon labeled Cadwell's dense aisles, 1.4 meter aisles, blind corners, under 1 second to react.
Same model, same confidence number, two completely different amounts of warning before it has to be right.

Renate wants that same free-roam mode live at Cadwell Grocery Distribution within twelve days, timed to a trade show demo, and to beat a rival who just announced their own "no geofencing" launch. Cadwell's aisles are narrow, 1.4 meters, with blind corners where two aisles meet. A robot rounding one of those corners has under a second between first seeing another robot and needing to know who goes first.

Knowledge spark: what does "the model is 91% sure" actually mean? The collision-avoidance model doesn't just decide who goes first, it also reports how sure it is. That number is supposed to match reality: if it says 90% sure a hundred times, it should be right about ninety of those times. When that match holds up, the number's worth trusting. When it doesn't, the model can sound just as confident while being wrong far more often, and nothing about its tone gives that away.

Here's the turn. The extra speed free-roam offers Cadwell isn't really what's on the table. What's actually on the table is a number nobody had checked in the one place it mattered. Draymark had quietly logged what the model would have decided at Cadwell's blind-corner merges for ten weeks, 3,200 of them, while the real marshal system still made the actual calls. Nobody had opened that log. When Zubin finally did, the model's "over 90% sure" calls were right only 79% of the time there, against 96% at the wide-open pilot site.

The model wasn't wrong about the merges. It was wrong about how sure it was, and that's the number a fast rollout was about to trust with nobody watching.
The choice I would take back Draymark's release process treated the pilot site's calibration number as the fleet-wide bar for turning on free-roam anywhere. That made sense in March, when Cadwell's blind-corner data barely existed and building a separate check for every aisle shape felt like slowing a promising pilot down for a difference that might not even matter. It mattered. I would take that back and require a geometry-specific calibration check before free-roam goes live on any layout that wasn't part of the original proof.

What I would leave alone: Cadwell's loading dock and staging yard are wide open, same shape as the pilot site, long sightlines, no blind corners. Free-roam there is already proven. Holding that up to wait on blind-corner data that has nothing to do with it would slow down the one part of the rollout that's already earned its yes.

The lesson: a model's confidence number is only as trustworthy as the ground it was measured on. You don't get to borrow a calibration number from one aisle shape and assume it holds in a narrower one, especially when the difference between the two is exactly the amount of time a robot has to react.

Now here is the same thing as a story

The short version above is what you'd actually say out loud in the room. Read this one for what it would have cost Draymark to find out the slow way.

The picking floor at Cadwell's regional distribution center gets loud around 4am, when the first trucks start loading and every aisle has two or three robots in it at once, gliding between racks of dairy crates and produce totes.

Zubin Kalantari has owned Draymark's pathfinding model for two years. He can read a merge log the way some people read a weather map, at a glance, for the shape of the thing rather than any one number.

The pilot went well enough that it stopped being a pilot in most people's heads. Free-roam ran for three months at a produce site with wide, easy aisles, and it just worked. Renate Odukoya, who runs field deployment, had been carrying that result into every client meeting since: zero collisions, eighteen percent faster, no lanes needed. She wasn't wrong to be proud of it.

Then a rival startup put out a demo video, robots weaving through a mock warehouse with no lane markings at all, and the caption said "fully autonomous, no geofencing, no waiting." Leadership wanted an answer fast. Renate picked Cadwell, the biggest account on the calendar, and set a twelve-day deadline tied to a trade show floor demo. Free-roam, fleet-wide, no lanes.

Hand sketched timeline diagram titled Twelve days at Draymark. Six milestones: wide-aisle pilot ships, zero collisions across 40,000 merges. A rival announces free-roam, no geofencing, same week. Renate sets a 12-day deadline, tied to the trade show floor. Sedgewick's incident, a robot clips a picker at a blind merge, this milestone emphasized. Zubin pulls Cadwell's shadow log, 3,200 blind merges never acted on. The no, with a number attached, the gap has to close to 5 points.
The rival's video and Renate's deadline landed the same week. What actually changed Zubin's mind arrived four days later, from a sister site nobody was watching that closely.

Four days into the twelve, word came in from Sedgewick Freight, a different Draymark client that had gone free-roam early under the same competitive pressure. A robot rounded a blind corner in a narrow aisle and clipped a picker's foot, not seriously, a bruise and a scare, but enough to trigger an incident review, a week with the whole fleet back on marshaled lanes, and a very uncomfortable call with Sedgewick's ops director.

Zubin didn't think Cadwell would follow the same script automatically. But he remembered something: months earlier, someone on his team had set up shadow logging at Cadwell, recording what the model would have decided at every blind-corner merge, while the real marshal system kept making the actual calls. It was meant to build a dataset for later. Nobody had ever gone back and read it.

He read it that afternoon. 3,200 blind-corner merges, ten weeks of them. He filtered for the calls where the model had said it was over 90% sure. At the pilot site, that confidence band was right 96% of the time, cleaner than the claim itself. At Cadwell, it was right 79% of the time. In the tightest corners, the ones with under a meter of visibility before the merge point, it dropped to 71%.

We weren't asking whether the robots could find their way through Cadwell. We were about to trust a number that had never once been checked at Cadwell.

I want to say the problem was that free-roam didn't work. It worked fine, at the site it was tested at. That's not really the story. Nobody at Draymark had a rule that said "check the confidence number on the actual layout before turning it loose." They had a feeling that a strong pilot result travels, and it mostly does, right up until the geometry underneath it changes enough that the model's confidence stops meaning what everyone assumed it meant.

Hand sketched comparison diagram titled What each of them is holding. Left panel, a person icon labeled Renate, field deployment, holds the deadline, the trade show floor, twelve days out. Right panel, a person icon labeled Zubin, navigation model, holds the shadow log, 3,200 real merges, never once looked at.
Neither of them was wrong about what they were holding. The deadline was real. So was the log. The question was which one got to decide first.

Zubin walked into Renate's office with the shadow log open on his laptop, not a warning, a number. He told her plainly: free-roam ships everywhere at Cadwell that looks like the pilot site, the loading dock, the staging yard, today, no delay. The dense aisles, the ones with blind corners, hold for eight weeks while the model gets recalibrated on real Cadwell data, with the old marshal system running underneath as backup the whole time. He named the number that would flip it: if the confidence gap at the tightest corners closes to under 5 points, he'd sign off on a phased pilot in one aisle before the eight weeks were even up.

Renate pushed back once, asked whether eight weeks would blow the trade show demo. Zubin said the open areas alone would still be a real demo, robots moving without lane markings through most of the floor, and that showing a rehearsed video of the blind corners would be a worse story than showing the real thing two months later. She agreed to split the rollout that same afternoon.

The recalibration work started that week: more blind-corner examples in the training data, a wider shadow-logging window, a bigger held-out sample to test against. Week zero, the gap at the tightest corners was 11 points. Week four, 8 points. Week eight, 4 points, under the line Zubin had set. Free-roam went live in Cadwell's dense aisles in week nine, marshaled fallback still running for the first two weeks as a safety net nobody ended up needing.

One version of this story ships free-roam everywhere on day twelve and spends the next month explaining an incident report to Cadwell's ops team. The other spends eight weeks closing an eleven-point gap and ships a version that's actually earned its confidence number. Same model, same deadline pressure, same trade show. The only thing that changed was whether the confidence number got checked against the ground it was actually going to run on.

What I'd tell myself, watching that shadow log sit unread for ten weeks: a number that's already sitting on your own server is not evidence until somebody actually reads it. It was there the whole time. It just hadn't been anybody's job yet.

PICK, for the day someone calls your no obstruction

Not a script for sounding confident. PICK only earns its keep if it makes you name, out loud, the one number that would actually change your mind, before anyone asks for it.

Hand sketched comparison diagram titled The asymmetry, drawn. Left panel, a gauge icon labeled Recalibrate first, about 120 hours over 8 weeks, on the existing schedule. Right panel, a question mark icon labeled Ship free-roam blind, Sedgewick's real bill, about 640 hours, a paused fleet, an incident report.
One path costs a scheduled eight weeks. The other already sent Draymark its real bill, from a different site, four days into the same deadline.
PPosition. The real distinction.
A no that names a specific, checkable reason, and a specific condition that would flip it, reads as calibrated judgment. A no that's vague, or unexplained, or just a bad feeling about new technology, reads as obstruction, because there's nothing in it anyone can go check.
This isn't about sounding brave in the meeting. It's about whether your objection points at something real that exists whether or not you say it out loud, in this case a confidence number that's never been tested on the actual aisle shape it's about to run in.
State the position before the story. A position that only appears after the incident already happened looks reverse-engineered from the ending.
IImpact. What's lost each way.
Say no with no evidence attached, and it burns your credibility to raise the next real concern, because it reads as reflexive caution. You get quietly routed around on the decisions that actually matter, which is worse for the product than being wrong once out loud.
Cave on every AI initiative to dodge the blocker label, and you ship things that fail for real, in public, which costs far more credibility than one uncomfortable "not yet" ever would.
Naming both losses is what stops this from collapsing into "just be brave," which isn't a decision either.
CCost asymmetry. The heart of it.
Saying no with a clearly named, checkable reason costs a moment of pushback in the room, cheap, forgotten by the next standup. A vague no costs your standing the next time you need to be heard. A reflexive yes costs a real failure, in Draymark's case an incident review, a paused fleet, and a client relationship that takes months to fully repair, if it repairs at all. Start from the cheap mistake. Only accept the expensive one once real evidence says it's actually safe to.
KKill criteria. The one test.
Does the confidence number, checked on the exact geometry it's about to run in, match the model's real accuracy within about 5 points. Zubin's team considered one shortcut instead of the eight-week recalibration: apply the wide-aisle confidence curve to Cadwell too, but scale every score down by a flat safety margin to be conservative. It got rejected. A safety margin invented without real blind-corner data is a guess wearing a number's clothes, and there was no way to know if the margin was even the right size without collecting the exact data the shortcut was trying to avoid collecting.
Hand sketched decision tree diagram titled Does the no hold, or is it just nerves. Root box, is free-roam ready for Cadwell's dense aisles, branching to four labeled conditions: no dense-aisle data, just a feeling, leading to not a real no yet. Tested on wide aisles only, leading to wrong geometry, unproven. Dense-aisle gap tested, still above 5 points, leading to hold, keep the fallback. Dense-aisle gap tested, under 5 points, leading to ship the phased pilot.
Four branches, one question: has the confidence number actually been checked on the ground it's about to run on, and does the gap clear the line.
Cost, by the numbers: recalibrating first versus what Sedgewick actually paid
640 hrs 320 hrs 0 120 hrs Recalibrate, before deciding 640 hrs Sedgewick's real bill, after shipping
Cheap, paid on a scheduleExpensive, paid after the fact
Recalibrating Cadwell's model on real blind-corner data cost about 120 hours over eight weeks, work Draymark could plan around. Sedgewick's actual incident, four days into the same kind of rollout, cost about 640 hours in review, a paused fleet, and client meetings, and that number already happened once.
The kill line, charted: confidence-versus-accuracy gap during recalibration
12 pts 6 pts 0 kill line: 5 pts 11 pts, week 0 8 pts, week 4 4 pts, cleared Week 0 Week 4 Week 8
Above the kill lineCleared the kill line
The team tracked the tightest blind corners specifically, not the average across all of Cadwell's aisles, because the average was already sitting close to the kill line and would have looked fine on a dashboard well before it was true.

The trade worth saying out loud: recalibrating on real blind-corner data before shipping cost eight real weeks nobody wanted to spend, against a twelve-day deadline that felt urgent for reasons that had nothing to do with Cadwell's actual aisles. That's slower, and it's worth paying, because the alternative, a robot that's sure of itself in exactly the spot it shouldn't be, only shows its true cost after someone's already standing too close to it.

And if you want to be sure it really works, try it somewhere else

Same four letters, a veterinary ER instead of a warehouse floor, and the missing check is a species, not a stretch of aisle.

Nyala Veterinary Emergency Network runs an AI triage model at intake: front-desk staff enter symptoms and vitals, and the model scores each animal's urgency to help route the sickest ones to a vet fastest. Mid flu season, with wait times climbing, someone proposes letting the model auto-discharge any animal it scores "low urgency" at 90% confidence or higher, no vet review, to free up staff for the animals that actually need one.

Hand sketched labeled parts diagram titled Where Nyala's auto-discharge gap sits. Center icon, a document labeled Auto-discharge decision. Four labels radiating out: dogs, 95 percent agreement at 90 percent plus confidence, cats, only 74 percent agreement in the same band, cats mask pain, undertrained in the data, no vet review sits behind that gap.
Same shape as Cadwell's aisles: a confidence number that was only ever proven on the majority case, about to be trusted on the minority one, unread.

Position: a no that names the exact subgroup the model was never checked on reads as judgment. A no about "AI in healthcare being risky" in general reads as obstruction, and it would be, since it doesn't point at anything specific. Impact: auto-discharge everything, and a genuinely sick cat gets sent home on a model that was mostly trained on dogs, since cats hide pain and distress far more than dogs do and made up only 18% of the training cases. Hold the whole initiative over general worry, and staff stay buried during flu season for no reason tied to actual risk. Cost asymmetry: pulling a held-out sample of cat cases and checking the model's calibration on them specifically costs a few days of a data scientist's time. A missed cat discharged as low urgency costs an animal's life and a lawsuit, not a symmetrical trade. Kill criteria: does the model's confidence match its real accuracy on cats specifically, within the same margin it already hits on dogs, about 95% agreement at the 90%-plus band. Today cats sit at 74%. Auto-discharge ships for dogs immediately, and holds for cats until that gap closes.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the number that would flip your no, before anything else gets said.
Cost: no time to run the full recalibration before a real deadline. Fine, but test the highest-risk subgroup first, cats or blind corners, not a bigger pile of the same easy cases that already passed.
The model got better, for real: say a new version scores near-perfect on the overall average. Still don't skip the subgroup check. A clean average can hide one bad slice nobody separated out.

Where people run it wrong.
They say no with a feeling instead of a number, and then can't explain what would ever change their mind.
They say yes just to avoid the blocker label, and quietly hope the gap doesn't matter this once.
They cite general risk instead of the specific untested condition, which makes every future no sound the same, whether it's real or not.

How to use it live. If you're ever pushed to approve something on the spot, buy yourself a second by asking one plain question out loud: "has this been checked on the exact case we're about to ship it into, or just on the case that made the pilot look good." That question is the whole method, asked instead of stated.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question about saying no without being called the blocker?
Tap to flip
ANSWER
PICK: state the real distinction between a calibrated no and a reflexive one, name what's lost on each side, find which mistake is cheap versus expensive, then give the one number that would flip the no.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Zubin Kalantari, who owns Draymark Robotics' pathfinding and collision-avoidance model for its warehouse robot fleet.
3 · THE POSITION
What's the real line Zubin draws between a good no and a bad one?
Tap to flip
ANSWER
A no that names a specific, checkable reason and a specific condition that would flip it reads as judgment. A no with no reason attached, or one nobody can check, reads as obstruction, and usually is.
4 · THE GAP
What's the two-number gap this whole answer turns on?
Tap to flip
ANSWER
At the pilot site, the model's over-90%-confident calls were right 96% of the time. At Cadwell's blind-corner merges, the same confidence band was right only 79% of the time, an 11-point gap nobody had checked before free-roam was proposed there.
5 · THE REVERSAL
What old decision would Zubin take back?
Tap to flip
ANSWER
Draymark treated the pilot site's calibration number as the fleet-wide bar for turning on free-roam anywhere, instead of requiring a geometry-specific check before every new aisle shape. That made sense before Cadwell's blind-corner data existed. It stopped making sense once the data showed a real gap.
6 · THE NUMBER
Fill in the blank: the shadow log had ___ blind-corner merges from ___ weeks. At over 90% confidence, the model was right ___ percent of the time at Cadwell, against ___ percent at the pilot site.
Tap to flip
ANSWER
3,200 merges, 10 weeks. 79 percent right at Cadwell, 96 percent right at the pilot site, at the same confidence band.
7 · THE KILL TEST
What's the one number that would flip Zubin's no into a yes?
Tap to flip
ANSWER
The confidence-versus-accuracy gap at Cadwell's tightest blind corners has to close to under 5 points, checked on real Cadwell data, not borrowed from the wide-aisle pilot.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of Cadwell's blind corners there?
Tap to flip
ANSWER
Nyala Veterinary Emergency Network's auto-discharge triage model. The role goes to cats, a subgroup the model was barely trained on and was never checked on directly, standing in for the aisle geometry the pilot data never covered.

Check yourself Score: 0 / 0

Multiple choice
1. What made Zubin's no at Cadwell a calibrated judgment rather than reflexive caution?
  • A. He felt uneasy about robots moving without lane markings.
  • B. He wanted more time before any change shipped, on principle.
  • C. He named the exact evidence gap, a confidence number never checked on Cadwell's blind corners, plus the specific number that would close it.
  • D. He didn't trust Renate's deadline.
Show hint
Check what the direct answer says separates a calibrated no from a felt one.
Show answer
C. A no built on a specific, checkable gap and a specific number to close it is what makes it judgment instead of nerves.
True or false
2. True or false: since the pilot site proved free-roam worked with zero collisions across 40,000 merges, that result was strong enough evidence to trust the same confidence numbers at Cadwell too.
  • True
  • False
Show hint
Think about what the pilot site's aisles had that Cadwell's didn't.
Show answer
False. The pilot site had wide, open aisles with long sightlines. Cadwell had narrow, blind-corner aisles with under a second of reaction time. A confidence number proven on one shape isn't automatically true on a very different one.
Fill in the blank
3. Sedgewick Freight's real incident, after shipping free-roam blind into a dense aisle, cost about ___ hours in review, a paused fleet, and client meetings. Recalibrating Cadwell's model first cost about ___ hours over ___ weeks.
Show hint
Look at the "cost by the numbers" chart in the PICK section.
Show answer
About 640 hours for Sedgewick's real bill. About 120 hours over 8 weeks to recalibrate Cadwell first. The gap between those two numbers is the cost asymmetry the whole no rests on.
Short answer, name the rejected alternative
4. What shortcut did Zubin's team consider instead of the full eight-week recalibration, and why did it lose?
Show hint
Look at the K, kill criteria, step of the framework recap.
Show answer
Model answer: Apply the wide-aisle confidence curve to Cadwell too, scaled down by a flat safety margin. It got rejected because a margin invented without real blind-corner data is a guess dressed up as a number, and there was no way to know if the margin was even the right size without collecting the exact data the shortcut was trying to skip.
Short answer, where it wouldn't matter
5. Name a place in Cadwell's own rollout where this exact scrutiny would NOT be needed, and say why.
Show hint
Look at "what I would leave alone" in the Let's learn section.
Show answer
Model answer: Cadwell's loading dock and staging yard are wide open, the same shape as the pilot site, with long sightlines and no blind corners. Free-roam is already proven there, so holding it up to wait on blind-corner data that has nothing to do with it would be waste, not caution.
Short answer, apply it yourself
6. Think of a time someone (or some AI feature) told you they were "pretty sure" about something. What would you have needed to see before trusting that confidence, rather than just the tone it was said in?
Show hint
Think about whether the confidence had ever actually been checked against the specific situation you were in.
Show answer
Model answer: A GPS app was "confident" about a shortcut through a neighborhood it had barely routed before. I'd have wanted to know how many times that exact route had actually been driven and checked, not just how sure the app sounded about it.
Before you close the answer
Why this works
Tests whether you understand that "saying no without being the blocker" is really a question about what makes a model's confidence number trustworthy, and where that trust runs out. Most candidates answer with generic diplomacy advice. The real judgment is knowing a model can sound exactly as sure in a place it's never been checked as in a place it has.
Follow-up traps
"Isn't this just you being afraid to be wrong in public?" Response: no, because the no comes with a specific number and a plan to re-test it in eight weeks, not an open-ended objection with no way to resolve it.

"What if 3,200 merges is too small a sample to trust?" Response: it's the honest limitation, which is exactly why the plan calls for a bigger, fresh 8-week recollection before making the final call, not for treating 3,200 as the last word either way.
If pressed
The real mechanism: the model's confidence head had been trained on a proxy label, a merge going through without a marshal override, and that label came almost entirely from wide-aisle merges. It learned to associate "clear camera frame" with high confidence regardless of how much sightline actually existed, which is a training-data imbalance, not a broken model.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more