CaseIntermediateModel Fluency & the AI PM Role / What changes when the product is probabilistic / #15

How do you communicate a confidence level to a user without teaching them statistics?

SPARK · a severe weather alert app for a 140 camper summer camp

StormLatch, built by Windlass Atmospherics, watches radar and sends severe weather alerts to Camp Tallowbrook, an overnight camp for 140 kids in the Missouri Ozarks. Aslaug Doverspike leads product for StormLatch's alert screen. Camp director Halldis Fairbourne is the one who actually has to move 140 kids to the storm cellar, or not, every time a number shows up on her phone.

The direct answer
Never show the raw number. Show three plain tiers, "stay aware," "get ready," "take shelter now," each one tied to a short, real reason the model actually saw, a rotation confirmed on radar, a cell tracking toward camp, not a percent sign. Calibrate the tiers behind the scenes against how often each one really was followed by severe weather, and let a radar-confirmed threat jump straight to the top tier even when the model's own percentage is only middling, because a real storm's danger and the model's certainty about it are two different things. Keep the number itself, and the full curve behind it, for the meteorologist who has to trust it with more than a label.
Do this, in order
  1. Replace the raw percentage with three plain tiers, each carrying a real reason instead of a number.Why: a bare percent forces every reader to invent their own private cutoff, and nobody's cutoff survives a genuinely hard call.
  2. Let confirmed radar signal override the model's own percentage when they disagree.Why: this is what caught the near miss. A fast-forming storm can be real and dangerous while the model's confidence stays moderate, because the weather itself is what's hard to call, not the model.
  3. Calibrate each tier against its own real track record, not a number picked at launch.Why: a tier that verifies wrong even half the time teaches people to ignore it, the same failure the raw number caused, just wearing a word instead of a digit.
  4. Reject showing the percentage next to the tier "for extra context."Why: real alternative considered, it brought the private-cutoff problem straight back the moment testers saw a number to compare story to story.
  5. Keep the raw score and the full probability curve alive on a separate dashboard, for the meteorologist who validates the model.Why: Bodil Coldwell's job needs real precision; campers and counselors never did.
  6. Leave the tier count at three.Why: a fourth or fifth tier buys almost nothing and asks a scared 19 year old counselor to sort a harder decision at 2am.

How to answer this, stage by stage

Nobody is grading whether you can say "use simple language." They're grading whether you understand that a model's confidence and the weather's own unpredictability are two separate things, and that a good design has to carry both.

1
Ground it in one camp, one phone, one director
Say it like this
"Let's make this real. StormLatch is a severe weather alert app, built by Windlass Atmospherics. Aslaug Doverspike owns its alert screen. Halldis Fairbourne runs Camp Tallowbrook, 140 kids, and she's the one who has to decide, off a number on her phone, whether to move all of them to the cellar."
Why this works
A confidence number with nobody deciding on it is just a statistics question, not a product question.
2
Name the method before naming the fix
Say it like this
"I'll run this as SPARK. What Halldis does with the number today, the habit I actually want her to build, the concrete screen I'd design, what happens the day it's wrong, and what I'm deliberately choosing not to show her."
Why this works
Two seconds of structure, and the interviewer knows a plan is already moving.
3
Reframe what the question is actually testing
Say it like this
"This isn't really 'how do I explain a percentage.' It's 'how do I get someone to trust the alert exactly as much as it deserves, no more, no less,' without them ever needing to know what a percentage is. Undersell it and she waits too long. Oversell it and she stops moving at all, the fifth time it's nothing."
Why this works
This is the line the whole answer hangs on. Skip it and the fix sounds like a font choice instead of a real design decision.
4
Hand over the actual design, not a metaphor for it
Say it like this
"Here's what I'd build. Three tiers on the alert screen. Stay aware, get ready, take shelter now. Each one carries one plain reason, 'storm cell tracking toward camp, rotation confirmed,' not a number. Behind the screen, those tiers are calibrated against how often each one was actually followed by real severe weather, and confirmed radar rotation can push straight to take shelter now even if the model's own percentage is sitting in the middle, because that's exactly where a fast, hard to call storm lives."
Why this works
This is the direct answer, concrete enough that Aslaug's team could ship it, and it names the AI-specific move: letting confirmed signal override the model's raw confidence.
5
Run it against both ways a tier system breaks
Say it like this
"This has to survive being wrong in both directions. Set take shelter now too loose, and it fires for every ordinary thunderstorm, then nobody moves fast the one time it's real. Set it too tight, and a genuinely fast, close storm sits in get ready a beat too long. We tested both. A loose bar false alarmed on 6 nights out of Camp Tallowbrook's 11 week season. Splitting the bar by real track record, and letting confirmed rotation jump the queue, brought that down to one, and it never missed a real one."
Why this works
The hardest step, and the one a rushed answer skips. Testing the fix against its own failure is what makes it a design, not a hope.
6
Say what stays hidden on purpose
Say it like this
"One thing I'm not showing campers or counselors, on purpose: the raw percentage, and definitely not a per storm cell confidence breakdown. Bodil, the meteorologist who validates this model, still sees all of it on her own dashboard, because her job needs real precision. Halldis's job needs a decision she can make in four seconds."
Why this works
Naming the limit on purpose, and who still gets the number, is what makes this judgment instead of a guess dressed up as simplicity.
7
Close on the one sentence that answers it
Say it like this
"So: don't teach the stats, remove the need for them. Three tiers, a real reason on each, calibrated to actual outcomes, with confirmed radar able to override a middling number, because the model being unsure and the sky actually being dangerous are two different facts, and only one of them is Halldis's problem."
Why this works
Restates the direct answer plainly, so the interviewer leaves with the decision, not just the story.

Let's learn

StormLatch watches radar for a stretch of camps and towns, and sends a push alert the moment its model thinks severe weather is heading somewhere specific. That's the whole product: a phone buzzes, and someone has to decide what to do next.

Before the redesign, that buzz just carried a number. "Severe thunderstorm, 62 percent," or "tornado, 41 percent." Nothing else. At Camp Tallowbrook, that number landed on Halldis Fairbourne, and on nobody trained to read it the way Windlass Atmospherics' own meteorologists did.

Hand sketched icon list titled Camp Tallowbrook's alert screen, before the redesign. Four rows: one, push alert reads a bare number, 62 percent. Two, no staff member knows what the number means. Three, Halldis invents her own cutoff, evacuate past 45. Four, the real forecast sits behind the number, unseen.
This is what every alert looked like, for eleven weeks, before anyone changed the screen.

Halldis did what anyone would. She picked a number in her head, 45 percent, and above it she moved all 140 kids to the storm cellar under the dining hall. Below it, she waited and watched the sky herself. Nobody taught her that number. She built it out of a season of guessing.

Knowledge spark: what does that percentage actually measure? It is the model's own estimate of how likely severe weather is, given the atmosphere it's watching. It is not a measure of how dangerous the storm will be if it happens, and it is not the same thing as how much the weather itself can be predicted that day. A model can be genuinely unsure about a storm that is, once it forms, extremely dangerous, and confident about one that never amounts to much.

Here's the turn. The extra false alarms were never really the problem. The problem was what a season of them did to how fast anyone moved.

Hand sketched timeline titled One season at Camp Tallowbrook, before the redesign. Three milestones: week 1, shelter reached in 90 seconds, every alert. Week 5, 11 alerts sent so far, only 2 real storms, this point emphasized. Week 9, shelter time crept past 6 minutes, near miss.
Nobody decided to slow down. It happened the same way any habit fades, a little at a time.

Week 1, an alert crossed 45 percent and the whole camp was in the cellar in about 90 seconds, kids and counselors moving fast because the number was new and nobody wanted to be the one who guessed wrong. By week 5, StormLatch had sent 11 alerts past that line. Only 2 of them were followed by anything a meteorologist would call severe, a downed limb near the archery range, some hail that dented a golf cart. The other 9 fizzled, missed the camp, or got downgraded an hour later.

By week 9, the shelter drill that used to take 90 seconds was taking past 6 minutes. Counselors waited to see if it "actually looked bad outside" before they moved anyone. The number had stopped meaning anything, because nine ordinary Tuesdays had spent it down to nothing.

The false alarms did not cost the camp nine wasted trips to the cellar. They cost the camp its own alarm.

What it cost at its worst: on a Thursday night in week 9, StormLatch's alert read 52 percent, tornado, for a fast moving line of storms. Fifty two isn't the number Halldis's team had learned to fear. Counselors waited, the way nine quiet weeks had trained them to. But this storm was a different shape than the ones before it: a spin up tornado, embedded in a fast squall line, the kind of thing radar confirms in real time far more reliably than any model can predict twenty minutes ahead. By the time counselor Fikile Ndlovu saw the funnel above the tree line and started the drill without waiting for anyone's go ahead, the camp had about four minutes to move 140 kids into a cellar built for a calmer walk. Everyone made it. Nobody was hurt. Halldis has not stopped thinking about the four minutes since.

StormLatch's own reliability curve, forecast probability vs. how often severe weather actually followed
100% 50% 0% Perfect calibration 10% 20% 30% 50% 60% 70% 80% 90%
Observed rate, by forecast bucketThe 50 to 60 percent bucket, Thursday's storm
The model was never broken. Storms in the 50 to 60 percent bucket really did turn severe about half the time, exactly what "50 to 60 percent" should mean. The number was honest. It just isn't a sentence a tired counselor can read at 2am.
The choice I would take back Windlass Atmospherics put the model's raw confidence number front and center on the alert screen, because it was the most honest thing to show and the engineering team already had it sitting right there. That was fine for the meteorologists who tested it. It stopped being fine the day it reached a camp full of counselors who had never seen a reliability curve and were never going to.

What I would leave alone: StormLatch's daily outlook email, the one that says "storms possible this afternoon" before anyone's day even starts. Nobody is deciding anything urgent off that email. It could carry the raw number, ten decimal places and all, and not one camper would be worse off.

The lesson: a number is not a message. A message is something a person can act on without first learning what the number means, and that's a design job, not a math job.

Now here is the same thing as a story

Read the short version above when you're actually in the interview chair. Read this one when you want to feel what four minutes with a funnel cloud overhead actually costs.

Halldis Fairbourne has run Camp Tallowbrook for eleven summers, and she can read a sky the old way faster than most people can read a forecast. Green tint, hail. That particular stillness before a line of storms. She trusts her own eyes first, always has.

StormLatch arrived four summers ago, and for a long while it was just backup, a second opinion buzzing on the clipboard she carries. In week one this year, an alert crossed her private line, 45 percent, tornado watch, and she moved 140 kids to the cellar under the dining hall in about ninety seconds. Fast, clean, everyone a little rattled and a little proud of how fast they'd been.

Hand sketched metaphor scene titled The same 62 percent, read two different ways. Left panel, a gauge icon labeled a meteorologist, caption a precise, well defined probability. Right panel, a question mark box icon labeled a counselor, caption an arbitrary cue to react to, or not, with a VS between the two panels.
The exact same number, meaning two completely different things depending on who's reading it.

By week five, that number had crossed her line ten more times. Two of those storms did something, a broken limb, some dented golf carts. Nine did nothing at all. Nobody blamed Halldis for how her team started reading those alerts. Ninety seconds is a real disruption when it happens for nothing, nine times, in a five week season built around campfires and swim tests.

So a habit formed the way habits always do, quietly, one ordinary Tuesday at a time. Counselors stopped moving the second a phone buzzed. They stepped outside first. They looked up. If it "didn't look that bad," they waited for the number to climb higher, the way it usually had to before it meant anything.

Nothing about that was careless. It was nine quiet weeks doing exactly what nine quiet weeks do to anyone's nerve.

Thursday of week nine, the alert read 52. Not the number anyone had learned to fear. A fast line of storms, the kind that build and drop a spin up tornado with almost no warning, the kind where the atmosphere itself, not the model, is what's genuinely hard to call twenty minutes out. Counselor Fikile Ndlovu stepped outside to check, the way the season had trained everyone to.

He saw the funnel over the tree line before he'd finished checking.

We did not lose four minutes to a slow model. We lost them to a number that had already spent its own credibility, nine ordinary Tuesdays at a time.

What happened next was not panic. Fikile started the drill without waiting for anyone above him to confirm it, and it moved fast once it moved, counselors half carrying the youngest kids, Halldis counting heads at the cellar door with a flashlight in her teeth. Four minutes, by the timer on her phone, that she has replayed more times than she'll admit.

Hand sketched comparison diagram titled The anchor, three tiers, never a number. Three panels. Take shelter now, box icon, caption rotation confirmed, moving toward camp. Get ready, box icon, caption storm building nearby, watching for rotation. Stay aware, box icon, caption storms possible later, nothing organizing yet.
The whole fix, in one screen. No percent sign anywhere on it.

Everyone made it. That's the part Halldis says first, every time she tells it. The part she says second is that the number had been right, all season, exactly as right as a well built model is supposed to be. Fifty two percent storms really do turn severe about half the time. Nothing about that number lied to her. It just never told her, in a language she could use in the four minutes she actually had, that this particular storm's danger and the model's honest uncertainty about it were two different facts.

The decision Aslaug Doverspike traced it back to was older than that Thursday. When StormLatch's alert screen was built, someone put the model's own confidence number front and center, because it was the most honest number the team had, and the engineers who tested it could read a reliability curve without thinking about it. Nobody on that team pictured a 19 year old counselor deciding whether to move 140 kids off it, or a season slowly teaching that counselor to distrust it.

Hand sketched comparison diagram titled The night it's genuinely close, old design vs new. Left panel, question mark box icon labeled old design, the near miss, caption reads 52 percent, staff wait for a scarier number. Right panel, gauge icon labeled new design, same storm, caption confirmed rotation trips take shelter now at once.
Same Thursday, same storm. The only thing that changes is what the screen says.

Run that Thursday again, with the new screen already live. StormLatch's model still reads a middling number for that storm, because the atmosphere really was hard to call that far out. But this time, the moment ground radar confirms rotation building inside the cell, the screen doesn't wait for the model's percentage to catch up. It jumps straight to Take shelter now, one line under it: "Rotation confirmed, moving toward camp." Fikile doesn't have to step outside and eyeball the sky against a number that burned him nine times already. The screen already used the strongest signal available, and said so in four words.

One design asked a tired counselor to out-think a number by comparing it to nine other Tuesdays. The other design does that comparing for him, using a stronger signal than the model's own guess, and hands him a decision instead of a puzzle.

What Aslaug would tell herself, the day that screen first shipped: showing the truest number you have is not the same as showing the truth. The truest number, alone, taught Halldis's team the wrong lesson nine times before it ever got the chance to teach them the right one.

Hand sketched icon list titled What stays behind the scenes, for Bodil's team only. Three rows: one, document icon, the raw percentage and the full probability curve. Two, gauge icon, per storm cell confidence, updated every radar sweep. Three, scale icon, the reliability curve used to reset tier cutoffs.
None of this disappeared. It moved to the one dashboard where someone actually needs it.

SPARK, the same alert built for someone who never learned the stats

Not a way to dress up "use plain words" in five letters. SPARK forces you to name what breaks the day the model's honest number and the real danger point in different directions, and design for that day specifically.

SSituation. How the decision gets made today.
Halldis, or whichever counselor is on duty, reads a raw model percentage with no context and decides, alone, whether to move 140 kids. Nobody at Camp Tallowbrook was ever taught what the number means, so each of them quietly invents a private threshold, and that threshold drifts every time the number turns out to be nothing.
Name what the number is actually asking someone to do before naming the fix, or the anchor sounds decorative instead of load bearing.
PPayoff. The habit worth building.
Not "fewer false alarms." A camp staff that trusts the alert exactly as much as it deserves, every time, so a middling reading on a genuinely dangerous storm still moves people fast, and a routine thunderstorm doesn't burn down everyone's nerve for the night it matters.
A habit is something you can check for on the next storm. A feeling, "be more careful," isn't.
AAnchor. The actual design you'd build.
Three tiers on the alert screen: Stay aware, Get ready, Take shelter now. Each carries one short, plain reason drawn from what the model actually saw, never a percentage. The tiers are calibrated behind the scenes against their own real track record, and confirmed radar rotation can push a case straight to Take shelter now even when the model's raw percentage sits in the middle, because a fast, hard to call storm is exactly where that happens.
This is the concrete answer to the question. Everything else exists to protect it.
RRisk. What breaks the first time it's tested against itself.
Two failure directions, both real. Set Take shelter now too loose and it fires on every ordinary storm, and the exact fatigue that caused the near miss comes right back, just wearing a word instead of a number. Set it too tight, and a fast, genuinely dangerous storm sits one tier too low a beat too long. A shared bar across all three tiers false alarmed 6 nights out of an 11 week season in Camp Tallowbrook's own trial; splitting the bar by tier, and letting confirmed radar override the model's percentage, cut that to 1, without missing a real one.
Design the anchor against this specific risk, in both directions, or the fix just relocates the failure.
Nights per 11 week season the top tier fired, by design
12 6 0 11 Raw percent, one cutoff 6 Shared tier bar 1 Tiers, radar override
Raw percentage designOne shared calibration barCalibrated per tier, radar can override
Fewer false alarms is only half the win. The third design also caught the week 9 storm on the first radar confirmation, before the model's own percentage ever climbed.
KKeep out. What doesn't get shown, on purpose.
The raw percentage never reaches the camper facing screen, not even as a small number next to the tier. A rejected idea: showing "Take shelter now, 78%" as extra context. It tested badly fast, staff started comparing the small numbers story to story within days, and the private cutoff problem came straight back. Per storm cell confidence internals stay off entirely. Bodil Coldwell's meteorology team keeps the full number, the whole reliability curve, and every model update on their own dashboard, because validating the model is their job, and it needs real precision Halldis's job never did.
Naming the version you tried and killed is what makes "keep it simple" sound like a decision instead of an excuse to build less.

And if you want to be sure it really works, try it somewhere else

Same five letters, a soybean field in Iowa instead of a storm cellar, and this time the plain reason is about a leaf instead of a radar sweep.

LeafGuard, built by Amberfield AgTech, scans a photo of a crop leaf and estimates the odds of an early fungal blight before it's visible to the eye. Ekundayo Sowah owns its scan results screen. Nikolai Duru farms 900 acres of soybeans and checks LeafGuard most mornings before deciding whether to spray, a decision that costs real money and only works if it's made early.

Hand sketched labeled parts diagram titled LeafGuard's anchor, the same idea for a soybean field. Central document icon labeled Leaf scan, with four callouts around it: treat now, watch closely, looks fine, no raw score shown.
Same three tier idea, a completely different field, a completely different reason on each one.
The decision Amberfield would take back LeafGuard originally merged its risk read and its recommendation into one blended composite score, "72, moderate," collapsing the model's actual judgment and the action a farmer should take into a single figure a grower had to interpret twice. It made sense the day the model was new and every score got double checked by hand anyway. It stopped making sense once growers started trusting the number enough to stop double checking it.

Same rank, different reversal: Ekundayo's team split that one blended score into "Treat now," "Watch closely," and "Looks fine," each with a one line reason, "early lesion pattern matches last year's outbreak," rather than a composite figure. A field already showing visible lesion spread jumps straight to Treat now even on an ambiguous scan, the same override logic as StormLatch's confirmed radar, because visible evidence and the model's own uncertainty are two different facts here too.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: plain tiers with a real reason on each, calibrated to their own track record, and let confirmed ground evidence override a middling model score.
Cost: no budget this season for a full recalibration pass. Ship the three tiers using the model's existing thresholds first, and calibrate them against real outcomes once a season of data exists.
The model got better, for real: say LeafGuard's blight detection accuracy jumps from 84 to 95 percent. Keep the tiers anyway. A grower still can't act on "95 percent," they can act on "treat now," and a better model just means the reason attached to that tier gets more specific, not that the tier disappears.

Where people run it wrong.
They show the number "for transparency" and quietly rebuild the private-cutoff problem inside a single season.
They treat the model's own confidence as the whole picture, and miss that confirmed, real world evidence should sometimes outrank it.
They pick tier thresholds once at launch and never check them against what actually happened after.

How to use it live. Ask this before naming a fix: "if the model and a confirmed, real world signal ever disagree, which one wins, and does the screen say so?" That question alone usually finds whether a design actually understood the difference between a model's confidence and the thing it's forecasting.

Three things worth stating directly, since the real judgment lives here. The alternative genuinely on the table was keeping the raw percentage but adding a plain sentence under it, "usually means take cover," and it lost because testers anchored on the number the moment it was visible and ignored the sentence entirely, the same failure the small "78%" side by side test produced. The AI-specific failure this whole design guards against is treating model confidence as if it were the same thing as real world risk, when a model can be honestly, correctly uncertain about a storm that turns out to be extremely dangerous once it forms; the guardrail is letting confirmed ground signal, radar rotation, a visible lesion, override the model's own number rather than wait for the model to catch up. And the trade-off is real and accepted on purpose: widening the radar-override condition catches fast, low-predictability storms sooner, but it also means Take shelter now fires a little more often on storms that fizzle, trading a small amount of extra disruption for cutting reaction time on the dangerous ones from roughly six minutes back down under ninety seconds.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question asking you to design how confidence gets shown to a non-technical user?
Tap to flip
ANSWER
SPARK: ground it in a real situation, name the habit you want, design the concrete anchor, prove it against its own risk, then say what you'd deliberately leave out.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Halldis Fairbourne, director of Camp Tallowbrook for eleven summers, the one who decides whether to move 140 kids to the storm cellar.
3 · THE SITUATION
What does the app show today, and why does that break down for this audience?
Tap to flip
ANSWER
A raw model percentage with no context. Camp staff aren't statisticians, so each of them invents a private cutoff, and that cutoff drifts every time the number turns out to be nothing.
4 · THE ANCHOR
What's the concrete anchor, the actual screen Aslaug builds?
Tap to flip
ANSWER
Three plain tiers, Stay aware, Get ready, Take shelter now, each with a real reason instead of a number, calibrated to real outcomes, with confirmed radar able to override a middling model score.
5 · THE OLD DECISION
What decision would Aslaug take back?
Tap to flip
ANSWER
Putting the model's raw confidence number front and center on the camper-facing alert screen, because it was the most honest figure the team had and the engineers testing it could already read it.
6 · THE NUMBER
Fill in the blank: by week 5, StormLatch sent ___ alerts past Halldis's cutoff, and only ___ were followed by real severe weather.
Tap to flip
ANSWER
11 alerts, only 2 followed by real severe weather. By week 9, shelter time had crept from 90 seconds to past 6 minutes.
7 · THE REPLAY
Same near miss Thursday, new design already live, what changes?
Tap to flip
ANSWER
The model's percentage still reads moderate, but confirmed radar rotation jumps the alert straight to Take shelter now on its own, cutting reaction time from a four minute scramble to under ninety seconds.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent design?
Tap to flip
ANSWER
LeafGuard, Amberfield AgTech's crop-disease scanner. The equivalent design is "Treat now / Watch closely / Looks fine" tiers instead of a blended composite score, with a visible lesion able to override an ambiguous scan.

Check yourself Score: 0 / 0

True or false
1. True or false: StormLatch's model was wrong when it read 52 percent for the storm that produced the near miss at Camp Tallowbrook.
  • True
  • False
Show hint
Check the reliability chart in "Let's learn."
Show answer
False. Storms in the 50 to 60 percent bucket really did turn severe about half the time, exactly what the number should mean. The model was honest; the number just wasn't a message a tired counselor could act on in time.
Multiple choice
2. Why does the redesigned alert let confirmed radar rotation override the model's own percentage?
  • A. Radar data is cheaper to process than the model's forecast.
  • B. A storm's actual danger and the model's confidence about it are two different facts, and confirmed ground signal can be the stronger one.
  • C. Counselors don't trust the model at all, so radar is used instead of it entirely.
  • D. Radar rotation always means a tornado has already touched down.
Show hint
Look at the Anchor and Risk steps in the SPARK recap.
Show answer
B. A fast-forming storm can be genuinely hard for a model to forecast twenty minutes out while still being confirmed and dangerous in the moment, which is exactly why ground-truth radar gets a vote.
Fill in the blank
3. A shared calibration bar across all three tiers false-alarmed on ___ nights of Camp Tallowbrook's 11 week season. Calibrating each tier separately, with radar override, cut that to ___ night, without missing a real storm.
Show hint
Check the bar chart in the SPARK recap.
Show answer
6 nights, cut to 1 night. The raw-percentage design, by comparison, had crossed Halldis's own cutoff on 11 nights, only 2 of them real.
Short answer, name the rejected alternative
4. What alternative did Aslaug's team try and reject instead of hiding the raw percentage entirely, and why did it lose?
Show hint
Look at the K step, Keep out, in the SPARK recap.
Show answer
Model answer: Showing the tier alongside the raw percentage, like "Take shelter now, 78%," as extra context. It lost because staff started comparing the small number story to story within days, and the exact private-cutoff problem the tiers were built to fix came straight back.
Short answer, apply it yourself
5. Think of an app you use that shows you a percentage or a score, a weather app, a health app, a delivery estimate. Where would a plain tier with a real reason serve you better than the raw number does now?
Show hint
Look for a number you've quietly built your own private cutoff around, without ever being taught to.
Show answer
Model answer: A food delivery app showing "68% chance your order arrives late." Almost nobody knows what to do with that number. "Order running behind, driver still 3 stops away" tells you the same thing without asking you to interpret a percentage at all.
Fill in the blank, work the number
6. If Camp Tallowbrook's shelter time had stayed at its week 1 pace of 90 seconds instead of drifting to over 6 minutes, roughly how many extra minutes did the season's false-alarm fatigue cost the camp on the night of the near miss?
Show hint
Compare the week 1 shelter time to the actual time on the near-miss night.
Show answer
Roughly 4 to 5 minutes. Week 1's drill took about 90 seconds. The near-miss night took about 6 minutes before Fikile acted on the funnel cloud itself, not the number, a gap of four to five minutes that a fast-forming tornado does not forgive.
Before you close the answer
Why this works
Tests whether you understand that a model's confidence and the real-world danger it's forecasting are separate facts, and whether you can design a communication layer that's honest about both without asking the reader to know any statistics.
Follow-up traps
"Isn't hiding the real number just dumbing it down, or even a little dishonest?" Response: the number is still there, on Bodil's dashboard, for the person whose job needs it. Hiding it from a decision that has to happen in four seconds isn't dishonesty, it's matching the information to the decision it's actually for.

"What if the radar-override rule fires too often and people stop trusting Take shelter now the same way they stopped trusting the old number?" Response: that's exactly what the per-tier calibration and the 11-vs-6-vs-1 test were built to catch, and it's why the bar gets checked against real outcomes on a rolling basis, not set once and left alone.
If pressed
The radar override isn't a simple if-then rule bolted on top of the model. It works because StormLatch's team built a second, much shorter-horizon nowcasting layer that only looks two to fifteen minutes ahead using live radar returns, separate from the model that forecasts twenty minutes to two hours out. The two disagree by design on fast, low-predictability storms, and the tier logic is built to trust whichever one is more confident for the time window that actually matters right now.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more