ConceptIntermediateDesigning for Uncertainty & Trust / Designing for failure and graceful degradation / #5
Describe three failure modes to design for before launch.
ORDER rank by what you can never win back, not by what's loudest
Thermara predicts which building HVAC units need service before they fail, using vibration and temperature sensors on each unit. Devendra Kulkarni has serviced rooftop units for nine years. A routine reliability audit, months before a planned expansion, found three separate ways Thermara could still go wrong, and the order they had to be handled in mattered more than any of them alone.
The direct answer
Design for three failure modes before launch: silent sensor gaps that produce false confidence, new equipment types the model was never calibrated on, and alert fatigue that makes technicians tune the system out. Build the sensor-gap detector first, since it has to exist before any alert can be trusted, but rank alert fatigue as the most urgent to prevent, because a technician who stops trusting the system is far harder to win back than a wrong reading is to fix.
Do this, in order
Design against alert fatigue as the top priority, even though it isn't built first.Why: a technician who decides the system cries wolf stops being a decision you can quietly reverse later.
Build sensor-gap detection before anything else technically.Why: every other failure mode depends on knowing whether the model's input was even real that day.
Stage new equipment types through a calibration pilot before trusting their readings.Why: a model trained mostly on one manufacturer's units will misread another one with confidence, not hesitation.
Test all three failure modes cheaply before launch, not just in production after.Why: a simulated sensor dropout and a held-out equipment type both cost a pilot week. A lost customer costs a lot more.
Leave the alert threshold alone for equipment types the model has already proven itself on.Why: this ranking only matters for the failure modes that are actually live risks, not every dial in the system.
How to answer this, stage by stage
Nobody is grading whether you can name three ways an AI system breaks. They're grading whether you can rank them by what actually can't be undone.
Stage 1
Scope it to one system, one launch
Say it like this
"I'll answer this using Thermara, a predictive maintenance tool for building HVAC units, and the three failure modes its own pre-launch audit actually found."
Why this works
Keeps "describe three failure modes" from turning into a generic checklist with no real stakes behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what all three failure modes threaten. Reversibility, which is hardest to undo. Dependency, what has to be built first. Evidence, what's cheap to test before launch. Rank, the final order, defended."
Why this works
Shows a real prioritization method instead of three unranked bullet points.
Stage 3
Name the outcome all three failure modes threaten
Say it like this
"All three threaten the same thing: whether a technician actually acts on what Thermara tells them. A model can be accurate and still fail completely if nobody trusts its output enough to act."
Why this works
Grounds the ranking in something real, instead of ranking by which failure sounds scariest.
Stage 4
Name the three failure modes plainly
Say it like this
"One: sensors go quiet and the model still outputs a confident 'fine' from stale data. Two: a new building's equipment type doesn't match what the model learned on, and it misreads it with false confidence. Three: too many low-priority alerts and technicians start dismissing them without reading them."
Why this works
This is the direct content of the answer, stated concretely instead of abstractly.
Stage 5
Rank by what's hardest to undo
Say it like this
"Alert fatigue ranks first, even though it isn't built first. A wrong reading gets fixed and forgotten. A technician who's decided the system cries wolf keeps ignoring it long after the bug is gone."
Why this works
This is the direct answer to the question, and it's the part most candidates skip by ranking on severity alone instead of reversibility.
Stage 6
Name the dependency order, separately
Say it like this
"Technically, sensor-gap detection has to be built first, since you can't design a trustworthy alert on top of data you don't know is real. New equipment calibration can run in parallel through a staged rollout."
Why this works
Shows you understand that build order and priority order aren't always the same thing.
Stage 7
Say what you'd test cheaply before launch
Say it like this
"I'd simulate a sensor dropout in a shadow-mode pilot, run a held-out equipment type through calibration before it goes live, and watch alert-dismissal rates in an internal beta before the wider rollout."
Why this works
Shows the ranking isn't just theoretical, it's testable before real customers are exposed to it.
Stage 8
Close on the one line
Say it like this
"Rank failure modes by what you can't win back, not by what looks scariest on a slide. A wrong reading is a bug. A technician who's stopped trusting you is a much longer project."
Why this works
Restates the direct answer in one breath, ready for a follow-up push.
Let's learn
Thermara watches vibration and temperature readings from sensors on rooftop HVAC units and tells a building's maintenance team which one needs service soon, before it actually fails.
Before Thermara, a technician like Devendra serviced units on a fixed schedule, every unit checked every ninety days whether it needed it or not, catching most real problems but wasting a lot of visits on units that were fine.
All three of these were found in the same pre-launch audit, months before any of them had actually happened to a real customer.
Now Thermara flags maybe fifteen units a month across a portfolio for service, cutting wasted visits by more than half.
Here's the turn: none of these three failure modes are about the model being wrong on purpose. They're about three different ways a genuinely working model can still lose the trust of the person meant to act on it, and losing that trust turned out to be a very different kind of damage than getting a single reading wrong.
False "fine" reads during a simulated 6-hour sensor dropout, by equipment age
Same dropout, same six hours. The model's confidence didn't drop on the unfamiliar equipment. It just kept guessing "fine" in a shape the audit hadn't expected.
At its worst, a compressor fails on a unit the model had confidently called fine for weeks, because a third of its sensors had gone quiet and nothing on Thermara's side ever said so out loud.
One of these you can patch overnight. The other one takes months of being right before anyone trusts you again.
The decision I would take back
Thermara's early onboarding materials told new facility teams the system was reliable enough to replace fixed-schedule checks entirely, written before anyone had tested it against a new equipment brand or a real sensor dropout. That made sense while the pilot customers all ran the same familiar equipment. It stopped making sense the moment Thermara expanded to buildings running gear the model had never actually been calibrated on.
What I would leave alone: equipment types the model has already proven itself on, over months of real service history, don't need a new calibration pilot every time. This ranking is about the live risks, not every dial in the system.
The lesson: the failure mode that costs you the most is rarely the one that breaks the loudest. It's the one that quietly changes what a person is willing to believe next time.
Now here is the same thing as a story
The short version above is what you'd say defending this ranking to Thermara's launch review board. Read this one for how the audit actually found all three.
Devendra can tell a compressor's about to go by the sound alone, a skill from nine years climbing onto rooftops with a service belt and a printed maintenance log. Thermara's rollout to his region promised to replace most of that guesswork with real sensor data.
Knowledge spark: what does "distribution shift" actually mean here?
A model learns patterns from the data it was trained on, which for Thermara meant thousands of hours from one manufacturer's rooftop units. A different brand of equipment vibrates differently, runs at different temperatures, and ages on a different curve. The model isn't wrong to be confident on the equipment it knows. It's wrong to stay just as confident on equipment it's never actually seen, and nothing about a normal confidence score tells you which situation you're in.
Three months before a planned expansion into a new region, Thermara's team ran a routine reliability audit, someone above the engineering team pulling ten recent alerts at random and checking each one against what actually happened on-site. It wasn't looking for anything specific. It found three things at once.
None of these three had caused a real failure yet. The audit found them because it went looking before launch, not after.
First: on a handful of units, a third of the vibration sensors had been dropping out intermittently for days, and Thermara's model, fed stale readings from before the dropout, kept outputting a calm "fine" the entire time. Second: a small pilot building running an unfamiliar equipment brand showed a false "fine" rate nearly ten times higher than the model's usual performance, with no drop in its stated confidence to match. Third, and the one that worried the team most: three technicians in a different region had quietly started ignoring Thermara's lowest-priority alerts entirely, after weeks of those alerts pointing to units that turned out to need nothing at all.
The model was never lying about any of these. It just didn't know what it didn't know, and nothing told the person reading the screen that it didn't.
Devendra had personally started doing this, quietly cross-checking the lowest-priority alerts against his own instinct before acting on them, a private workaround nobody on the product team knew existed until the audit went looking.
Alert fatigue isn't the rarest failure. It's the one sitting furthest into "hard to undo," which is exactly why it ranks first.
With the redesigned launch plan, sensor-gap detection now flags stale data explicitly, on-screen, before any reading built from it reaches a technician. New equipment types run through a two-week shadow calibration before their alerts count as trustworthy. And Thermara now tracks each technician's dismissal rate on its lowest-priority alerts as its own number, escalating for review the moment it climbs, instead of only ever showing up in an audit three months later.
Sensor-gap detection had to come first, technically. But the alert-fatigue risk it enables catching was the one that mattered most to prevent.
The old plan asked technicians to simply trust the model more with each launch. The new one earns that trust back, one verified reading at a time, before it ever asks for it.
We built Thermara to sound confident, on purpose, so a busy technician wouldn't have to parse a hedge on every single alert. It took an audit three months early to see that confidence without a way to check it just teaches people, eventually, to stop believing it at all.
ORDER, in one screenNot a lecture on testing plans. ORDER is what forces you to rank by what you can't undo, not by what's easiest to picture.
O
Outcome. What all three failure modes are competing to protect.
Whether a technician actually acts on what Thermara tells them. A model can be accurate and still fail if nobody trusts it enough to act.
Grounds the whole ranking in something real, instead of ranking by which failure sounds scariest.
R
Reversibility. Which is hardest to undo.
Alert fatigue is the hardest. A wrong reading gets patched and forgotten. A technician who's decided the system cries wolf keeps ignoring it long after the bug is gone.
The hardest step and the real answer to the question: ranking by what you can't win back, not by severity alone.
D
Dependency. What has to be built first.
Sensor-gap detection has to exist before any alert can be trusted. New equipment calibration can run in parallel through a staged rollout.
Separates the build order, which is forced by reality, from the priority order, which is a judgment call.
E
Evidence. What's cheap to test before launch.
A simulated sensor dropout, a held-out equipment type, and an internal beta's dismissal rate, all testable before a real customer is exposed to any of the three.
Shows the ranking is checkable in a pilot, not just argued from a whiteboard.
R
Rank. The final order, defended.
Alert fatigue first to guard against, sensor-gap detection first to build, new equipment calibration staged in parallel.
A real, defensible order instead of three equally weighted bullet points.
Every one of the three failure modes gets tested on purpose, weeks before a real customer ever sees the system live.
The recap, one line per letter: outcome is whether a technician actually acts on what they're told, reversibility ranks alert fatigue first since trust doesn't rebuild on its own, dependency puts sensor-gap detection first technically, evidence is testing all three cheaply before launch, and rank is the final order, defended by what can't be undone.
And if you want to be sure it really works, try it somewhere elseSame five letters, a hospital's sepsis early-warning system instead of a rooftop HVAC unit. A different set of three failure modes, the same ranking logic.
Halvorsen General runs an early-warning model that flags patients showing signs of developing sepsis. Nurse Aline Okonjo has worked the medical ward for twelve years. Mapped onto ORDER: the outcome is whether a nurse actually escalates a flagged patient, not just whether the model's own accuracy looks good on paper. The three candidate failure modes here are a monitor that silently stops transmitting vitals, a rare patient presentation the model was never trained to recognize, and false-positive fatigue from too many low-confidence flags on stable patients.
Reversibility ranks the same way it did for Thermara: a nurse who's decided the alert "always cries wolf" is far harder to win back than a single missed reading is to patch, so false-positive fatigue ranks first to guard against even though monitor-dropout detection, structurally, has to be built first. The evidence step is identical in shape too, a simulated monitor dropout, a held-out rare presentation, and a nursing-floor beta tracking dismissal rates, all cheap to run before the system reaches a real patient.
Swap "technician" for "nurse," and the same asymmetry between a fixable bug and a lost habit still holds.
Nurse response time to flagged patients, before and after dismissal-rate tracking
Response time crept up for months before tracking existed to catch it, then recovered only partway once it did, evidence that trust, once spent, doesn't fully return on its own.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "rank by what can't be undone, not by what's scariest, and alert fatigue almost always wins that ranking," and stop.
Cost: there's no time to build all three guardrails before this launch. Say so honestly, and start with dismissal-rate tracking, since it's the cheapest early-warning signal for the failure mode that's hardest to reverse.
The model gets better, for real: if Thermara's underlying accuracy genuinely improves and false reads become rare, that's still not a reason to drop sensor-gap detection, a rarer gap is exactly the one nobody will think to double-check anymore.
Where people run it wrong.
They rank failure modes by which one sounds most dramatic in a review meeting, not by which one actually can't be undone.
They treat build order and priority order as the same list, when dependency and reversibility often disagree.
They wait for a real incident to discover a failure mode an audit could have found for free, months earlier.
How to use it live. When someone asks you to name failure modes to design for, ask yourself one question before ranking them: if this happens and gets fixed, does the person on the other end trust the system again automatically, or do they have to be won back? Rank whichever one requires winning back first.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "name failure modes to design for" question?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Rank by what's hardest to undo, not by what sounds scariest.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Devendra Kulkarni, an HVAC field technician for nine years, who quietly started cross-checking Thermara's lowest-priority alerts against his own instinct.
3 · THE THREE FAILURE MODES
What are the three failure modes this answer names?
Tap to flip
ANSWER
Silent sensor gaps producing false confidence, new equipment types the model was never calibrated on, and alert fatigue from too many low-priority flags.
4 · THE RANK
Which failure mode ranks first to guard against, and why?
Tap to flip
ANSWER
Alert fatigue. A wrong reading gets patched and forgotten. A technician who's decided the system cries wolf keeps ignoring it long after the underlying bug is gone.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling new facility teams at onboarding that Thermara was reliable enough to replace fixed-schedule checks entirely, before it had been tested against a new equipment brand or a real sensor dropout.
6 · THE NUMBER
Fill in the blank: during the simulated dropout test, unfamiliar equipment showed a false "fine" rate of about ___ percent, versus 4 percent on familiar equipment.
Tap to flip
ANSWER
About 38 percent. Nearly ten times higher, with no drop in the model's stated confidence to match.
7 · THE DEPENDENCY
Which failure mode's fix has to be built first, technically, even though it isn't ranked first?
Tap to flip
ANSWER
Sensor-gap detection. No alert can be trusted until the system knows whether the data behind it was even real that day.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and how does the ranking logic transfer?
Tap to flip
ANSWER
Halvorsen General's sepsis early-warning system. False-positive fatigue ranks first to guard against, the same way alert fatigue did for Thermara, even though monitor-dropout detection has to be built first.
Check yourself Score: 0 / 0
Multiple choice
1. Why does alert fatigue rank first to guard against, even though sensor-gap detection has to be built first?
A. Alert fatigue is the most common of the three failure modes.
B. A wrong reading gets fixed and forgotten, but a technician who's stopped trusting the system takes far longer to win back.
C. Sensor-gap detection is too expensive to build first.
D. Regulations require alert fatigue to be addressed before anything else.
Show hint
Look at the Reversibility step.
Show answer
B. The ranking is built on what can't be undone, not on frequency or build cost.
True or false
2. True or false: the redesigned launch plan requires every equipment type, including ones already proven over months of service, to go through a new calibration pilot before every launch.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. Equipment types the model has already proven itself on over real service history don't need a fresh calibration pilot every time.
Fill in the blank
3. Fill in the blank: the pre-launch audit pulled ___ recent alerts at random and checked each one against what actually happened on-site.
Show hint
Look at the story section describing the audit.
Show answer
Ten. A small, cheap sample was enough to surface all three failure modes at once, months before a real customer hit any of them.
Short answer, apply it yourself
4. Think of an alert or notification system you've personally started ignoring. Was it because it was often wrong, or because there were just too many of them?
Show hint
Ask whether you'd trust it again immediately if the underlying accuracy improved tomorrow, or whether the habit of ignoring it would stick around anyway.
Show answer
Model answer: Most people can name a notification they now ignore on reflex, even for things they'd technically want to know about, the same sticky habit alert fatigue describes.
Short answer, name the reversal
5. What old decision does this answer take back, and why did it make sense when Thermara first launched to its early customers?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Telling new customers the system was reliable enough to fully replace fixed-schedule checks. It made sense while every pilot customer ran the same familiar equipment.
Short answer, where it wouldn't matter
6. Name a part of Thermara's system where this three-failure-mode ranking genuinely doesn't need to apply.
Show hint
Look at the fifth bullet in the priority list.
Show answer
Model answer: Equipment types with months of proven, real service history behind them. The live risk this ranking addresses doesn't apply once a type has already earned its trust.
Before you close the answer
Why this works
Tests whether you can rank real risks by what actually can't be undone, instead of listing three scary-sounding failure modes with no order behind them.
Follow-up traps
"Isn't the sensor-gap failure objectively worse, since it could cause an actual equipment failure?" Response: a single missed equipment failure is genuinely serious, but it's also a fixable, one-time cost. A workforce that's quietly stopped trusting the system compounds every day after, on every alert, not just the ones related to the original bug.
"How would you even measure alert fatigue before it becomes a real habit?" Response: track each technician's dismissal rate on low-priority alerts as its own number, and escalate for review the moment it climbs, instead of waiting for the next audit to notice.
If pressed
Thermara's redesigned confidence score is now computed separately from its sensor-completeness score, and the alert only shows a "fine" reading at full confidence when both scores clear their own threshold, so a stale-data situation can never borrow confidence from a model that's actually just missing information.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.