CaseAdvancedShipping & Model Lifecycle / Incident management for AI products / #12
What is the communication plan when the incident is publicly visible?
Incident management for AI products
The direct answer
Once an incident is already visible outside the company, the first public statement names only what is confirmed, right now, and what is being done about it. It states plainly what is not yet known, and it commits to a named time for the next update. It never names a cause in that first message, because in hour one the only source for a cause is the model's own unchecked account of itself.
Do this, in order
Lock the first statement to four lines: confirmed scope, action taken, what is not known, named next update time.Why: this is the one design decision the whole plan hangs on. Skip it and every later choice is a guess made under pressure.
Never let the model's own incident summary supply the public cause line.Why: a confidence score is not a check. It is exactly what turns an incomplete diagnosis into a public claim.
Name a real, confirmed scope number, even a wide one, instead of waiting for the exact final count.Why: waiting for precision costs more trust than an honest "at least this many, still confirming" said early.
Commit to a named next update time and hit it, even with an unfinished answer.Why: a promise with a clock on it survives new facts. A promise that it is fixed does not.
Leave out the cause, any "this can't happen again" line, and any account level detail.Why: none of it is needed to be believed in hour one, and all of it is exactly what hour four proves wrong.
Require an independent reconciliation check, not the model's own confidence, before a cause ever goes public.Why: this is the guardrail that stops the next incident's first draft from making the same overclaim.
The whole plan is really this one document. Four lines, no fifth line, and a cause is never one of them.
How to answer this, stage by stage
Nobody is grading whether you can say "we'd be transparent." They are grading whether you can hand over a locked design for the exact sentence that goes out first, before anyone knows how the story ends. Seven moves get you there.
1
Scope it to one real product, one real person
Say it like this
"Let's ground this. Fairmarch Wealth runs an engine called Cambrix that auto rebalances about two hundred and ten thousand accounts overnight. Denice Corliss runs trust and communications there. She's the one who has to write the first public line once an incident is already outside the building."
Why this works
Grounds the answer in a real number and a real seat before any framework talk starts.
2
Say what the plan is actually for, before any story
Say it like this
"Here's how I'd frame it up front. A communication plan for a public incident isn't a press release template. It's a rule that stops the first message from saying more than the team can still stand behind six hours later."
Why this works
States the reframe before a single detail, so the interviewer hears a design decision, not a PR instinct.
3
Name the anchor, the one decision the whole plan hangs on
Say it like this
"The anchor is this. The first public statement is only allowed to say two things: a confirmed number, and an action already taken. It is never allowed to name a cause, because the only source for a cause in hour one is the model's own incident summary, and nobody has checked that yet."
Why this works
This is the line an interviewer is listening for. It turns a vague value like "be transparent" into an inspectable rule.
4
Show the day it breaks, and why
Say it like this
"Here's the risk this has to survive. An earlier version of Denice's plan let Cambrix's own root cause line straight into the nine a.m. statement. It said isolated, already patched. Five hours later, reconciliation found two more affected funds the drift detector had missed. Isolated and patched were still sitting on the page, and a reporter quoted them next to the new complaints."
Why this works
Names the specific way an AI incident's comms plan fails: trusting the model's own account of itself as if it were checked.
5
Name what you considered and ruled out
Say it like this
"I'd also say what I ruled out. Staying quiet until the cause is fully confirmed doesn't work either. Once an incident is already public, silence reads as proof something is being hidden. The fix isn't speed or silence. It's narrowing what the fast version is allowed to claim."
Why this works
A rejected alternative shows a real judgment call, not the first idea that came to mind.
6
Say what's deliberately left out, and why that's safe
Say it like this
"The first statement leaves out the cause, any promise it can't happen again, and any account level detail. None of that is needed to be believed in hour one. All of it is exactly what a later hour tends to prove wrong."
Why this works
Shows judgment instead of a wish list. Leaving something out on purpose is a design choice, not a gap.
7
Show the replay, and close on the count
Say it like this
"Run the same day with the redesigned plan. First statement goes out forty four minutes after the thread crosses five hundred replies, naming at least four thousand one hundred and twenty accounts confirmed so far, and a promise to update by four p.m. whether or not the cause is known. At two ten, when two more funds turn up, they fold into that four p.m. update instead of contradicting a claim nobody made. Same bad day. No retraction."
Why this works
Closes on a number and a clock someone could check, not a promise to "communicate better" next time.
Let's learn
Say a wealth management company builds an engine that moves your money between funds overnight, on its own, to keep your account matching the risk level you picked. Call the company Fairmarch Wealth and call the engine Cambrix. It handles automatic rebalancing for about two hundred and ten thousand accounts, every week, while the account holder sleeps.
Right now, when something like this goes wrong and customers notice before the company says a word, someone drafts a statement fast, usually inside an hour. But fast has a cost nobody counts. In two of the last three incidents like this one this year, that fast draft named a cause nobody had actually confirmed yet, because the cause was sitting right there in the tool that watches Cambrix for problems, and it looked confident.
Knowledge spark: what is a drift detector's flag line?
Cambrix has a watcher built to catch its own mistakes. It flags any account whose real mix of funds has moved more than two percentage points away from the target the customer picked. Checked against a labeled set of past corporate action mistakes, it catches about ninety seven percent of the confirmed ones. Not all of them. The other three percent are exactly the kind that slip through quiet, not because nobody built a check, but because every check has a line under which it stops looking.
With a locked plan, the statement still goes out inside an hour. It just never claims a cause. It names a confirmed number, an action already taken, and a time it will say more.
Without a locked plan, the fastest person in the room becomes the plan, and whatever the model says about itself rides straight into the public line.
Here is the turn. The extra minutes it takes to write the disciplined version are not the real risk. The real risk is a sentence that a person copies onto a public page in hour one and then has to unsay in hour four. Speed was never the problem. An overclaim was.
The danger was never how long the first message took to write. It was whether the sentence on the page was still true four hours later.
At its worst, this costs far more than the original mistake. A wrong early claim, once it gets contradicted, reads as proof of a cover up even when nobody was covering anything up. That pulls in press coverage, regulator questions, and a story that isn't about the incident anymore, it's about the statement.
What the company could honestly say, at two moments in the same day
The number grew by 290 accounts once two more funds were folded in. Nothing on the page needed correcting, because the first statement never claimed to be final.
The choice I would take back. When Cambrix's own incident summary tool shipped, someone on the trust team suggested wiring its output straight into the public statement draft, to save Denice the ten minutes it usually took to write one by hand. That merged two steps that used to be separate: the internal summary, and the public claim. Collapsing them removed the pause where a person would normally ask, has anyone outside the model actually checked this yet.
The decision that mattered
Letting the model's own incident summary flow straight into the public draft, with no independent check in between. It made sense the week it shipped, since the tool had been right on every smaller incident before this one. It also meant the one thing that could have caught an incomplete diagnosis, a human asking "checked by what," was gone.
What I would leave alone. Internal engineering write ups can name a cause the moment engineers believe they have found it. That is not what is broken, and slowing that down would only make the real fix take longer. An incident quiet enough that nobody outside the company ever notices does not need this locked plan either. A short private note to the accounts affected is enough.
The lesson. A first public statement is not a report. It is a promise to keep updating, not a promise of an answer.
Now here is the same thing as a story
The short version is above. Read on if you want to feel what one confident sentence, and forty four minutes, actually cost.
Denice Corliss has run trust and communications at Fairmarch Wealth for four years. Give her a spike in support tickets and she can tell inside a minute whether it is a real incident or a Monday grumble about a slow app. She built that instinct on small, honest things: a login outage, a late statement, a display bug that made a balance look wrong for an hour.
For most of those four years, Cambrix's own incident summary tool made her fast. It watched every system Cambrix touched and, when something broke, it wrote a plain paragraph explaining what happened. She would read it, check it against the engineering channel, and build her public line around it. It was right, again and again, on the small stuff.
So she started checking it less. First she skimmed the engineering channel instead of reading it closely. Then, for the smaller incidents, she stopped checking at all. She would just paste the tool's own summary line into her draft, because it had been right nine times running, and nine times feels like a rule.
Nobody decided to stop checking. Nine correct answers in a row just started to feel like a promise the tenth one never made.
One Monday, support tickets about "my portfolio looks wrong" started climbing around 8:02am. The volume looked routine, the kind that usually meant a display glitch. At 8:55am, a customer posted a screenshot of their account to IndexTalk, a large investing forum, showing a bond fund gone and a much heavier stock position in its place. Denice pulled Cambrix's incident summary the way she always did. It reported, in a clean and confident paragraph, "isolated batch classification error in overnight corporate action processing, patch already deployed." She built her statement around that line and posted it at 9:04am, nine minutes after the first screenshot, with no defined trigger telling her whether this counted as public yet. It just felt public, so she moved.
At 9:29am the forum thread crossed five hundred replies. Now it was unmistakably public, and Denice's statement, already live for twenty five minutes, called the problem isolated and already patched.
Same discovery, six hours later. One design breaks a promise. The other one keeps it, because it never made the promise the first design made.
At 2:10pm, Fairmarch's reconciliation team, running its slower, independent trade log check, found two more bond funds caught in the same misclassified corporate action. Cambrix's own drift detector had never flagged them. Their drift sat at 1.8 percentage points, just under its 2 point flag line. Isolated was now wrong. Patched was now wrong. At 2:31pm, a reporter with about forty thousand followers quoted Denice's own 9:04am line, word for word, next to a fresh customer screenshot.
The dollar cost of putting 4,410 accounts back where they belonged came to about $340,000 in reversed trades and tax adjustments over one weekend. That part was survivable. What actually hurt was three finance publications running a version of "robo advisor understated its own mistake," and a state regulator's office opening an informal inquiry that took eleven weeks to close, over the contradiction, not the original error.
The two extra funds were never the real problem. The real problem was a sentence Denice had already promised nobody would need to unsay.
The old decision, told as a memory of a meeting. When the incident summary tool shipped eighteen months earlier, someone suggested wiring it straight into the comms template to save time. Everyone in that room had watched it be right, incident after incident. Nobody was wrong about what the data showed them that week.
The replay, run the same Monday forward with the redesigned plan already in place. Same tickets at 8:02am. Same screenshot at 8:55am. Same five hundred replies at 9:29am, which is now the defined trigger, not a feeling. At 10:13am, forty four minutes later, the first statement goes out: at least 4,120 accounts confirmed affected so far, reconciliation still confirming the final count. All automatic rebalancing tied to corporate action events, paused company wide as of 9:41am. We do not yet know the root cause. We will publish an update by 4:00pm ET today, whether or not we have a complete answer. At 2:10pm, the same two funds turn up. They fold straight into the 4:00pm update Denice already promised. Nothing on the page needs walking back.
What I would tell my past self, back in that first meeting: a tool that is right nine times running has not earned your trust. It has earned a check that catches the tenth time, before the public does.
SPARK: designing the statement before the story breaks
This is a design question wearing a crisis costume. Nothing has broken yet when the plan gets written, so SPARK runs forward, building the anchor before the failure it has to survive.
S
Situation. Who is this for, and how does it get handled today, without the plan?
One person, one real desk, not a policy on a shelf.
Denice Corliss, four years running trust and communications at Fairmarch Wealth. Today, without a locked plan, whoever is fastest in the room drafts a line straight from the ops channel, often lifting the model's own incident summary wholesale because it is sitting right there and it is usually right.
P
Payoff. What habit should this build, and what should it let her stop doing?
The habit is the product. The time saved is downstream of it.
The habit: separate what is confirmed from what is still being checked, every time, without deciding it fresh under pressure. What she gets to stop doing: reaching for one reassuring line that makes the company sound already in control.
A
Anchor. What is the one design decision everything else hangs on?
The hardest step, and the actual answer to the question.
The first public statement is locked to four lines only: a confirmed scope number, an action already taken, an explicit line on what is not known yet, and a named clock for the next update. No fifth line, and a cause is never one of the four.
R
Risk. What breaks the first time this is wrong?
Not "accuracy drops." What the reader does next.
An earlier version let the model's own root cause line into that first statement. It said isolated, already patched. Five hours later, two more affected funds surfaced that the drift detector had missed, and the earlier line sat there, contradicted, screenshotted next to new complaints.
K
Keep out. What does the plan deliberately leave out of the first message?
Judgment, not a wish list.
No named cause, no promise it cannot happen again, no account level detail, no internal team named. None of it is needed to be believed in hour one, and all of it is exactly what a later hour proves wrong.
And if you want to be sure it really works, try it somewhere else
Solden Veterinary Labs runs Farrowgate Read, an AI tool that scans X-rays sent in by partner clinics and flags likely findings for a radiologist to review before the pet's own vet sees the result.
S. Sekou Girard, clinical communications lead at Solden, four years writing every notice that goes out to referring clinics when something goes wrong with a batch of scans. P. The habit: never let the first notice claim a diagnosis is right or wrong, only that a batch is back in front of a person. What he stops doing: promising the software already caught its own mistake before a radiologist has actually looked again. A. A different anchor than Fairmarch's, because the failure shape is different. Instead of a scope number, the anchor here is a guaranteed action: "every flagged image from the affected update is back in front of a licensed radiologist before any client hears a final word," posted before the exact count of affected images is even final. Clinics waiting on a result need the action first, not the number. R. An early note said "the false flags were limited to one clinic's batch," pulled straight from Farrowgate Read's own change log. Within a day, the same pattern reached three more clinics the change log had not caught. K. The notice never speculates on which flagged cases are false and which are real, because some pet owners in that same batch had genuine findings mixed in, and guessing which is which in public is exactly the overclaim the plan exists to prevent.
Same method, different anchor
At Fairmarch the anchor was a number and a clock. At Solden it is a guaranteed re-review, because nobody waiting on a scan result wants a count, they want to know a person is looking again. Different shape, same rule underneath: never let the system under question narrate its own explanation before an independent check backs it up.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the anchor, the four locked lines, and the one rule that a cause is never in the first message.
Cost: the comms team is one person covering three products, with no time to write custom lines under pressure. Pre-draft the four line skeleton for every product now, before any incident, and fill in only the numbers live.
The model got better, for real: suppose Cambrix's drift detector gets tuned from catching 97 percent of confirmed misclassifications to 99.9 percent. The plan does not relax. A detector being usually right is not the same as this specific claim being checked, and the public statement still waits for the independent reconciliation, not the model's improved confidence.
Where people run it wrong.
They let whoever is fastest in the room decide what to say, instead of running the same four lines every time.
They wait for a fully confirmed cause before saying anything at all, and by the time they publish it reads like a cover up.
They treat the model's own confidence in its summary as if it were an independent check, instead of a lead that still needs verifying.
How to use it live. Say the reframe before any story: "the point of a communication plan for a public incident isn't to sound like everything's under control in hour one. It's to make sure nothing said in hour one has to be walked back in hour four." That buys the room to name the actual anchor, instead of reciting "we would be transparent and act quickly."
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
What framework is this, and what makes it fit a design question like this one?
Tap to flip
ANSWER
SPARK. It runs forward, designing the plan against a failure that has not happened yet, instead of running backward from a failure that already did.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Denice Corliss, Head of Trust and Communications at Fairmarch Wealth, who has written every public incident notice the company has sent in four years.
3 · THE PAYOFF
What habit does the plan build in her, and what does it let her stop doing?
Tap to flip
ANSWER
Separating confirmed facts from what is still being checked, every time. She stops reaching for one reassuring line that claims more control than the team actually has.
4 · THE ANCHOR
What is the one locked decision the whole plan hangs on?
Tap to flip
ANSWER
The first statement is capped at four lines: confirmed scope, action taken, what is not known, and a named next update time. A cause is never one of the four.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Wiring Cambrix's own incident summary tool straight into the public draft, with no independent check, to save about ten minutes. It made sense because the tool had been right on every smaller incident before this one.
6 · THE NUMBER
Fill in the blank: the first statement named at least ___ accounts confirmed so far. By the 4:00pm update that grew to ___.
Tap to flip
ANSWER
4,120. Then 4,410, once two more affected funds were folded in.
7 · THE REPLAY
Same bad Monday, redesigned plan, what changes?
Tap to flip
ANSWER
The first statement goes out 44 minutes after the visibility trigger, naming only scope and action. When two more funds turn up at 2:10pm, they fold into the already promised 4:00pm update instead of contradicting a claim nobody made.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different anchor. Which product, and which anchor?
Tap to flip
ANSWER
Farrowgate Read, an X-ray triage tool at Solden Veterinary Labs. The anchor is a guaranteed re-review commitment instead of a scope number.
Check yourself Score: 0 / 0
Fill in the blank
1. The first public statement is locked to four lines: a confirmed ___, an ___ already taken, what is ___ yet, and a named ___ time.
Show hint
Check the Anchor step in the SPARK recap, and the diagram right after the priority list.
Show answer
Scope; action; not known; next update. A cause never gets a fifth line, no matter how confident the model's own summary sounds.
True or false
2. True or false: the redesigned plan protects Fairmarch Wealth because it makes Denice write the statement faster than before.
True
False
Show hint
Compare the two posting times in the story: 9:04am for the old plan against 10:13am for the new one.
Show answer
False. The new plan is actually slower to post, 44 minutes after the trigger instead of 9. It protects the company by narrowing what the fast version is allowed to claim, not by making it faster.
Multiple choice
3. Why does this answer reject letting Cambrix's own incident summary tool draft the first public statement, even though it is usually accurate?
A. Because the tool is too slow to run during a live incident.
B. Because using it would cost too much in compute during a spike.
C. Because it treats the model's own self report about its own failure as a confirmed fact, before anyone outside the model has checked it.
D. Because legal requires every public statement to be signed by a named human.
Show hint
Look at the Risk step in the SPARK recap, and what happened at 2:10pm in the story.
Show answer
C. A confidence score is not a check. The tool being right nine times before does not make its tenth answer verified.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the meeting memory about the incident summary tool, not a dial someone could just turn up.
Show answer
Model answer: Wiring Cambrix's incident summary tool straight into the public draft, with no independent check, to save Denice about ten minutes of writing. It made sense when it shipped because the tool had been right on every smaller incident before this one, so nobody in that meeting was wrong about what the data showed them that week.
Short answer, apply it yourself
5. Think of a product you use that could have a public incident someday. What is one line you would ban your own first public statement from saying, before anything has even gone wrong?
Show hint
Look for the kind of claim that sounds reassuring today and gets contradicted the moment more facts arrive.
Show answer
Model answer: A food delivery app that uses a model to guess which orders are likely to arrive late. If a bad batch of guesses ever went public, the first statement should never say "this only affected one city," because a city-level claim is exactly the kind of scope number that a model's own regional breakdown might get wrong first.
Short answer, the number question
6. If reconciliation had found the two extra affected funds within the first thirty minutes, before the 10:13am statement even went out, would the plan's "what we do not know yet" line still be needed? Show the reasoning.
Show hint
Think about whether finding two more funds proves there is not a third one still hiding under the drift detector's 2 point cutoff.
Show answer
Yes, still needed. Finding two more affected funds does not confirm there is not a third. The drift detector's own flag line only catches drift over 2 percentage points, so the same kind of misclassification could still be sitting under that line, unconfirmed, anywhere else in the portfolio lineup. Declaring the incident fully scoped would still be an overclaim.
Before you close the answer
Why this works
Tests whether a communication plan, to you, is a claims discipline exercise and not a tone exercise. With a model in the loop, the real risk in hour one is not sounding cold, it's stating the model's own unverified diagnosis as if a person had already checked it.
Follow-up traps
"Isn't withholding the cause the same as being evasive?" Response: no, because in hour one there genuinely is no confirmed cause, only the model's own unchecked guess. Stating a guess as fact is the actual evasion, dressed up as confidence.
"What if leadership wants to reassure customers faster than this plan allows?" Response: the plan already ships in under an hour, at 44 minutes in the replay. What it refuses is one specific overclaim, not speed. The tradeoff was never time, it was which sentence goes out.
If pressed
The independent check that has to clear before any cause goes public is a trade log reconciliation, joining the corporate action feed against actual executed trades, not Cambrix's own confidence score. On this incident that reconciliation did not clear until 2:10pm, more than five hours after the model's own incident summary had already offered a confident, incomplete answer.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.