Artifact critiqueIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #5

Write the three sentences you would use to reset expectations after an overpromised launch date.

LEAD · why the reset message needed a number, not a mood, tested on Bramblehurst's drafting model Longstitch

Bramblehurst turns a podcast episode into a finished blog post: upload the audio, and Longstitch, its drafting model, writes the headline, the subheads, and the pull quotes. Nydia Kirchgessner owns Auto-Publish, the mode that skips the human review screen and posts straight to a creator's blog. She promised it for March 3rd on a public roadmap call. Two days out, Cynwrig Achenbach, who runs the model's evals, told her the real number: 61 percent, not the 92 it needed. What she sent next did more damage than the missed date ever could have.

The direct answer
Reset expectations with three sentences, not a mood: the real, specific reason for the delay, a new date tied to a real number already moving, and what changes if that date slips too. Never send a vague or an extra-cheerful message instead, it does more damage than the missed date itself. A missed date costs a launch. A reset with no real reason and no real evidence costs whether your word can be trusted the next time you give one.
Do this, in order
  1. Send three sentences: the real reason, the new date's evidence, and the fallback if it slips again.Why: this is the whole shape of the answer, just said in order instead of all at once.
  2. Name the specific eval gap or capability gap, never a phrase like "unforeseen challenges."Why: a vague reason reads as spin, and spin gets punished twice as hard as an honest number would.
  3. Tie the new date to a number that is already moving, not a fresh guess.Why: a second guessed date is the exact mistake the first one made, just wearing a later month.
  4. Watch the opt-out rate in the first 48 hours, not the renewal number six weeks out.Why: by the time renewal drops, the real damage already happened weeks earlier and quietly.
  5. Never send the message as a single blast written to end an awkward five minutes faster.Why: that is exactly how a reset turns into corporate wording that buries the real reason.
  6. Skip the full three-sentence ritual for a small, low-stakes slip.Why: a copy edit sliding three days does not need an eval number and a fallback clause, only a real broken promise does.

How to answer this, stage by stage

Nobody is grading whether you can apologize well. They're grading whether you can produce the actual three sentences, out loud, with a real reason and a real number inside them.

1
Scope it to one real product and person
Say it like this
"Let's ground this in one real case. Bramblehurst turns podcast episodes into blog posts. Nydia Kirchgessner owns Auto-Publish, the mode that skips the review screen and posts straight to a creator's blog. I'll answer using her real launch, not the idea of a reset message in general."
Why this works
Naming a real product and person stops "reset expectations" from turning into a vague list of etiquette tips.
2
Say your structure out loud
Say it like this
"I'll run this as LEAD. Link, what a reset message actually has to protect. Early signal, the number that tells you in two days whether it worked, weeks before renewal does. Abuse, how these messages get written badly. Decision, the three sentences themselves."
Why this works
Two seconds of structure tells the interviewer you have a method, not a vibe you're improvising live.
3
Answer the literal question first, in one line
Say it like this
"Short version: three sentences. The real reason, specific enough to name the actual gap. The new date, tied to a real number already moving. And what changes if that date slips too."
Why this works
This is the actual answer, said plainly, before any story. The interviewer shouldn't have to wait for it.
4
Reframe why anyone's actually asking this
Say it like this
"This isn't really 'can you write an apology.' It's 'do you understand that a reset message is itself a claim,' because a vague one gets checked against reality just as hard as the date it's replacing. Maybe harder, since people are already primed to distrust you."
Why this works
Shows the interviewer you understand the trap under the question, not just a script for answering it.
5
Give the decision, committed
Say it like this
"So here's what I'd actually send. Not a mood. Not an apology with no content in it. A reason I can name, a date I can defend with a number, and a real answer to 'what if you're wrong again.'"
Why this works
This is the direct answer, spoken with no hedge in it, right before the actual words follow.
6
Read the actual three sentences, word for word
Say it like this
"Something like this. 'Auto-Publish is not shipping on March 3, because on jargon-heavy episodes, like founder and deep-tech interviews, Longstitch still invents a wrong company name, wrong number, or wrong credit in about 4 of every 10 drafts, and we will not publish that straight to your blog with no one reading it first. We are now targeting April 21, the week our jargon accuracy score is on track to clear 92 percent two runs in a row; it sat at 61 percent this week and has climbed about 4 points every week for the last five weeks. If we are not there by April 21, Auto-Publish ships that week anyway, but only for plain-language shows, and jargon-heavy episodes stay in manual review until the number clears, instead of us picking a fourth date.'"
Why this works
This is what a real reset message sounds like. Saying it out loud, not describing it, is what proves you could actually write one under pressure.
7
Prove it with the real 48 hours, numbers first
Say it like this
"Here's what actually happened. The first reset, the vague one, went to 40 beta customers on March 1. Within 48 hours, 14 of them opted out, thirty five percent. The second reset, built from the eval number, went to the remaining 26 a few days later. Within 48 hours, only 2 opted out, eight percent. Same missed date under both messages. Different message, very different math."
Why this works
Two real percentages, days apart, beat any amount of talk about tone or empathy.
8
Name the abuse, say what you'd leave alone, then close
Say it like this
"Someone on the growth team floated just picking a safer, further-out date and not explaining why, to close the conversation faster. I'd turn that down, a second guessed date is the same mistake wearing a new outfit. I also wouldn't run this full three-sentence process for a minor copy tweak sliding by three days, that just needs 'out Thursday instead of Tuesday.' So: name the real reason, back the new date with a real number, say what changes if you're wrong again, and a reset message finally does what it's supposed to do. Hold."
Why this works
Naming a bad idea you'd turn down, and a place you wouldn't bother, is what proves this is judgment, not a script.

Let's learn

What do you actually say when the date you promised out loud is not going to happen?

Bramblehurst takes a podcast episode and turns it into a blog post. Upload the audio, and Longstitch, its drafting model, reads the transcript and writes a headline, a few subheads, some pull quotes, and a short summary for search. Every draft lands on a review screen first. A creator reads it, fixes anything odd, and hits publish. That screen has caught every invented fact Longstitch has ever produced, for four straight years.

Knowledge spark: what is a capability gap? A specific thing a model cannot yet do reliably, not a general "it's not perfect" complaint. Longstitch can summarize a plain conversation well. It cannot yet tell, on a jargon-heavy show, a company name or number it actually heard from one that just sounds plausible.

Auto-Publish removes that screen. The draft posts straight to the creator's blog the moment an episode finishes processing, no one reads it first. For plain-talk shows, story podcasts, advice shows, Longstitch is safe enough to skip the screen; it only has to describe what was said. Jargon-heavy shows are a different problem: founder interviews, funding round recaps, deep-tech explainers, where guests talk fast and drop company names, dollar figures, and acronyms with no time to spell anything out. When Longstitch does not quite catch one, it does not leave a blank. It fills the gap with something that sounds right.

Hand sketched flow diagram titled Where Longstitch invents things. Four connected boxes in a row: Fast jargon episode, Term unclear in audio, Model fills the gap, highlighted in orange as the key step, and Wrong detail goes out.
The model does not leave a blank. It fills the gap with something that sounds right, and the sentence reads the same either way.

Bramblehurst built a jargon eval set to catch exactly this: real jargon-heavy episodes, checked by hand for any invented company name, wrong number, or the wrong person given credit for something. The bar to ship Auto-Publish safely on those shows was 92 percent of drafts coming back clean, on two separate eval runs in a row, so one lucky run could not fake it. Nydia Kirchgessner, who owns Auto-Publish, promised customers a launch date in January, on a public roadmap call: March 3rd. Growth wanted a date on the slide. A date is what she gave them.

Two days before launch, Cynwrig Achenbach, who runs the eval, told her the real number. 61 percent. Not 92. Longstitch was still inventing a wrong company name, figure, or credit in about 4 of every 10 jargon-heavy drafts, the exact thing the review screen existed to catch.

Knowledge spark: what is confident wrongness? When a model states a wrong answer with the same tone as a right one. Longstitch does not write "I'm not sure about this number." It writes the wrong number in the same clean sentence it writes a real one in, so a reader has no way to tell them apart just by looking.

Here is the turn. That 61 percent was not really the emergency. Every team building on a model misses an eval bar sometimes, it is a normal, fixable problem with a normal fix: more training data, more time. The real emergency, Nydia built herself, in the ninety seconds it took to write two vague sentences instead of three honest ones.

Two days out, with no time and a public date already loose in the world, she sent all 40 Auto-Publish beta customers one line: "We're continuing to fine-tune Auto-Publish to make sure it meets our quality bar, and we'll share an update soon." No real reason. No real date. It read exactly like what it was, a sentence written to end an awkward conversation fast.

Hand sketched comparison diagram titled How a reset message gets gamed. Left panel, a document icon labeled The message sent, caption soft wording, no real reason given. Right panel, a question mark box icon labeled The real reason, caption 61 percent not 92, never said out loud.
The message got sent. The real reason never made it into a single sentence of it.
We did not lose the beta list because Longstitch missed its bar. We lost it in the ninety seconds it took to write two sentences that said nothing.

Within 48 hours, 14 of the 40 opted out of the beta. Several had already told their own readers, or their own bosses, that Auto-Publish was landing March 3rd, and now they had nothing to point to except a company that had gone quiet. One posted in a podcasting community: "Bramblehurst just ghosted their own launch date. Second time this quarter." It was the first time. Support tickets marked frustrated tripled that same week.

Hand sketched metaphor scene titled Two reset messages, unequal weight. Left, a plain square labeled The vague email, caption two sentences, no reason, no date. Right, a small scale icon labeled What it cost, caption 14 of 40 customers gone in 2 days.
The vague email looked like the lighter thing to send. It was not the lighter thing that happened next.

What that cost at its worst: if the second message had made the same mistake, the whole beta program would likely have folded before Auto-Publish shipped to anyone at all. Not because the model was behind, models fall behind schedule constantly, but because nobody would have been left who still believed a date coming out of Bramblehurst.

Nydia and Cynwrig sat down with the eval history that already existed: five weeks of real, climbing scores, sitting on a chart nobody had pulled into a customer message. That data was not new. It had been climbing the whole time, uninvolved in a single promise anyone had made.

Hand sketched timeline titled The eval score, climbing the whole time. Six marks along a line: Week 1, 41 percent. Week 2, 45 percent. Week 3, 49 percent. Week 4, 53 percent. Week 5, 57 percent. Week 6, highlighted in teal, 61 percent, vague reset sent here.
The evidence for the new date was not new. It had been climbing on a chart nobody had opened in a week.

They wrote three sentences instead, sent to the remaining 26 beta customers a few days later:

The reset that actually shipped "Auto-Publish is not shipping on March 3, because on jargon-heavy episodes, like founder and deep-tech interviews, Longstitch still invents a wrong company name, wrong number, or wrong credit in about 4 of every 10 drafts, and we will not publish that straight to your blog with no one reading it first. We are now targeting April 21, the week our jargon accuracy score is on track to clear 92 percent two runs in a row; it sat at 61 percent this week and has climbed about 4 points every week for the last five weeks. If we are not there by April 21, Auto-Publish ships that week anyway, but only for plain-language shows, and jargon-heavy episodes stay in manual review until the number clears, instead of us picking a fourth date."
Hand sketched labeled parts diagram titled The anatomy of a reset message that holds. A central document icon labeled Reset message, with three callouts: The real reason, The new date, evidence-based, and What changes if it slips again.
Three parts, every time. Leave one out and it goes back to being a mood with a date attached.

Within 48 hours of that message, only 2 of the 26 opted out, 8 percent instead of 35. At the 6-week renewal point, 22 of the remaining 24 converted to paid, 92 percent, against Bramblehurst's normal beta-to-paid rate of 75 percent company-wide.

Cumulative beta opt-out rate, by day after each reset message
50% 25% 0% 35% at 48 hrs 41% 8% at 48 hrs 9% Day 0 Day 2 Day 10
Vague reset (March 1)Evidence-based reset (early March)
Both curves had already told the whole story by day 2, five and a half weeks before anyone's renewal date arrived.
Beta-to-paid renewal rate at 6 weeks: company average vs. this cohort
100% 50% 0% 75% Company average 92% Auto-Publish cohort
Company averageThis cohort
The cohort that got the honest message did not just recover. It renewed at a higher rate than launches that never slipped at all.
Hand sketched comparison diagram titled Which clock rings first. Left panel, a gauge icon labeled Opt-out rate, caption known within 48 hours. Right panel, a gauge icon labeled Renewal rate, caption known 6 weeks later.
One of these clocks rings in two days. The other one waits six weeks to say what the first one already knew.
The choice I would take back Back in January, in the meeting where the roadmap got set, the team gave customers a fixed calendar date, March 3rd, before the eval bar had ever been cleared once. That made sense in the room. A date is what the VP of Growth wanted to say on the call, and a bar with no date attached felt too soft to announce. It stopped making sense the moment the model's own number said the date was not real yet, and nobody had a plan for what to say if that happened, because the promise had never been conditional on anything.

What I would leave alone: not every slipped date needs three sentences and an eval number behind it. A small copy change sliding by three days does not need a fallback clause, it needs "out Thursday instead of Tuesday." The full version is for a promise that was public, specific, and tied to whether Bramblehurst's word means anything the next time it makes one.

The lesson: a reset message is not neutral. It is its own claim, made at the exact moment customers are already primed to check your word against reality. Write it vague to dodge five awkward minutes, and you do not skip the hard conversation. You just move it two days later with a credibility problem stapled to it.

Now here is the same thing as a story

The short version above is what you'd actually say in an interview. Read this one for the two days in March that decided whether Bramblehurst's beta list survived.

Nydia Kirchgessner had run four product launches before Auto-Publish, and she had a rule for all of them: never say a date out loud until the number behind it is real. Ask her about any of those launches, and she could tell you the exact eval score behind whatever she'd just promised, because she'd pulled it herself that morning.

The roadmap call in January went well. She stood up in front of 40 beta customers and said Auto-Publish, no review screen, straight to your blog, would ship March 3rd. The room lit up. For three weeks after that, checking in on Longstitch's progress was the best part of her Tuesdays. Cynwrig Achenbach's updates always had a climbing number in them, and a climbing number toward a stated date reads exactly like progress.

Week one, she pulled the raw eval score herself before every check-in. Week two, she asked Cynwrig for it instead of pulling it herself. Week three, she stopped asking for the number at all, and just asked, "still on track for March 3rd?" Mostly, he kept saying, and mostly sounded close enough.

Then it was a Tuesday, two days before launch, at 4pm.

Cynwrig messaged her, nothing dramatic, no meeting called. "We're at 61. Not going to make 92 by Friday." Six words, sent between two other Slack threads.

Nydia had ninety minutes before she had to say something to 40 people who'd already told their own readers, and in some cases their own bosses, that March 3rd was real. She did the fast thing. She wrote two sentences that sounded careful and said nothing: "We're continuing to fine-tune Auto-Publish to make sure it meets our quality bar, and we'll share an update soon." She sent it at 5:42pm and closed her laptop.

By Thursday morning, 14 of the 40 had opted out of the beta. One had posted about it in a podcasting community she wasn't even a member of, until someone forwarded it to her: "Bramblehurst just ghosted their own launch date. Second time this quarter." It was the first time.

The missed date cost Longstitch two extra months. The vague email cost Nydia fourteen customers in two days, and only one of those numbers was ever really about the model.

I want to say the problem was that the model wasn't ready. It wasn't. But that's not really the story. A missed eval bar is a normal Tuesday for any team building on a model, fixable with time and more training runs. What wasn't fixable in the same way was 40 people who'd been told nothing, twice, by a company that had asked for their trust back in January.

The decision Nydia would take back sits further behind that Tuesday. It sits in the January meeting where the roadmap got set, marker in the VP of Growth's hand, a room that wanted a date it could put on a slide. Nydia gave them one. Nobody in that room, herself included, asked what they'd say if the date and the number ever disagreed, because in January, agreeing felt like the only outcome anyone was planning for.

Friday morning, she and Cynwrig sat down with the eval history on a shared screen instead of a Slack message. Five weeks of real numbers were already there, climbing on their own the whole time, about 4 points a week. Nobody had ever put a customer-facing date next to that line before.

They wrote three sentences this time, and Nydia read them out loud twice before sending, to hear whether they'd sound honest coming from someone else. The real reason: jargon-heavy episodes, invented details, about 4 in 10 drafts, no auto-publish while that number stood. The new date: April 21st, the week the climb was on track to reach 92 twice in a row, if the last five weeks kept holding. What changes if it slips again: Auto-Publish ships anyway that week, for the shows it's actually ready for, and the hard ones wait in manual review instead of getting a fourth date.

Within 48 hours, 2 of the remaining 26 opted out. Eight percent, not thirty five. Six weeks later, 22 of 24 renewed, ninety two percent, three points above Bramblehurst's normal rate for launches that go out exactly on time.

Same missed date, both times. One message cost fourteen people in two days. The other one, sent about the exact same bad news, built a cohort that stuck around better than a launch that never slipped at all.

What I'd tell myself, back in that January meeting: the date was never really the promise. The bar was. I promised the wrong one first, and it took a Tuesday at 5:42pm to find out.

LEAD, so the reset message survives being checked against reality

Not a way to prove Nydia was careless. LEAD is what forces you to say what a reset message should actually track, and to catch a bad one before it costs more than the missed date did.

LLink. The real outcome a reset message actually has to protect.
Not whether the inbox goes quiet for a day. What actually matters is whether burned beta customers still convert to paying customers once Auto-Publish is finally real. A vague message can buy a quiet Thursday and still lose the cohort by the time renewal comes due.
Bramblehurst's real number was the 6-week renewal rate. The vague reset put the whole beta on a path that would have wrecked it long before the model was even the problem anymore.
EEarly signal. What moves in days, weeks before renewal does.
The share of beta customers who opt out within 48 hours of a reset message. Renewal takes six weeks to show up. The opt-out number is back before the weekend.
35 percent opted out in 48 hours after the vague reset. 8 percent did after the evidence-based one. The 6-week renewal gap, 75 versus 92, just confirmed what the 48-hour number had already said.
Hand sketched comparison diagram titled Which clock rings first. Left panel, a gauge icon labeled Opt-out rate, caption known within 48 hours. Right panel, a gauge icon labeled Renewal rate, caption known 6 weeks later.
The same picture again, on purpose. It's the whole reason to watch the 48-hour number instead of waiting on renewal.
AAbuse. How a reset message gets written badly.
Two ways, and Nydia's team found both. Bury the real reason in soft, corporate wording, "continuing to fine-tune to make sure it meets our quality bar," so nobody can point to what's actually wrong. Or swing hard the other way and hand out a brand-new date with nothing behind it, just to end the awkward conversation faster.
The vague version cost 14 customers in 2 days. A second guessed date, if April 21 had been picked without the eval trend behind it, would have cost most of the rest the next time it slipped too.
Hand sketched comparison diagram titled How a reset message gets gamed. Left panel, a document icon labeled The message sent, caption soft wording, no real reason given. Right panel, a question mark box icon labeled The real reason, caption 61 percent not 92, never said out loud.
The same picture again, on purpose. It's the exact shape of the abuse this letter exists to catch.
DDecision. What actually goes in the message.
Three sentences, every time. The real reason, specific enough to name the actual gap. The new date, tied to a number already moving, not a fresh guess. And what happens if that date slips too, so the message can be checked instead of just believed.
Read the exact words in the walkthrough and the story above. The shape is reason, evidence, fallback, every time, regardless of what the delay actually is.

The recap, one line per letter: link the message to whether burned customers actually renew, not whether the inbox goes quiet today. The early signal is the 48-hour opt-out number, because it's back before the weekend and renewal takes six weeks. Name both abuses plainly, corporate wording that hides the real reason, and a second guessed date standing in for evidence. And the decision is what makes it real: a reason, a number, and a fallback, all three, every time.

Two things worth saying outright, since the real judgment sits here. Bramblehurst's growth team floated saying nothing at all, quietly pushing the beta window back and hoping nobody noticed until Auto-Publish actually shipped. That would have made things worse, not better, since more than a dozen customers had already told other people the March 3rd date in public; silence reads as hoping to get away with it, which is a worse story than an honest miss. The AI-specific failure worth naming by name is confident wrongness, sometimes called hallucination: Longstitch does not leave a blank when it is unsure, it fills the gap with a wrong company name or number stated in the exact same tone as a correct one, so a reader can't spot the fake one just by reading the sentence. The guardrail that catches it is the jargon eval bar itself, 92 percent, two runs in a row, checked by hand against real jargon-heavy episodes, with those shows routed to manual review until the number clears, no matter how good Longstitch looks on everything else. And the tradeoff was real and taken on purpose: holding jargon-heavy shows back from Auto-Publish costs Bramblehurst part of its fully hands-off pitch to its highest-value customers, the deep-tech and venture podcasters who wanted the review screen gone the most. That's a real cost. It's smaller than putting a made-up statistic on a real company's blog under that company's own name, which is what skipping the bar would have risked.

And if you want to be sure it really works, try it somewhere else

Same four letters, a field-service company instead of a content tool, and this time the jargon that trips the model is refrigerant leaks, not company names.

Plenum builds Same-Visit Report, an AI tool that drafts a client-facing diagnostic summary the moment an HVAC technician finishes a job, from the technician's shorthand notes and photos, and sends it to the client without the technician reading it first. Radomira Szabados owns its rollout, and hit a close cousin of Nydia's exact problem two months after a similar public launch date.

Same-Visit Report's kickoff promise was June 10th. On single-system faults, a broken capacitor, a dirty filter, the model was safe to auto-send from week one. On rare, multi-system faults, a refrigerant leak paired with a control-board fault code, it sometimes wrote a specific part number or diagnosis it hadn't actually confirmed, in about a third of those visits. The eval bar for auto-sending those reports was 90 percent clean, two runs in a row. Ten days before June 10th, the real number was 58.

Hand sketched timeline titled Same trap, a different job site. Five marks along a line: June 10, launch promised. Vague note sent, 12 of 30 crews switch it off. Eval check, highlighted in orange, 58 percent not 90. Evidence reset sent, real reason, real date, real fallback. August 4, target, tied to the climb.
Same shape of mistake, a different kind of jargon. This time it's refrigerant leaks and fault codes instead of company names.

Radomira's first instinct matched Nydia's exactly: a short, vague note to the 30 field crews signed up for the beta, "still polishing accuracy, more soon." Within a week, 12 of the 30 crews had quietly turned Same-Visit Report off on their own tablets and gone back to writing reports by hand, the same early-signal shape as Bramblehurst's opt-outs, just measured in a switch flipped instead of an account closed.

The decision Radomira took back matched Nydia's too: Plenum had promised a calendar date before the eval bar had ever been cleared once, because "give us a date" felt more concrete in the launch meeting than a bar with no date attached. Once she rewrote the message with the same three-sentence shape, a real reason (multi-system faults, wrong part numbers, about a third of the time), a real new date (August 4th, tied to the eval climbing about 3 points a week from 58), and a real fallback (auto-send ships for single-system faults on June 10th as planned, multi-system faults stay in manual review until the number clears), only 3 of the remaining 18 crews turned the feature off in the following week.

Mapped onto LEAD, the shape holds. The link is whether crews trust the tool enough to keep using it past the first bad week, not whether the group chat goes quiet. The early signal is the share of crews who quietly disable it within a week, ready fast, while the real outcome, client callback complaints about a wrong diagnosis, takes 30 days to show up in the data. The abuse Radomira caught herself making was the same one, soft wording standing in for a real reason. And her decision matched: a reason, a number, and a fallback, sized to what the job actually was.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the real reason, give a new date tied to a real number, say what changes if it slips again.
Cost: there's no budget this quarter to keep running the eval weekly. Whoever owns the model pulls it by hand every other week instead, a slower real number still beats a fast guessed one.
The model got better, for real: say Longstitch's next version halves its invented-detail rate overnight. The eval bar and the two-runs-in-a-row rule stay in place anyway, because one great run is still one run, and a single lucky score is not the same thing as a bar that means something.

Where people run it wrong.
They write the reason in language soft enough that nobody can actually name what's broken.
They swap a second guessed date for a real number, because it feels like faster relief than waiting for the evidence.
They skip the fallback sentence entirely, so the new date ends up just as unconditional as the first one was.

How to use it live. When an interviewer hands you a question like this, buy yourself a second by asking out loud what the actual number behind the delay is. That question is the whole answer in miniature, and it tells the interviewer you're about to build the message from evidence, not from tone.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that asks for the literal words to reset a broken promise?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Built for metric questions, it forces a reset message to name a real reason and a real number, not just a tone.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Nydia Kirchgessner, who owns Auto-Publish at Bramblehurst, and Cynwrig Achenbach, who runs the eval that told her the real number two days before launch.
3 · THE LINK
What should a reset message actually protect?
Tap to flip
ANSWER
Whether burned customers still renew once the real launch happens, not whether the inbox goes quiet for a day.
4 · THE EARLY SIGNAL
What moved first here, and what took six weeks to catch up?
Tap to flip
ANSWER
The 48-hour opt-out rate: 35 percent after the vague reset, 8 percent after the evidence-based one. Renewal, 75 versus 92 percent, only confirmed it six weeks later.
5 · THE OLD DECISION
What decision would Nydia take back?
Tap to flip
ANSWER
Promising a fixed calendar date, March 3rd, in a public roadmap call, before the eval bar behind it had ever been cleared once.
6 · THE NUMBER
Fill in the blank: within 48 hours of the vague reset, ___ of 40 beta customers opted out. After the evidence-based reset, only ___ of the remaining 26 did.
Tap to flip
ANSWER
14 (35 percent); 2 (8 percent). Same missed date, two very different messages.
7 · THE REPLAY
Same bad Tuesday, three honest sentences instead, what changes?
Tap to flip
ANSWER
Opt-outs drop from 35 percent to 8 percent in 48 hours. Six weeks later, 22 of 24 renew, 92 percent, three points above Bramblehurst's normal on-time rate.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the parallel problem there?
Tap to flip
ANSWER
Same-Visit Report, an HVAC diagnosis tool run by Radomira Szabados at Plenum. There, a vague note about "polishing accuracy" led 12 of 30 crews to quietly switch the feature off, until an evidence-based reset with the same three-sentence shape cut that to 3 of 18.

Check yourself Score: 0 / 0

Fill in the blank
1. Within 48 hours of the vague reset email, ___ of Bramblehurst's 40 beta customers opted out. After the evidence-based reset, only ___ of the remaining 26 did.
Show hint
Look at the numbers right after each message in Let's learn.
Show answer
14 (35 percent); 2 (8 percent). Same missed date under both messages, very different outcomes.
Multiple choice
2. Why did Nydia's first reset message do more damage than the missed March 3rd date itself?
  • A. It named the wrong engineer as responsible for the delay.
  • B. It gave no real reason and no real date, so customers who had already told others about March 3rd were left with nothing to point to.
  • C. It was sent too late at night for most people to see it.
  • D. It offered a refund Bramblehurst could not actually afford.
Show hint
Check the Abuse step in the framework recap.
Show answer
B. The message spent trust with nothing to show for it. A real reason and a real number would have given customers something to check instead of something to distrust.
True or false
3. True or false: the fix here was mainly to sound warmer and more apologetic in the second message.
  • True
  • False
Show hint
Check the Abuse step, it names "over-apologetic" as its own way to get this wrong.
Show answer
False. The fix was adding a specific reason and a real number. Sounding warmer without adding either is one of the exact traps named in the abuse step. Tone was never the missing part.
Short answer, name the reversal
4. What old decision would Nydia take back, and why did it make sense the first time she made it?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Promising a fixed calendar date, March 3rd, in the January meeting where the roadmap got set, instead of a bar-conditioned promise. It made sense because a hard date is what the room wanted to hear, and nobody had planned for what to say if the date and the real number ever disagreed.
Short answer, apply it yourself
5. Think of a product you use that has slipped a promised date or feature. What would the three honest sentences to reset your trust in it actually need to say?
Show hint
Name a real reason, a number-backed date, and a fallback if that date slips too.
Show answer
Model answer: A budgeting app once promised bank-sync fixes "this month" for three months running. The honest version would name the real reason (one bank's own data format changed and broke matching for about 1 in 5 accounts), a new date tied to a real number (once matching accuracy on that bank clears 95 percent in testing), and a fallback (manual entry stays free and supported if that date slips again too).
Short answer, work the number
6. If Longstitch's jargon accuracy had climbed only 1 point a week instead of 4, would April 21 still have been a defensible date to put in sentence two? Why or why not?
Show hint
Work out how many weeks it would take to go from 61 to 92 at 1 point a week, then compare that to how far away April 21 actually was.
Show answer
No. At 1 point a week it would take about 31 weeks to clear 92 from 61, nowhere near April 21, which was roughly 7 weeks out. Sentence two would have been another guess wearing a number, exactly what it was written to avoid.
Before you close the answer
Why this works
Tests whether you can compress trust recovery into three checkable sentences instead of a mood, and whether you understand that a model failing on jargon-heavy audio is a specific, nameable capability gap you can cite, not a vague excuse to wave at.
Follow-up traps
"Isn't naming a 4-in-10 error rate publicly just handing customers a stick to beat you with?" Response: No, the delay already told them something's wrong. Naming the real number is what makes the new date checkable, and a vague reason gets assumed to be worse than the real one anyway.

"What if the eval stalls and April 21 slips too, doesn't the fallback just become the next broken promise?" Response: No, because the fallback isn't another date, it's a structural change, ship the safe subset, keep the risky one in review, that doesn't depend on the model improving any further, so it can't itself slip.
If pressed
The two-runs-in-a-row rule in sentence two exists because a single eval run on a sample this size carries enough randomness that Longstitch could clear 92 percent once by luck while its real rate sits closer to 85. Requiring it twice in a row is what turns the number into evidence instead of a good roll.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more