Write the three sentences you would use to reset expectations after an overpromised launch date.
Bramblehurst turns a podcast episode into a finished blog post: upload the audio, and Longstitch, its drafting model, writes the headline, the subheads, and the pull quotes. Nydia Kirchgessner owns Auto-Publish, the mode that skips the human review screen and posts straight to a creator's blog. She promised it for March 3rd on a public roadmap call. Two days out, Cynwrig Achenbach, who runs the model's evals, told her the real number: 61 percent, not the 92 it needed. What she sent next did more damage than the missed date ever could have.
- Send three sentences: the real reason, the new date's evidence, and the fallback if it slips again.Why: this is the whole shape of the answer, just said in order instead of all at once.
- Name the specific eval gap or capability gap, never a phrase like "unforeseen challenges."Why: a vague reason reads as spin, and spin gets punished twice as hard as an honest number would.
- Tie the new date to a number that is already moving, not a fresh guess.Why: a second guessed date is the exact mistake the first one made, just wearing a later month.
- Watch the opt-out rate in the first 48 hours, not the renewal number six weeks out.Why: by the time renewal drops, the real damage already happened weeks earlier and quietly.
- Never send the message as a single blast written to end an awkward five minutes faster.Why: that is exactly how a reset turns into corporate wording that buries the real reason.
- Skip the full three-sentence ritual for a small, low-stakes slip.Why: a copy edit sliding three days does not need an eval number and a fallback clause, only a real broken promise does.
How to answer this, stage by stage
Nobody is grading whether you can apologize well. They're grading whether you can produce the actual three sentences, out loud, with a real reason and a real number inside them.
Let's learn
What do you actually say when the date you promised out loud is not going to happen?
Bramblehurst takes a podcast episode and turns it into a blog post. Upload the audio, and Longstitch, its drafting model, reads the transcript and writes a headline, a few subheads, some pull quotes, and a short summary for search. Every draft lands on a review screen first. A creator reads it, fixes anything odd, and hits publish. That screen has caught every invented fact Longstitch has ever produced, for four straight years.
Auto-Publish removes that screen. The draft posts straight to the creator's blog the moment an episode finishes processing, no one reads it first. For plain-talk shows, story podcasts, advice shows, Longstitch is safe enough to skip the screen; it only has to describe what was said. Jargon-heavy shows are a different problem: founder interviews, funding round recaps, deep-tech explainers, where guests talk fast and drop company names, dollar figures, and acronyms with no time to spell anything out. When Longstitch does not quite catch one, it does not leave a blank. It fills the gap with something that sounds right.
Bramblehurst built a jargon eval set to catch exactly this: real jargon-heavy episodes, checked by hand for any invented company name, wrong number, or the wrong person given credit for something. The bar to ship Auto-Publish safely on those shows was 92 percent of drafts coming back clean, on two separate eval runs in a row, so one lucky run could not fake it. Nydia Kirchgessner, who owns Auto-Publish, promised customers a launch date in January, on a public roadmap call: March 3rd. Growth wanted a date on the slide. A date is what she gave them.
Two days before launch, Cynwrig Achenbach, who runs the eval, told her the real number. 61 percent. Not 92. Longstitch was still inventing a wrong company name, figure, or credit in about 4 of every 10 jargon-heavy drafts, the exact thing the review screen existed to catch.
Here is the turn. That 61 percent was not really the emergency. Every team building on a model misses an eval bar sometimes, it is a normal, fixable problem with a normal fix: more training data, more time. The real emergency, Nydia built herself, in the ninety seconds it took to write two vague sentences instead of three honest ones.
Two days out, with no time and a public date already loose in the world, she sent all 40 Auto-Publish beta customers one line: "We're continuing to fine-tune Auto-Publish to make sure it meets our quality bar, and we'll share an update soon." No real reason. No real date. It read exactly like what it was, a sentence written to end an awkward conversation fast.
Within 48 hours, 14 of the 40 opted out of the beta. Several had already told their own readers, or their own bosses, that Auto-Publish was landing March 3rd, and now they had nothing to point to except a company that had gone quiet. One posted in a podcasting community: "Bramblehurst just ghosted their own launch date. Second time this quarter." It was the first time. Support tickets marked frustrated tripled that same week.
What that cost at its worst: if the second message had made the same mistake, the whole beta program would likely have folded before Auto-Publish shipped to anyone at all. Not because the model was behind, models fall behind schedule constantly, but because nobody would have been left who still believed a date coming out of Bramblehurst.
Nydia and Cynwrig sat down with the eval history that already existed: five weeks of real, climbing scores, sitting on a chart nobody had pulled into a customer message. That data was not new. It had been climbing the whole time, uninvolved in a single promise anyone had made.
They wrote three sentences instead, sent to the remaining 26 beta customers a few days later:
Within 48 hours of that message, only 2 of the 26 opted out, 8 percent instead of 35. At the 6-week renewal point, 22 of the remaining 24 converted to paid, 92 percent, against Bramblehurst's normal beta-to-paid rate of 75 percent company-wide.
What I would leave alone: not every slipped date needs three sentences and an eval number behind it. A small copy change sliding by three days does not need a fallback clause, it needs "out Thursday instead of Tuesday." The full version is for a promise that was public, specific, and tied to whether Bramblehurst's word means anything the next time it makes one.
The lesson: a reset message is not neutral. It is its own claim, made at the exact moment customers are already primed to check your word against reality. Write it vague to dodge five awkward minutes, and you do not skip the hard conversation. You just move it two days later with a credibility problem stapled to it.
Now here is the same thing as a story
The short version above is what you'd actually say in an interview. Read this one for the two days in March that decided whether Bramblehurst's beta list survived.
Nydia Kirchgessner had run four product launches before Auto-Publish, and she had a rule for all of them: never say a date out loud until the number behind it is real. Ask her about any of those launches, and she could tell you the exact eval score behind whatever she'd just promised, because she'd pulled it herself that morning.
The roadmap call in January went well. She stood up in front of 40 beta customers and said Auto-Publish, no review screen, straight to your blog, would ship March 3rd. The room lit up. For three weeks after that, checking in on Longstitch's progress was the best part of her Tuesdays. Cynwrig Achenbach's updates always had a climbing number in them, and a climbing number toward a stated date reads exactly like progress.
Week one, she pulled the raw eval score herself before every check-in. Week two, she asked Cynwrig for it instead of pulling it herself. Week three, she stopped asking for the number at all, and just asked, "still on track for March 3rd?" Mostly, he kept saying, and mostly sounded close enough.
Then it was a Tuesday, two days before launch, at 4pm.
Cynwrig messaged her, nothing dramatic, no meeting called. "We're at 61. Not going to make 92 by Friday." Six words, sent between two other Slack threads.
Nydia had ninety minutes before she had to say something to 40 people who'd already told their own readers, and in some cases their own bosses, that March 3rd was real. She did the fast thing. She wrote two sentences that sounded careful and said nothing: "We're continuing to fine-tune Auto-Publish to make sure it meets our quality bar, and we'll share an update soon." She sent it at 5:42pm and closed her laptop.
By Thursday morning, 14 of the 40 had opted out of the beta. One had posted about it in a podcasting community she wasn't even a member of, until someone forwarded it to her: "Bramblehurst just ghosted their own launch date. Second time this quarter." It was the first time.
I want to say the problem was that the model wasn't ready. It wasn't. But that's not really the story. A missed eval bar is a normal Tuesday for any team building on a model, fixable with time and more training runs. What wasn't fixable in the same way was 40 people who'd been told nothing, twice, by a company that had asked for their trust back in January.
The decision Nydia would take back sits further behind that Tuesday. It sits in the January meeting where the roadmap got set, marker in the VP of Growth's hand, a room that wanted a date it could put on a slide. Nydia gave them one. Nobody in that room, herself included, asked what they'd say if the date and the number ever disagreed, because in January, agreeing felt like the only outcome anyone was planning for.
Friday morning, she and Cynwrig sat down with the eval history on a shared screen instead of a Slack message. Five weeks of real numbers were already there, climbing on their own the whole time, about 4 points a week. Nobody had ever put a customer-facing date next to that line before.
They wrote three sentences this time, and Nydia read them out loud twice before sending, to hear whether they'd sound honest coming from someone else. The real reason: jargon-heavy episodes, invented details, about 4 in 10 drafts, no auto-publish while that number stood. The new date: April 21st, the week the climb was on track to reach 92 twice in a row, if the last five weeks kept holding. What changes if it slips again: Auto-Publish ships anyway that week, for the shows it's actually ready for, and the hard ones wait in manual review instead of getting a fourth date.
Within 48 hours, 2 of the remaining 26 opted out. Eight percent, not thirty five. Six weeks later, 22 of 24 renewed, ninety two percent, three points above Bramblehurst's normal rate for launches that go out exactly on time.
Same missed date, both times. One message cost fourteen people in two days. The other one, sent about the exact same bad news, built a cohort that stuck around better than a launch that never slipped at all.
What I'd tell myself, back in that January meeting: the date was never really the promise. The bar was. I promised the wrong one first, and it took a Tuesday at 5:42pm to find out.
LEAD, so the reset message survives being checked against reality
Not a way to prove Nydia was careless. LEAD is what forces you to say what a reset message should actually track, and to catch a bad one before it costs more than the missed date did.
The recap, one line per letter: link the message to whether burned customers actually renew, not whether the inbox goes quiet today. The early signal is the 48-hour opt-out number, because it's back before the weekend and renewal takes six weeks. Name both abuses plainly, corporate wording that hides the real reason, and a second guessed date standing in for evidence. And the decision is what makes it real: a reason, a number, and a fallback, all three, every time.
Two things worth saying outright, since the real judgment sits here. Bramblehurst's growth team floated saying nothing at all, quietly pushing the beta window back and hoping nobody noticed until Auto-Publish actually shipped. That would have made things worse, not better, since more than a dozen customers had already told other people the March 3rd date in public; silence reads as hoping to get away with it, which is a worse story than an honest miss. The AI-specific failure worth naming by name is confident wrongness, sometimes called hallucination: Longstitch does not leave a blank when it is unsure, it fills the gap with a wrong company name or number stated in the exact same tone as a correct one, so a reader can't spot the fake one just by reading the sentence. The guardrail that catches it is the jargon eval bar itself, 92 percent, two runs in a row, checked by hand against real jargon-heavy episodes, with those shows routed to manual review until the number clears, no matter how good Longstitch looks on everything else. And the tradeoff was real and taken on purpose: holding jargon-heavy shows back from Auto-Publish costs Bramblehurst part of its fully hands-off pitch to its highest-value customers, the deep-tech and venture podcasters who wanted the review screen gone the most. That's a real cost. It's smaller than putting a made-up statistic on a real company's blog under that company's own name, which is what skipping the bar would have risked.
And if you want to be sure it really works, try it somewhere else
Same four letters, a field-service company instead of a content tool, and this time the jargon that trips the model is refrigerant leaks, not company names.
Plenum builds Same-Visit Report, an AI tool that drafts a client-facing diagnostic summary the moment an HVAC technician finishes a job, from the technician's shorthand notes and photos, and sends it to the client without the technician reading it first. Radomira Szabados owns its rollout, and hit a close cousin of Nydia's exact problem two months after a similar public launch date.
Same-Visit Report's kickoff promise was June 10th. On single-system faults, a broken capacitor, a dirty filter, the model was safe to auto-send from week one. On rare, multi-system faults, a refrigerant leak paired with a control-board fault code, it sometimes wrote a specific part number or diagnosis it hadn't actually confirmed, in about a third of those visits. The eval bar for auto-sending those reports was 90 percent clean, two runs in a row. Ten days before June 10th, the real number was 58.
Radomira's first instinct matched Nydia's exactly: a short, vague note to the 30 field crews signed up for the beta, "still polishing accuracy, more soon." Within a week, 12 of the 30 crews had quietly turned Same-Visit Report off on their own tablets and gone back to writing reports by hand, the same early-signal shape as Bramblehurst's opt-outs, just measured in a switch flipped instead of an account closed.
The decision Radomira took back matched Nydia's too: Plenum had promised a calendar date before the eval bar had ever been cleared once, because "give us a date" felt more concrete in the launch meeting than a bar with no date attached. Once she rewrote the message with the same three-sentence shape, a real reason (multi-system faults, wrong part numbers, about a third of the time), a real new date (August 4th, tied to the eval climbing about 3 points a week from 58), and a real fallback (auto-send ships for single-system faults on June 10th as planned, multi-system faults stay in manual review until the number clears), only 3 of the remaining 18 crews turned the feature off in the following week.
Mapped onto LEAD, the shape holds. The link is whether crews trust the tool enough to keep using it past the first bad week, not whether the group chat goes quiet. The early signal is the share of crews who quietly disable it within a week, ready fast, while the real outcome, client callback complaints about a wrong diagnosis, takes 30 days to show up in the data. The abuse Radomira caught herself making was the same one, soft wording standing in for a real reason. And her decision matched: a reason, a number, and a fallback, sized to what the job actually was.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the real reason, give a new date tied to a real number, say what changes if it slips again.
Cost: there's no budget this quarter to keep running the eval weekly. Whoever owns the model pulls it by hand every other week instead, a slower real number still beats a fast guessed one.
The model got better, for real: say Longstitch's next version halves its invented-detail rate overnight. The eval bar and the two-runs-in-a-row rule stay in place anyway, because one great run is still one run, and a single lucky score is not the same thing as a bar that means something.
Where people run it wrong.
They write the reason in language soft enough that nobody can actually name what's broken.
They swap a second guessed date for a real number, because it feels like faster relief than waiting for the evidence.
They skip the fallback sentence entirely, so the new date ends up just as unconditional as the first one was.
How to use it live. When an interviewer hands you a question like this, buy yourself a second by asking out loud what the actual number behind the delay is. That question is the whole answer in miniature, and it tells the interviewer you're about to build the message from evidence, not from tone.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the eval stalls and April 21 slips too, doesn't the fallback just become the next broken promise?" Response: No, because the fallback isn't another date, it's a structural change, ship the safe subset, keep the risky one in review, that doesn't depend on the model improving any further, so it can't itself slip.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Managing stakeholder expectations and AI hype
- #1 Your CEO saw a demo on social media and wants that feature in six weeks. Structure your response.
- #2 How do you set expectations about AI capability without sounding like you are blocking?
- #3 Describe the difference between a demo and a product, using a concrete example.
- #4 Your board asks why competitors ship AI features faster. Prepare your answer.
- #6 How do you handle a sales team that has already sold a capability you do not have?
- #7 What is the honest way to describe your AI feature's limitations in marketing copy?