How does probabilistic output change the shape of your support and escalation process?
Haventide watches a traveler's connecting flights and shows one number: how likely each leg is to land on time. Peregrina Osgar trusted a 94 percent call on the one trip that couldn't slip, her sister's rehearsal dinner. Cadfael Wrensfield took her call afterward and found his ticket queue had no box for what had actually happened to her. This is what has to change about a support process once the thing it's defending isn't a bug. It's a percentage that was built to be wrong sometimes, on purpose.
- Split every miss into two escalation lanes at intake: tracked pattern, or genuine one-off.Why: this is the decision the rest of the answer works out, and nothing downstream matters if this split doesn't exist first.
- Send an automatic, visible follow-up to the specific person the moment the true outcome is known.Why: Peregrina's report went into a feedback form with no reply, so she had nothing to trust the next time Haventide showed her a number.
- Tag every miss ticket with the model's confidence and whatever condition was active at prediction time.Why: this is what turns "we said 94 percent" into something searchable, instead of a sentence buried in free text nobody can query.
- Track accuracy by condition, not only the one headline number.Why: Haventide's blended 91 percent never moved while the weather-flagged segment sat near 58, invisible underneath it.
- Give agents a real tracked case number for a known cluster, separate from a one-off goodwill credit.Why: Cadfael had one tool for two different situations, and only one of them should ever reach the model team.
- Leave the credit-and-apology flow alone for genuine one-offs.Why: not every miss is a pattern; routing all of them through model review would drown the real clusters in noise.
How to answer this, stage by stage
Nobody is grading whether you can list good customer-service habits. They're grading whether you can point at the exact place a support process built for bugs breaks on a call that was never a bug.
Let's learn
Before Haventide existed, a traveler with a tight connection had exactly one method: stand near the gate, refresh the airport's flight board, and guess. About 1 in 6 tight connections got missed with no warning at all, because nothing told a passenger which flight was actually shaky until it was too late to do anything about it.
Haventide changed that. It reads live flight data and gives every leg of a trip one number: how likely it is to land close enough to make the next flight. Below a set bar, it offers a backup, an earlier flight, a different connection, sometimes for a small fee. Missed connections on flagged legs fell by more than two-thirds in Haventide's first year, and travelers who took the backup it offered saved themselves, on average, about three hours of scrambling from a gate agent's desk. Across every flight it scores, Haventide correctly calls on-time or delayed about 91 percent of the time. For most trips, on most days, it just works.
The number is honest, on average, across every flight it's ever scored. What it doesn't do is change when today isn't average. On a normal day at a busy connecting hub, Haventide's shown confidence tracks reality closely: it says 94, the real outcome lands near 92. On a day with an active winter weather advisory at that same hub, the true on-time rate for regional connections drops to about 58 percent, and the app still shows the exact same 94, because the number is built from a static, per-route baseline that isn't checked against live weather-advisory data at the moment it's shown.
Here's the turn. Those six wrong calls in every hundred were never the real problem. That's what a 94 percent tool is supposed to produce, on purpose, some of the time. The real problem is what happens after. Nobody's process could tell the difference between "this was always going to happen sometimes" and "this keeps happening on the same kind of day." So every one of those misses got the exact same shrug, a credit, and an apology, whether it was pure bad luck or a recurring pattern nobody had named yet.
What it costs at its worst: those weather-flagged misses kept stacking up underneath one blended 91 percent that never moved, because nobody's ticket form had a field to catch them. Peregrina Osgar's inbound flight, on the one trip that couldn't slip, was one of them. What she lost wasn't a connection. It was a rehearsal dinner, and any reason to trust the next number Haventide showed her.
What I would leave alone: for long layovers, ninety minutes or more, with no active advisory anywhere near the route, the current blended number is fine exactly as it is. A miss there costs a traveler a longer wait in a terminal, not a missed event, and there's real slack to recover without any of this machinery.
The lesson: a support process built for deterministic bugs assumes every complaint has a fix waiting at the other end of it. A probabilistic feature's support process has to do something a bug queue never had to: tell a pattern from a one-off, before anyone picks up the phone.
Now here is the same thing as a story
Read the short version above when you're mid-interview and the clock is running. Read this one when you want to actually feel what twelve minutes cost.
Peregrina Osgar flies to see her family about nine times a year, always through a connecting hub, always on a budget that means whatever fare is cheapest that month. She used to be good at this the old way: printed boarding passes, a phone alarm set for boarding time, her own eyes on the departure board the second she cleared the jet bridge.
On her first trip with Haventide open on her phone, she still checked the board herself at every gate, out of habit more than doubt. By her fourth trip, she glanced at Haventide's number once at the gate and trusted it. By her ninth, she'd stopped checking anything at all. The number had never once let her down, so there was nothing left to double-check.
Her tenth trip was the one that mattered. Her sister's rehearsal dinner, a Friday evening in Portland, a fifty-two-minute layover through a hub with an active winter storm advisory that morning. Haventide showed her 94 percent likely on time. It also offered a backup, an earlier connection through the same hub, forty dollars more. She looked at the 94, looked at the forty dollars, and kept her original booking. Ninety-four felt like a promise, not a probability.
Her inbound flight ran sixty-four minutes late, de-icing backed up behind a dozen other regional jets doing the same thing. She reached her connecting gate twelve minutes after the door had already closed.
She opened the app right there at the gate and tapped "report this." It disappeared into a general feedback form. No case number. No reply, not that day, not the next week.
Cadfael Wrensfield has worked Haventide's support line for two years, and he's good at it. Duplicate charges, stuck bookings, a card that got declined mid-checkout, he clears those in minutes because there's always a fix on the other end of the ticket. When Peregrina called, he did everything right by the book. Pulled her record, confirmed the delay, apologized, and issued a seventy-five dollar travel credit, the same resolution he gives every missed-connection call, because it's the only lever his queue gives him.
What Cadfael couldn't do was tell Peregrina whether her miss was rotten luck or something that keeps happening. He had no field on his screen for "was there a weather advisory active," no way to see that this was the eighty-somethingth ticket that month carrying the exact same shape. Nobody had ever built that field, because a bug ticket doesn't need one. A bug either happened once or it happened to everyone; either way there's a report to file and a patch to point at. A well-calibrated miss doesn't work that way. It happened exactly as often as it was designed to, to someone specific, on a day that looked identical to every other day on the screen.
The decision Haventide's model lead, Frideswide Tench, would take back reaches back to the week the confidence score first shipped. In that launch review, someone asked whether the number should shift in real time against live weather-advisory data. The answer was that a live lookup would add real latency, and weather-exposed connections were rare enough not to be worth it. Fine, everyone agreed, and moved to the next line item on the list.
Run the same storm again, with two changes in place. Every missed-connection ticket gets tagged automatically with the model's confidence and whether an advisory was active. Peregrina's ticket matches a known, tracked cluster the moment it's filed, so it routes to the model team with a case number instead of vanishing into a feedback form, and Haventide sends her a follow-up that afternoon: "We told you 94. This one didn't go your way. Here's what we found, and what we're checking." She still misses the flight, because nothing catches every case. But she knows, within hours, that someone saw it, and that it wasn't just her.
One design treats every miss as a fresh mystery a support agent has to solve alone, on the phone, with no data. The other tells you, the moment the ticket is filed, which misses were already known about.
What I'd tell myself, back in that launch review: a number that's right ninety-four times out of a hundred still needs a way of admitting, out loud, which days it already knows are shakier than the rest.
GUARD: the model wasn't broken, so the ticket had nowhere to go
Not a checklist for spotting risk in general. GUARD forces you to say, specifically, who's stuck holding a miss that nobody's process was built to explain.
And if you want to be sure it really works, try it somewhere else
Same five letters, a county licensing office instead of an airline app, and the missing gap is a business's renewal instead of a rehearsal dinner.
Braelock County's business licensing office runs Clearway, a tool that reads an incoming permit-renewal application and scores it for compliance risk, low enough to fast-track approval, or high enough to send to a manual inspection. Marged Vantwist owns a small bakery that renewed its license two weeks after Braelock County adopted a new fire-safety code covering commercial kitchens. Clearway fast-tracked her renewal at a shown 92 percent low-risk score, the same number it had been giving bakeries like hers for years, because the model had been trained on years of data from before the new code took effect and hadn't yet learned what the rule change actually meant for a kitchen like hers.
Same rank, different lever, mapped straight onto GUARD: the groups are Marged, who never sees Clearway's score at all, and Dovile Cindercreek, the clerk who works the fast-track queue Clearway sorts for her. The harm concentrates on businesses fast-tracked in the weeks right after a code change, not spread across every renewal. Marged has no way to know a score decided her renewal needed no second look, and no channel to ask, it just reads to her as a licensing office that approved her, then fined her months later. The fix is the same shape: flag every fast-tracked business in an affected category for a quick manual recheck once a code change ships, instead of waiting for the next routine inspection cycle. Detecting it means watching violation notices cluster in the sixty days after any code change, the same leading signal that would have caught Haventide's gap early.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip to it: two lanes at intake, a tracked pattern versus a genuine one-off, and a follow-up that reaches the person before they ask.
Cost: no budget for a live weather-advisory lookup this quarter. Ship the ticket-tagging and the two lanes first, since that costs process discipline, not new infrastructure, then add the live lookup once budget allows.
The model got better, for real: say Haventide's accuracy climbs to 97 percent. Keep the two lanes anyway. A smaller miss still deserves a way to be told apart from a pattern, because the harm on the traveler who hits it hasn't gotten any smaller.
Where people run it wrong.
They watch one blended accuracy number and call the rollout a success.
They let every missed-connection ticket get the same resolution, whether it's a one-off or the fortieth ticket that month shaped exactly like it.
They treat "we'd probably notice a pattern eventually" as a plan, instead of pricing what waiting for that pattern actually costs the next person who hits it.
How to use it live. Ask the routing question before naming a fix: "when a miss gets reported, does anything connect it to every other miss that looked like it, or does each one start from zero?" That question alone tells you whether the harm is contained or just quietly repeating.
Three things worth stating directly, since the real judgment sits here. The alternative Haventide could have taken instead of a live, condition-aware confidence score was simply lowering the shown number network-wide any time a weather advisory existed anywhere in the system. That was rejected: it would have hedged the roughly 95 percent of trips nowhere near the advisory, triggering needless backup-flight prompts and burning trust exactly where the number was already right. The AI-specific failure worth naming is silent segment-level miscalibration under distribution shift: a confidence score generated as a static, per-route average will read as accurate in aggregate while being systematically overconfident in the one condition, an active weather advisory, that a traveler most needs it to be honest about. The guardrail is tracking accuracy by condition, not just headline accuracy, and tagging every miss ticket with confidence-plus-condition so a cluster is queryable instead of buried in free text. And the trade-off is real: a live weather-advisory lookup at inference time costs roughly 150 milliseconds of added latency and genuine engineering work, accepted on purpose, because the alternative is a static number that silently fails on exactly the days it mattered most to get right.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't you just raise the confidence bar so the app is more cautious overall?" Response: That was considered and rejected. It would hedge the roughly 95 percent of trips nowhere near a weather advisory, for no real gain, and burn trust exactly where the number is already right.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on What changes when the product is probabilistic
- #1 Name three product decisions that change when a feature's output is probabilistic rather than deterministic.
- #2 A traditional feature either works or has a bug. Explain why that framing breaks for an LLM feature.
- #3 What does 'correct' mean for a summarization feature? Give a definition your engineering team could test against.
- #4 QA files a bug that reads: the model gave a wrong answer once. How do you triage it?
- #5 Explain the difference between a defect and an acceptable error rate to a non-technical executive.
- #6 Why can you not write an acceptance criterion like 'the output must be accurate' for a generative feature?