ConceptIntermediateShipping & Model Lifecycle / Rollout strategy and phased launches / #18

Describe the on-call arrangement you would want during an AI rollout.

The direct answer
Staff someone who can tell a wrong answer from a right one on the primary on-call rotation, before the rollout starts, not after something goes wrong. Standard infra on-call comes second, for when the tool itself breaks. A path to the model's vendor comes last, for the rare case the problem is in the model and not in how you built with it.
The ranking, by what breaks first if skipped
  1. Put someone who can judge the output, not just the uptime, on primary on-call before day one.Why: this is the whole decision. Skip it and there is nobody on the rotation who can tell a fluent wrong answer from a right one.
  2. Rule out running the rollout on the standard infra rotation alone, even though it already exists.Why: reversibility. A wrong answer that reaches a customer cannot be unsent. Who is on secondary this week can be swapped by Friday.
  3. Confirm the primary on-call person has actually reviewed the tool's old cases, before scheduling a single shift.Why: dependency. Nothing about the rollout is safe to start until this is true.
  4. Check the on-call roster against the list of people who reviewed those old cases, cheaply, before committing the schedule.Why: evidence. It is the cheapest thing to check first, and it is often the one nobody checks.
  5. Rank quality judgment first, infra on-call second, and the vendor escalation last.Why: each one only helps with a different kind of failure. The slowest, least controllable one should never be the first call.

How to answer this, stage by stage

Seven moves. The trap in this question is describing a normal rotation with an AI label stuck on it, when the real question is who can tell a good answer from a bad one.

1
Ground it in one real product
Say it like this
"Let me make this concrete. Say Ostrom Cloud ships Synopsis, a feature inside its incident tool Signalboard that drafts what happened, why, and what fixed it, right after a major incident closes. Sanaa Bhargava is running the rollout, Joaquin Merrick carries the standard on-call pager, and Livia Holmberg built the two hundred old incidents Synopsis was checked against before anyone trusted it. I will answer against that."
Why this works
Grounds an open question in one real system, so the on-call ranking that follows is not a guess in the air.
2
Name your method before you use it
Say it like this
"I would use ORDER here. Rank the candidate on-call setups by what is actually hardest to undo if it is missing, not by which rotation already happens to exist."
Why this works
Signals a plan up front, so the answer reads as a method, not a habit repeated from the last rollout.
3
Say what the question is actually testing
Say it like this
"This is not really asking who answers the page fastest. It is asking who, once they pick up, can tell whether Synopsis just told a customer something false."
Why this works
Separates a scheduling question from the real judgment call the interviewer is testing.
4
Give the ranked answer straight
Say it like this
"Put someone who can judge the output on primary, before day one. Standard infra on-call covers actual outages, second. A path to the model vendor sits last, for when the problem is genuinely theirs to fix."
Why this works
This is deliverable zero, said out loud, in the order that actually matters.
5
Show what has to be true before anything else
Say it like this
"None of this works if the person on primary has not already reviewed those two hundred old incidents. At Ostrom that is Livia. Joaquin can restart Synopsis in his sleep. He cannot tell you if the paragraph it just wrote is true."
Why this works
Shows the order is not arbitrary. One thing has to be real before the rest of the plan is worth trusting.
6
Name what is hardest to take back
Say it like this
"The hardest thing to undo is a wrong explanation that already reached a customer. Who is on secondary this week, you can swap by Friday. A customer who already read the wrong reason for their own outage, you cannot un-tell them."
Why this works
Names the one gap that turns a normal staffing choice into something you cannot walk back.
7
Back it with the numbers and close on the rule
Say it like this
"Here is what it looked like without this. A wrong root cause sat live on the status page for forty minutes before Livia happened to catch it, off duty, in a Slack thread. With her staffed on primary, the same draft gets caught in about six minutes, before it is ever posted. So: quality judgment on call first, since nothing else works without it. Infra on-call second, for when the tool actually breaks. The vendor escalation last, because that path is slow and it is not ours to speed up."
Why this works
Ends on the literal ranking the question asked for, backed by a number instead of just asserted.

Let's learn

Synopsis is a box that appears on screen the second a major incident closes at Ostrom Cloud, already holding a paragraph that explains what happened, why, and what fixed it.

Before Synopsis, whoever closed the incident wrote that paragraph by hand. It took about forty minutes, on top of already being awake at three in the morning for a dozen incidents most months. Without it, half the write-ups from the middle of the night never got finished at all, because forty minutes is a long time to stay sharp once the fire is out.

Knowledge spark: what is an eval set? A pile of old cases with the right answer already written down next to them. Ostrom built one from two hundred past incidents, so anyone reviewing a new Synopsis draft could check it against cases where the real cause is already known, instead of guessing whether it sounds right.

Sanaa Bhargava, the product manager running the rollout, did not open Synopsis to every incident at once. She started it on internal write-ups only, the notes engineers read to each other, never shown to a customer. For the first two weeks it went well. Drafts came back in under a minute. Engineers read every one closely, checked it against the actual logs, and it kept holding up.

Hand-sketch dependency diagram, three boxes connected by arrows left to right: Someone who reviewed the old incidents joins the rotation, circled in amber, then Synopsis goes live on real incidents, then A wrong draft gets caught in minutes, not weeks, showing the order this has to happen in.
Someone who already knows what a wrong draft looks like has to be on the rotation before the rest of this chain is true.

So the team moved Synopsis to draft the public status page update too, the paragraph customers read while an outage is happening. That is where it stopped being a convenience and started being something someone had to answer for.

Minutes to write an incident summary, by hand versus with Synopsis
40 min 1 min Written by hand Drafted by Synopsis
Before SynopsisAfter Synopsis
Forty minutes saved is real. It says nothing about whether the minute-long draft is actually true, and nobody on the standard rotation was ever asked to check that.
We did not need Synopsis to stop getting things wrong. We needed someone on call who could tell when it did.

Here is what that costs at its worst. A wrong explanation on a public status page is not a bug ticket. It is something a customer already read and believed, while deciding whether to trust Ostrom with their own outage.

The choice I would take back Ostrom staffed Synopsis's rollout the same way it staffs every rotation: whoever is next on the standing infra on-call schedule owns it. That was fine for every earlier feature, because being on call had only ever meant being able to restart something. It stopped being fine the moment being on call also had to mean being able to tell if a paragraph was lying.

What I would leave alone. When Synopsis fails outright, times out, throws an error, produces nothing, that is a normal outage. Joaquin's job does not change at all. Restarting a broken tool and judging a fluent, confident, wrong paragraph are not the same failure, and only one of them needs a new kind of person on call.

The lesson. An on-call rotation built to answer "is it up" does not automatically answer "is it right." Those are two different jobs, and we only ever staffed for the first one.

Now here is the same thing as a story

The short version is above. Keep reading if you want to feel why a page that looked handled still cost forty minutes.

Joaquin Merrick can read a stack trace the way some people read a menu. He has carried Ostrom's platform pager for three years. Give him a service that is down and he will have it back inside minutes, and he will tell you exactly why it fell over before the postmortem doc is even open.

When Synopsis started drafting the public status page too, Joaquin read every one of its early drafts against the raw logs before approving it. That took an extra five minutes on top of the outage itself, but the drafts kept coming back clean, sentence after sentence, incident after incident.

So he stopped reading the whole thing. He started skimming the first two lines. Then he started approving on sight, because it always read cleanly and it always matched whatever he could already see with his own eyes, the service was down, now it is back up.

Then came a Tuesday that did not look different from any other.

Hand-sketch comparison: on the left, a green door swinging both ways labeled Who is on secondary infra on-call this week, captioned change it by Friday, costs nothing. On the right, a red-orange door bolted shut labeled A wrong root cause already read by a customer, captioned found after it posted, cannot be unread. A VS sits between the two panels.
One of these you can change next week. The other one, caught late, has already reached someone.

A major outage closed, and Synopsis drafted the public explanation: the cause was a database failover taking longer than expected. It read well. It matched what customers had seen, things went slow, then things came back. Joaquin approved it and it posted.

The real cause was a feature flag, mis-scoped weeks earlier, that had silently turned off rate limiting for one region. The fix that shipped was a flag revert, not a failover change of any kind. Nobody caught the mismatch for thirty two minutes, until Livia Holmberg, off duty, not on the rotation at all, saw the status page mentioned in a customer support Slack thread and noticed it did not match the deploy log she had glanced at out of habit. It took her eight more minutes to confirm it and get it pulled. Forty minutes, start to finish, with a wrong reason for their own outage sitting in front of every customer who read that page.

He was not careless. A wrong-but-fluent explanation looks exactly like a right one from the outside, and nobody had ever asked Joaquin to be the person who could tell the difference.

Here is what he did next. He stopped trusting any Synopsis draft he could not check himself, and since he could not check the ones that mattered, he started paging Livia, at any hour, for anything that read even slightly odd. Eleven times over the next two weeks, most of them nothing at all, just a phrase that struck him as off. There was no setting between "trust it because it reads well" and "wake Livia up for everything," because nobody had ever built him one. He was never meant to make that call alone.

We did not lose forty minutes twice. We lost something bigger: an on-call system for the thing that actually mattered, quality, that ran entirely on one person's unscheduled goodwill, with no rotation, no shift, and no end date.

A month before rollout, in the meeting where the on-call plan got scoped, someone had asked whether Synopsis needed anything different from the standard rotation. The standard rotation already existed, already had a schedule, already had a pager assigned. Nobody wrote down that the whole plan would lean on whoever happened to be up that week knowing something none of them had been trained to know.

I would go back and put Livia on primary for those first two weeks, not as the person Joaquin calls when he is unsure, as the person who reads it first. Same near miss, same Tuesday, same wrong draft about a database failover. This time Livia checks it against the deploy log in about six minutes and holds it before it ever reaches the status page. Joaquin, on secondary, never gets paged that day at all, because the service itself never goes down.

One design put a person on call who could tell the lights were back on. The other put a person on call who could tell whether the sentence explaining it was true.

What I would tell myself, back in that scoping meeting: a rotation is not just about who answers the page. It is about what they can actually tell you once they pick up.

ORDER, for ranking who actually has to be on call first

GUARD would fit if this were about who gets harmed and cannot push back. This is a straight ranking of on-call roles by what is hardest to undo if the wrong one is missing, which is ORDER's job.

O, outcome. Every candidate on-call setup competes for one thing: a wrong-but-fluent Synopsis draft gets caught inside minutes, by someone who can actually tell, instead of being mistaken for business as usual or missed entirely.
R, reversibility. The hardest thing to undo is a wrong explanation a customer has already read. Who carries secondary infra on-call this week is trivial to swap. A status page a customer already trusted, and already acted on, is not.
D, dependency. Nothing about the rollout is safe to start until someone who can judge Synopsis's output, not just its uptime, is arranged on the rotation. Standard infra on-call can restart Synopsis in seconds. It cannot tell you whether the paragraph it just wrote is true.
E, evidence. Cheap to check before day one: pull the on-call roster and compare it against the list of people who reviewed the two hundred old incidents Synopsis was checked against. At Ostrom that took Sanaa one afternoon, and the overlap came back at zero.
R, rank. Quality judgment first, on primary. Standard infra on-call second, for actual outages. A fallback path to the model vendor last, for the rare case the fault sits inside the model itself, since that path is slow and outside the team's own control.
Minutes a wrong draft stayed live, without a quality judge on primary versus with one
32 min, unnoticed 8 min, confirm and pull 40 min total 6 min, caught before posting No quality judge on primary Livia on primary
Live, unnoticedConfirm and pullCaught before it posts
Same wrong draft, same Tuesday. The only thing that changed between forty minutes of exposure and zero was who was already scheduled to read it first.
The check that keeps this ranking honest If the standard on-call rotation could already tell a wrong draft from a right one, this would not need to rank first. It ranks first because the gap stayed invisible right up until the exact failure it was built to catch, a confident, wrong sentence, actually showed up.

Same order, a recycling facility instead of a cloud platform

Cardew Recycling runs a camera tool over its sorting line that flags contaminated loads before they get baled and shipped to a buyer, so nobody finds out about the problem after a truck has already left.

O. Every version of Cardew's on-call setup protects one thing: a wrongly cleared load gets caught before it is baled, not after a buyer rejects the whole shipment.
R. A contaminated bale that has already shipped under a clean label is the hardest thing to undo. It costs a rejected shipment and a buyer relationship, against a five-minute delay at the sorting line.
D. Nothing about the rollout matters until someone who can tell a false clearance from a real one is on call. Facility maintenance on-call can fix a jammed conveyor. It cannot judge a camera flag.
E. Cheap to check: whether the on-call roster includes anyone who reviewed the tool's set of past contamination photos, the ones with the right call already written down next to them.
R. Same order. Sorting lead Kaelan Amaresh, who reviewed those old photos, goes on primary first. Facility maintenance on-call goes second, for the conveyor itself. The camera vendor's support line goes last.

Swap the trigger and it still runs

  • Leadership wants Synopsis rolled out to every team faster after a good quarter. The order does not move. Speed makes putting a quality judge on primary matter more, not less, since there is less time to catch a mismatch the slow way.
  • Synopsis's underlying model gets noticeably more accurate. Does not reorder either. A better model still cannot report on its own rare wrong answer. Someone still has to be there to notice.
  • Ostrom offers Synopsis to a partner company running its own smaller incident desk. Does not reorder. The rule protects the same thing no matter whose customers are reading the status page.

Where people run it wrong

  • Treating "the on-call engineer restarted it fine" as proof the rotation is working, when restarting a tool and judging its output are different jobs.
  • Using the one person who can judge quality as an emergency contact instead of a scheduled primary, which quietly turns her judgment into unpaid, unscheduled overtime.
  • Checking the roster's skill fit once at launch and never rechecking it once the model, or the team, changes.

How to use it live

Say the outcome out loud before naming a role. "Every choice about who is on call protects one thing: whether we catch a wrong-but-confident answer before it reaches someone outside the team." Then ask what the standard on-call person in that rotation could actually tell you if the output looked slightly off. If the honest answer is not much, that is the gap you rank first.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits ranking on-call roles for an AI rollout, and why not GUARD?
Tap to flip
ANSWER
ORDER, for ranking candidate on-call roles by what is hardest to undo if the wrong one is missing. GUARD fits questions about who is harmed and cannot push back, not a straight staffing sequence.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Sanaa Bhargava, the product manager running Synopsis's rollout at Ostrom Cloud, working with on-call engineer Joaquin Merrick and eval-set builder Livia Holmberg.
3 · THE HABIT
What did Joaquin stop doing once Synopsis's drafts kept reading well?
Tap to flip
ANSWER
He stopped checking each draft against the real logs, and started approving whatever read cleanly and matched what he could already see for himself.
4 · THE SWITCH
What is the two-setting switch in Joaquin's story, with no middle?
Tap to flip
ANSWER
Trust the draft because it reads well, or page Livia for absolutely everything, because he cannot tell which flagged thing is real. Nobody had ever built him a calibrated middle setting.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Staffing Synopsis's rollout on the standing infra on-call schedule, same as every other feature. It made sense because being on call had only ever meant being able to restart something.
6 · THE NUMBER
The wrong draft stayed live on the status page for ___ minutes before Livia caught it.
Tap to flip
ANSWER
40 minutes, 32 unnoticed plus 8 to confirm and pull it. The number the whole ranking rests on.
7 · THE REPLAY
Same near miss, Livia staffed as primary on-call from day one. What changes?
Tap to flip
ANSWER
She checks the draft against the deploy log in about 6 minutes and holds it before it ever posts. Zero minutes of a wrong explanation reach a customer, instead of 40.
8 · THE TRANSFER
Section 4 runs ORDER again on a different product. Which one, and what plays the role of the standard infra rotation there?
Tap to flip
ANSWER
Cardew Recycling's load-contamination camera tool. Facility maintenance on-call plays that role, since it can fix a jammed conveyor but cannot judge a camera flag.

Check yourself Score: 0 / 0

Fill in the blank
1. The wrong root-cause draft stayed live on Ostrom's public status page for ______ minutes before Livia caught it.
Show hint
It is the number the entire ranking argument rests on.
Show answer
40 minutes. Thirty two before anyone noticed, eight more to confirm it was wrong and pull it down.
Multiple choice
2. Which on-call role should be staffed first for Synopsis's rollout, according to this answer?
  • A. Whoever is fastest to reach on the standard infra schedule
  • B. Someone who reviewed the old incidents and can tell a wrong-but-fluent draft from a right one
  • C. The most senior engineer on the team, since seniority means good judgment
  • D. No one extra. Synopsis's own confidence score is enough on its own
Show hint
Ask who, once they pick up the page, can actually judge the draft itself.
Show answer
B. Only someone who already knows what a real miss looks like can tell a fluent wrong answer from a right one before it reaches a customer.
True or false
3. True or false: because Joaquin could restart Synopsis and close the incident, the on-call rotation was working as designed. Say why.
  • True
  • False
Show hint
Ask whether restarting a tool and judging what it wrote are the same skill.
Show answer
False. Restarting Synopsis and judging whether its draft was true are different jobs. Joaquin could do the first. Nobody had staffed the rotation for the second.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when Synopsis's rollout was first scoped?
Show hint
Look for the decision that let the on-call schedule get built fast, not the one that made the numbers look good.
Show answer
Model answer: Staffing Synopsis's rollout on the standing infra on-call schedule, because that schedule already existed and already had a pager assigned. It made sense when on-call had only ever needed to answer whether something was up, not whether its output was true.
Short answer, apply it yourself
5. Pick an AI product you use yourself. If it gave you a confident, wrong answer at an inconvenient hour, who around you could actually tell it was wrong, and how would they know?
Show hint
Look for someone who has seen enough real failures to recognize one, not just someone available to answer a phone.
Show answer
Model answer: "My bank's chat assistant once told me a fee was refundable when it wasn't. A front-line support rep who has actually seen the refund policy exceptions could tell in seconds. A rep who has only ever escalated tickets couldn't, because knowing the policy and reading a fluent sentence are different skills."
Short answer, the number question
6. If the near miss had happened on day one of the rollout instead of day three, before Livia had seen any real Synopsis drafts yet, would her six-minute catch time still hold? Why or why not?
Show hint
The dependency step is about having reviewed the old cases, not just being present.
Show answer
Model answer: Probably not as fast. Her six-minute catch relied on already knowing Synopsis's drafting patterns from reviewing the two hundred old incidents, not just being on the rotation. Being present without having done that review would not satisfy the dependency step, only being present after doing it would.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more