…
AI Product Case Questions

You lead the ChatGPT 6 rollout. How do you launch it?

A worked answer to a real AI PM interview question: you lead the launch of a major new ChatGPT model. How do you roll it out?

Transcript

Read the full transcript (1,908 words)

[INTERVIEWER] You lead the ChatGPT 6 rollout. How do you launch it? You're leading the ChatGPT 6 rollout. How do you launch it? Here's the trap, and it's a big one: treating this like a marketing date, a day you flip the switch and send the tweet. A frontier model launch is not a date. It's a controlled ramp with a kill switch.

The strong answer never turns it on for everyone at once. It stages exposure user group by user group, it watches safety and cost as primary metrics sitting right next to quality, and it can roll the whole thing back in minutes if something goes wrong. If you say launch it and monitor, you've already failed. What they're testing is whether you understand that shipping a frontier model is a risk management exercise, not a launch party.

A more capable model is also a more capable model at doing harm, at burning GPU budget, at melting your serving capacity. So the interviewer wants to see you build the guardrails before the ramp, stage the exposure, and hold the line when one segment misbehaves. Walk out of this able to describe a real ramp with thresholds and a rollback, and you can answer any launch question in the book.

Let me build it. First, clarify, because ChatGPT 6 could mean a few things. Is it a new base model behind the same product, or genuinely new capabilities, longer context, agents, tool use? What's the surface: consumer ChatGPT, the API, or both? And who bears the risk here: is the pressure on safety, on latency, on cost per query, or on raw serving capacity?

I'll make my assumptions explicit. I'll assume a new frontier model, materially more capable, going to both consumer ChatGPT and the API, and carrying meaningfully higher inference cost. And I'll state the goal of the rollout itself, which is not get it out fast. It's ship the capability without a safety incident, without a latency regression, and without a capacity meltdown.

Name that goal, because it decides every tradeoff after it. You can't ramp against nothing, so before the first user sees it, I define the metrics. Success first. Task success rate. Human preference, meaning the win rate of ChatGPT 6 responses against ChatGPT 5 in blind comparisons. Retention. And thumbs up rate. Then the guardrails, and each one gets a hard threshold and a named owner, not a vibe.

Unsafe output rate, so jailbreaks and policy violations. Hallucination and factuality on the eval set. p95 latency. Cost per query. Error and timeout rate. And refusal rate, because over refusing is a real regression too, a model that won't answer safe questions is a broken model even if it's perfectly safe. Here's the discipline that matters: I write the rollback thresholds down now, in the calm before launch, so nobody's trying to negotiate them at two in the morning in the middle of an incident.

Negotiating thresholds under pressure usually leads to compromised standards. Now the gates, before a single external user touches it. The full offline eval suite runs, capabilities and safety benchmarks both. Red teaming, plus an external safety review. And a load test at projected peak queries per second to confirm you've actually got the capacity and to price the cost per query at scale.

This is where a frontier lab differs from a normal product team: the safety eval and the red team sign off are a gate, a hard stop, not an optional extra you skip when the deadline's tight. And stage zero is dogfooding: your own staff use it internally first, because they'll find the embarrassing failures before the public does. Here's the heart of it, the ramp.

Internal staff, then trusted testers or alpha, then a one percent canary of production, then five percent, twenty five, fifty, and finally a hundred. And you hold at each step long enough to actually read the guardrails, which is hours to a day or two, and longer for safety signals because rare events need volume before they show up in the data.

A couple of technical details that separate a real answer from a vague gesture. One, route users by a stable hash of their user ID, so a given person gets a consistent experience and isn't flipping between models mid conversation. Two, and this is the one people forget, hold back a control group on ChatGPT 5 at every single stage.

That way it's better is a measured A/B delta against a live control, not just drift you're eyeballing week over week. At each gate, I promote only if two things are true together: success metrics are up or flat, and every guardrail is inside its threshold. Both, not either. And crucially, I segment the read by platform, by geography, by language, and by use case.

Because a model can be brilliant in English and unsafe in another language, or perfect on chat and broken on tool use, and if you only look at the aggregate, that regression hides inside a healthy looking average. Every ramp needs an exit. I want a feature flag, or a model routing switch, that reverts traffic back to ChatGPT 5 in minutes.

And I want it tested before launch, actually pulled in a drill, not theorised in a doc. Then I define what trips it automatically versus what needs a human call. Automatic triggers: unsafe output rate over threshold, a p95 latency blowout, an error spike. Those page and can flip on their own. Judgement calls, like a subtle quality complaint, go to a human.

And the rollback has to be a config change, not a redeploy, because a redeploy takes too long when users are getting harmed right now. Relying on a new code deployment for rollback means you lack a true kill switch. Last part of the frame, and it's the frontier specific bit people skip. A frontier launch is often GPU capacity bound, so I gate the ramp on serving headroom, not just on quality.

If cost per query is high, I consider paid tiers or rate limits first, so the economics don't blow up before you've learned anything. And I stagger the announcement with the ramp, so the marketing demand doesn't front run the capacity and crash the service on day one. Then after launch: I keep monitoring the long tail safety reports that only surface at scale, I hold the control group for a couple of weeks to measure the true retention lift, and I write the launch retro so the next model ramp is smoother.

Let me run the whole thing concretely. ChatGPT 6, higher capability, roughly 1.4 times the inference cost of ChatGPT 5. Pre launch: the eval suite passes, the red team finds two classes of jailbreak, both get patched and retested, and the load test confirms twenty percent capacity headroom at projected peak. Good, we ramp. Staff, then a thousand alpha testers, then the one percent canary.

At one percent, blind preference is fifty eight percent win against ChatGPT 5, thumbs up is up four percent, both great. But the unsafe output rate is sitting at 0.9 percent against a half percent threshold, and it's concentrated in one language segment. So I hold. I do not promote, even though the quality numbers look fantastic, because a guardrail is red.

We ship a safety filter tuned for that language, remeasure, and unsafe output drops to 0.3 percent, back under threshold. Now I resume: five percent, then twenty five, watching p95 latency, which stays under the three second cap, and cost per query. At twenty five percent a latency regression shows up, but only on the API tool use path, not on chat.

So I hold that one surface at twenty five while consumer chat proceeds all the way to a hundred. Kill switch tested, control group retained throughout. The full ramp completes over roughly two weeks. And the headline: nothing shipped to everyone on day one, and the one segment that misbehaved got caught and fixed before it ever reached scale. That holding decision, letting chat proceed while pinning the broken surface, is the exact judgement they're listening for.

The interviewer will often press the business against the caution: you're being so careful you'll get scooped, a competitor ships their frontier model while you're still at five percent. Real tension, and you need a straight answer. The point of the staged ramp is speed with a floor, not slowness for its own sake. Each hold is hours to a day, not weeks, so the whole ramp still completes in roughly two weeks, and the control is what lets you move fast, because you can promote confidently the moment the guardrails are green instead of waiting and guessing.

The thing I won't trade away for speed is the safety gate, because a frontier lab that ships a jailbreak to a hundred million people doesn't save the time it cut, it loses far more cleaning up the incident and the trust. The second push you should expect: your unsafe output rate looks fine in aggregate, why keep segmenting by language?

Because aggregate safety is a lie of averages. A model that's ninety nine point nine percent safe overall can be dangerously unsafe in one language that's five percent of traffic, and that five percent is still millions of real people. An aggregate green metric hiding a red segment is exactly what segmented reads prevent, proving you understand safety at scale.

Here's what makes them lean in. First, you treated safety and cost as primary ramp guardrails, right alongside quality, which is the frontier lab signal. Most people watch only quality and retention. Second, you held a control group at every stage, so it's better was a measured delta and not a feeling. Third, you described a tested, rapid rollback with pre agreed automatic trip conditions, and you showed the discipline to hold the ramp when one segment failed while letting the rest keep going.

That last one, partial hold instead of all or nothing, is a senior move. The traps that sink people. One, jumping to launch it and monitor with no stages, no thresholds, no rollback. Jumping straight to launching and monitoring without stages, thresholds, or a rollback is just wishing for the best. Two, watching only quality and retention while ignoring latency, cost per query, and unsafe output rate, which on a frontier model are the metrics that actually bite you.

And three, reading the ramp only in aggregate, so a per language or per surface safety regression hides inside a healthy average and reaches millions of people before anyone notices. You must segment every gate because averages can hide critical failures. So the whole picture. Clarify what's shipping and to which surface. Define success and guardrail metrics with hard thresholds before you ramp.

Gate on eval, red team, and load test. Then ramp staff to one percent to five to twenty five to fifty to a hundred, holding a control group at every stage, promoting only when success is up or flat and every guardrail is green, and segmenting by platform, geo, language, and use case. Keep a tested, rapid kill switch the whole way.

Carry this into the room: stage the exposure, gate every promotion on success metrics up and guardrails green, keep a control group and a tested kill switch, and never flip a frontier model on for everyone at once.

Keep learning