What guardrail metrics would you track? One flags an issue. Now what?
Transcript
Read the full transcript (1,598 words)
[INTERVIEWER] What guardrail metrics would you track? One flags an issue. Now what? Consider what guardrail metrics you would track and what you actually do if one of them flags an issue. Here's why guardrails exist at all. You can win your primary metric and still be quietly wrecking the product. Queries per user goes up because your answers got worse and people keep re-asking.
So the strong answer names the specific counter metrics for an AI answer engine. More importantly, it has a runbook for the moment one trips. Contain first, diagnose second, fix third. In that order. This is a two part question, and the second half is where it's won. Naming guardrails is table stakes. What the interviewer really wants is the on call reflex.
When your factuality metric spikes at three in the afternoon and users are getting hurt right now, do you flail, or do you stop the bleeding before you understand it? That instinct to contain the blast radius before you diagnose is what separates a PM who's actually shipped an AI product from one who's only read about it. Let me build both halves.
We need a quick foundation first. Every North Star can be gamed, or bought at a cost you didn't intend to pay. A guardrail, or counter metric, is the thing you refuse to sacrifice while you're chasing the primary goal. For Perplexity, the primary metric might be queries per user, or answer engagement. The guardrails are what protect answer quality, trust, cost, and speed while you push that primary number up.
Say that framing first, because it explains why the whole set exists. Now for the set itself. I tie every single one to an answer engine specifically, rather than generic quality and latency. Factuality, or hallucination rate, is the share of answers with an unsupported or flat out wrong claim. You measure this on a sampled eval set and through user report clicks.
For a product replacing search, that's guardrail number one. Citation quality is the share of answers where the cited source actually supports the claim. This means real grounding, not just any link slapped underneath. Latency at p95 matters because an answer engine competes on speed and a slow answer is a lost user. Cost per query is vital because LLM inference is expensive and runaway cost per query quietly kills your unit economics.
Safety covers the harmful content rate, policy violations and unsafe instructions. User trust signals include the thumbs down rate, report rate, and the regenerate or go search Google instead rate, which is the user telling you the answer failed. Finally, coverage, or refusal rate, because over refusing real queries is its own regression. Naming those product specific ones is what shows you actually understand this product.
Then we have a short but strict rule. Every guardrail gets a threshold and an owner. A guardrail with no threshold is just decoration. You only know it flagged if you defined in advance what flagging means. For example, the hallucination report rate must stay under one percent of answers. If it crosses that line, it pages the on call PM and engineer.
Without a defined threshold, you simply cannot answer the second half of this question. Okay, so the factuality guardrail spikes. Here's the loop, and the order is the point. First, contain. If the harm is live and rising, cut the blast radius right now, before you understand anything. Throttle or feature flag off the affected surface, or route to a safer fallback.
That could be a more conservative model, or just showing the sources instead of a synthesised answer. Stop the bleeding first. Second, ask if it's real or a measurement artifact. Did a logging change or a new report button just inflate the rate artificially? Rule out the artifact before you declare a quality crisis, because half the time the metric is lying.
Third, isolate. Segment the spike by query category, like news, medical, or code. Check which model version, which language, which geography, and since exactly when. Then correlate against the release timeline. Did a model swap, a prompt change, or a retrieval index update ship right before the spike? An AI quality regression almost always traces back to one specific change.
Fourth, fix at the right layer. If a model or prompt change caused it, roll that back. If retrieval is surfacing bad sources, fix the index or raise the grounding threshold. If it's a genuinely hard class of query, add a guardrail to refuse or hand off on low confidence, rather than shipping a confident wrong answer. Fifth, verify and restore.
Confirm the metric is back to baseline on the sampled eval before you ramp the surface up again. Sixth, communicate. Tell the stakeholders what tripped, what you contained, and the fix. Then decide whether users who saw bad answers deserve a correction. I'll make this real with an example. The primary metric is answers per weekly active user, trending nicely up after a new synthesis model shipped on Tuesday.
Then on Wednesday, the hallucination report guardrail jumps from 0.6 percent to 2.1 percent of answers. This is past the one percent threshold, and the on call team gets paged. Contain first. I feature flag the new model back to conservative synthesis for the affected query classes within the hour. Next, check if it's real or an artifact. The report button hasn't changed, and a sampled eval confirms a genuine factuality drop, so it's real.
Isolate the issue. The spike concentrates in recent news queries, and it started exactly when the Tuesday model shipped. The root cause is that the new model synthesises confidently over thin or stale retrieval on breaking news. It makes things up when the sources are weak. Fix at the right layer by raising the grounding threshold and forcing a show sources without synthesis fallback whenever retrieval confidence is low for news.
Remeasure to see the report rate back down to 0.7 percent. Restore the model with the new guardrail in place, keep the news query fallback, and write the retro. Here's the kicker. The primary metric of answers per user was up all week the entire time. The guardrail is the only thing that stopped a slow motion trust disaster hiding underneath a healthy looking headline number.
That's exactly what guardrails are for. The interviewer will often press on the tension. They will ask if containing the surface hurts your primary metric, since you're throttling the thing users came for. Yes, it does briefly, and that's the right trade. So I'd say it plainly. I'll accept a short, measured dip in answers per user to stop shipping wrong answers, because a trust hit compounds and an engagement dip recovers.
A user who gets one confidently wrong medical answer doesn't just churn. They tell people the product lies, and that's far more expensive than an hour of conservative fallbacks. Containing is buying time at a known, small cost to avoid an unknown, large one. The second push you should expect is about the guardrail itself being the problem. They might ask if your one percent threshold is just too tight and you're paging people over nothing.
That's a real failure mode called alert fatigue, where a badly tuned guardrail cries wolf until the on call team stops trusting it. So thresholds aren't set once and forgotten. I'd calibrate them against historical baseline plus normal variance. This uses the same sizing discipline as any diagnosis, so the threshold fires on genuine signal and not on the daily noise.
A guardrail that pages every Tuesday is worse than no guardrail, because it trains the team to ignore the one page that actually matters. Naming alert fatigue as a risk, and calibration as the fix, shows you've lived on the other end of a pager and not just drawn the dashboard. Here's what makes them lean in. First, you named answer engine specific guardrails like hallucination rate, citation grounding, and cost per query, instead of a generic quality and latency list that could apply to anything.
Second, you contained the blast radius before diagnosing, which is what a real on call PM does when people are getting hurt in real time. Third, you isolated the spike to a query class and tied it to a specific release, rather than sitting there guessing at causes. Precision under pressure is the signal. Watch out for the traps. Trap one is listing guardrails with no thresholds, which means the idea of flagging an issue has no meaning since there's nothing to flag against.
Trap two is jumping straight to a fix with no containment and no check that the spike is even real. This means you either ship over an artifact or leave users exposed while you investigate. Trap three is treating a quality regression as some mysterious act of God instead of tracing it to the change that shipped right before it.
In AI products, the regression almost always has a Tuesday deploy behind it. Find the deploy. So here's the shape. Define what a guardrail is, name the product specific set, give every one a threshold and an owner, and then run the trip runbook. That means you contain, confirm it's real, isolate to a change, fix at the right layer, verify, and communicate.
Carry this into the room. Guardrails are the metrics you refuse to sacrifice. Each needs a threshold and an owner. When one trips you contain the blast radius, confirm it's real, isolate to a change, fix at the right layer, and then restore.