Design an AI agent for a streaming service
Transcript
Read the full transcript (1,745 words)
[INTERVIEWER] Design an AI agent for a streaming service. An agent is not a chatbot. Say that to yourself before you open your mouth, because the strong answer to this question defines four precise things, and a chatbot only has one of them. The agent's job. The tools it can call. The stop conditions that end its loop. And the guardrails that keep it inside its lane.
Miss the stop conditions, and you've designed a system that either loops forever or takes actions it should never take. So today is about those four things, in order, and about the edges of the design more than its centre. What's Sierra testing? Sierra builds customer-facing AI agents for a living, so they want to know if you actually understand what an agent is, mechanically.
Do you know it's a loop over tools with a policy, not just a clever text box? Do you think about when it should stop, what it's forbidden to do, and when it hands off to a human? By the end of this you'll have a four-part frame, job, tools, stops, guardrails, that you can run on any "design an agent for X" question in any interview.
First, ninety seconds. Which agent are we even building? A viewer-facing "what should I watch" and account-help agent, or an internal ops agent for the support team? Take viewer-facing, because it's the richer design and gives you more to talk about. Then ask what it can actually do on the user's behalf. Recommend, play content, manage the watchlist, handle account and billing, apply parental controls?
And confirm the boundary: it acts inside one account, with that account's permissions, and anything that touches money or account settings needs an explicit confirmation. Establishing that boundary up front is itself a signal that you think about agents safely. Now, the job, in one sentence. The agent helps a viewer find something worth watching right now, and handles the small account tasks that get in the way, without a human agent and without a menu tree.
But here's the part people forget: the scope boundary is as important as the scope. So say what it doesn't do. It doesn't make refunds beyond a set limit. It doesn't change the plan tier without explicit confirmation. And it doesn't recommend around parental controls. A good agent design is defined by its edges, and naming what it won't do shows more maturity than a long list of what it will.
An agent is its tools plus a policy for calling them, so let's list them, and give each one a permission and a cost. Catalogue search, query by mood, length, genre, cast. User-state read, watch history, watchlist, current subscription, parental settings. Playback control, start, resume, add to list. Account actions, update payment, change plan, cancel, each one behind a confirmation.
And a policy and knowledge lookup for support questions, grounded and cited, exactly like the RAG designs we've covered. Now here's the framing that impresses. Each tool has a permission and a cost. Reads are cheap and safe. Writes like playback and watchlist are reversible, so they're medium risk. And account and billing writes are high-stakes, so they're gated behind explicit user confirmation.
Sorting the tools by reversibility and stakes is how you show you'd design this responsibly. This is the part that separates agent designers from chatbot designers, so slow down here. The loop ends when, one, the user's goal is met, something's playing or the task is done. Two, the user says stop or changes topic. Three, the agent hits a confirmation gate and is now waiting on the human, billing, cancellation.
Four, the agent has tried a bounded number of tool calls without making progress, a step budget, so it can never loop forever. And five, the request is out of scope and it hands off to a human agent. Name the step budget explicitly, say something like "at most six tool calls per turn before it must either answer or escalate." That single number, a hard cap on the loop, is often the thing that makes an interviewer sit up, because it proves you know an agent is a loop that needs a brake.
Now the guardrails, the hard limits. Parental controls are a hard filter on every recommendation and every playback tool, not a polite suggestion the model can reason its way around. Money and account changes require explicit confirmation and get logged. The agent grounds its support answers in cited policy and refuses to invent, the same no-source-no-answer rule from the RAG questions, because a made-up refund policy is a real problem.
It respects a spend and action budget, so a runaway loop can't rack up tool calls or take the same action ten times. And it always offers a human handoff for anything it can't resolve or isn't permitted to do. Guardrails aren't features you add if there's time. They're the walls the loop runs inside. Let me go deep on the loop itself, because "an agent is a loop over tools" is easy to say and harder to design well.
Walk one turn. The user says something. The agent forms a plan, decides which tool to call first, calls it, reads the result, and decides the next move. And every hop through that loop, you're spending latency, money, and risk. So the policy that governs the loop needs real rules. First, prefer cheap reads before expensive writes, and never take a write action you haven't confirmed the user actually wants.
Second, cache within a turn, if you already read the watch history this turn, don't read it again, that's a wasted call against your step budget. Third, and this is the one people miss, handle tool failure gracefully. The catalogue search times out, or the billing API returns an error, what does the agent do? It doesn't silently give up and it doesn't blindly retry ten times.
It retries once, and if it still fails, it tells the user plainly and offers the human handoff. And fourth, the agent has to know when it's genuinely stuck versus just needing one more call, which is why the step budget exists as a hard backstop underneath the smarter logic. Designing the loop's policy, not just listing the tools, is what proves you understand that an agent is a control system, not a prompt.
Metrics. Your North Star is successful task completion per session, something got watched, or the account task actually finished, without needing a human agent. Guardrails: escalation rate, and read it carefully, too high means the agent is under-scoped and useless, too low might mean it's overreaching and doing things it shouldn't. Wrong-action rate, any account or playback action the user immediately undoes.
And a hard-zero target on parental-control violations, because that's a category where one failure is a headline. For feedback, undos, escalations, and "not this" recommendation rejections all train the recommender and the tool-calling policy, so the agent's judgement improves with use. And watch one efficiency metric alongside the outcome ones: tool calls per completed task. If that number is creeping up, the agent is getting less decisive, taking six calls to do what used to take two, and that's both a cost problem and a latency problem the user feels as sluggishness.
A good agent gets more efficient over time, not just more capable, so you track both. Let me ground it. A viewer says "find me something for twenty minutes, nothing scary, kids are around." The agent calls catalogue search filtered to an eighteen-to-twenty-five-minute runtime, applies the parental-control filter as a hard constraint, reads watch history to avoid repeats, and returns three options with one-line reasons.
The user picks one, the agent starts playback, and the loop stops. Notice, no account tool was touched, so no confirmation was needed, the autonomy matched the low stakes. Now the contrast. If instead the user said "cancel my subscription," the agent would confirm explicitly, state the effective date and any retention offer, and only then act, and log it.
Same agent, completely different level of caution, because the stakes are different. Your target: seventy percent of sessions completed without a human handoff, zero parental-control violations, and wrong-action rate under three percent. And if the interviewer pushes with "what if the user asks it to do two things at once, watch something and also dispute a charge?" you've got it: the agent handles the low-stakes recommendation immediately, and routes the billing dispute to the confirmation-gated path, telling the user it's started both.
It doesn't let a complex request collapse its stop conditions. Here's what makes them lean in. First, that you defined the job, the tools with their permissions and costs, and the stop conditions as three separate, explicit things, instead of blurring them into "it helps users with stuff." Second, the step budget, a bounded loop, plus confirmation gates on high-stakes writes, because that's the mechanical heart of a safe agent.
And third, parental controls as a hard filter and a designed human handoff, which shows you think about the edges and the failure cases, not just the demo where everything goes right. Now the traps. The first trap is designing a chatbot and calling it an agent. It answers questions beautifully but it has no tools, no actions, and no stop conditions, so it can't actually do anything.
The second trap is letting the agent take high-stakes actions, cancel, change plan, refund, with no confirmation, which is how you get an agent that empties someone's account by accident. And the third trap is no bounded loop, so the agent can spin on tool calls forever with no exit, burning money and never answering. Name the step budget and you've dodged all three.
So let's assemble it. You pick viewer-facing and set the account boundary. You define the job in one sentence, and just as importantly, what it won't do. You list the tools, each with a permission and a cost, sorted by reversibility. You define the stop conditions, including an explicit step budget. You set the guardrails, parental controls as a hard filter, confirmation on money, grounded support answers, human handoff.
And you measure task completion with a hard zero on parental violations. Carry this one line into the room. An agent is a job plus tools plus stop conditions plus guardrails: name what it does, what it can call, when it must stop, and the hard limits it can never cross. ---