GenAI System Design Tradeoffs: Interview Questions
The tradeoff side of GenAI system design: latency, token cost, choosing a model, build vs buy vs fine-tuning, rollouts, model migrations and incidents. AI engineers meet these in system design rounds and AI product managers in strategy rounds. They cover the decisions, not architecture diagrams or code. 162 questions across 7 subtopics, each with a model answer.
LLM Latency and UX Tradeoffs
Why latency targets are set at the 95th percentile, when a fast weaker model beats a slow strong one, and why voice interfaces need stricter limits. A common system design topic for AI engineers and PMs.
- What is a latency budget and how would you allocate one across a RAG pipeline?
- Why do you set latency targets at the 95th percentile rather than the mean?
- Describe how streaming changes perceived latency without changing actual latency.
- At what point does latency stop mattering and quality take over?
- How would you decide between a fast weak model and a slow strong one for autocomplete?
- Explain the UX options available when a response will take 30 seconds.
- Describe the latency requirements for a voice interface and why they are stricter.
- How do reasoning models complicate a latency budget?
LLM Cost Modeling and Unit Economics
How to work out what an AI feature costs: cost per interaction from token counts, how RAG changes the cost structure, and monthly cost at scale. Expect calculation questions in AI PM and AI engineer interviews.
- Build the cost-per-interaction model for a feature with a 2,000-token prompt and a 500-token response.
- What cost drivers exist for an AI feature beyond model tokens?
- Explain how a RAG pipeline's cost structure differs from a single model call.
- How does prompt caching change your unit economics, and when does it not help?
- Model the monthly cost of a feature used by 50,000 users averaging 12 interactions each.
- What is the cost impact of moving from a single call to a five-step agent?
- Describe how you would find the most expensive one percent of your traffic.
- Explain the gross margin problem for AI products with flat-rate pricing.
Choosing an LLM: Model Selection
How to pick a model for a product: quality against price, why benchmark leaderboards are a weak guide, what a bigger context window really buys, and how to compare models on your own use case. Asked of AI PMs and AI engineers.
- What are the five dimensions a PM should compare models on before a product decision?
- Why is benchmark leaderboard position a weak input to model selection?
- Describe how you would build a model bake-off for your specific use case.
- How do you weigh a model that is 20 percent better and three times more expensive?
- Explain when a smaller, faster model beats the frontier model in a product.
- What does context window size actually buy you in product terms?
- How do you evaluate a model's reliability on structured output for your schema?
- Describe a test set you would build to choose between two candidate models.
Build vs Buy vs Fine-Tuning Decisions
When to use an API, when to buy a product and when fine-tuning is worth it: data needs, cost predictability, reversibility and lock in. A favourite question for AI PMs and a common system design discussion for AI engineers.
- Lay out the decision tree for build, buy, prompt, retrieve or fine-tune.
- What conditions make buying a vendor solution the right call for an AI feature?
- Explain the strategic risk of building your core differentiator on a third-party model.
- When does fine-tuning pay for itself relative to a better prompt?
- Your vendor's pricing changes annually. How does that affect a build-versus-buy analysis?
- Describe the total cost of ownership of a self-hosted open model versus an API.
- What is the switching cost of a fine-tuned model versus a prompted one?
- A vendor offers 90 percent of what you need. How do you evaluate closing the last 10 percent yourself?
AI Rollout Strategy and Phased Launches
How to launch an AI feature safely: canary percentages, phased rollouts to large user bases, and why passing evals is not a reason to ship to everyone on day one. Asked of AI PMs and AI engineers.
- Design the rollout plan for an AI feature going to two million users.
- What percentage would you start a canary at, and how do you decide?
- Explain the difference between a feature flag rollout and a model rollout.
- What metrics gate each stage of a phased rollout?
- How do you choose which users go first?
- Describe the rollback criteria you would set before launch.
- How long should each rollout phase last, and what determines it?
- Explain how rollout differs when quality varies by user segment.
LLM Model Migration and Version Changes
What to do when the model behind your product changes: why a better model can still be a bad migration, planning for a deprecation, and what to tell users when behaviour shifts.
- Your provider deprecates the model behind your main feature in 60 days. Write the plan.
- How do you test a replacement model against the behaviour users have come to expect?
- Explain why a strictly better model can still be a bad migration.
- What should you tell users when model behaviour changes underneath them?
- Describe a dual-running strategy for a model migration.
- How do you handle customers who tuned their prompts to the old model?
- What contractual commitments should you avoid making about model behaviour?
- Design the regression suite you would run before any model swap.
AI Incident Management
What counts as an incident for an AI feature, how to define severity for quality problems, and how to choose between switching a feature off and degrading it. Relevant to AI PMs and AI engineers on call.
- What counts as an incident for an AI feature but not for a normal one?
- Write the severity definitions for AI quality incidents.
- Your model starts producing offensive output. Describe the first hour.
- How do you triage an incident where the code is fine and the model is the problem?
- What is the AI equivalent of a rollback, and when is it not available?
- Describe the on-call runbook entry for a sudden quality drop.
- How do you decide whether to disable a feature or degrade it during an incident?
- Explain how you would investigate a complaint you cannot reproduce.