How would you build an LLM inference pipeline?
Transcript
Read the full transcript (1,505 words)
[INTERVIEWER] How would you build an LLM inference pipeline? "How would you build an LLM inference pipeline?" This question rewards structure. You need to do three things. First, walk the request path: tokenize, batch, prefill and KV cache, decode, and stream. Second, name the levers: batching, quantization, caching, and speculative decoding. Finally, land the whole thing on the single tradeoff that governs everything, which is latency versus throughput versus cost, all pulled by one dial, batch size.
Do those three in that order, and you sound like you have actually run inference. Let me walk you through it. For a PM, this is really testing whether you understand that these infrastructure knobs are product decisions in disguise. Time to first token is a conversion lever. Cost per token is your gross margin. If you can connect the plumbing to the P&L, you have answered the question they are actually asking.
So let's get the plumbing right first, then tie it back. Walk the path. A request comes in. The prompt, plus parameters like max tokens and temperature, hits an API gateway that handles authentication and rate limiting. Then you tokenize. The text becomes token IDs using something like BPE, or byte pair encoding. Next is scheduling and batching. Requests queue up and get batched together onto the GPU.
Here is the first big idea, which is continuous, or in flight, batching. Instead of static batches where everyone waits for the slowest request, you add and evict requests every single step. This keeps the GPU full. That is the single biggest throughput lever in the whole stack. Then comes prefill. The entire prompt runs through the model in one forward pass, building the KV cache for every prompt token and producing that first output token.
Prefill is compute bound. It scales with how long your prompt is, and it sets your time to first token, your TTFT. Then comes decode, and this is a different regime entirely. Generation is autoregressive, one token at a time, and each step reads the KV cache and appends to it. Decode is memory bandwidth bound, not compute bound, and it sets your inter token latency and your tokens per second.
That prefill versus decode distinction is worth stating out loud because they have different bottlenecks and different fixes. Now consider the KV cache itself, which is the thing they will probe you on. It stores the keys and values for all the previous tokens so you never have to recompute them. It grows with sequence length times batch size. It is the main constraint on GPU memory, and it is usually what limits how big a batch you can run.
PagedAttention, which is the trick in vLLM, pages that cache like an operating system pages memory. This cuts fragmentation and lets you fit a bigger batch in the same GPU. Then you stream. Tokens go back over server sent events as they are generated, so the perceived latency is your TTFT, not the full completion time. Finally, you stop on an end of sequence token or by hitting max tokens, and you return and log.
Now for the levers, which are the knobs you turn to hit your targets. First is quantization. You take the weights from FP16 down to FP8, INT8, or INT4. Smaller weights mean less memory, which means a bigger batch, which makes it faster and cheaper. The cost is a small quality drop. Concretely, FP8 on an H100 roughly doubles throughput versus FP16 with only minor quality loss, which is why it has become a default.
Second is KV cache optimisation. This is PagedAttention plus prefix or prompt caching. You reuse the KV for a shared prefix, like your system prompt or your RAG context, across many requests, and skip prefilling it again every time. If 10,000 requests share the same long system prompt, you compute it once. Third is continuous batching, the throughput workhorse I already mentioned.
Fourth is speculative decoding, which is quite clever. A small, fast draft model proposes several tokens ahead, the big model verifies them all in one pass, and any accepted tokens skip a full forward pass each. That gives roughly two times lower latency with no quality change. The no quality change part matters because the verification step keeps the output distribution mathematically identical to the big model running alone.
Beyond those, you have model parallelism, tensor and pipeline parallel for models too big for one GPU, FlashAttention kernels for faster attention maths, distillation or a smaller mixture of experts model for raw cost, and a semantic cache for repeated queries. Consider the tradeoff that governs the whole thing. This is where you show you understand it as a system.
Latency versus throughput versus cost, all pulled by batch size. Turn the batch size up, and GPU utilisation goes up, throughput goes up, and cost per token goes down. That is all good. But every individual request now waits longer in the batch, so latency goes up. That is the tension. Interactive chat, where a human is waiting, optimises for TTFT and inter token latency.
This means smaller batches, speculative decoding, and streaming. Bulk or offline jobs, where nobody is watching in real time, optimise for throughput and cost, which means you max out the batch. You tune to the SLA. You do not pick one setting and use it for both because the two workloads want opposite ends of the same dial. Here is the PM translation, because everything I just described looks like infrastructure but reads as product.
TTFT is a conversion and retention lever. If the model takes 3 seconds to start talking, people leave. Cost per token is literally your gross margin on the feature. The batch size dial trades one directly against the other. The division of labour is clean. The PM sets the latency SLA and the cost ceiling, which are the business constraints, and the infrastructure team tunes batching and quantization to hit both.
You own the targets, and they own the knobs. But you have to know the knobs exist to set targets that are actually achievable. Let me make it concrete. A chat product on a Llama 70B class model, running on H100s, served with vLLM. You go FP8 on the weights, continuous batching, prefix caching on the shared system prompt and the RAG context, and speculative decoding with a small 7B draft model, all streaming over server sent events.
Your SLA is p95 time to first token under 400 milliseconds, and at least 40 tokens per second, which is comfortably faster than a person reads. Now watch the stack build up. Continuous batching plus PagedAttention has been reported at up to around 24 times the throughput of naive HuggingFace Transformers serving. FP8 adds roughly another two times. Speculative decoding roughly halves your latency on top.
Those levers stack, and they are not mutually exclusive. Separately, your offline evaluation jobs run on a different pool entirely, one tuned for high batch and cost, not for latency, because nobody is waiting on those. It is the same model with two completely different serving configurations because the two workloads have different SLAs. That is the answer that shows you get it.
What makes them lean in? First, you walked the actual path. You distinguished prefill from decode and you knew what the KV cache is for, instead of hand waving. Second, you separated TTFT from throughput and you knew which levers move which. Speculative decoding is for latency, and batching is for throughput. Third, you landed on the batch size tradeoff and you tied it to both an SLA and to gross margin.
That is the move that turns an infrastructure answer into a product answer. That connection is the whole reason a PM gets asked this. The traps are easy to fall into on a technical question like this. Trap one is listing buzzwords like quantization, vLLM, and speculative decoding, with no request path connecting them. This makes it clear you have heard the words but never traced a request.
Trap two is missing that prefill and decode are different regimes with different bottlenecks, one compute bound and one memory bound. Trap three is optimising throughput on an interactive product, or latency on a batch job. This happens whenever you never bothered to set an SLA. The SLA is the thing that tells you which way to turn the dial.
Let us recap. The path is tokenize, batch, prefill into the KV cache, decode, and stream. The levers are continuous batching for throughput, quantization for memory and speed, caching to skip repeated work, and speculative decoding for latency. The governing tradeoff is that batch size trades latency against throughput and cost, so you tune it to the SLA. Interactive wants small, and batch wants large.
The one line to carry in is this. Tokenize, batch, prefill into the KV cache, decode, and stream. The levers are batching, quantization, caching, and speculative decoding, with batch size trading latency against throughput and cost.