…
AI Product Case Questions

What's the effect of adjusting an LLM's context window size?

A worked answer to a real AI PM interview question: what changes when you adjust an LLM's context window size?

Transcript

Read the full transcript (980 words)

[INTERVIEWER] What's the effect of adjusting an LLM's context window size? Changing the context window size seems straightforward. The easy answer, the one that gets you a polite nod and nothing else, is that bigger means it can hold more. True, but shallow. The strong answer says bigger isn't free, and it isn't always better. It costs more on every call, it runs slower, and here's the kicker, the model doesn't reliably use the middle of it anyway.

Let me show you why that last part matters most. This is a knowledge check, but it's testing whether you'd feel the cost of a design choice in the unit economics. A product manager who thinks a bigger window is a free upgrade will happily ship something that quietly burns money and gets worse over time. By the end of this section, you'll be able to reason about window size as a lever you tune, not just a number you max out.

Let's define it cleanly first. The context window is the maximum number of tokens the model can attend to in a single request, and that's input and output combined. A token is roughly four characters, or about three-quarters of a word. For reference points you can drop in the room, GPT-4o sits around 128,000 tokens, Claude around 200,000 with some variants at a million, and Gemini 1.5 goes up to a million or more.

Think of it as the working memory for that one specific request. Nothing carries over between requests unless you explicitly put it back in. Consider what growing the window actually buys you. Real things. Longer documents fit in one go. More retrieved chunks fit, so your retrieval augmented generation can pull in more context. More chat history fits, so the conversation remembers further back.

And more few-shot examples fit, so you can steer it harder. That upside is genuine, and you should say so before you pivot to the catch. Don't argue that big windows are useless. Argue that they're not free. Here's the catch, and there are three parts to it. First, cost. You pay per input token, every single call. And in the vanilla case, attention scales roughly quadratically with sequence length, so filling a 128,000 token window is dramatically more expensive than a 4,000 token prompt.

Second, latency. A longer prompt means a longer prefill pass, which means your time to first token goes up, and the user waits longer to see anything. And third, the one people miss, quality. There's a well-known result called Lost in the Middle from Liu and colleagues in 2023. Models recall facts placed at the very start or the very end of a long context far better than facts buried in the middle.

So a bigger window does not mean the model actually reads all of it well. Dump in a load of irrelevant text and you dilute the signal, you don't strengthen it. The product conclusion writes itself. Don't stuff the whole corpus in just because it fits. Retrieve and rank the most relevant chunks first. Put your most important content at the edges of the prompt, the start and the end, where the model actually attends.

And cap your history rather than letting it grow forever. Window size is a cost and quality lever you tune to the job, not a free capacity upgrade you crank to the top. Let me put numbers on it, because that's what makes it land. Say a frontier model charges about two dollars fifty per million input tokens. A 128,000 token prompt is then roughly thirty-two cents in input alone, per call.

Now run that on every request in a support product doing 100,000 calls a day, and you're at around 32,000 dollars a day just on input. A tight 4,000 token RAG prompt, by contrast, is about a cent a call. That's exactly thirty-two times cheaper. And here's the part that surprises people. The small prompt often scores higher, because the model isn't hunting through 128,000 tokens for the one fact that matters.

On buried fact tasks, that lost in the middle effect can drop accuracy by twenty points or more versus just putting the fact at the top. So the stuffed window cost you thirty-two times the money and lost you accuracy. Both directions, worse. Interviewers are looking for three specific signals. You gave the full three-way tradeoff, cost, latency, and quality, not just that it holds more.

You knew about the lost in the middle effect, which means you don't assume the model reads the whole window evenly. And you drew the actual product conclusion, retrieve and rank rather than stuff, and backed it with cost math. That last bit, the math, is what tells them you'd feel this in the profit and loss statement. The traps are quick.

Trap one is thinking bigger is always better. No. It costs more, runs slower, and can recall worse. Trap two is forgetting that output tokens share the window with input, so your generation budget eats into your context budget. And trap three is giving the answer with no cost math at all, so the interviewer can't tell whether you'd ever notice this on the bill.

Numbers are what separate a product manager from someone who just read a blog post. Let's recap. The context window is the per-request working memory in tokens. Growing it fits more in, which is genuinely useful. But it costs more, runs slower, and thanks to the lost in the middle effect, it doesn't use all of that space well. So you retrieve and rank rather than stuff, and you park the important stuff at the edges.

The one line to carry in is that a bigger context window fits more but costs more, runs slower, and gets lost in the middle, so retrieve the right chunks and don't stuff the whole corpus.

Keep learning