…
AI Product Case Questions

What's your understanding of the RAG framework?

A worked answer to a real AI PM interview question: explain your understanding of the RAG framework.

Transcript

Read the full transcript (1,214 words)

[INTERVIEWER] What's your understanding of the RAG framework? "What's your understanding of RAG?" RAG is retrieval augmented generation. In one line, you fetch relevant external text at query time and drop it into the prompt, so the model answers from that text instead of from its own memory. It's an open book exam, not a closed book one. The weak answer stops at "it looks stuff up." The strong answer walks the whole pipeline and names exactly where quality gets capped.

That last part is what separates you. Let me walk it. This is testing whether you understand the system well enough to know why answers go wrong and where to fix them. Because when a RAG system gives a bad answer, most people blame the model. Usually it's the retrieval. By the end of this you'll know the two phases cold, and you'll know which knob to turn when it breaks.

There are two phases. The first is indexing, and it happens offline, ahead of time. Step one, you ingest the documents, whatever your source of truth is. Step two, you chunk them, typically two hundred to five hundred tokens each, with a little overlap between neighbours so you don't slice a sentence cleanly in half. Chunking strategy matters more than people expect.

It isn't just a minor detail. Step three, you embed each chunk with an embedding model, something like text-embedding-3-small, which turns a chunk into a vector. That is a list of numbers, say fifteen hundred and thirty-six dimensions, that captures its meaning. Step four, you store those vectors in a database like Pinecone or pgvector, alongside metadata for filtering and access control.

That's the index. It sits there ready. The second phase is query, and it runs live per request. Step five, you embed the user's question with the same embedding model you used on the chunks, so they live in the same space. Step six, you retrieve the top K nearest chunks by cosine similarity, with K around five to twenty.

The better setups don't stop at plain vector search. They go hybrid, using dense vector search for meaning plus sparse keyword search like BM25 for exact terms, and then a cross encoder reranker to reorder the shortlist by real relevance. Step seven, you assemble the prompt with your system instruction, the retrieved chunks, and the user's question. Step eight, you generate.

The model answers grounded in those chunks, with citations. Step nine, optionally, you verify faithfulness before you return the answer, checking the output is actually supported by the chunks you retrieved. Now why do we do all this instead of just fine tuning the knowledge in? A few reasons matter to a product. Answers stay current because you update the index, not the model weights, so new information is live the moment you add the document.

Answers are grounded and citable, which directly cuts hallucination, because now there's a source behind every claim. Retrieval can respect access control, meaning you only fetch what this specific user is allowed to see, which fine tuning can't do. It's also far cheaper than fine tuning when what you need is knowledge, not a change in behaviour. For knowledge, reach for RAG.

For behaviour and tone, that's when you'd fine tune. Here's the tradeoff, and it's the sentence that makes you sound like you've built one. RAG adds retrieval latency and a chunk of infrastructure, the vector DB, the embedding pipeline, all of it. But the real catch is that its quality is capped by retrieval. If the retriever misses the right chunk and doesn't put it in front of the model, then no amount of model quality can save the answer.

The best model in the world can't reason about a passage it never received. Garbage in, grounded garbage out. Let me make it concrete. Picture an internal HR policy assistant sitting over ten thousand documents. You chunk at four hundred tokens with fifty tokens of overlap. You embed with text-embedding-3-small into pgvector. You retrieve the top twenty by hybrid search, then rerank down to the top five, and you prompt the model to answer only from these passages and cite the document ID.

Now here's the number that decides everything, which is retrieval recall. Specifically, how often the right passage actually lands in your top five. Say it's only there eighty percent of the time. That means twenty percent of your answers are ungrounded no matter how good your LLM is, because the source they needed simply wasn't in the prompt. Now you add the reranker, and recall at five climbs from around eighty percent to about ninety two percent.

That one change, the reranker, is the single biggest quality move in the entire system. Not a better model, but better retrieval. The follow up a good interviewer reaches for is RAG versus fine tuning, and when do you pick which. The clean line is RAG for knowledge, fine tuning for behaviour. If the problem is that the model doesn't know something, like current facts, your internal docs, or this customer's data, that's RAG.

This is because you can update the index the moment the fact changes and you never retrain. If the model doesn't act right, with the wrong tone or format, that's fine tuning, because you're changing behaviour, not knowledge. They often combine, where you fine tune for the house style and use RAG for the live facts. The other push is asking what you actually change when your retrieval keeps missing.

The order is to fix chunking first, because bad chunks poison everything downstream, then add hybrid search, then add the reranker, and then tune K. You don't reach for a bigger model, because as I said, the model can't reason about a passage it never received. So what makes them lean in? First, you covered both phases, indexing and query, not just the generation step, so it's clear you understand the whole machine.

Second, you know that retrieval quality caps the system, and you said you'd measure context recall, which is the metric that actually predicts whether the thing works. Third, you know hybrid search plus reranking, which means you're past the naive top K version everyone describes. Those three signals together read as someone who has debugged a real RAG pipeline. The traps are predictable.

Trap one is describing generation but skipping chunking, embedding, and the vector store, which leaves out the half of the system that usually breaks. Trap two is assuming a better LLM fixes bad retrieval. It cannot. If the chunk isn't in the prompt, the model never sees it, full stop. Trap three is treating chunk size and K as trivial details, when they're actually the main tuning knobs you'll spend most of your time on.

So, let's recap. Indexing offline means you ingest, chunk, embed, and store. Query online means you embed the question, retrieve top K, rerank, ground the prompt, generate with citations, and optionally verify. RAG earns its place because answers stay current, grounded, access controlled, and cheap. The whole thing is capped by retrieval. The one line to carry in is that RAG is chunk, embed, retrieve top K, rerank, ground the prompt, cite, and its quality is capped by retrieval, so measure context recall first.

Keep learning