Are you familiar with RLHF? What do you know about DPO?
Transcript
Read the full transcript (1,178 words)
[INTERVIEWER] Are you familiar with RLHF? What do you know about DPO? "Do you know RLHF? What about DPO?" Both are preference tuning methods. Both do the same job, which is making a model prefer the kind of output humans actually like. RLHF learns a separate reward model and then optimises against it with reinforcement learning. DPO skips the reward model and the RL loop entirely, and optimises straight on the preference pairs.
The weak answer explains RLHF and then treats DPO as a vague synonym. The strong answer says what specific problem each one solves. Let me give you that. You're almost never going to pick the training algorithm yourself, as that's down to the research team. So why do they ask? Because it tests whether you understand where alignment quality comes from, which decides where your actual control lies as a PM.
Get this and you'll know the one thing you actually own in this process, and it's not the algorithm. So set the context first. A model gets built in stages. Pretraining teaches it to predict the next token over a huge pile of text. Supervised fine tuning then teaches it to imitate good demonstrations. But after both of those, the model still doesn't reliably prefer answers that are helpful, harmless, and the kind people actually want.
That final step, teaching it human preference, is preference optimisation. And that's the job both RLHF and DPO are doing. Same goal, two routes. Naming the shared problem is the framing that makes you sound like you understand it rather than just recite it. A good answer is really hard to write down as a labelled target. What's the single correct response to "cheer me up"?
There isn't one. But here's the thing humans can do easily. Show a person two answers and they'll happily tell you which is better. So both methods learn from pairwise comparisons, A versus B, which is preferred, not from gold standard answers. That's the trick underneath both of them. RLHF, the classic recipe, is three steps. You start from the supervised fine tuned model.
Then step one, you train a reward model. You collect a big pile of human comparisons, A versus B, and you train a model to predict which one a human would prefer and output that as a scalar reward, a single number for how good. Step two, you optimise with reinforcement learning, usually PPO. You fine tune the language model to maximise that reward, but with a KL penalty that keeps it close to where it started, so it doesn't drift off into gibberish or learn to game the reward model.
What problem does it solve? It turns fuzzy human preference into a concrete, trainable signal. This is how InstructGPT and ChatGPT were aligned. The cost is that it's complex and unstable. You've got multiple models in the loop at once, namely the policy, a value model, the reward model, and the reference, plus a finicky RL step that's genuinely hard to get right.
Then DPO, from Rafailov and colleagues in 2023. The insight is genuinely elegant. They proved the RLHF objective has a closed form solution, which means you can optimise the policy directly on the preference pairs with a simple classification style loss, using the starting model as a reference. No separate reward model. No PPO. One training stage instead of the whole apparatus.
So the problem DPO solves is precisely the instability, cost, and complexity of RLHF. It's simpler, it's more stable, and it's a lot cheaper to reproduce. That's why so many teams reached for it the moment it landed. Declaring a winner is a mistake, because that's the giveaway. Name the tradeoff. DPO is simpler, but it's offline. It trains on a fixed, frozen dataset of preference pairs, and it never samples fresh outputs from the current policy as it learns.
The online exploration in RLHF, plus that explicit reward model, can capture things a frozen dataset just misses, as the model learns from its own new mistakes. That's exactly why frontier teams still use RLHF, or online variants of these ideas, rather than treating DPO as the final word. Both are in active use. It's not a settled win, and saying that shows you actually understand the state of the field.
And here's the part that's actually yours. You rarely pick the algorithm, but you own the preference data and the evals. Alignment quality is bounded by the coverage and quality of your comparison labels. If your labels don't cover a scenario, the model won't be aligned for it, and no algorithm fixes that. So cheaper alignment, which is what DPO buys you, means you can preference tune on your own domain data faster and more often.
The true advantage is in the data. That's the line to remember. Grounding this in a real scenario helps. Say you're aligning a support model to the tone of your brand and your refusal policy. With RLHF, you collect, say, fifty thousand human A and B comparisons, you train a reward model on them, then you run PPO. That's two model training stages, plus a separate reward model you have to build and maintain.
With DPO, you take those exact same fifty thousand comparison pairs and you train the policy directly on them, which removes the reward model and the whole RL stage. Roughly, you've halved the number of moving parts and the compute in your alignment pipeline. And the original DPO paper showed this matching or beating PPO based RLHF on summarisation and dialogue preference win rate.
Same result, half the machinery. That's the case for it, and also the reason it spread so fast. Hiring managers lean in when you do three things. First, you stated the shared problem, learning from comparisons, before you contrasted the two methods, which shows structure. Second, you know the actual insight of DPO, that it removes the reward model and the RL loop, not just the hand wave that it's simpler.
And third, you named the real tradeoff, that DPO is offline, instead of crowning a winner. Those three read as genuine understanding rather than a memorised summary. The traps are quick. Trap one, explaining RLHF fluently and then treating DPO as a fuzzy synonym for it. Trap two, saying DPO is always better, when it's simpler but offline only and RLHF still earns its place.
And trap three, the PM specific one, missing that your advantage is in the preference data, not the algorithm choice. To recap the whole picture. Both make the model prefer what humans like, and both learn from pairwise comparisons. RLHF trains a reward model, then optimises against it with RL and a KL penalty. DPO drops both of those and trains straight on the pairs, simpler and more stable but offline.
And your job is the data. The core takeaway is that both learn from human comparisons, RLHF trains a reward model then optimises with RL, while DPO drops both and trains on the pairs directly, making it simpler but offline.