Fine-Tuning LLMs
Fine-tuning LLMs is the process of continuing the training of a pre-trained large language model on a smaller, task-specific dataset so that its behaviour changes permanently. It adapts a general model to a particular task, output format, vocabulary or style without training a new model from the beginning.
- Pre-trained base: It starts from a model that has already learned language and code from a large general corpus.
- Task dataset: It uses hundreds to thousands of input and output examples that demonstrate the desired behaviour.
- Weight updates: It changes the model's parameters, unlike prompt engineering, which changes only the input.
- Parameter-efficient methods: Techniques such as LoRA train a small set of additional weights instead of the whole model.
- Separate from knowledge updates: It teaches behaviour and format more reliably than it teaches new facts.
For example, a team can fine-tune a model on its own reviewed pull requests so that generated diff summaries follow the team's exact format and terminology.
Key Characteristics of Fine-Tuning LLMs
- Supervised fine-tuning: The model learns from labelled pairs, where each input has a reference output written or approved by people.
- Full fine-tuning: Every parameter is updated, which requires substantial GPU memory and produces a complete copy of the model.
- LoRA adapters: Low-rank adaptation freezes the original weights and trains small added matrices, which reduces memory and storage requirements.
- Preference tuning: Methods such as reinforcement learning from human feedback adjust the model using rankings of better and worse answers.
- Data quality dominance: A small, consistent and accurate dataset usually outperforms a large, noisy one.
- Model access requirement: LLM fine-tuning needs open weights or a provider that offers tuning for its hosted models, as described in open source vs closed source LLMs.
How to Fine-Tune an LLM
- Define the task: Specify the exact input and the expected output, such as a code diff and a concise two-sentence summary.
- Collect examples: Gather representative inputs and reviewed outputs, then remove duplicates, empty records and confidential information.
- Format the dataset: Convert each example into the format the training tool expects, commonly JSON Lines with chat messages.
- Split the data: Reserve a validation set that the model never trains on, so that improvement can be measured objectively.
- Train: Run several passes, called epochs, in which the model predicts the reference outputs and its weights are adjusted to reduce errors.
- Evaluate: Compare the tuned model with the original base model on the validation set and on realistic manual tests.
- Deploy and monitor: Serve the tuned model and continuously track output quality, because data distributions and requirements change over time.
Derivation of the LoRA Update
LoRA (low-rank adaptation) rests on one observation, namely that the change fine-tuning makes to a weight matrix can be approximated by the product of two narrow matrices. The original matrix remains frozen, and training updates only those two narrow matrices.
- : the frozen pre-trained weight matrix, which has rows and columns.
- : the adapted matrix that the tuned model applies during inference.
- : the two small trainable matrices, whose product has exactly the same dimensions as .
- : the LoRA rank, the shared inner dimension of and , which is typically far smaller than and .
- : a scaling constant that determines how strongly the learned update influences the original weights.
In other words, the tuned layer equals the original layer plus a correction, and that correction is constructed from a small number of column vectors in and row vectors in . Counting the parameters that training must learn reveals the saving:
Worked example. Each attention projection in a 7-billion-parameter model such as Llama 2 7B is a 4096 × 4096 matrix, which is a realistic size for this calculation. With :
- Full fine-tuning: trainable parameters.
- LoRA: trainable parameters.
- Ratio: , so LoRA trains 256 times fewer parameters, approximately 0.39 percent of the complete layer.
- Miniature verification: For a 3 × 3 identity matrix with , , and , the update modifies only rows 1 and 3, where the entries of are nonzero.
The program below calculates the miniature update and then compares the parameter counts for several different values of the rank.
# LoRA update W' = W + (alpha / r) * B @ A on a tiny 3 x 3 layer with rank r = 1.
def matmul(X, Y):
return [[sum(X[i][t] * Y[t][j] for t in range(len(Y))) for j in range(len(Y[0]))] for i in range(len(X))]
W = [[1, 0, 0], [0, 1, 0], [0, 0, 1]] # frozen pre-trained weights (d = k = 3)
B = [[1], [0], [2]] # d x r, trained
A = [[1, -1, 0]] # r x k, trained
alpha, r = 2, 1
delta = matmul(B, A)
W_new = [[W[i][j] + alpha / r * delta[i][j] for j in range(3)] for i in range(3)]
print("B @ A =", delta)
print("W' =", [[int(x) for x in row] for row in W_new])
# Trainable parameters for one 4096 x 4096 projection matrix.
d = k = 4096
full = d * k
print(f"\nfull fine-tuning: {full:,} parameters")
for rank in (4, 8, 16, 64):
lora = rank * (d + k)
print(f"LoRA r={rank:<3}: {lora:>9,} parameters ({lora / full:.2%} of full, {full // lora}x fewer)")B @ A = [[1, -1, 0], [0, 0, 0], [2, -2, 0]]
W' = [[3, -2, 0], [0, 1, 0], [4, -4, 1]]
full fine-tuning: 16,777,216 parameters
LoRA r=4 : 32,768 parameters (0.20% of full, 512x fewer)
LoRA r=8 : 65,536 parameters (0.39% of full, 256x fewer)
LoRA r=16 : 131,072 parameters (0.78% of full, 128x fewer)
LoRA r=64 : 524,288 parameters (3.12% of full, 32x fewer)- Rank sets the budget: Doubling the LoRA rank doubles the number of trainable parameters, yet even a rank of 64 represents only about 3 percent of the full matrix.
- B starts at zero: The matrix is initialised with zeros, so the adapted weights initially equal the original weights and training begins from the base model's behaviour.
- Merging: After training, the scaled product can be added permanently into , so the tuned model introduces no additional inference latency.
- QLoRA: A related method stores the frozen weights in 4-bit precision and trains the LoRA matrices on top, which reduces GPU memory requirements even further.
In practice, the LoRA rank is the principal tuning decision, because a higher rank can capture a larger behavioural change but requires more memory and produces a larger adapter file.
Example: Preparing Fine-Tuning Data in Python
The program below converts reviewed diff summaries into chat-format training records, drops invalid examples and creates a validation split.
import json
# Team examples: a diff summary written the way reviewers want it
examples = [
("def add(a, b): return a+b -> def add(a: int, b: int) -> int: return a + b",
"Adds type hints to add(). No behaviour change."),
("retries = 3 -> retries = 5",
"Raises retry count from 3 to 5. Check timeout budget."),
("except Exception: pass -> except ValueError as e: log.warning(e)",
"Narrows the except clause and logs the error instead of hiding it."),
("", "Empty diff."),
]
SYSTEM = "Summarise the diff in one or two sentences for a code reviewer."
records, skipped = [], 0
for diff, summary in examples:
if not diff.strip(): # drop examples with no input
skipped += 1
continue
records.append({"messages": [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": diff},
{"role": "assistant", "content": summary},
]})
# Hold back the last record to check the tuned model on unseen data
train, validation = records[:-1], records[-1:]
jsonl = "\n".join(json.dumps(r) for r in train) # contents of train.jsonl
print(f"kept {len(records)}, skipped {skipped}")
print(f"train {len(train)}, validation {len(validation)}")
print("roles:", [m["role"] for m in train[0]["messages"]])
print("target:", train[0]["messages"][2]["content"])kept 3, skipped 1
train 2, validation 1
roles: ['system', 'user', 'assistant']
target: Adds type hints to add(). No behaviour change.- Data preparation is most of the work: Training itself is a single command or API call, but the examples decide what the model learns.
- Assistant message as target: The model is trained to reproduce the assistant content, so every summary must already be correct.
- Scale in practice: A real dataset contains hundreds of records or more, and training runs on GPUs through a training library or a provider's tuning service.
Applications of Fine-Tuning LLMs
- Consistent output formats: Pull request summaries, commit messages or JSON responses that always follow one template.
- Internal code conventions: Completions that follow an organisation's frameworks, naming conventions and architectural patterns.
- Domain vocabulary: Legal, medical or financial terminology that general models handle poorly.
- Classification: Categorising bug reports, log entries or support tickets into predefined labels.
- Smaller replacement models: Training small language models to match a larger model on one narrow task.
Advantages
- Shorter prompts: Instructions and examples move into the weights, which reduces tokens and latency per request.
- Consistent behaviour: Output style and format become more reliable than with prompting alone.
- Lower serving cost: A tuned small model can replace a larger general model for a specific task.
- Private customisation: Open-weight models can be tuned entirely on internal hardware.
Limitations
- Poor fit for changing facts: Frequently updated information is better supplied at query time through RAG, as compared in RAG vs fine-tuning.
- Data effort: High-quality labelled examples are expensive and time-consuming to produce and review.
- Catastrophic forgetting: Narrow training can degrade general abilities that the base model previously had.
- Maintenance: Each new base model version requires the entire tuning process to be repeated and re-evaluated.
- Hallucination remains: A tuned model can still produce confident errors, as covered in LLM hallucination.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What does fine-tuning change in an LLM?
Frequently Asked Questions
Fine-tuning vs prompting: which should be tried first?
Prompting should usually be tried first, because it needs no training data and changes take effect immediately. Fine-tuning is worth its cost when prompts cannot make the output format or style consistent enough.
How much data is needed to fine-tune an LLM?
Narrow tasks such as a fixed output format can improve with a few hundred high-quality examples. Broader behaviour changes need more data, and quality matters more than quantity.
What is LoRA fine-tuning?
LoRA, or low-rank adaptation, freezes the original model weights and trains small additional matrices. It needs far less GPU memory than full fine-tuning and produces a small adapter file.
Can closed models such as GPT or Claude be fine-tuned?
Some providers offer fine-tuning for selected hosted models through their platforms, while others do not. Availability changes over time, so the provider's current documentation is the reference.
Does fine-tuning add new knowledge to a model?
Fine-tuning can add some knowledge, but it is more reliable for teaching behaviour, style and format. Facts that change often are better supplied through retrieval at query time.
Related Articles
- RAG vs Fine-TuningRAG vs fine-tuning compared: what each changes, cost, freshness and accuracy trade-offs, when to use each, and one API docs example handled both ways.
- Small Language Models (SLM)Learn what small language models are, how quantization and distillation shrink them, where SLMs are used, with a Python memory estimate for local models.
- Open Source vs Closed Source LLMsCompare open source vs closed source LLM options on data control, cost, customisation and setup, with a comparison table and a pull request example.
- What is Prompt EngineeringLearn what prompt engineering is, how to design a prompt step by step, its key characteristics, uses and limits, with a Python code review prompt example.
- What is a Large Language Model (LLM)Learn what a large language model is, how an LLM predicts the next token, its key characteristics, uses and limits, with a Python example and diagram.
- How LLMs WorkLearn how LLMs work from training to inference: tokenization, embeddings, attention layers and decoding, with a Python trace of one forward pass.