…
Skip to content
Topics
On this page

Fine-Tuning LLMs

Fine-tuning LLMs is the process of continuing the training of a pre-trained large language model on a smaller, task-specific dataset so that its behaviour changes permanently. It adapts a general model to a particular task, output format, vocabulary or style without training a new model from the beginning.

  • Pre-trained base: It starts from a model that has already learned language and code from a large general corpus.
  • Task dataset: It uses hundreds to thousands of input and output examples that demonstrate the desired behaviour.
  • Weight updates: It changes the model's parameters, unlike prompt engineering, which changes only the input.
  • Parameter-efficient methods: Techniques such as LoRA train a small set of additional weights instead of the whole model.
  • Separate from knowledge updates: It teaches behaviour and format more reliably than it teaches new facts.
How fine-tuning turns a base LLM into a tuned modelA pre-trained base model and a train.jsonl file of reviewed pull request summaries both feed a training step, which adjusts the model weights over several epochs. The result is a tuned model that writes diff summaries in the team's format. A held-back validation set is used to check the tuned model. The flow is illustrative.Base modelpre-trained weightsTeam examplestrain.jsonlTrainingadjusts weightsTuned modelsummaries in team formatchecked on held-back validation set
How fine-tuning turns a base LLM into a tuned model

For example, a team can fine-tune a model on its own reviewed pull requests so that generated diff summaries follow the team's exact format and terminology.

Key Characteristics of Fine-Tuning LLMs

  • Supervised fine-tuning: The model learns from labelled pairs, where each input has a reference output written or approved by people.
  • Full fine-tuning: Every parameter is updated, which requires substantial GPU memory and produces a complete copy of the model.
  • LoRA adapters: Low-rank adaptation freezes the original weights and trains small added matrices, which reduces memory and storage requirements.
  • Preference tuning: Methods such as reinforcement learning from human feedback adjust the model using rankings of better and worse answers.
  • Data quality dominance: A small, consistent and accurate dataset usually outperforms a large, noisy one.
  • Model access requirement: LLM fine-tuning needs open weights or a provider that offers tuning for its hosted models, as described in open source vs closed source LLMs.

How to Fine-Tune an LLM

  1. Define the task: Specify the exact input and the expected output, such as a code diff and a concise two-sentence summary.
  2. Collect examples: Gather representative inputs and reviewed outputs, then remove duplicates, empty records and confidential information.
  3. Format the dataset: Convert each example into the format the training tool expects, commonly JSON Lines with chat messages.
  4. Split the data: Reserve a validation set that the model never trains on, so that improvement can be measured objectively.
  5. Train: Run several passes, called epochs, in which the model predicts the reference outputs and its weights are adjusted to reduce errors.
  6. Evaluate: Compare the tuned model with the original base model on the validation set and on realistic manual tests.
  7. Deploy and monitor: Serve the tuned model and continuously track output quality, because data distributions and requirements change over time.

Derivation of the LoRA Update

LoRA (low-rank adaptation) rests on one observation, namely that the change fine-tuning makes to a weight matrix can be approximated by the product of two narrow matrices. The original matrix remains frozen, and training updates only those two narrow matrices.

  • : the frozen pre-trained weight matrix, which has rows and columns.
  • : the adapted matrix that the tuned model applies during inference.
  • : the two small trainable matrices, whose product has exactly the same dimensions as .
  • : the LoRA rank, the shared inner dimension of and , which is typically far smaller than and .
  • : a scaling constant that determines how strongly the learned update influences the original weights.

In other words, the tuned layer equals the original layer plus a correction, and that correction is constructed from a small number of column vectors in and row vectors in . Counting the parameters that training must learn reveals the saving:

Worked example. Each attention projection in a 7-billion-parameter model such as Llama 2 7B is a 4096 × 4096 matrix, which is a realistic size for this calculation. With :

  1. Full fine-tuning: trainable parameters.
  2. LoRA: trainable parameters.
  3. Ratio: , so LoRA trains 256 times fewer parameters, approximately 0.39 percent of the complete layer.
  4. Miniature verification: For a 3 × 3 identity matrix with , , and , the update modifies only rows 1 and 3, where the entries of are nonzero.

The program below calculates the miniature update and then compares the parameter counts for several different values of the rank.

Python
# LoRA update W' = W + (alpha / r) * B @ A on a tiny 3 x 3 layer with rank r = 1.
def matmul(X, Y):
    return [[sum(X[i][t] * Y[t][j] for t in range(len(Y))) for j in range(len(Y[0]))] for i in range(len(X))]

W = [[1, 0, 0], [0, 1, 0], [0, 0, 1]]  # frozen pre-trained weights (d = k = 3)
B = [[1], [0], [2]]                    # d x r, trained
A = [[1, -1, 0]]                       # r x k, trained
alpha, r = 2, 1

delta = matmul(B, A)
W_new = [[W[i][j] + alpha / r * delta[i][j] for j in range(3)] for i in range(3)]
print("B @ A =", delta)
print("W'    =", [[int(x) for x in row] for row in W_new])

# Trainable parameters for one 4096 x 4096 projection matrix.
d = k = 4096
full = d * k
print(f"\nfull fine-tuning: {full:,} parameters")
for rank in (4, 8, 16, 64):
    lora = rank * (d + k)
    print(f"LoRA r={rank:<3}: {lora:>9,} parameters ({lora / full:.2%} of full, {full // lora}x fewer)")
Output
B @ A = [[1, -1, 0], [0, 0, 0], [2, -2, 0]]
W'    = [[3, -2, 0], [0, 1, 0], [4, -4, 1]]

full fine-tuning: 16,777,216 parameters
LoRA r=4  :    32,768 parameters (0.20% of full, 512x fewer)
LoRA r=8  :    65,536 parameters (0.39% of full, 256x fewer)
LoRA r=16 :   131,072 parameters (0.78% of full, 128x fewer)
LoRA r=64 :   524,288 parameters (3.12% of full, 32x fewer)
LoRA adds a low-rank product B times A to a frozen weight matrix and trains 256 times fewer parametersA large square labelled W, 4096 by 4096, is frozen. It is added to alpha over r times the product of two trained matrices: B, a tall thin matrix of 4096 by 8, and A, a wide thin matrix of 8 by 4096. The thin matrices are drawn wider than their true proportion so they stay visible. Two bars compare trainable parameters for this one layer: 16,777,216 for full fine-tuning and 65,536 for LoRA at rank 8, a bar about 0.4 percent as long. The layer size is illustrative and matches the worked example.W (frozen)4096 × 4096+α / rB: 4096 × 8A: 8 × 4096B and A are trainedFull fine-tuning: 16,777,216 trainable parametersLoRA, r = 8: 65,536 trainable (256x fewer)
LoRA adds a low-rank product B times A to a frozen weight matrix and trains 256 times fewer parameters
  • Rank sets the budget: Doubling the LoRA rank doubles the number of trainable parameters, yet even a rank of 64 represents only about 3 percent of the full matrix.
  • B starts at zero: The matrix is initialised with zeros, so the adapted weights initially equal the original weights and training begins from the base model's behaviour.
  • Merging: After training, the scaled product can be added permanently into , so the tuned model introduces no additional inference latency.
  • QLoRA: A related method stores the frozen weights in 4-bit precision and trains the LoRA matrices on top, which reduces GPU memory requirements even further.

In practice, the LoRA rank is the principal tuning decision, because a higher rank can capture a larger behavioural change but requires more memory and produces a larger adapter file.

Example: Preparing Fine-Tuning Data in Python

The program below converts reviewed diff summaries into chat-format training records, drops invalid examples and creates a validation split.

Python
import json

# Team examples: a diff summary written the way reviewers want it
examples = [
    ("def add(a, b): return a+b  ->  def add(a: int, b: int) -> int: return a + b",
     "Adds type hints to add(). No behaviour change."),
    ("retries = 3  ->  retries = 5",
     "Raises retry count from 3 to 5. Check timeout budget."),
    ("except Exception: pass  ->  except ValueError as e: log.warning(e)",
     "Narrows the except clause and logs the error instead of hiding it."),
    ("", "Empty diff."),
]

SYSTEM = "Summarise the diff in one or two sentences for a code reviewer."
records, skipped = [], 0
for diff, summary in examples:
    if not diff.strip():  # drop examples with no input
        skipped += 1
        continue
    records.append({"messages": [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": diff},
        {"role": "assistant", "content": summary},
    ]})

# Hold back the last record to check the tuned model on unseen data
train, validation = records[:-1], records[-1:]
jsonl = "\n".join(json.dumps(r) for r in train)  # contents of train.jsonl
print(f"kept {len(records)}, skipped {skipped}")
print(f"train {len(train)}, validation {len(validation)}")
print("roles:", [m["role"] for m in train[0]["messages"]])
print("target:", train[0]["messages"][2]["content"])
Output
kept 3, skipped 1
train 2, validation 1
roles: ['system', 'user', 'assistant']
target: Adds type hints to add(). No behaviour change.
  • Data preparation is most of the work: Training itself is a single command or API call, but the examples decide what the model learns.
  • Assistant message as target: The model is trained to reproduce the assistant content, so every summary must already be correct.
  • Scale in practice: A real dataset contains hundreds of records or more, and training runs on GPUs through a training library or a provider's tuning service.

Applications of Fine-Tuning LLMs

  • Consistent output formats: Pull request summaries, commit messages or JSON responses that always follow one template.
  • Internal code conventions: Completions that follow an organisation's frameworks, naming conventions and architectural patterns.
  • Domain vocabulary: Legal, medical or financial terminology that general models handle poorly.
  • Classification: Categorising bug reports, log entries or support tickets into predefined labels.
  • Smaller replacement models: Training small language models to match a larger model on one narrow task.

Advantages

  • Shorter prompts: Instructions and examples move into the weights, which reduces tokens and latency per request.
  • Consistent behaviour: Output style and format become more reliable than with prompting alone.
  • Lower serving cost: A tuned small model can replace a larger general model for a specific task.
  • Private customisation: Open-weight models can be tuned entirely on internal hardware.

Limitations

  • Poor fit for changing facts: Frequently updated information is better supplied at query time through RAG, as compared in RAG vs fine-tuning.
  • Data effort: High-quality labelled examples are expensive and time-consuming to produce and review.
  • Catastrophic forgetting: Narrow training can degrade general abilities that the base model previously had.
  • Maintenance: Each new base model version requires the entire tuning process to be repeated and re-evaluated.
  • Hallucination remains: A tuned model can still produce confident errors, as covered in LLM hallucination.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What does fine-tuning change in an LLM?

Frequently Asked Questions

Fine-tuning vs prompting: which should be tried first?

Prompting should usually be tried first, because it needs no training data and changes take effect immediately. Fine-tuning is worth its cost when prompts cannot make the output format or style consistent enough.

How much data is needed to fine-tune an LLM?

Narrow tasks such as a fixed output format can improve with a few hundred high-quality examples. Broader behaviour changes need more data, and quality matters more than quantity.

What is LoRA fine-tuning?

LoRA, or low-rank adaptation, freezes the original model weights and trains small additional matrices. It needs far less GPU memory than full fine-tuning and produces a small adapter file.

Can closed models such as GPT or Claude be fine-tuned?

Some providers offer fine-tuning for selected hosted models through their platforms, while others do not. Availability changes over time, so the provider's current documentation is the reference.

Does fine-tuning add new knowledge to a model?

Fine-tuning can add some knowledge, but it is more reliable for teaching behaviour, style and format. Facts that change often are better supplied through retrieval at query time.