Modern AI Engineering

Fine-Tuning an LLM: When to Do It and How It Works (LoRA)

By Modern AI Engineering · · 8 min read

Fine-tuning an LLM means taking a model that is already trained and training it a little more on your own examples, so that it behaves the way you need without long instructions. It is worth doing when you need a consistent style, format or narrow skill that prompting cannot hold. It is the wrong tool for adding facts.

Today most fine-tuning uses LoRA, a method that trains a small set of extra weights and leaves the original model untouched. This article explains when to fine-tune, how the process runs and how LoRA works.

What is fine-tuning?

A pre-trained model has general abilities. It can write, summarise and follow instructions. Fine-tuning continues its training on a smaller set of examples that show one specific behaviour. Each example is a pair: an input and the output you want for it.

The mechanism is the same as in any training. The model produces an output. A loss function measures how far it is from the target. Backpropagation works out how each weight should change, and the weights move a little. After enough examples the new behaviour is built into the model.

The key point is that fine-tuning changes the weights. Prompting and RAG change only the input.

When should you fine-tune an LLM?

Try a good prompt first, then retrieval if the model lacks knowledge. Fine-tune when those are not enough and you can name the failure. The table shows the usual cases.

SituationFine-tune?Reason
You need a fixed tone or house styleOften yesStyle is a behaviour, and examples teach it well.
You need a strict output format every timeOften yesTraining makes the format the default.
You want a small model to do one job wellYesA tuned small model can replace a costly large one for that job.
Your prompt is very long and sent on every requestMaybeTuning can move instructions and examples into the weights.
The model lacks your company factsNoUse RAG. Facts in weights are hard to update and to trace.
The facts change every weekNoYou would have to train again each time.
You have only a handful of examplesNoPut them in the prompt as few-shot examples.

How does fine-tuning an LLM work, step by step?

The steps are the same whichever tool you use.

  1. Define the task and write down how you will judge success.
  2. Build an evaluation set and measure the base model with your best prompt. This is the baseline you must beat.
  3. Collect training examples as pairs of input and ideal output. Quality matters more than quantity.
  4. Hold some examples back. Never test on data the model trained on.
  5. Choose a base model and a method, usually LoRA.
  6. Train, and watch the loss on the held-back examples to spot overfitting.
  7. Compare the tuned model with the baseline on the evaluation set.
  8. Read real outputs, not only scores. Then deploy and keep monitoring.

What is LoRA and why is it used?

Full fine-tuning updates every weight in the model. For a large model that needs a lot of GPU memory, and the result is a complete new copy of the model for each task.

LoRA stands for low-rank adaptation. It freezes the original weights and adds two small matrices next to each of the large weight matrices you choose to adapt. Only the small matrices are trained. Multiplied together, they form an update of the same shape as the large matrix, and that update is added to it.

The idea behind it is that the change a task requires is simple compared with the full model, so it can be captured with far fewer numbers. A setting called the rank controls how many. A higher rank gives the update more capacity and costs more memory.

The benefits are practical. Training needs much less memory. The result, called an adapter, is a small file. You can keep one base model and swap adapters for different tasks, or merge an adapter into the base weights so it adds no delay at run time.

What is QLoRA?

LoRA cuts the number of weights you train, but the frozen base model still has to sit in memory. QLoRA goes one step further. It loads the base model in a quantized form, which stores each weight with fewer bits, and trains LoRA adapters on top.

This lowers memory use enough that a model of moderate size can be tuned on a single GPU. The trade is that quantization loses a little precision and training can be slower. If you have the hardware, compare a QLoRA run with a plain LoRA run on your evaluation set.

Other ways to adapt a model

LoRA is the common choice, but it is not the only one.

  • Prefix tuning trains a short sequence of vectors that is placed before the input, and leaves the model frozen.
  • Knowledge distillation trains a small model to copy the outputs of a large one, so the small model can do the task at lower cost.
  • Preference tuning trains on pairs of better and worse answers, to move the model towards what people prefer.
  • Hosted fine-tuning, where a provider trains a private version of its model on the examples you upload.

Common fine-tuning mistakes

Most failed fine-tuning projects fail for ordinary reasons.

  • No baseline. Without one you cannot tell whether tuning helped.
  • Poor data. The model learns every flaw in the examples, including inconsistent formats.
  • Fine-tuning to teach facts. The model may still state them wrongly, and updating them means training again.
  • Overfitting. With too many passes over a small dataset, the model repeats the training examples and handles new inputs badly.
  • Forgetting. Heavy tuning on a narrow task can weaken general abilities the model had before.
  • Mismatched format. The prompt format used in training must match the one used in production.

How to learn fine-tuning

Fine-tuning rests on ideas from deep learning: loss, gradient descent, backpropagation and overfitting. If those are clear, LoRA is a short step. If they are not, start there.

Module 8 of the AI Engineering Bootcamp, Teaching and Shaping Models, has eleven lessons. They cover fine-tuning, LoRA, prefix tuning, knowledge distillation, continual learning, and the methods used to align models with human preferences, such as RLHF and DPO. Quantization, which QLoRA depends on, has a lesson in the module on serving models.

Frequently asked questions

What is fine-tuning in simple terms?

It is extra training of an existing model on your own examples, so that it learns a specific behaviour. The weights of the model change as a result.

How much data do I need to fine-tune an LLM?

It depends on the task and the model. A narrow format task needs far fewer examples than a broad skill. Start small with clean examples, measure, and add more only if the results call for it.

Is LoRA as good as full fine-tuning?

For many tasks it comes close at a fraction of the cost. Full fine-tuning can do better when the task differs a lot from what the model already knows.

Should I fine-tune or use RAG?

Use RAG when the model lacks knowledge. Fine-tune when it lacks a behaviour, such as a style or format. Many products use both.

Can I fine-tune a model on my own computer?

Small models can be tuned with LoRA or QLoRA on a single GPU with enough memory. Larger models need rented hardware or a hosted fine-tuning service.

Learn it properly: the AI Engineering Bootcamp

More articles