RAG vs Fine-Tuning vs Prompt Engineering: Which to Use When
By Modern AI Engineering · · 7 min read
Use prompt engineering first. Add RAG when the model lacks knowledge, such as your documents or recent facts. Use fine-tuning when the model lacks a behaviour, such as a fixed style or format, and a prompt cannot get it there reliably. The three are not competitors. They change different things, and many products use all of them.
The quick test is to ask what is missing. If the model does not know something, that is a knowledge problem and RAG fits. If the model knows enough but does not act the way you need, that is a behaviour problem, and prompting or fine-tuning fits.
What is prompt engineering?
Prompt engineering means changing what you send to the model. You write clear instructions, give examples of good answers, set a format, and break a hard task into steps. The model itself does not change.
It is the fastest and cheapest of the three. You can try an idea in minutes. It is also the most limited. A prompt cannot give the model facts it never saw unless you paste those facts in, and every word you add is paid for on every request.
Today the term context engineering is often used for the wider job: deciding everything that goes into the context window, including instructions, retrieved documents, tool results and conversation history.
What is RAG (retrieval-augmented generation)?
RAG adds a search step before the model answers. Your documents are split into chunks, and each chunk is turned into an embedding and stored. When a question arrives, the system finds the chunks most related to it and places them in the prompt. The model then answers from that text.
RAG suits knowledge that is private, large or changing. To update what the system knows, you update the documents. No training is needed. It also lets you show sources, so a user can check where an answer came from.
Its weak point is retrieval quality. If the right chunk is not found, the model cannot use it. Most of the work in a RAG system goes into chunking, search and reranking, not into the model call.
What is fine-tuning?
Fine-tuning continues training a pre-trained model on your own examples, so its weights change. After fine-tuning, the model behaves differently without being told to in the prompt.
It works well for behaviour: a consistent tone, a strict output format, a narrow task done the same way every time, or making a small model good at one job so you can stop paying for a large one. Methods such as LoRA train only a small set of extra weights, which makes fine-tuning much cheaper than updating the whole model.
It is a poor way to add facts. Facts trained into weights are hard to update and hard to trace, and the model can still state them wrongly. It also needs a set of good examples, time to train and a way to measure whether the result is better.
RAG vs fine-tuning vs prompt engineering: comparison table
The table sums up the practical differences.
| Prompt engineering | RAG | Fine-tuning | |
|---|---|---|---|
| What changes | The input | The input, with retrieved text | The model weights |
| Best for | Instructions, format, simple tasks | Private or changing knowledge | Consistent behaviour and style |
| Adds new facts | Only what you paste in | Yes, from your documents | Poorly |
| Effort to start | Low | Medium | High |
| Updating | Edit the prompt | Update the documents | Train again |
| Can cite sources | No | Yes | No |
| Main risk | Long, fragile prompts | Poor retrieval | Bad or too little training data |
How do you choose between them?
Go through these steps in order and stop as soon as the result is good enough.
- Write a clear prompt with a few examples and test it on real inputs. Many problems end here.
- Look at the failures. If the model lacks facts, or the facts change, add RAG.
- If answers are still wrong, check retrieval before blaming the model. Is the right chunk being found?
- If the model has the facts but the style, format or reasoning is still inconsistent, collect examples of ideal outputs and consider fine-tuning.
- Whatever you choose, build an evaluation set first, so you can tell whether each change helped.
When should you combine RAG and fine-tuning?
Combining them is common, because they solve different problems. A support assistant might use RAG to pull the current help articles, a fine-tuned model to answer in the company voice and always return the same structure, and a prompt that sets the rules for the conversation.
A sensible order is to get prompting and RAG working first. Add fine-tuning only when you can point to a clear, repeated failure that the first two do not fix. Fine-tuning before that often means paying for training to solve a problem that a better prompt or better retrieval would have solved.
Common mistakes
A few errors come up again and again.
- Fine-tuning to teach facts. Use retrieval for facts.
- Blaming the model for a retrieval failure. Read what was actually retrieved.
- Stuffing everything into the prompt. Long contexts cost more, and details in the middle are easier for the model to miss.
- Chunking documents without thought. Chunks that cut a table or a sentence in half lose meaning.
- Skipping evaluation. Without a test set you cannot compare the three approaches fairly.
Frequently asked questions
Is RAG better than fine-tuning?
Neither is better in general. RAG is better for adding knowledge that is private or changes often. Fine-tuning is better for changing how the model behaves. Many systems use both.
Does RAG stop hallucinations?
It reduces them when the right documents are retrieved, because the model can answer from text in front of it. It does not remove them. Poor retrieval or unclear sources still lead to wrong answers.
Can prompt engineering replace fine-tuning?
Often, yes. A clear prompt with good examples solves many tasks. Fine-tuning becomes worth it when you need consistent behaviour at scale that prompts cannot hold.
Is fine-tuning expensive?
It costs more than prompting, because you need training data, compute and evaluation. Methods such as LoRA lower the cost a lot by training only a small number of extra weights.
Do I need a vector database for RAG?
Usually, but not always. Vector search is the common choice. Keyword search, hybrid search and approaches without embeddings can also retrieve the right text, depending on the data.