AI Engineering · 6 MIN
RAG or fine-tuning: which should you use?
Use RAG when the model needs your changing facts, fine-tuning when it needs a consistent behavior. A decision framework with a comparison table and a test order.
Use retrieval-augmented generation (RAG) when the model needs facts it does not have, especially facts that change or that differ by user. Use fine-tuning when the model has the knowledge but needs to behave in a consistent way, such as a fixed format, tone, or narrow classification. Most teams should try prompting first, add RAG for knowledge gaps, and fine-tune last. This post gives the decision rules, a comparison table, and the order in which to test them.
- RAG supplies knowledge at request time. Fine-tuning adjusts behavior through training. They solve different problems.
- OpenAI's optimization guide treats prompting as the starting point and fine-tuning as the step after prompting falls short.
- RAG wins when data changes often, must be traceable to a source, or varies by permission.
- Fine-tuning fits stable, narrow tasks with many labeled examples, or cases where a small model must replace a large one at scale.
- RAG quality depends mostly on retrieval, so measure retrieval separately from generation.
- Nactore builds the eval set first, then picks the lightest technique that clears it, so the choice rests on measurements.
What is the actual difference between RAG and fine-tuning?
RAG retrieves relevant passages from your own data at request time and places them in the prompt, so the model answers from them. The idea comes from the original RAG paper, which combined a generator with an external document index. Updating knowledge means updating the index, not the model.
Fine-tuning continues training a model on your examples so its weights change. It teaches patterns of behavior: output format, style, domain phrasing, how to handle a class of inputs. It is a poor way to store facts, because the knowledge is baked in, hard to update, and hard to trace to a source. (Parameter-efficient methods such as LoRA make training cheaper, but they do not change this distinction.)
A useful shorthand: RAG is giving the model an open book. Fine-tuning is training the model's habits.
How do you decide?
Ask what is actually failing.
| If the failure is... | The usual fix |
|---|---|
| Model lacks your private or recent facts | RAG |
| Answers must cite a source or be auditable | RAG |
| Different users may see different data | RAG with permission filters |
| Output format or style is inconsistent despite good prompts and examples | Fine-tuning (or stricter structured output first) |
| A narrow high-volume task is too costly on a large model | Fine-tune a smaller model |
| Prompt is huge because of many examples | Fine-tuning, or prompt caching first |
| Model reasoning is weak on the task | A better model or different reasoning settings |
The last row matters. Neither technique fixes a model that cannot do the task. If it fails with the right context in the prompt, switch or upgrade the model.
What should you try first?
OpenAI's model optimization guide lays out a flow that holds up in practice: build evaluation tests first, optimize prompts with those evals in place, and consider fine-tuning only if needed. It lists the conditions where fine-tuning pays off: handling more varied inputs than fit in a prompt, cutting token cost and latency at scale, and training smaller models for specific tasks. That guide also notes OpenAI is winding down its fine-tuning platform, so check current provider availability before planning around it.
A practical test order:
- Prompt with examples. Clear instructions, a few examples, a strict output schema.
- Add retrieval if the failures are missing facts.
- Improve retrieval (chunking, hybrid search, reranking) before touching the model.
- Fine-tune only if a measured behavior gap remains and you have enough clean labeled examples.
Run every step through the same eval set. See why AI features need evals before production.
Why does RAG quality come down to retrieval?
When RAG answers badly, the cause is usually that the right passage never reached the prompt. Two published findings are worth knowing.
First, retrieval can be improved without touching the model. Anthropic's contextual retrieval technique prepends chunk-specific context before indexing, and the company reports that combining it with keyword search and reranking cut top-20 retrieval failures by 67 percent in its tests. Treat that as evidence the approach is worth testing on your data, not as a number you will reproduce.
Second, more context is not always better. The Lost in the Middle paper found that models use information at the start and end of a long context more reliably than information in the middle. Sending fewer, better passages often beats sending many.
Measure retrieval on its own. Build 50 to 100 real questions, label which passage answers each, and track how often that passage lands in the top results. If this number is low, no prompt change will rescue the answers.
What about doing both?
Combining them is legitimate and common once the basics work. A team might fine-tune a small model to follow a strict output format and call style, and use RAG to feed it current facts. The key is that each piece is justified by a measured gap, not added because it sounds sophisticated.
What does each approach cost to run and maintain?
| Factor | RAG | Fine-tuning |
|---|---|---|
| Updating knowledge | Re-index changed documents | Retrain on new data |
| Source traceability | Natural (retrieved passages can be cited) | Weak |
| Per-request cost | Higher input tokens from retrieved context | Often lower prompts, may use a smaller model |
| Upfront work | Pipeline, index, retrieval evals | Data labeling, training runs, regression checks |
| Main failure mode | Wrong or missing passage | Overfitting, stale behavior, regressions after provider changes |
| Permissions | Filter at retrieval time | Cannot be separated per user after training |
Prompt caching can also shrink the cost of long, stable prompts, which sometimes removes the case for fine-tuning. See LLM cost control and prompt caching.
Frequently asked questions
Can fine-tuning teach the model our company knowledge?
It can shift style and surface familiarity, but it is unreliable for exact facts and hard to keep current. Use retrieval for facts you need to be right and traceable.
Is RAG just a vector database?
No. A vector index is one component. Production systems often combine embeddings with keyword search, reranking, metadata filters, and permission checks.
How much data does fine-tuning need?
It depends on the task and provider. The requirement is clean, representative labeled examples and a held-out eval set. Check your provider's current guidance.
When is long context better than RAG?
When the whole corpus fits in the window and cost is acceptable, putting it all in the prompt (with caching) can be simpler. It stops working as data grows or when permissions differ by user.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.