Your model does not know your product, your data, or your domain — and you have two fundamentally different ways to fix that: fine-tuning, which rewrites the model's weights, and RAG, which hands the model documents at query time. Teams routinely pick the wrong one, then pay for it in wasted GPU bills or brittle pipelines. This article explains what each technique actually changes inside the system, gives you a decision framework based on the kind of knowledge you need, and compares the real costs.
What fine-tuning actually does
Fine-tuning continues training: you run gradient descent on your own examples, and the model's weights change. The knowledge moves into the parameters. After fine-tuning, the model behaves differently with no extra context in the prompt — it has absorbed patterns from your data.
Full fine-tuning updates every weight, which for a multi-billion-parameter model means serious GPU memory. In practice most teams use LoRA (Hu et al., 2021): instead of updating the huge weight matrix W, you learn a small low-rank update ΔW = A·B, where A and B are tiny matrices. If W is 4096×4096 (16.7M parameters), LoRA with rank 8 learns two matrices totalling about 65K parameters — a 250× reduction — and adds the update back at inference time. You get most of the benefit at a fraction of the cost.
What fine-tuning is genuinely good at: behaviour, not facts. It teaches the model your desired style, tone, output format, and domain conventions — “answer like a support agent,” “always respond with valid JSON,” “follow this diagnostic procedure.” It bakes in how to act.
What it is bad at: injecting new factual knowledge. Research by Gekhman and colleagues (2024) showed that fine-tuning on facts the model did not already know can actually increase hallucinations — the model learns to sound confident about the new domain without reliably learning the facts. If your goal is “know these 10,000 documents,” fine-tuning is the wrong-shaped tool.
What RAG actually does
RAG changes nothing inside the model. The weights stay frozen; the knowledge lives outside, in your document store, and the relevant passages are pasted into the prompt at query time. The model reads them the way a person reads an open book during an exam.
This gives RAG three structural advantages. First, freshness: update a document and the answers change immediately — no retraining. Second, attribution: you can show exactly which passages the answer came from, which matters for anything compliance-adjacent. Third, reversibility: remove a document and the model instantly “forgets” it, which is nearly impossible with fine-tuned weights.
The price: every query pays for retrieval plus a longer prompt (more tokens, more latency, more cost), and the system is only as good as its retriever. RAG is also poor at teaching behaviour — no document will reliably make the model adopt your house style across all outputs.
The decision framework
Forget “which is better.” Ask what kind of knowledge you are adding:
- Does the knowledge change? Prices, policies, docs, news — anything updated more often than you want to retrain — points to RAG. Stable conventions (a coding style, a brand voice) point to fine-tuning.
- Is it facts or behaviour? “What is our refund policy?” is a fact: RAG. “Write refund emails in our voice” is behaviour: fine-tuning.
- Is it long-tail? Rare, specific facts (obscure error codes, internal project names) are exactly what retrieval handles well and what fine-tuning memorizes poorly.
- Do you need citations? If users or auditors must see sources, RAG wins outright — fine-tuned weights cannot point at their sources.
- How much example data do you have? Fine-tuning needs hundreds to thousands of quality examples; with fifty examples you will mostly overfit. RAG needs documents, not examples.
- Latency budget? RAG adds a retrieval hop and a fatter prompt to every request. Fine-tuning adds zero inference latency — the knowledge is already in the weights.
A useful rule of thumb: RAG for knowing, fine-tuning for behaving. Most production systems that need both end up using both, which brings us to cost.
Cost and effort, honestly compared
Fine-tuning costs concentrate up front and repeatedly. You need a curated dataset (the expensive part — data quality dominates outcomes), GPU time for training (LoRA on a 7B model fits on a single 24GB card; full fine-tuning of large models does not), an evaluation set to prove you improved things, and a retraining cadence every time the knowledge drifts. The failure mode is silent: the model gets slightly worse at things it used to do — catastrophic forgetting — and you only catch it with evals.
RAG costs concentrate in ongoing operations. You need the vector database and embedding pipeline, chunking and ingestion code, retrieval tuning (this is where the real engineering goes), and you pay per-query in tokens and latency forever. The failure mode is visible: bad retrieval produces obviously wrong answers, which is actually easier to debug than a quietly degraded fine-tune.
Roughly: fine-tuning is a capital expense with maintenance; RAG is an operating expense with tuning. Small teams with changing data almost always find RAG cheaper to start and easier to iterate on.
When you need both
The techniques compose. The common production pattern is fine-tune for behaviour, RAG for knowledge: teach the model your domain's reasoning style and output format with LoRA, then feed it fresh facts via retrieval. A medical assistant might be fine-tuned to follow diagnostic protocols and hedge appropriately, while RAG supplies the latest treatment guidelines.
Research is also blending them deliberately. RAFT (Retrieval-Augmented Fine-Tuning, Zhang et al., 2024) fine-tunes the model specifically to be good at reading retrieved documents — including training it to ignore irrelevant ones. That targets RAG's weakest link (the model misusing context) with fine-tuning's strength (behaviour shaping). It is a good illustration of the principle: use each tool for what it is shaped for.
Failure modes to plan for
Fine-tuning fails when the dataset is too small (memorization instead of generalization), when new facts are forced in (confident hallucinations, per Gekhman et al.), when training runs too long (the model forgets general capabilities), and when the world changes (stale weights, full retrain required).
RAG fails when retrieval misses (the right document exists but is not in the top-k — add hybrid keyword search), when documents contradict each other (the model picks one arbitrarily — deduplicate and date your corpus), when the context gets stuffed (key facts lost in the middle — retrieve less, re-rank more), and when the corpus is junk (no retrieval method rescues bad documents).
Notice the asymmetry: RAG's failures are usually fixable without touching the model. Fine-tuning's failures require another training run. That alone decides it for many teams.
A quick way to prototype the decision
If you are unsure, run a two-weekend experiment instead of debating. Weekend one: build a minimal RAG pipeline over your documents with an off-the-shelf embedding model and a hosted vector database, and score its answers on 50 representative questions. Weekend two: fine-tune a small open model with LoRA on a few hundred examples of desired behaviour, and score the same questions. The scoreboard usually settles the argument: if RAG answers correctly whenever retrieval hits, invest in retrieval quality; if the answers are right but the form is consistently wrong — wrong tone, wrong structure, missing steps — that is the signal that fine-tuning the behaviour will pay off.
Putting it together
The choice is not about sophistication — it is about the shape of your knowledge. Changing facts with sources to cite: RAG. Stable behaviour, style, and format: fine-tuning. Both: fine-tune the behaviour, retrieve the facts. Start with RAG unless you can articulate exactly which behaviour fine-tuning would teach; it is cheaper to prototype, easier to debug, and trivially reversible. Graduate to fine-tuning when your evals show the model consistently mishandling the form of its answers rather than their content.
Further reading
- Hu, E. J., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685
- Lewis, P., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401
- Gekhman, Z., et al. (2024). Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? arXiv:2404.10182
- Zhang, T., et al. (2024). RAFT: Adapting Language Model to Domain Specific RAG. arXiv:2403.10131






