Use RAG when answers must come from documents that change, need citing or differ per user; fine-tune only for a fixed format, tone or classification behaviour that prompting cannot hold. On October 2026 prices, RAG over 100,000 documents costs roughly $125 to $275 a month, while one tuning run over the same text costs $1,800 or more and must be repeated when content changes. OpenAI is winding down self-serve fine-tuning, so new tuning projects mean Google or open models.
RAG vs fine-tuning is a choice between changing what the model can see and changing how it behaves: use RAG when answers must come from documents that change or need citing, and fine-tune only when you need a fixed format, tone or classification behaviour that prompting cannot hold. In 2026 the default is RAG, partly on cost and partly because self-serve fine-tuning is getting harder to buy: OpenAI is winding its platform down, which leaves Google's Gemini tuning and open models as the main routes.
Built from the vendors' documentation and pricing pages, checked October 2026: OpenAI API pricing and deprecations, Google's supervised fine-tuning docs and Agent Platform pricing, Anthropic's pricing docs, Supabase pricing and its compute sizing guide, and the pgvector README. Every cost below is arithmetic on those list prices with stated inputs, not a measured bill.
If you only need the model prices themselves, they are in our LLM API pricing comparison, the hub for this topic. This page is about the decision that comes before picking a model: where your knowledge and your rules should live.
RAG vs fine-tuning: the short answer
Ask two questions. Does the knowledge change? Does the model's default behaviour need to change? The answers put you in one of four boxes, and only one of them is fine-tuning alone.
- Pick RAG if the answer lives in documents: policies, product data, contracts, tickets, a knowledge base. It stays current, it can cite, and it respects who is allowed to see what.
- Pick fine-tuning if the job is a stable behaviour with many examples: classify, extract, or write in one strict format, at a volume where long prompts hurt.
- Pick neither if the material fits in the prompt. That is more common than the debate suggests.
Fine-tuning vs RAG, job by job
Most teams are not choosing a technology; they are trying to get one job done. Find the job in the first column.
| The job | With RAG | With fine-tuning | Pick |
|---|---|---|---|
| Answer from documents that change | Update the index and the next answer is current | Retrain for every change | RAG |
| Show where an answer came from | Returns the chunks it used, so you can cite them | No source to point to | RAG |
| Keep each customer's data separate | Filter rows at query time, for example with row level security | Weights cannot be partitioned per user | RAG |
| Hold a fixed format, tone or house style | Possible with examples in every prompt, paid for on every call | Learns it from examples | Fine-tune, after few-shot prompting fails |
| Classify or extract into fixed labels | Not a retrieval problem | A listed use case in Google's tuning docs | Fine-tune, or few-shot |
| Cut latency | Adds a search step and a longer prompt | Shorter prompts; Google lists lower latency among the benefits | Fine-tune, if volume justifies the work |
| Cut cost per request | Adds retrieved tokens to every call | Shorter prompts, but tuned Gemini 3 models bill at 1.5x | Usually neither: see the sums below |
| Small corpus that rarely changes | Works, but is more machinery than needed | Wrong tool | Neither: put it in the prompt and cache it |
The per-user row deserves a second look if you sell to businesses. With RAG on Postgres, the same row level security that protects your tables protects the chunks, so a question can only retrieve what that user may read. A fine-tuned model has no such boundary: whatever went into training can come out for anyone.
The cost row surprises people. Fine-tuning is often sold as a way to shorten prompts, and it does, but Google bills tuned Gemini 3 models at 1.5 times the base rate. Illustrative sums on Gemini 3.1 Flash-Lite ($0.25 input, $1.50 output per million tokens): a request with 1,500 tokens of instructions, 500 of user input and 300 of output costs about $95 per 100,000 requests on the base model. Tune the instructions down to 100 tokens and the same traffic costs about $90 at the tuned rate. The markup eats the saving, so tune for behaviour, not to save tokens.
On latency, no honest number fits every stack, so this page gives none. The direction is clear enough: RAG adds a vector search and a longer prompt to every request, and a tuned model with a short prompt removes input tokens. Whether either matters depends on your model and volume, which is a thing to measure on your own traffic.
Where you can still fine-tune in 2026
The practical question has changed since most RAG-versus-tuning advice was written. On May 7, 2026, OpenAI told developers it was changing the availability of its self-serve fine-tuning platform, and the pricing page now says the platform is being wound down.
- May 7, 2026Closed to newcomers
Organisations that had never run fine-tuning can no longer create jobs.
- Jul 2, 2026Closed to lapsed users
Job creation ends for organisations with no inference on a fine-tuned model in the previous 60 days.
- Jan 6, 2027No new jobs for anyone
Active existing customers lose the ability to create fine-tuning jobs.
- After thatExisting models keep answering
Inference on a fine-tuned model continues until its base model is deprecated.
Source: OpenAI API deprecations page, checked October 2026
So a new project cannot start on OpenAI fine-tuning at all, and an existing one has a date on it. Here is what the main vendors document today.
| Provider | Status, October 2026 | Training price | Inference on the tuned model |
|---|---|---|---|
| OpenAI | Self-serve fine-tuning is being wound down. Closed to organisations that had not used it; new jobs end for everyone on Jan 6, 2027 | Legacy list only, e.g. gpt-4.1-mini at $5.00 per 1M training tokens | Existing tuned models keep serving until the base model is deprecated |
| Supervised fine-tuning on Gemini Enterprise Agent Platform for Gemini 3.5 Flash, Gemini 3.1 Flash-Lite and older Gemini 2.5 models | $10 per 1M training tokens (3.5 Flash); $3 per 1M (3.1 Flash-Lite) | 1.5x the base price for Gemini 3 and later | |
| Anthropic | No fine-tuning endpoint in the Claude API docs for current models. The one guide Anthropic publishes covers Claude 3 Haiku on Amazon Bedrock, from 2024 | Not listed | Not listed |
| Open models | Tune on Google's platform (Gemma 3, Llama and Qwen 3 models are listed) or on your own GPUs | $0.67 per 1M (Llama 3.1 8B) to $6.72 per 1M (Llama 3.3 70B) on Google's platform | Depends on where you deploy the result |
Google lists tuning prices per 1,000 tokens ($0.01 and $0.003); they are converted to per-million here. Training tokens are the tokens in your dataset multiplied by the number of epochs.
That table is a risk assessment as much as a price list. A RAG pipeline is a database and some prompts; it can move to another model without retraining anything. A tuned model is tied to one base model on one platform, and you have just watched a major platform withdraw the service. Weigh that before you invest in a training set.
The cost model: 1,000 and 100,000 documents
Here are the sums for two corpus sizes. Every input is illustrative, so swap in your own:
- Documents: 2,000 tokens each on average (about 1,500 words), split into 500-token chunks, so four vectors per document.
- Embeddings: OpenAI's text-embedding-3-small at $0.02 per million tokens, 1,536 dimensions by default. (Google's Gemini Embedding 2 lists $0.20.)
- Vector store: pgvector on a Supabase Pro project, $25 a month including $10 of compute credit. pgvector stores each vector in 4 bytes per dimension plus 8.
- Traffic: 10,000 questions a month, five chunks retrieved for each, so 2,500 extra input tokens per question.
- Answering model: $2 per million input tokens, the list price of both Claude Sonnet 5.5 and GPT-6.1 Sol.
- Tuning comparison: three epochs over the same text on Google's platform, as if you tried to put the knowledge into the weights.
| Cost line | 1,000 documents | 100,000 documents |
|---|---|---|
| RAG: embed the corpus once | $0.04 | $4.00 |
| RAG: vectors to store | 4,000 (about 25 MB) | 400,000 (about 2.5 GB) |
| RAG: Supabase plan and compute | $25 a month (Micro compute is covered by the plan) | $75 a month at 384 dimensions; $225 at 1,536 |
| RAG: retrieved context, 10,000 questions a month | $50 | $50 |
| RAG: monthly total | About $75 | About $125 to $275 |
| Tuning: one run, 3 epochs, Gemini 3.1 Flash-Lite | $18 | $1,800 |
| Tuning: one run, 3 epochs, Gemini 3.5 Flash | $60 | $6,000 |
| Tuning: when the documents change | Another run | Another run |
| Tuning: inference | 1.5x base price on every request | 1.5x base price on every request |
RAG bars are the one-off embedding cost plus the first month of hosting and retrieved context. Tuning bars are a single training run, before any inference. Inputs are illustrative. Source: Arithmetic on OpenAI, Supabase, Anthropic and Google list prices, checked October 2026
Three things stand out once the numbers are side by side.
- Embedding is a rounding error. Four dollars for 200 million tokens. Nobody should choose an architecture to avoid the embedding bill.
- Hosting is the real RAG cost at scale, and vector size drives it. Supabase's sizing guide benchmarks 500,000 vectors of 1,536 dimensions on an XL instance ($210 a month) and the same count at 384 dimensions on a Medium ($60). OpenAI's embedding models accept a
dimensionsparameter to shorten vectors, which is the cheapest lever in the whole table. - RAG cost follows questions; tuning cost follows content. The retrieved-context line is $50 at both sizes because it depends on traffic. The tuning line grows a hundredfold with the corpus and repeats on every update.
To build the store, our pgvector on Supabase guide has the schema, index and search function in runnable SQL, and pgvector vs Qdrant vs Pinecone prices the alternatives at a million vectors. For agents rather than documents, see our comparison of AI agent memory options. If your documents live on websites, the Firecrawl tutorial covers turning pages into clean text before you embed them.
The do-both pattern
Mature products often end up with both, each doing the half it is good at. Retrieval carries the facts. The tuned model carries the house style: the output schema, the tone, when to refuse, how to cite.
- 01A user asks a question
Nothing about the request needs to mention format or rules.
- 02Retrieve
Embed the question and pull the closest chunks, filtered to what this user may see.
- 03The tuned model reads both
Training taught it the format, tone and refusals. The chunks supply today's facts.
- 04Answer with sources
Citations point at retrieved chunks, so a wrong answer can be traced to its source.
- 05Maintain each half alone
Documents change: re-embed them. Behaviour changes: retrain. Neither forces the other.
Build the retrieval half first. It works with any model, it exposes whether a behaviour gap really exists once the facts are right, and the logged questions and corrected answers it produces are exactly the examples a later tuning run needs. Google adds one practical note for tuned models that support thinking: set the thinking level to minimal, because the tuned model learned the task without it.
RAG or fine-tuning for a business: the order to try things
Go from cheap and reversible to expensive and sticky. Each step either solves the problem or tells you precisely what is still missing.
- 1Write the prompt properly
Clear instructions and three to five worked examples. Many "we need fine-tuning" problems end here.
- 2Put the documents in the prompt if they fit
A small, stable corpus can ride in the context window with prompt caching. No index, no retrieval misses.
- 3Add RAG when the corpus outgrows the prompt
Or when it changes weekly. Chunk, embed, store, retrieve. Start in Postgres with pgvector if you already run Postgres.
- 4Build a test set before changing anything else
50 to 100 real questions with known good answers. Score retrieval and answers separately.
- 5Fine-tune only for a gap that survives
Wrong format, drifting tone or weak classification after steps 1 to 4, on a platform that will still offer tuning next year.
- 6Keep RAG after tuning
A tuned model still needs today's facts. Tuning replaces prompt instructions, not the index.
- Answers must reflect documents that change
- Users need to see the source
- Data is per customer and must stay separated
- You want to switch model provider later without retraining
- You have documents but no labelled examples
- The task is classification, extraction or one fixed output format
- You have hundreds of good input and output examples
- Prompts are long and volume is high enough to matter
- The desired behaviour will be stable for months
- Your platform still offers tuning: Google, or open models
Step one is where most of the value sits, and our prompt engineering guide covers it. For steps two and three on a real product, the AI SaaS Builder program has a lesson that builds RAG on Supabase with pgvector and the Claude API, starting with the question this page ends on: whether you need retrieval at all, or can send the documents whole and cache them.
RAG vs fine-tuning: FAQ
What is the difference between RAG and fine-tuning?
RAG, retrieval-augmented generation, leaves the model unchanged and adds a search step: your documents are split into chunks, embedded and stored, and the closest chunks are placed in the prompt at question time. Fine-tuning trains the model further on example inputs and outputs so its default behaviour changes. RAG changes what the model can see; fine-tuning changes how it responds. They solve different problems and can be combined.
Is RAG cheaper than fine-tuning?
For knowledge, yes, on published October 2026 prices. Embedding 100,000 documents of 2,000 tokens with OpenAI's text-embedding-3-small costs about $4 once, and hosting the vectors on Supabase runs roughly $75 to $225 a month depending on vector size. Training on the same 200 million tokens for three epochs costs $1,800 on Gemini 3.1 Flash-Lite and has to be repeated whenever the documents change. These are illustrative sums, not measured bills.
Can you still fine-tune OpenAI models in 2026?
Only if you were already doing it. OpenAI's deprecations page says that since May 7, 2026, organisations that had not previously run fine-tuning cannot create jobs, and that from January 6, 2027 active existing customers will not be able to create new ones either. Models already fine-tuned keep working for inference until their base model is deprecated. New projects should plan around Google's Gemini tuning or open models.
Should a business use RAG or fine-tuning?
Most business cases, such as answering from policies, product docs, contracts or a knowledge base, are RAG cases: the content changes, people want to see the source, and access differs by user. Fine-tuning suits narrower jobs with stable rules and plenty of examples, such as routing tickets into fixed categories or writing in a strict house format. Start with a good prompt, add RAG when documents are involved, and tune last.
Can I use RAG and fine-tuning together?
Yes, and it is a common end state for mature products. Retrieval supplies current facts from your documents, while a tuned model supplies consistent format, tone and refusal behaviour. The two are maintained separately: when documents change you re-embed them, and when the desired behaviour changes you retrain. Build the RAG half first, because it works with any model and shows you whether a tuning gap really exists.
Does fine-tuning teach a model my documents?
Not in a way you can rely on for exact answers. Supervised fine-tuning learns patterns from input and output examples; Google's documentation lists output formats, specific behaviours and classification among its use cases, not document lookup. A tuned model cannot cite a source, cannot forget a deleted document and cannot keep one customer's data apart from another's. For facts that must be right and current, retrieve them at question time.
When do I need neither RAG nor fine-tuning?
When the documents fit in the prompt. If your whole reference set is a few dozen pages, send it with every request and use prompt caching so repeat reads are billed at a fraction of the input price. On Claude Sonnet 5.5, for example, a cached read lists $0.10 per million tokens against $2 uncached, checked October 2026. There is no index to maintain and no retrieval step to miss the right passage.
Decided on RAG? Build it on a stack you own.
AI SaaS Builder, included in All Access, builds retrieval on Supabase with pgvector and the Claude API, then covers caching, cost control and shipping the product around it, alongside the other three programs, live coaching and the private community.
Not sure which box you are in?
Describe the job and your documents in the free Discord and get a second opinion before you spend on either route.