Skip to main content

RAG vs Fine-Tuning in 2026: A Decision Guide With Costs

RAG vs fine-tuning in 2026: which to pick by job, a cost model for 1,000 and 100,000 documents from published prices, and where tuning is still offered.

Founder of IImagined.ai

Published
Oct 11, 2026
Reading time
12 min read
Quick answer

Use RAG when answers must come from documents that change, need citing or differ per user; fine-tune only for a fixed format, tone or classification behaviour that prompting cannot hold. On October 2026 prices, RAG over 100,000 documents costs roughly $125 to $275 a month, while one tuning run over the same text costs $1,800 or more and must be repeated when content changes. OpenAI is winding down self-serve fine-tuning, so new tuning projects mean Google or open models.

RAG vs fine-tuning is a choice between changing what the model can see and changing how it behaves: use RAG when answers must come from documents that change or need citing, and fine-tune only when you need a fixed format, tone or classification behaviour that prompting cannot hold. In 2026 the default is RAG, partly on cost and partly because self-serve fine-tuning is getting harder to buy: OpenAI is winding its platform down, which leaves Google's Gemini tuning and open models as the main routes.

Built from the vendors' documentation and pricing pages, checked October 2026: OpenAI API pricing and deprecations, Google's supervised fine-tuning docs and Agent Platform pricing, Anthropic's pricing docs, Supabase pricing and its compute sizing guide, and the pgvector README. Every cost below is arithmetic on those list prices with stated inputs, not a measured bill.

If you only need the model prices themselves, they are in our LLM API pricing comparison, the hub for this topic. This page is about the decision that comes before picking a model: where your knowledge and your rules should live.

RAG vs fine-tuning: the short answer

Ask two questions. Does the knowledge change? Does the model's default behaviour need to change? The answers put you in one of four boxes, and only one of them is fine-tuning alone.

Two questions, four answers
Behaviour must change
Fine-tune, but only after a prompt with worked examples has failed to hold the format or tone.
Both. RAG supplies the facts; a tuned model supplies the format and tone.
Default behaviour is fine
Neither. A good prompt, with the documents in context if they fit.
RAG. Index the documents and retrieve at question time.
Knowledge is stable
Knowledge changes
  • Pick RAG if the answer lives in documents: policies, product data, contracts, tickets, a knowledge base. It stays current, it can cite, and it respects who is allowed to see what.
  • Pick fine-tuning if the job is a stable behaviour with many examples: classify, extract, or write in one strict format, at a volume where long prompts hurt.
  • Pick neither if the material fits in the prompt. That is more common than the debate suggests.

Fine-tuning vs RAG, job by job

Most teams are not choosing a technology; they are trying to get one job done. Find the job in the first column.

The jobWith RAGWith fine-tuningPick
Answer from documents that changeUpdate the index and the next answer is currentRetrain for every changeRAG
Show where an answer came fromReturns the chunks it used, so you can cite themNo source to point toRAG
Keep each customer's data separateFilter rows at query time, for example with row level securityWeights cannot be partitioned per userRAG
Hold a fixed format, tone or house stylePossible with examples in every prompt, paid for on every callLearns it from examplesFine-tune, after few-shot prompting fails
Classify or extract into fixed labelsNot a retrieval problemA listed use case in Google's tuning docsFine-tune, or few-shot
Cut latencyAdds a search step and a longer promptShorter prompts; Google lists lower latency among the benefitsFine-tune, if volume justifies the work
Cut cost per requestAdds retrieved tokens to every callShorter prompts, but tuned Gemini 3 models bill at 1.5xUsually neither: see the sums below
Small corpus that rarely changesWorks, but is more machinery than neededWrong toolNeither: put it in the prompt and cache it

The per-user row deserves a second look if you sell to businesses. With RAG on Postgres, the same row level security that protects your tables protects the chunks, so a question can only retrieve what that user may read. A fine-tuned model has no such boundary: whatever went into training can come out for anyone.

The cost row surprises people. Fine-tuning is often sold as a way to shorten prompts, and it does, but Google bills tuned Gemini 3 models at 1.5 times the base rate. Illustrative sums on Gemini 3.1 Flash-Lite ($0.25 input, $1.50 output per million tokens): a request with 1,500 tokens of instructions, 500 of user input and 300 of output costs about $95 per 100,000 requests on the base model. Tune the instructions down to 100 tokens and the same traffic costs about $90 at the tuned rate. The markup eats the saving, so tune for behaviour, not to save tokens.

On latency, no honest number fits every stack, so this page gives none. The direction is clear enough: RAG adds a vector search and a longer prompt to every request, and a tuned model with a short prompt removes input tokens. Whether either matters depends on your model and volume, which is a thing to measure on your own traffic.

Where you can still fine-tune in 2026

The practical question has changed since most RAG-versus-tuning advice was written. On May 7, 2026, OpenAI told developers it was changing the availability of its self-serve fine-tuning platform, and the pricing page now says the platform is being wound down.

OpenAI self-serve fine-tuning: the wind-down
  1. May 7, 2026
    Closed to newcomers

    Organisations that had never run fine-tuning can no longer create jobs.

  2. Jul 2, 2026
    Closed to lapsed users

    Job creation ends for organisations with no inference on a fine-tuned model in the previous 60 days.

  3. Jan 6, 2027
    No new jobs for anyone

    Active existing customers lose the ability to create fine-tuning jobs.

  4. After that
    Existing models keep answering

    Inference on a fine-tuned model continues until its base model is deprecated.

Source: OpenAI API deprecations page, checked October 2026

So a new project cannot start on OpenAI fine-tuning at all, and an existing one has a date on it. Here is what the main vendors document today.

ProviderStatus, October 2026Training priceInference on the tuned model
OpenAISelf-serve fine-tuning is being wound down. Closed to organisations that had not used it; new jobs end for everyone on Jan 6, 2027Legacy list only, e.g. gpt-4.1-mini at $5.00 per 1M training tokensExisting tuned models keep serving until the base model is deprecated
GoogleSupervised fine-tuning on Gemini Enterprise Agent Platform for Gemini 3.5 Flash, Gemini 3.1 Flash-Lite and older Gemini 2.5 models$10 per 1M training tokens (3.5 Flash); $3 per 1M (3.1 Flash-Lite)1.5x the base price for Gemini 3 and later
AnthropicNo fine-tuning endpoint in the Claude API docs for current models. The one guide Anthropic publishes covers Claude 3 Haiku on Amazon Bedrock, from 2024Not listedNot listed
Open modelsTune on Google's platform (Gemma 3, Llama and Qwen 3 models are listed) or on your own GPUs$0.67 per 1M (Llama 3.1 8B) to $6.72 per 1M (Llama 3.3 70B) on Google's platformDepends on where you deploy the result

Google lists tuning prices per 1,000 tokens ($0.01 and $0.003); they are converted to per-million here. Training tokens are the tokens in your dataset multiplied by the number of epochs.

That table is a risk assessment as much as a price list. A RAG pipeline is a database and some prompts; it can move to another model without retraining anything. A tuned model is tied to one base model on one platform, and you have just watched a major platform withdraw the service. Weigh that before you invest in a training set.

The cost model: 1,000 and 100,000 documents

Here are the sums for two corpus sizes. Every input is illustrative, so swap in your own:

  • Documents: 2,000 tokens each on average (about 1,500 words), split into 500-token chunks, so four vectors per document.
  • Embeddings: OpenAI's text-embedding-3-small at $0.02 per million tokens, 1,536 dimensions by default. (Google's Gemini Embedding 2 lists $0.20.)
  • Vector store: pgvector on a Supabase Pro project, $25 a month including $10 of compute credit. pgvector stores each vector in 4 bytes per dimension plus 8.
  • Traffic: 10,000 questions a month, five chunks retrieved for each, so 2,500 extra input tokens per question.
  • Answering model: $2 per million input tokens, the list price of both Claude Sonnet 5.5 and GPT-6.1 Sol.
  • Tuning comparison: three epochs over the same text on Google's platform, as if you tried to put the knowledge into the weights.
Cost line1,000 documents100,000 documents
RAG: embed the corpus once$0.04$4.00
RAG: vectors to store4,000 (about 25 MB)400,000 (about 2.5 GB)
RAG: Supabase plan and compute$25 a month (Micro compute is covered by the plan)$75 a month at 384 dimensions; $225 at 1,536
RAG: retrieved context, 10,000 questions a month$50$50
RAG: monthly totalAbout $75About $125 to $275
Tuning: one run, 3 epochs, Gemini 3.1 Flash-Lite$18$1,800
Tuning: one run, 3 epochs, Gemini 3.5 Flash$60$6,000
Tuning: when the documents changeAnother runAnother run
Tuning: inference1.5x base price on every request1.5x base price on every request
First-month cost at 100,000 documents, by route
RAG, 384-dimension vectors
$129
RAG, 1,536-dimension vectors
$279
One tuning run, Gemini 3.1 Flash-Lite
$1,800
One tuning run, Gemini 3.5 Flash
$6,000

RAG bars are the one-off embedding cost plus the first month of hosting and retrieved context. Tuning bars are a single training run, before any inference. Inputs are illustrative. Source: Arithmetic on OpenAI, Supabase, Anthropic and Google list prices, checked October 2026

Three things stand out once the numbers are side by side.

  1. Embedding is a rounding error. Four dollars for 200 million tokens. Nobody should choose an architecture to avoid the embedding bill.
  2. Hosting is the real RAG cost at scale, and vector size drives it. Supabase's sizing guide benchmarks 500,000 vectors of 1,536 dimensions on an XL instance ($210 a month) and the same count at 384 dimensions on a Medium ($60). OpenAI's embedding models accept a dimensions parameter to shorten vectors, which is the cheapest lever in the whole table.
  3. RAG cost follows questions; tuning cost follows content. The retrieved-context line is $50 at both sizes because it depends on traffic. The tuning line grows a hundredfold with the corpus and repeats on every update.

To build the store, our pgvector on Supabase guide has the schema, index and search function in runnable SQL, and pgvector vs Qdrant vs Pinecone prices the alternatives at a million vectors. For agents rather than documents, see our comparison of AI agent memory options. If your documents live on websites, the Firecrawl tutorial covers turning pages into clean text before you embed them.

The do-both pattern

Mature products often end up with both, each doing the half it is good at. Retrieval carries the facts. The tuned model carries the house style: the output schema, the tone, when to refuse, how to cite.

RAG for facts, a tuned model for behaviour
  1. 01
    A user asks a question

    Nothing about the request needs to mention format or rules.

  2. 02
    Retrieve

    Embed the question and pull the closest chunks, filtered to what this user may see.

  3. 03
    The tuned model reads both

    Training taught it the format, tone and refusals. The chunks supply today's facts.

  4. 04
    Answer with sources

    Citations point at retrieved chunks, so a wrong answer can be traced to its source.

  5. 05
    Maintain each half alone

    Documents change: re-embed them. Behaviour changes: retrain. Neither forces the other.

Build the retrieval half first. It works with any model, it exposes whether a behaviour gap really exists once the facts are right, and the logged questions and corrected answers it produces are exactly the examples a later tuning run needs. Google adds one practical note for tuned models that support thinking: set the thinking level to minimal, because the tuned model learned the task without it.

RAG or fine-tuning for a business: the order to try things

Go from cheap and reversible to expensive and sticky. Each step either solves the problem or tells you precisely what is still missing.

Cheapest first
  1. 1
    Write the prompt properly

    Clear instructions and three to five worked examples. Many "we need fine-tuning" problems end here.

  2. 2
    Put the documents in the prompt if they fit

    A small, stable corpus can ride in the context window with prompt caching. No index, no retrieval misses.

  3. 3
    Add RAG when the corpus outgrows the prompt

    Or when it changes weekly. Chunk, embed, store, retrieve. Start in Postgres with pgvector if you already run Postgres.

  4. 4
    Build a test set before changing anything else

    50 to 100 real questions with known good answers. Score retrieval and answers separately.

  5. 5
    Fine-tune only for a gap that survives

    Wrong format, drifting tone or weak classification after steps 1 to 4, on a platform that will still offer tuning next year.

  6. 6
    Keep RAG after tuning

    A tuned model still needs today's facts. Tuning replaces prompt instructions, not the index.

The decision in one screen
Choose RAG if
  • Answers must reflect documents that change
  • Users need to see the source
  • Data is per customer and must stay separated
  • You want to switch model provider later without retraining
  • You have documents but no labelled examples
Choose fine-tuning if
  • The task is classification, extraction or one fixed output format
  • You have hundreds of good input and output examples
  • Prompts are long and volume is high enough to matter
  • The desired behaviour will be stable for months
  • Your platform still offers tuning: Google, or open models

Step one is where most of the value sits, and our prompt engineering guide covers it. For steps two and three on a real product, the AI SaaS Builder program has a lesson that builds RAG on Supabase with pgvector and the Claude API, starting with the question this page ends on: whether you need retrieval at all, or can send the documents whole and cache them.

RAG vs fine-tuning: FAQ

What is the difference between RAG and fine-tuning?

RAG, retrieval-augmented generation, leaves the model unchanged and adds a search step: your documents are split into chunks, embedded and stored, and the closest chunks are placed in the prompt at question time. Fine-tuning trains the model further on example inputs and outputs so its default behaviour changes. RAG changes what the model can see; fine-tuning changes how it responds. They solve different problems and can be combined.

Is RAG cheaper than fine-tuning?

For knowledge, yes, on published October 2026 prices. Embedding 100,000 documents of 2,000 tokens with OpenAI's text-embedding-3-small costs about $4 once, and hosting the vectors on Supabase runs roughly $75 to $225 a month depending on vector size. Training on the same 200 million tokens for three epochs costs $1,800 on Gemini 3.1 Flash-Lite and has to be repeated whenever the documents change. These are illustrative sums, not measured bills.

Can you still fine-tune OpenAI models in 2026?

Only if you were already doing it. OpenAI's deprecations page says that since May 7, 2026, organisations that had not previously run fine-tuning cannot create jobs, and that from January 6, 2027 active existing customers will not be able to create new ones either. Models already fine-tuned keep working for inference until their base model is deprecated. New projects should plan around Google's Gemini tuning or open models.

Should a business use RAG or fine-tuning?

Most business cases, such as answering from policies, product docs, contracts or a knowledge base, are RAG cases: the content changes, people want to see the source, and access differs by user. Fine-tuning suits narrower jobs with stable rules and plenty of examples, such as routing tickets into fixed categories or writing in a strict house format. Start with a good prompt, add RAG when documents are involved, and tune last.

Can I use RAG and fine-tuning together?

Yes, and it is a common end state for mature products. Retrieval supplies current facts from your documents, while a tuned model supplies consistent format, tone and refusal behaviour. The two are maintained separately: when documents change you re-embed them, and when the desired behaviour changes you retrain. Build the RAG half first, because it works with any model and shows you whether a tuning gap really exists.

Does fine-tuning teach a model my documents?

Not in a way you can rely on for exact answers. Supervised fine-tuning learns patterns from input and output examples; Google's documentation lists output formats, specific behaviours and classification among its use cases, not document lookup. A tuned model cannot cite a source, cannot forget a deleted document and cannot keep one customer's data apart from another's. For facts that must be right and current, retrieve them at question time.

When do I need neither RAG nor fine-tuning?

When the documents fit in the prompt. If your whole reference set is a few dozen pages, send it with every request and use prompt caching so repeat reads are billed at a fraction of the input price. On Claude Sonnet 5.5, for example, a cached read lists $0.10 per million tokens against $2 uncached, checked October 2026. There is no index to maintain and no retrieval step to miss the right passage.

All Access · all four programs · $99/mo

Decided on RAG? Build it on a stack you own.

AI SaaS Builder, included in All Access, builds retrieval on Supabase with pgvector and the Claude API, then covers caching, cost control and shipping the product around it, alongside the other three programs, live coaching and the private community.

Start All Access — $99/mo →30-day money-back guarantee
Free · no signup

Not sure which box you are in?

Describe the job and your documents in the free Discord and get a second opinion before you spend on either route.