Checked October 2026, mid-tier models from OpenAI (GPT-6.1 Sol) and Anthropic (Claude Sonnet 5.5) both list $2 input and $10 output per million tokens, small models like GPT-6 Luna and Claude Haiku 5.5 list $0.10 and $0.50, and open models on Vertex AI go lower still. Batch processing halves the price on every major provider. Budget on tokens per finished task, not on the headline rate.
In an LLM API pricing comparison for October 2026, the mid tier is a tie: OpenAI's GPT-6.1 Sol and Anthropic's Claude Sonnet 5.5 both cost $2 per million input tokens and $10 per million output tokens, while the cheapest small models (GPT-6 Luna, Claude Haiku 5.5) cost $0.10 and $0.50. Google's Gemini 3.8 Flash sits between them at an introductory $0.75 and $3.75 until the end of 2026, and hosted open models such as gpt-oss-120b go as low as $0.09 and $0.36.
Prices checked October 2026 on OpenAI API pricing, Anthropic's Claude pricing docs, the Gemini Developer API pricing page, Google Cloud Vertex AI pricing (Gemini Pro, Flash-Lite and the open models) and Mistral API pricing. All prices are USD per million tokens, standard (non-batch) tier.
This page is the hub for API costs on iimagined.ai. It holds one table we re-check every month, a formula you can run on paper, a worked example and the things the table leaves out. Per-provider deep dives sit underneath it, starting with our Claude API pricing breakdown.
LLM API pricing comparison table (October 2026)
| Model | Provider | Input / 1M | Output / 1M | Notes |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | Flagship; prompts over 272K tokens cost more |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | Top Claude tier; 1M context at standard price |
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 | Cache reads at 0.05x input |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | Up to 200K prompt; $4 / $18 above that | |
| Claude Sonnet 5.5 | Anthropic | $2.00 | $10.00 | Mid tier |
| GPT-6.1 Sol | OpenAI | $2.00 | $10.00 | Mid tier; cached input $0.10 |
| Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | Cached input $0.15 |
| Gemini 3.8 Flash | $0.75 | $3.75 | Introductory to Dec 31, 2026; $1.50 / $7.50 from Jan 1, 2027 | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Vertex AI global endpoint | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | Text, image and video input | |
| Mistral Small 4 | Mistral | $0.15 | $0.60 | Cached input $0.015 |
| Claude Haiku 5.5 | Anthropic | $0.10 | $0.50 | Prompts up to 100K tokens; $0.50 / $2.50 above |
| GPT-6 Luna | OpenAI | $0.10 | $0.50 | Cached input $0.01 |
Two patterns stand out. First, output is four to eight times the price of input on every model in the table, so the length of answers drives most bills. Second, the newest generation cut some prices: Claude Opus 5.5 lists $4 and $20, below the $5 and $25 of Claude Opus 5, and Anthropic made Claude Sonnet 5's introductory $2 and $10 its standard price.
Open models: what hosted inference costs
Open-weight models can be self-hosted, but most teams start on a managed endpoint. These are Google Cloud's Vertex AI list prices for partner open models, checked October 2026:
| Open model (Vertex AI) | Input / 1M | Output / 1M |
|---|---|---|
| DeepSeek-V3.2 | $0.56 | $1.68 |
| Llama 4 Maverick | $0.35 | $1.15 |
| Qwen3-235B-A22B-Instruct-2507 | $0.22 | $0.88 |
| gpt-oss-120b | $0.09 | $0.36 |
| gpt-oss-20b | $0.07 | $0.25 |
Other hosts price the same weights differently, and self-hosting swaps token prices for GPU hours, so treat this as one reference point rather than the floor. If you are weighing an open model for multilingual work, our Qwen 3 guide covers what the family is good at.
How to run an LLM API cost comparison
We have not shipped an interactive calculator for this yet, so here is the method it will use. It takes ten minutes with a spreadsheet.
- 1Count requests per day
Use logs if you have them; otherwise estimate from users times actions per user.
- 2Measure average tokens per request
Input = system prompt + context + user message + tool definitions. Output includes reasoning tokens.
- 3Multiply out to a month
Monthly input tokens = requests per day x input tokens x 30. Same for output.
- 4Apply the price per million
Cost = input millions x input price + output millions x output price.
- 5Apply discounts you will really use
Batch roughly halves eligible traffic; cached input is billed at a fraction of the input price.
- 6Test on real prompts
Run 50 to 100 real requests per candidate model and replace your estimates with measured token counts.
Worked example: one workload across eight models
Illustrative inputs, not data from a real customer: a support assistant handling 2,000 requests a day, with 1,500 input tokens and 400 output tokens per request. Over 30 days that is 90 million input tokens and 24 million output tokens. At October 2026 standard list prices, with no batch or caching:
Illustrative workload. Cost = 90 x input price + 24 x output price. Source: provider pricing pages, checked October 2026
The spread is a factor of 100 between the top and bottom bars for the same traffic. That is why model choice matters more than any single optimisation. The usual pattern is routing: send classification, extraction and short replies to a small model, and escalate only the hard cases to a mid or top tier. If half of this example's traffic could run as a nightly batch job, that half would cost about 50% less again.
Cheapest LLM API: it depends on the task, not the row
- The task has one right answer (classify, extract, route)
- Outputs are short and checked by code
- You can retry cheaply on failure
- Volume is high and latency matters
- A wrong answer costs a refund or a client
- The task needs multi-step reasoning or code
- Long documents must be read in one pass
- A human would otherwise review every output
For the provider-specific detail, use the spokes: Claude pricing across plans and API, ChatGPT API workflows and their real costs, and Claude vs ChatGPT for the quality side of the trade-off. To get keys, see getting an OpenAI API key and getting a Claude API key.
Batch, caching and free tiers by provider
- OpenAI. Batch and Flex are listed at half the standard rate (GPT-6.1 Sol drops to $1 and $5). Prompts over 272K input tokens are billed at higher long-context rates for the whole request. Regional data-residency endpoints add 10% on newer models.
- Anthropic. The Batch API takes 50% off input and output. Claude 4.6 and later models (except Haiku 5.5) include the full 1M-token context window at standard price. Pinning US-only inference adds a 1.1x multiplier. New accounts get a small amount of free credit to test.
- Google. Batch and Flex are half price on the paid tier. The Gemini Developer API has a free tier for testing, and Google states free-tier content is used to improve its products while paid-tier content is not.
- Mistral. The pricing page lists Standard, Batch and Priority tiers, with cached input billed at a tenth of the input price on the models in our table.
What the price per token does not tell you
Caching is the other big lever. OpenAI lists cached input for GPT-6.1 Sol at $0.10 against $2 for fresh input. Anthropic bills cache reads at a tenth of the input price on most models (lower on Opus 5.5) and charges a premium to write the cache. Google lists Gemini 3.8 Flash cached input at $0.075 plus an hourly storage fee. If your assistant re-sends the same long system prompt, a knowledge base or a tool list, caching often saves more than switching model.
LLM API pricing in 2026: what changes next
- Nov 21, 2026OpenAI GPT-5.6 Sol promo
OpenAI says promotional pricing runs at least through this date.
- Dec 31, 2026Gemini Flash intro pricing ends
Gemini 3.8, 3.7 and 3.6 Flash move from $0.75 / $3.75 to $1.50 / $7.50 on January 1, 2027.
- MonthlyThis table is re-checked
We update prices against the official pages and note changes here.
Source: OpenAI and Google pricing pages, checked October 2026
If you are pricing a product built on these APIs, your margin depends on this table. Our AI SaaS build guide covers how to set plans above your token cost, and the AI SaaS Builder program walks through routing, caching and usage limits in a real app.
LLM API pricing: FAQ
Which LLM API is cheapest in 2026?
On list price per million tokens, the cheapest options checked in October 2026 are open models hosted on Vertex AI, such as gpt-oss-20b at $0.07 input and $0.25 output, followed by small proprietary models like GPT-6 Luna and Claude Haiku 5.5 at $0.10 input and $0.50 output. Cheapest per token is not cheapest per finished task: a small model that needs retries can cost more than a mid-tier model that gets it right once.
How do I compare LLM API costs fairly?
Estimate monthly input and output tokens for your real request mix, multiply each by the model's per-million price, and add them. Then adjust for batch (about half price on all four major providers), cached input and long-context surcharges. Finally, run the same 50 to 100 real prompts through each candidate and count tokens, because tokenizers differ and the same text can produce different token counts.
Is OpenAI or Anthropic cheaper?
At matched tiers they are close. Checked October 2026, GPT-6.1 Sol and Claude Sonnet 5.5 both list $2 input and $10 output per million tokens, and GPT-6 Luna and Claude Haiku 5.5 both list $0.10 and $0.50. At the very top, Claude Fable 5.1 and GPT-6 Astra both list $10 and $50, while Claude Opus 5.5 sits below at $4 and $20. Caching rules and long-context pricing decide most real bills.
Does the Gemini API have a free tier?
Yes. Google's Gemini Developer API pricing page lists a free tier for models such as Gemini 3.8 Flash, with free input and output tokens under rate limits, then paid pricing per million tokens. Free-tier use is meant for testing and small projects, and Google states that free-tier data may be used to improve its products, so production and client data belong on the paid tier.
How much does batch processing save?
Anthropic, OpenAI, Google and Mistral all list batch processing at roughly half the standard token price, checked October 2026. The trade-off is time: batch jobs are asynchronous, and Google targets a 24-hour turnaround. Use it for anything that does not need an answer while a user waits, such as nightly summaries, bulk classification, content drafts or evaluation runs.
Why is my real API bill higher than the price table suggests?
Usually output tokens, retries and hidden input. Reasoning and thinking tokens bill as output on most providers, tool definitions and system prompts are re-sent on every call, and conversation history grows each turn. Anthropic also notes its newer tokenizer produces about 30% more tokens for the same text. Measure tokens per finished task, not per request, before you budget.
Build on these APIs without guessing the bill.
All Access includes the AI SaaS Builder program, with the other three programs, live coaching and the private community, in one subscription.
Comparing models for a build?
Join the free Discord and ask what other builders are paying per task on the models they run.