The Gemini API free tier gives free input and output tokens on certain models, including Gemini 3.8 Flash and the Flash-Lite models, under per-project limits on requests per minute, input tokens per minute and requests per day. Google does not publish those numbers per model in its docs; your project's real limits are in Google AI Studio. A 429 comes in three codes, and only two of them are fixed by retrying. Paid starts with a $5 prepayment.
The Gemini API free tier gives you free input and output tokens on certain models (Gemini 3.8 Flash, Gemini 3.6 Flash, the Flash-Lite models, Gemini Embedding 2 and Gemma 4 among them), capped by three per-project limits: requests per minute, input tokens per minute and requests per day. Google's rate-limits page does not list those numbers per model; it sends you to Google AI Studio, where the limits for your own project are shown.
Checked October 2026 against Google's rate limits page (last updated October 9, 2026), the Gemini Developer API pricing page (October 9, 2026), the billing guide, the API errors reference, the troubleshooting guide, the deprecations page and the Gemini API additional terms. Google changes limits and prices often, so treat every number here as dated.
That last point shapes this whole page. Any table of "free tier RPM per model" you find on a blog is a snapshot of somebody's project on some past date. What Google does publish is which models are free, how the three limits behave, what each error code means and what each paid tier costs to reach. Those are below, in that order, with the place to read your own numbers. For how Gemini's paid prices sit beside OpenAI and Anthropic, see our LLM API pricing comparison, the hub for this topic.
What the Gemini API free tier includes
The free tier is a list of models, not the whole catalogue. Google's pricing page describes it as limited access to certain models, free input and output tokens, Google AI Studio access, and content used to improve its products. Each model then has its own table with a Free Tier column that says either "Free of charge" or "Not available".
| Model | Model code | Free tier (standard) | Paid tier, input / output per 1M tokens |
|---|---|---|---|
| Gemini 3.8 Flash | gemini-3.8-flash | Free of charge | $0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50 |
| Gemini 3.6 Flash | gemini-3.6-flash | Free of charge | $0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50 |
| Gemini 3.5 Flash-Lite | gemini-3.5-flash-lite | Free of charge | $0.30 / $2.50 |
| Gemini 3.1 Flash-Lite | gemini-3.1-flash-lite | Free of charge | $0.25 / $1.50 (audio input $0.50). Shutdown date May 7, 2027 |
| Gemini 3 Flash Preview | gemini-3-flash-preview | Free of charge | $0.50 / $3.00 (audio input $1.00) |
| Gemini Embedding 2 | gemini-embedding-2 | Free of charge | $0.20 text input |
| Gemma 4 | see model card | Free of charge | No paid tier listed |
| Gemini 3.1 Pro Preview | gemini-3.1-pro-preview | Not available | $2.00 / $12.00 up to 200K-token prompts, $4.00 / $18.00 above |
| Gemini Nano Banana 2.1 (images) | gemini-nano-banana-2.1 | Not available | $1.50 input; $7.50 text and $30.00 image output |
| Veo 3.1 (video) | veo-3.1-generate-preview | Not available | From $0.05 per second (Lite, 720p) |
Source: Gemini Developer API pricing page, standard tier, USD, checked October 2026. The page lists more models (Live, TTS, Transcribe, Robotics) with their own rows.
Four things in that page are easy to miss:
- The Pro model is paid only. Gemini 3.1 Pro Preview shows "Not available" in the free column. If a tutorial tells you to call it with a free key, expect an error, not a slow response.
- Batch and Flex are paid only. The half-price Batch and Flex rows show "Not available" on the free tier for the Flash models, so the free tier cannot be stretched with overnight jobs.
- Gemini 2.5 models still list a free tier, but Google's deprecations page says access to them is now limited to users who have actively used them before, and tells new projects to use Gemini 3.5 Flash-Lite or 3.8 Flash.
- Grounding with Google Search is inconsistent on the page. The Gemini 3 model tables list it as "Not available" on the free tier (with a note that it can be tested in AI Studio), while the tools table lists a free daily allowance for Flash and Flash-Lite. Test it on your own key before you design around it.
How the Gemini API rate limit works on the free tier
Every request is measured three ways, and going over any one of them is enough to fail it. Google's example: with an RPM limit of 20, the 21st request in a minute errors even if you are nowhere near the token or daily limit.
- RPM, requests per minute. How many calls, regardless of size.
- TPM, tokens per minute. Input tokens only. A long system prompt or a pasted document can trip this while your request count looks harmless.
- RPD, requests per day. Resets at midnight Pacific time, not 24 hours after your first call.
- 01Your app sends a request
Any API key in the project. They all share one allowance.
- 02Google counts it three ways
Requests this minute (RPM), input tokens this minute (TPM) and requests today (RPD), for the model you called.
- 03Under every limit: a response
Capacity is still not promised: Google says actual capacity may vary from the specified limits.
- 04Over any one limit: HTTP 429
With a code: rate_limit_exceeded, too_many_requests or quota_exceeded.
- 05Per-minute code: back off
Wait 1s, 2s, 4s, 8s with jitter, then retry. The official SDKs retry transient errors by default.
- 06Daily code: stop
Queue the work until the reset, call a different model, or move the project to Tier 1.
Three rules from the same page explain most "why am I limited" threads. Limits vary by model, and some models add their own dimension, such as images per minute for the image models or tokens per day. Experimental and preview models have tighter limits than stable ones, which matters because several free models carry "Preview" in their name. And the limits belong to the project, not to the key.
Gemini free tier RPM: where to find your real numbers
Your project's RPM, TPM and RPD are in Google AI Studio, not in the docs. The rate-limits page says limits "depend on a variety of factors (such as your usage tier) and can be viewed in Google AI Studio", and that they update automatically as your tier and account status change. So the lookup takes two minutes and beats any third-party table.
- 1Open the rate limit page in Google AI Studio
aistudio.google.com/rate-limit, signed in with the account that owns the key. The docs link to it as "View your active rate limits in AI Studio".
- 2Make sure you are on the right project
Limits are per project. A key created in another project has its own numbers.
- 3Find the model you actually call
Write down its RPM, TPM and RPD with today's date. Preview models are usually the tight ones.
- 4Check the tier on the Projects page
The Billing Tier column shows the project's tier and whether billing is set up.
- 5Compare with real traffic
Dashboard, then Usage, shows what the project has been sending. Look at the peak minute, not the daily average.
- 6Re-check after any change
A new model, a tier upgrade or a Google update all move the numbers. Read the page again instead of trusting an old screenshot.
Once you have the three numbers, the planning is arithmetic. Illustrative inputs: if your model shows 10 RPM and your job is 600 documents, the fastest clean run is one request every six seconds, about an hour, provided 600 also fits under the RPD figure. If it does not, the job spans two Pacific days or it needs a paid tier. Do this sum before you write the loop, not after the first 429.
Gemini API 429 errors: what each code means
A 429 is three different problems wearing one status code, and only two of them are fixed by retrying. Google's API errors reference, written for the Interactions API, gives each its own code. The rate-limits page and troubleshooting guide use an umbrella status name for the same condition, 429 RESOURCE_EXHAUSTED, so you will see both in logs and forum posts.
| Code | HTTP | What Google says it means | What to do |
|---|---|---|---|
| rate_limit_exceeded | 429 | You exceeded the per-minute or per-second request or token limit. | Back off exponentially and retry. Smooth bursts with a queue. |
| too_many_requests | 429 | Too many requests in a short period of time. | Same fix: backoff with jitter, and cap how many calls run at once. |
| quota_exceeded | 429 | You exceeded your daily quota. | Stop retrying. Wait for the reset, use another model, or move to a paid tier. |
| payment_required | 402 | Your Prepay credit balance is depleted. | Add credits or turn on auto-reload. Retrying cannot succeed. |
| service_unavailable | 503 | The service is temporarily overloaded or down. | Back off and retry. This is Google's capacity, not your quota. |
| permission_denied | 403 | Your API key does not have permission for this resource. | Check the key and the project. Do not retry. |
The two non-429 rows are there because they get misread as rate limiting. A 402 means a paid project ran out of prepaid credit: the billing guide says every key on that billing account stops at once when the balance reaches zero. A 503 is Google being overloaded. Neither is solved by lowering your request rate.
Paid projects have one more way to earn a 429: a spend-based limit, measured over a rolling 10-minute window ($10 on Tier 1, $50 on Tier 2, $200 on Tier 3, checked October 2026). It does not apply to the free tier.
- Retry only transient errors: 429, 408 and 5xx
- Never retry 400, 402 or 403. They fail the same way every time
- Start around 1 second and double the wait: 1s, 2s, 4s, 8s
- Add random jitter so parallel workers do not all retry in the same instant
- Set a maximum number of attempts, then queue the job or fail loudly
- Read the code first: quota_exceeded means wait for the daily reset, not retry
- Do not wrap your own loop around an SDK that already retries
The last item saves real quota. Google says its client SDKs retry transient errors with exponential backoff by default; the Python SDK, for example, retries up to four times with a first delay of about one second and a ceiling of 60 seconds. Put a hand-written retry loop around that and one failing call can become a dozen or more requests against the same per-minute limit.
How to stay inside the free tier longer
Most free-tier 429s come from bursts and bloated prompts, not from real volume. These levers are ordered by how little work they take.
- Queue instead of bursting. Send requests at a steady pace just under your RPM rather than firing a batch in parallel. A loop with a fixed delay is enough for a script.
- Cut input tokens. TPM counts input. Trim the system prompt, send the paragraph instead of the document, and stop re-sending the whole conversation when the last few turns will do. Our prompt engineering guide covers writing tighter prompts.
- Combine small jobs. RPD counts requests, not items. Classifying 20 short reviews in one prompt uses one request; classifying them one at a time uses 20.
- Use the smallest model that passes your test. Limits vary by model. Keep the larger Flash model for the steps that need it and send routine work to a Flash-Lite model.
- Cache your own answers. If two users ask the same question, the second answer should come from your database, not from another API call.
- Have a fallback. When the daily code arrives, route to a second model or queue the job, and tell the user it is delayed rather than showing a raw error.
None of this turns the free tier into production hosting. It buys you a longer, calmer prototype. If you are building a product on an LLM API, our AI SaaS Builder program covers the production side on the Claude API: errors, rate limits and fallbacks, cost control with caching, and a per-user quota on every AI route so one customer cannot spend the whole allowance.
When to pay: the tiers and the jump math
Moving off the free tier costs a $5 prepayment and usually takes effect at once (some accounts are offered a pay-afterwards Postpay plan instead). You link a billing account to the project in AI Studio, buy at least $5 of credit, and the project becomes Tier 1. From there, tiers rise with cumulative spend on the billing account and time since the first payment.
| Usage tier | Qualification | Billing tier cap | Spend rate limit per 10 minutes |
|---|---|---|---|
| Free | Active project or free trial | N/A | N/A |
| Tier 1 | Set up and link an active billing account | $250 | $10 |
| Tier 2 | Paid $100 + 3 days from first successful payment | $2,000 | $50 |
| Tier 3 | Paid $1,000 + 30 days from first successful payment | $20,000 - $100,000+ | $200 |
Source: Gemini API rate limits and billing pages, checked October 2026. The Tier 2 and 3 spend counts all Google Cloud services on the billing account, not only the Gemini API.
- Day 0Free tier
An active project. Free tokens on certain models, tight per-project limits, content used to improve Google products.
- Same dayTier 1
Link a billing account and prepay at least $5. Google says this upgrade typically takes effect instantly.
- $100 paid + 3 daysTier 2
Automatic once the billing account qualifies. Later upgrades take effect within 10 minutes.
- $1,000 paid + 30 daysTier 3
The top self-serve tier. Beyond it, Google has a rate limit increase form for paid projects and offers no promise of approval.
Source: Gemini API rate limits page, checked October 2026
What Google does not print is the Tier 1 RPM per model: like the free numbers, it shows on your AI Studio rate limit page once the project upgrades. What it does print is the price, so the money side can be worked out in advance. Illustrative inputs: 1,000 requests a day, each with 2,000 input tokens and 500 output tokens, is 60 million input and 15 million output tokens over 30 days.
Arithmetic from list prices: input millions x input price + output millions x output price. The workload is an illustrative input, not measured traffic. Source: Gemini Developer API pricing page, standard tier, checked October 2026
Scale it to your own case by dividing. A side project at a tenth of that volume, 100 requests a day, costs about $3.75 a month on Gemini 3.1 Flash-Lite. Two things push the real bill above the sum: thinking tokens are billed as output, and Gemini 3.8 Flash doubles in price on January 1, 2027, so budget on the later figure. The cheapest bar has an expiry too: Google lists a shutdown date of May 7, 2027 for Gemini 3.1 Flash-Lite, with Gemini 3.5 Flash-Lite as the replacement. For the same workload priced on other providers, use the tables in our LLM API pricing comparison and the Claude API pricing breakdown.
Three billing details catch people after the upgrade. Prepaid credits expire 12 months after purchase and are non-refundable unless you switch to a Postpay plan. When the balance hits zero, requests fail with a 402 until you add credit, so turn on auto-reload with a monthly limit if anything depends on the key. And the billing guide warns that usage can overrun the balance during roughly ten minutes of billing delay, which matters for long batch or agent jobs.
Free tier or paid: how to decide
Pay when a failed request costs more than the tokens would have. For a learning project that threshold never arrives. For anything with users it arrives on the first day, and there is a second reason that has nothing to do with speed: data.
Google's terms say content sent through unpaid quota is used to improve its products, that human reviewers may read it, and that you should not submit sensitive, confidential or personal information to the unpaid services. The same terms say you may use only paid services when making your app available to users in the European Economic Area, Switzerland or the United Kingdom. A customer's support ticket is exactly the kind of content that belongs on the paid tier.
- You are learning the API or building a prototype
- Nothing sensitive, confidential or personal goes into the prompts
- A failed request costs you nothing but a rerun
- The only user is you
- You can wait for the midnight Pacific reset
- A real user sees the error when you hit a 429
- Prompts contain customer or client data
- Your app serves users in the EEA, Switzerland or the UK
- You need Gemini 3.1 Pro Preview, Batch or the image models
- Working around the limit has already cost you more than $5 of time
If you are still choosing a provider rather than a tier, the setup guides for the other two are here: getting an OpenAI API key and getting a Claude API key. For what to build once the key works, start with how to build an AI SaaS in 2026.
Gemini API free tier: FAQ
Is the Gemini API free?
Yes, for certain models. Checked October 2026, Google's pricing page lists free input and output tokens on the free tier for Gemini 3.8 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite, Gemini Embedding 2 and Gemma 4, among others. Gemini 3.1 Pro Preview, the image models and Veo 3.1 show Not available on the free tier. Free use is capped by per-project rate limits, and your content can be used to improve Google products.
What are the Gemini API free tier rate limits?
Three limits apply per project: requests per minute (RPM), input tokens per minute (TPM) and requests per day (RPD). Google's rate-limits page, last updated October 9, 2026, does not list the numbers per model. It says limits depend on factors such as your usage tier and can be viewed in Google AI Studio, so the reliable figure is the one on your own project's rate limit page.
Why am I getting a 429 error from the Gemini API?
A 429 means the project went over a limit. Google's error reference lists three codes with that status: rate_limit_exceeded for per-minute or per-second request or token limits, quota_exceeded for the daily quota, and too_many_requests for a burst in a short period. Back off and retry for the first and third. For the daily quota, waiting for the reset at midnight Pacific time or moving to a paid tier are the fixes.
Does creating more API keys increase the Gemini free tier limit?
No. Google states that rate limits are applied per project, not per API key, and that keys inherit the tier limits and billing status of their project. Every key in a project draws on the same RPM, TPM and RPD allowance. The supported ways to get more capacity are to link billing and move to Tier 1, or to request a rate limit increase once you are on a paid tier.
How much does it cost to move from the Gemini free tier to paid?
The minimum is a $5 prepayment, which moves the project to Tier 1 once a billing account is linked (checked October 2026). After that you pay per token: Gemini 3.1 Flash-Lite, for example, lists $0.25 per million input tokens and $1.50 per million output tokens. Unused prepaid credits expire after 12 months, and requests fail with a 402 error when the balance reaches zero.
Does Google use free tier Gemini API data for training?
Google's Gemini API terms say content sent through unpaid quota is used to provide, improve and develop Google products and machine learning technologies, and that human reviewers may read and annotate it after it is disconnected from your account and API key. The terms tell you not to submit sensitive, confidential or personal information to the unpaid services. On the paid tier, the pricing page lists Used to improve our products as No.
When does the Gemini API daily limit reset?
Requests per day quotas reset at midnight Pacific time, according to Google's rate-limits page. That is a fixed clock time, not a rolling 24 hours from your first request, so a project that uses up its daily allowance in the morning stays blocked until the reset. Per-minute limits recover within the minute, which is why backoff helps with those and does nothing for the daily cap.
Past the free tier? Build the product the quota was for.
AI SaaS Builder, included in All Access, takes an AI feature to production on the Claude API (errors, rate limits, fallbacks and caching), with Supabase, deployment and Stripe billing around it, alongside the other three programs, live coaching and the private community.
Hit a limit you cannot explain?
Bring the error code and your model name to the free Discord and compare notes with other people building on LLM APIs.