Fix Gemini API Error 429 Rate Limit Exceeded
Updated 9/22/2026
Getting a 429 Resource Exhausted error while building with the Gemini API can halt your application entirely. This error indicates that you have exceeded the rate limits allocated to your API key or project. Google enforces strict rate limits on both free-tier and pay-as-you-go billing plans to prevent server abuse.
Here is how to identify why your requests are being throttled and the exact steps you need to take to fix the error.
Why Gemini Returns Error 429
When using the Gemini API (via Google AI Studio or Vertex AI), rate limits are measured in three ways: 1. RPM (Requests Per Minute): The maximum number of API calls you can make in 60 seconds. 2. RPD (Requests Per Day): The total limit allowed in a 24-hour cycle (applicable primarily to the free tier). 3. TPM (Tokens Per Minute): The volume of raw data (input prompt text + generated output response) processed per minute.
If you exceed any of these thresholds, the Gemini API gateway instantly rejects your connection and returns a 429 status code.
Follow these steps to eliminate Gemini API rate limit errors in your application.
Step 1: Implement Exponential Backoff in Your Code
If your code sends bursts of traffic, standard sequential requests will trigger a 429 error. Instead of failing immediately when a request is rejected, your code should wait and try again using a strategy called exponential backoff.
Here is how to implement basic exponential backoff in Python when calling the Gemini API:
`python import time import google.generativeai as genai from google.api_core import exceptions
genai.configure(api_key="YOUR_GEMINI_API_KEY") model = genai.GenerativeModel('gemini-1.5-flash')
def generate_with_retry(prompt, max_retries=5): delay = 1.0 # Start with a 1-second delay for attempt in range(max_retries): try: response = model.generate_content(prompt) return response.text except exceptions.ResourceExhausted as e: if attempt == max_retries - 1: raise e print(f"Rate limit hit. Retrying in {delay} seconds...") time.sleep(delay) delay *= 2 # Double the wait time for the next attempt `
By doubling the delay after every rejected request, you allow the rate-limit window to clear before spamming the server again.
Step 2: Check Your Quotas in Google AI Studio
If you get 429 errors on your very first request of the day, your project might have reached its daily limit, or Google may have temporarily restricted your API key due to policy issues.
- Open Google AI Studio and sign in.
- Navigate to your Plan & Billing dashboard or view your key details.
- Check your active tier. The free tier of Gemini 1.5 Flash has a limit of 15 RPM, 1,500 RPD, and 32,000 TPM.
- If your daily active requests have reached 1,500, your API key will reject all requests until the daily reset time (UTC midnight).
Step 3: Optimize Your Token Usage (TPM)
Sometimes you are not exceeding your requests per minute (RPM), but you are hitting your tokens per minute (TPM) limit. Large prompts, extensive system instructions, or long conversational histories consume massive token volumes.
- Trim input prompts: Strip out unnecessary context, boilerplate instructions, or repetitive examples in your system prompts.
- Limit output tokens: Use the generation_config parameter to set a lower max_output_tokens value. This stops the model from writing excessively long responses that eat into your TPM.
- Clear chat history: If you are building a chatbot, prune older turns of the conversation history from your input payload to reduce the cumulative token overhead per request.
Step 4: Upgrade to Pay-As-You-Go Billing
If your application requires higher throughput, the free tier limits will not suffice. Upgrading your account lifts the low limits and converts your usage into a paid structure.
- Go to the Google AI Studio Billing section.
- Click Set up billing and link a valid Google Cloud Platform (GCP) billing account.
- Once billing is active, your limits are scaled up dramatically (for example, Gemini 1.5 Flash scales up to 1,000 RPM and 4,000,000 TPM with pay-as-you-go pricing).
When to Escalate
If you are on a paid tier, have implemented exponential backoff, and still experience persistent 429 errors that fail to clear after several minutes, check the official Google Cloud Status Dashboard or the Vertex AI status page. A sudden spike in 429 errors across multiple projects can indicate a regional infrastructure outage on Google's side. If systems are healthy, you must submit a Quota Increase Request directly through your Google Cloud Console's IAM & Admin Quotas page.