Tickd.ai
API errors

Gemini API 429 Error: How to Fix Rate Limits

Updated 10/9/2026

A 429 Too Many Requests error indicates that your application has exceeded its allocated rate limit or quota on the Gemini API. This occurs when you send more requests per minute (RPM) or tokens per minute (TPM) than your current tier allows.

While the Gemini API offers a generous free tier, its strict rate limits can easily be triggered during development, multi-threaded tasks, or sudden traffic spikes. Follow these practical steps to troubleshoot, bypass, and resolve Gemini API 429 errors.

1. Verify Your Current Quota Limits Before modifying your codebase, identify which specific rate limit you are hitting. Gemini limits are split into Requests Per Minute (RPM), Requests Per Day (RPD), and Tokens Per Minute (TPM).

  1. Open the Google AI Studio console.
  2. Navigate to Plan and Billing or look at your project settings.
  3. If you are on the Free Tier, your limits are strictly capped (typically 15 RPM, 1,500 RPD, and 1 million TPM on models like Gemini 1.5 Flash).
  4. Check if your current application flow is exceeding these limits. For example, loop-based API calls without delays will trigger a 429 instantly.

2. Implement Exponential Backoff in Your Code If your code makes simultaneous or rapid consecutive calls, you must handle 429 errors gracefully using an retry mechanism with exponential backoff. This prevents your script from crashing and retries the request after a calculated delay.

Here is a clean Python example using the official google-generativeai SDK and the tenacity library to handle rate limits automatically:

`python import time import google.generativeai as genai from google.api_core import exceptions

genai.configure(api_key="YOUR_GEMINI_API_KEY") model = genai.GenerativeModel('gemini-1.5-flash')

def generate_with_retry(prompt): max_retries = 5 delay = 2 # initial delay in seconds for attempt in range(max_retries): try: response = model.generate_content(prompt) return response.text except exceptions.ResourceExhausted as e: if attempt == max_retries - 1: raise e print(f"Rate limit hit (429). Retrying in {delay} seconds...") time.sleep(delay) delay *= 2 # Double the wait time for the next attempt `

3. Enable Context Caching for Large Prompts If you are sending massive system instructions, documents, or video files with every request, you might be hitting the Tokens Per Minute (TPM) limit rather than the Request Per Minute (RPM) limit.

To fix this: * Use Gemini Context Caching for repetitive, heavy contexts (e.g., a 100k token reference document used across multiple prompts). * Caching reduces the active token overhead on subsequent API calls, preventing you from hitting the TPM ceiling while also reducing API costs.

4. Upgrade to a Pay-As-You-Go Plan If your application requires more than 15 requests per minute, you must transition from the free tier to the pay-as-you-go tier in Google AI Studio.

  1. Go to the Google AI Studio dashboard.
  2. Click on Set up billing or Upgrade to Pay-As-You-Go.
  3. Link a valid Google Cloud Billing account.
  4. Once billing is active, your limits will scale up significantly (e.g., up to 360 RPM and 4 million TPM on paid tiers, depending on the model chosen).

5. Implement Client-Side Rate Limiting Rather than relying on the API to reject your requests, control the flow of outbound calls directly inside your application infrastructure.

  • Implement a Queue: Use a queue system (like Redis, Celery, or BullMQ) to throttle outgoing requests to the Gemini API.
  • Introduce Artifical Delays: If processing a batch of inputs, add a simple sleep timer between iterations (e.g., time.sleep(4) to guarantee you stay under 15 RPM on the free tier).

When to escalate If you have already upgraded to a Pay-As-You-Go account and are still receiving persistent 429 errors despite keeping your traffic below standard limits, your project might have a custom quota restriction placed on it.

Log in to the Google Cloud Console, navigate to IAM & Admin > Quotas & System Limits, and search for the "Generative Language API" limits. If you require higher capacity for enterprise needs, you can submit a formal Quota Increase Request directly through that dashboard.

While you're here

Tickd is more than troubleshooting — these three are free and take seconds.

Agent BuilderDesign your own AI agent and export it to ChatGPT, Claude, Gemini or Grok.Build one free