Fix Gemini API Resource Exhausted Error (429)
Updated 10/7/2026
The RESOURCE_EXHAUSTED error is one of the most common issues developers face when integrating Google Gemini models via Google AI Studio or Vertex AI. This error corresponds to an HTTP 429 status code (Too Many Requests) or a gRPC status code 8.
It indicates that your application has exceeded its allocated quota limit. This could be your Requests Per Minute (RPM), Requests Per Day (RPD), or Tokens Per Minute (TPM) limit. Here is how to diagnose and resolve this issue step by step.
1. Verify Your API Tier and Rate Limits Before modifying your code, you must determine which rate limits you are hitting. Google AI Studio offers a Free Tier and a Pay-as-you-go Tier, both of which have distinct constraints.
- Free Tier Limits: Currently, models like Gemini 1.5 Flash have strict default limits on the free tier (e.g., 15 RPM, 1 million TPM, and 1,500 RPD). If your application sends requests concurrently or processes large batches of documents, you will quickly trigger the RESOURCE_EXHAUSTED error.
- Pay-as-you-go Tier: Upgrading to a paid tier increases your RPM limits significantly and removes the daily limit (RPD), charging you only for the tokens you actually consume.
To check your tier and usage: 1. Open Google AI Studio. 2. Navigate to the Settings (gear icon) or Billing section. 3. Check your active billing plan. If you are on the Free Tier, link a Google Cloud billing account to transition to the pay-as-you-go tier and instantly unlock higher limits.
2. Implement Exponential Backoff in Your Code If your application occasionally spikes in usage, you can handle the `RESOURCE_EXHAUSTED` error gracefully by writing retry logic with exponential backoff. This prevents your code from crashing when it hits a temporary RPM or TPM limit.
Here is an example of how to implement exponential backoff using Python and the google-generativeai SDK:
`python import time import google.generativeai as genai from google.api_core import exceptions
genai.configure(api_key="YOUR_API_KEY") model = genai.GenerativeModel('gemini-1.5-flash')
def generate_text_with_retry(prompt, max_retries=5, initial_delay=2): delay = initial_delay for attempt in range(max_retries): try: response = model.generate_content(prompt) return response.text except exceptions.ResourceExhausted as e: if attempt == max_retries - 1: raise e print(f"Rate limit reached. Retrying in {delay} seconds...") time.sleep(delay) delay *= 2 # Double the wait time for the next attempt `
3. Monitor Your Quotas in Google Cloud Console If you are using Vertex AI or have linked your AI Studio project to Google Cloud, you can track exactly which quota limit is being triggered.
- Go to the Google Cloud Console.
- In the search bar, type "Quotas & System Limits" and select it.
- Filter the Service by Generative Language API (for AI Studio keys) or Vertex AI API.
- Look for metrics like Generate Content requests per minute or Generate Content tokens per minute.
- Check the chart to identify whether you are hitting token limits (TPM) or request limits (RPM).
If you are on a paid billing account and your production traffic naturally exceeds the default paid quotas, you can click the checkbox next to the limit you want to change, click Edit Quotas, and submit a request for an increase.
4. Optimize Token Consumption and Payload Size If you are hitting the Tokens Per Minute (TPM) limit, the issue is likely the size of your prompts rather than the frequency of your calls. You can optimize your token usage with the following practices:
- Trim System Instructions: Keep your system instructions concise. Long, repetitive instructions consume tokens on every single API request.
- Limit Context History: In multi-turn chat applications, prune the chat history. Do not pass the entire conversation history back to the model if only the last few turns are relevant.
- Lower max_output_tokens: Restrict the maximum response length by setting the max_output_tokens parameter in your generation configuration.