How to Fix Claude API 429 Rate Limit Error
Updated 9/11/2026
The Claude API returns an HTTP 429 Too Many Requests status code when your application exceeds its allocated usage limits. Anthropic enforces limits on three distinct metrics: Requests Per Minute (RPM), Tokens Per Minute (TPM), and Requests Per Day (RPD).
When your API keys hit any of these thresholds, the API blocks subsequent requests and throws an error. This guide provides concrete, actionable steps to handle, mitigate, and resolve these rate limits in your application.
1. Inspect the Rate Limit Headers Before changing your code, identify which limit you are hitting. Anthropic includes real-time rate limit metrics in the HTTP response headers of every API call. Write code to inspect these values when debugging:
- anthropic-ratelimit-requests-limit: The total number of requests permitted per minute.
- anthropic-ratelimit-requests-remaining: How many requests you have left in the current minute.
- anthropic-ratelimit-requests-reset: The time remaining until your request counter resets (formatted as ISO 8601 duration/time).
- anthropic-ratelimit-tokens-limit: The total number of tokens permitted per minute.
- anthropic-ratelimit-tokens-remaining: The number of tokens remaining before you are throttled.
- anthropic-ratelimit-tokens-reset: The time remaining until your token pool resets.
Log these headers in your application's error-handling block to determine if your bottleneck is request volume (RPM) or payload size (TPM).
2. Implement Exponential Backoff with Jitter Do not immediately retry failed requests at fixed intervals, as this creates a "thundering herd" problem that keeps your keys blocked. Instead, use exponential backoff with a randomized delay (jitter).
Here is how to implement backoff handling in Python using the tenacity library:
`python import anthropic from tenacity import retry, stop_after_attempt, wait_random_exponential
client = anthropic.Anthropic(api_key="your_api_key")
@retry( wait=wait_random_exponential(min=1, max=60), stop=stop_after_attempt(5), retry=lambda e: isinstance(e, anthropic.RateLimitError) ) def generate_text_with_backoff(prompt): response = client.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=1024, messages=[{"role": "user", "content": prompt}] ) return response `
If you are using Node.js/TypeScript, you can build a manual retry loop:
`javascript async function callClaudeWithRetry(prompt, retries = 5, delay = 1000) { try { return await client.messages.create({ model: "claude-3-5-sonnet-20241022", max_tokens: 1024, messages: [{ role: "user", content: prompt }], }); } catch (error) { if (error.status === 429 && retries > 0) { // Calculate randomized exponential delay const jitter = Math.random() * 1000; const nextDelay = delay * 2 + jitter; console.warn(Rate limited. Retrying in ${nextDelay.toFixed(0)}ms...); await new Promise(resolve => setTimeout(resolve, nextDelay)); return callClaudeWithRetry(prompt, retries - 1, delay * 2); } throw error; } } `
3. Reduce Token Consumption Per Request If you are hitting TPM (Tokens Per Minute) limits, optimizing your prompts and configurations will prevent throttling:
- Reduce max_tokens: Set the max_tokens parameter only as high as necessary. Do not default to the maximum allowed length if you only need short responses.
- Prune System Prompts: Avoid passing large, static instructions or repetitive documentation in every request. Use system prompts efficiently.
- Truncate Chat History: When sending multi-turn conversation histories, remove older turns. Only send the last 3 to 5 exchanges to preserve context without bloating token usage.
- Use Prompt Caching: If you send large, static context blocks (such as API documentation or reference files) repeatedly, implement Anthropic's Prompt Caching feature. This reduces token count and speeds up processing time.
4. Upgrade Your Account Tier Anthropic API accounts are divided into tiers based on payment history. New accounts (Tier 1) have highly restrictive limits.
To increase your limits systematically: 1. Navigate to the Anthropic Console. 2. Go to the Billing section. 3. Add funds to your balance. Your tier upgrades automatically as you meet lifetime spend requirements (e.g., Tier 2 requires a $40 lifetime deposit, Tier 3 requires $100).
Updating your payment tier takes effect within minutes, instantly multiplying your RPM and TPM allocations.
When to escalate If you have upgraded your account tier but your limits have not refreshed within an hour, or if you are on Tier 4/5 and require custom enterprise limits, navigate to the **Anthropic Console**, click the Help button (question mark icon), and submit a "Rate Limit Increase Request" with your estimated usage metrics.
Quick fixes
- Claude is down or not loading
- Claude Pro billing or payment problem
- Can't sign in to Claude