Tickd.ai
API errors

How to Fix Claude API 429 Rate Limit Exceeded

Updated 9/16/2026

The HTTP 429 Too Many Requests error occurs when your application exceeds Anthropic’s rate limits. The Claude API enforces these limits on three distinct vectors: Requests Per Minute (RPM), Tokens Per Minute (TPM), and Tokens Per Day (TPD).

When you cross any of these thresholds, the API blocks subsequent requests and returns a rate_limit_error. This guide covers how to identify which limit you are hitting and how to configure your application to handle and prevent these blocks.

1. Diagnose Your Rate Limit Type Before changing code, determine which exact limit is being tripped. Check the error payload returned by the Anthropic API. It typically contains a detailed message indicating whether you exceeded your RPM, TPM, or TPD allocation.

Additionally, inspect the HTTP response headers in your failed API calls. Anthropic provides real-time limit tracking headers: * anthropic-ratelimit-requests-limit: Your maximum allowed requests per minute. * anthropic-ratelimit-requests-remaining: The number of requests you have left in the current minute window. * anthropic-ratelimit-requests-reset: The UTC timestamp when your request count resets. * anthropic-ratelimit-tokens-limit and anthropic-ratelimit-tokens-remaining: Your current TPM status.

If your remaining count is consistently near zero, your application's concurrency is too high for your current Tier.

2. Implement Exponential Backoff with Jitter Do not immediately retry failed requests in a loop. Doing so creates a "thundering herd" problem that keeps your API key rate-limited. Instead, implement exponential backoff with randomized jitter.

This algorithm increases the waiting time between retries exponentially and adds a small amount of random delay to prevent synchronized retries.

Here is a conceptual Python implementation using the official anthropic SDK, which has built-in retry mechanisms that you can customize:

`python import anthropic import time import random

client = anthropic.Anthropic( api_key="your_api_key", max_retries=5 # The SDK automatically uses exponential backoff )

To implement manual backoff for custom pipelines: def call_with_backoff(prompt, max_attempts=5): base_delay = 1.0 # start with 1 second for attempt in range(max_attempts): try: message = client.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=1024, messages=[{"role": "user", "content": prompt}] ) return message except anthropic.RateLimitError as e: if attempt == max_attempts - 1: raise e # Calculate exponential delay with jitter delay = (base_delay * (2 ** attempt)) + random.uniform(0, 1) print(f"Rate limited. Retrying in {delay:.2f} seconds...") time.sleep(delay) ```

3. Reduce Token Consumption per Request If you are hitting TPM (Tokens Per Minute) limits, you can prolong your budget by optimizing your payload size. The TPM metric counts both input tokens (prompt, system prompt, history) and output tokens.

  • Trim Chat History: Do not send the entire history of a long session. Keep only the last 3-5 message exchanges or generate a periodic summary of the conversation to feed into the prompt context.
  • Compress System Prompts: Avoid overly verbose instructions. Consolidate rules and remove repetitive examples.
  • Set Lower max_tokens: Limit the maximum response length to only what is necessary for your use case.

4. Use Client-Side Queueing and Concurrency Limits If your application processes background tasks or bulk jobs (such as document processing), do not run them concurrently without a throttle. Implement a worker pool or a token-bucket queue.

Limit the number of active concurrent workers sending requests to the API. For example, if your RPM limit is 50, configure your worker queue to process a maximum of 40 tasks per minute to leave a safe margin for transient spikes.

5. Upgrade Your Usage Tier Anthropic determines rate limits based on your account's tier, which is tied to your lifetime deposit history. New accounts start in Tier 1 with lower limits (typically 20,000 to 40,000 TPM).

To raise your limits automatically: 1. Log in to the Anthropic Console. 2. Navigate to Billing. 3. Deposit funds to transition to a higher tier. Tier 2 is unlocked at a $40 cumulative deposit, Tier 3 at $200, and Tier 4 at $400. Once the payment clears, your RPM and TPM limits scale up immediately.

When to Escalate If you are at Tier 4 or Tier 5 and require enterprise-level scale that exceeds standard tier limits, do not rely on code-level workarounds. Navigate to the **Limits** section in your Anthropic Console and click "Request Limit Increase" to submit a formal capacity request to the Anthropic team.

Quick fixes

  • Claude is down or not loading
  • Claude Pro billing or payment problem
  • Can't sign in to Claude

While you're here

Tickd is more than troubleshooting — these three are free and take seconds.

Agent BuilderDesign your own AI agent and export it to ChatGPT, Claude, Gemini or Grok.Build one free