Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build a Local Fallback Proxy with Python to Route LLM Requests Between Claude and OpenAI

Don't let API outages sink your application. Learn how to build an ultra-fast, async Python proxy that automatically routes requests from Anthropic to OpenAI during rate limits or downtimes.

Updated 10/5/2026

Relying on a single AI provider for a production-grade application is architectural self-sabotage. APIs fail. Rate limits are hit. Rate-limit increases are ignored during peak hours. If your entire system grinds to a halt because Anthropic's API throws a 503 error, your users won't blame Anthropic—they will blame you.

To build resilient AI systems, you need a local gateway proxy. This sits between your main application code and the LLM providers, intercepting outgoing calls, handling routing, and instantly falling back to an alternative model if the primary provider drops the ball.

In this guide, we are going to build a high-performance, asynchronous fallback proxy in Python using FastAPI and HTTPX. It will default to sending prompts to /platforms/claude (the gold standard for complex reasoning) and fall back seamlessly to /platforms/openai if Anthropic drops or rate-limits our requests.

The Design Philosophy

To keep our proxy fast, we need to respect two design rules: 1. Low Overhead: We cannot introduce latency. We must use asynchronous requests to avoid blocking the main event loop. 2. Payload Normalisation: Claude and OpenAI use slightly different structures for their messaging endpoints. Our proxy will accept a simplified unified JSON format and translate it on the fly depending on which provider is currently serving the request.

If you find yourself needing to master the precise payload terms used by both providers, check out our /glossary for detailed explanations of structured outputs and context windows.

Step 1: Project Setup

Let’s create a new project directory and install the necessary async libraries.

`bash mkdir llm-fallback-proxy cd llm-fallback-proxy python3 -m venv venv source venv/bin/activate

pip install fastapi uvicorn httpx pydantic `

We will use fastapi for our local API, uvicorn to run the server, and httpx to handle our outbound asynchronous requests to Anthropic and OpenAI.

Step 2: Designing the Unified Payload Schema

We want our main application code to write to a single structured endpoint. We will define a clean, standard request format using pydantic. Create a file named schemas.py:

`python from pydantic import BaseModel from typing import List, Dict, Optional

class MessageItem(BaseModel): role: str # 'user' or 'assistant' or 'system' content: str

class LLMRequest(BaseModel): messages: List[MessageItem] temperature: Optional[float] = 0.7 max_tokens: Optional[int] = 1024 `

Step 3: Translating Requests for Claude and OpenAI

Because Anthropic handles system messages differently than OpenAI (passing them as a top-level parameter rather than an element in the messages array), we need translator utilities.

Create a file named translators.py:

`python from typing import Dict, Any, List from schemas import LLMRequest

def to_anthropic_payload(req: LLMRequest) -> Dict[str, Any]: # Pull out system messages system_message = "" anthropic_messages = [] for m in req.messages: if m.role == "system": system_message = m.content else: # Anthropic expects 'assistant' instead of 'system' role = "assistant" if m.role == "assistant" else "user" anthropic_messages.append({"role": role, "content": m.content}) payload = { "model": "claude-3-5-sonnet-20241022", "messages": anthropic_messages, "max_tokens": req.max_tokens, "temperature": req.temperature, } if system_message: payload["system"] = system_message return payload

def to_openai_payload(req: LLMRequest) -> Dict[str, Any]: openai_messages = [] for m in req.messages: openai_messages.append({"role": m.role, "content": m.content}) return { "model": "gpt-4o", "messages": openai_messages, "max_tokens": req.max_tokens, "temperature": req.temperature, } `

Step 4: Writing the Fallback Gateway

Now, let’s build the API Gateway in main.py. This will attempt to call Claude first. If a connection timeout, rate-limiting status (429), or server error (5xx) occurs, it catches the exception and routes the request to OpenAI instead.

We will load our API keys from the local environment. If you hit persistent authentication bugs during configuration, remember to consult the Anthropic support guide or the OpenAI support hub to verify your billing status and active key scopes.

`python import os import logging import httpx from fastapi import FastAPI, HTTPException, status from schemas import LLMRequest from translators import to_anthropic_payload, to_openai_payload

logging.basicConfig(level=logging.INFO) logger = logging.getLogger("fallback-proxy")

ANTHROPIC_API_KEY = os.getenv("ANTHROPIC_API_KEY") OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")

if not ANTHROPIC_API_KEY or not OPENAI_API_KEY: raise ValueError("Please ensure ANTHROPIC_API_KEY and OPENAI_API_KEY are configured in environment variables.")

app = FastAPI(title="LLM Fallback Proxy")

Configure HTTPX clients with sensible timeouts (e.g., 15s to react quickly to hangs) limits = httpx.Limits(max_keepalive_connections=5, max_connections=10) timeout = httpx.Timeout(15.0, connect=5.0)

@app.post("/v1/chat/completions") async def route_llm_request(payload: LLMRequest): async with httpx.AsyncClient(limits=limits, timeout=timeout) as client: # --- TRY CLAUDE FIRST --- try: logger.info("Attempting to route to Anthropic (Claude 3.5 Sonnet)... ") anthropic_data = to_anthropic_payload(payload) response = await client.post( "https://api.anthropic.com/v1/messages", headers={ "x-api-key": ANTHROPIC_API_KEY, "anthropic-version": "2023-06-01", "content-type": "application/json" }, json=anthropic_data ) # If Anthropic throws a rate limit or server error, trigger fallback if response.status_code in [429, 500, 502, 503, 504]: logger.warning(f"Anthropic returned status {response.status_code}. Initiating fallback routing...") raise httpx.HTTPStatusError("Anthropic failure", request=response.request, response=response) response.raise_for_status() data = response.json() # Normalise the response back to a clean standard string return { "provider": "anthropic", "text": data["content"][0]["text"], "usage": { "input_tokens": data["usage"]["input_tokens"], "output_tokens": data["usage"]["output_tokens"] } } except (httpx.HTTPError, httpx.TimeoutException) as exc: logger.error(f"Primary LLM call failed: {str(exc)}. Routing to OpenAI fallback...") # --- FALLBACK TO OPENAI --- try: openai_data = to_openai_payload(payload) response = await client.post( "https://api.openai.com/v1/chat/completions", headers={ "Authorization": f"Bearer {OPENAI_API_KEY}", "Content-Type": "application/json" }, json=openai_data ) response.raise_for_status() data = response.json() return { "provider": "openai", "text": data["choices"][0]["message"]["content"], "usage": { "input_tokens": data["usage"]["prompt_tokens"], "output_tokens": data["usage"]["completion_tokens"] } } except Exception as backup_exc: logger.critical(f"Both LLM providers failed. Backup error: {str(backup_exc)}") raise HTTPException( status_code=status.HTTP_502_BAD_GATEWAY, detail="All configured upstream AI backends are currently unreachable." ) `

Step 5: Testing Your Resilient Proxy

Set your environment variables and boot the proxy locally:

`bash export ANTHROPIC_API_KEY="your-anthropic-key" export OPENAI_API_KEY="your-openai-key" uvicorn main:app --reload --port 8080 `

Now, test the proxy using curl or any HTTP client:

`bash curl -X POST http://127.0.0.1:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "messages": [ {"role": "system", "content": "You are a short, cheeky assistant."}, {"role": "user", "content": "What is the capital of the UK?"} ], "temperature": 0.2 }' `

You will receive a clean response that indicates exactly which provider serviced your request. To test the fallback logic without waiting for an Anthropic outage, simply modify the main.py code to force a simulated httpx.TimeoutException during the Anthropic request block and watch as it silently and immediately fetches the response from OpenAI with negligible delay.

pythonfastapiclaude-3-5openai-api

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.