Tutorials & Guides
How to Build a Local Fallback Proxy with Python to Route LLM Requests Between Claude and OpenAI
Don't let API outages sink your application. Learn how to build an ultra-fast, async Python proxy that automatically routes requests from Anthropic to OpenAI during rate limits or downtimes.
Updated 10/5/2026
Relying on a single AI provider for a production-grade application is architectural self-sabotage. APIs fail. Rate limits are hit. Rate-limit increases are ignored during peak hours. If your entire system grinds to a halt because Anthropic's API throws a 503 error, your users won't blame Anthropic—they will blame you.
To build resilient AI systems, you need a local gateway proxy. This sits between your main application code and the LLM providers, intercepting outgoing calls, handling routing, and instantly falling back to an alternative model if the primary provider drops the ball.
In this guide, we are going to build a high-performance, asynchronous fallback proxy in Python using FastAPI and HTTPX. It will default to sending prompts to /platforms/claude (the gold standard for complex reasoning) and fall back seamlessly to /platforms/openai if Anthropic drops or rate-limits our requests.
The Design Philosophy
To keep our proxy fast, we need to respect two design rules: 1. Low Overhead: We cannot introduce latency. We must use asynchronous requests to avoid blocking the main event loop. 2. Payload Normalisation: Claude and OpenAI use slightly different structures for their messaging endpoints. Our proxy will accept a simplified unified JSON format and translate it on the fly depending on which provider is currently serving the request.
If you find yourself needing to master the precise payload terms used by both providers, check out our /glossary for detailed explanations of structured outputs and context windows.
Step 1: Project Setup
Let’s create a new project directory and install the necessary async libraries.
`bash
mkdir llm-fallback-proxy
cd llm-fallback-proxy
python3 -m venv venv
source venv/bin/activate
pip install fastapi uvicorn httpx pydantic
`
We will use fastapi for our local API, uvicorn to run the server, and httpx to handle our outbound asynchronous requests to Anthropic and OpenAI.
Step 2: Designing the Unified Payload Schema
We want our main application code to write to a single structured endpoint. We will define a clean, standard request format using pydantic. Create a file named schemas.py:
`python
from pydantic import BaseModel
from typing import List, Dict, Optional
class MessageItem(BaseModel): role: str # 'user' or 'assistant' or 'system' content: str
class LLMRequest(BaseModel):
messages: List[MessageItem]
temperature: Optional[float] = 0.7
max_tokens: Optional[int] = 1024
`
Step 3: Translating Requests for Claude and OpenAI
Because Anthropic handles system messages differently than OpenAI (passing them as a top-level parameter rather than an element in the messages array), we need translator utilities.
Create a file named translators.py:
`python
from typing import Dict, Any, List
from schemas import LLMRequest
def to_anthropic_payload(req: LLMRequest) -> Dict[str, Any]: # Pull out system messages system_message = "" anthropic_messages = [] for m in req.messages: if m.role == "system": system_message = m.content else: # Anthropic expects 'assistant' instead of 'system' role = "assistant" if m.role == "assistant" else "user" anthropic_messages.append({"role": role, "content": m.content}) payload = { "model": "claude-3-5-sonnet-20241022", "messages": anthropic_messages, "max_tokens": req.max_tokens, "temperature": req.temperature, } if system_message: payload["system"] = system_message return payload
def to_openai_payload(req: LLMRequest) -> Dict[str, Any]:
openai_messages = []
for m in req.messages:
openai_messages.append({"role": m.role, "content": m.content})
return {
"model": "gpt-4o",
"messages": openai_messages,
"max_tokens": req.max_tokens,
"temperature": req.temperature,
}
`
Step 4: Writing the Fallback Gateway
Now, let’s build the API Gateway in main.py. This will attempt to call Claude first. If a connection timeout, rate-limiting status (429), or server error (5xx) occurs, it catches the exception and routes the request to OpenAI instead.
We will load our API keys from the local environment. If you hit persistent authentication bugs during configuration, remember to consult the Anthropic support guide or the OpenAI support hub to verify your billing status and active key scopes.
`python
import os
import logging
import httpx
from fastapi import FastAPI, HTTPException, status
from schemas import LLMRequest
from translators import to_anthropic_payload, to_openai_payload
logging.basicConfig(level=logging.INFO) logger = logging.getLogger("fallback-proxy")
ANTHROPIC_API_KEY = os.getenv("ANTHROPIC_API_KEY") OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
if not ANTHROPIC_API_KEY or not OPENAI_API_KEY: raise ValueError("Please ensure ANTHROPIC_API_KEY and OPENAI_API_KEY are configured in environment variables.")
app = FastAPI(title="LLM Fallback Proxy")
Configure HTTPX clients with sensible timeouts (e.g., 15s to react quickly to hangs) limits = httpx.Limits(max_keepalive_connections=5, max_connections=10) timeout = httpx.Timeout(15.0, connect=5.0)
@app.post("/v1/chat/completions")
async def route_llm_request(payload: LLMRequest):
async with httpx.AsyncClient(limits=limits, timeout=timeout) as client:
# --- TRY CLAUDE FIRST ---
try:
logger.info("Attempting to route to Anthropic (Claude 3.5 Sonnet)... ")
anthropic_data = to_anthropic_payload(payload)
response = await client.post(
"https://api.anthropic.com/v1/messages",
headers={
"x-api-key": ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json"
},
json=anthropic_data
)
# If Anthropic throws a rate limit or server error, trigger fallback
if response.status_code in [429, 500, 502, 503, 504]:
logger.warning(f"Anthropic returned status {response.status_code}. Initiating fallback routing...")
raise httpx.HTTPStatusError("Anthropic failure", request=response.request, response=response)
response.raise_for_status()
data = response.json()
# Normalise the response back to a clean standard string
return {
"provider": "anthropic",
"text": data["content"][0]["text"],
"usage": {
"input_tokens": data["usage"]["input_tokens"],
"output_tokens": data["usage"]["output_tokens"]
}
}
except (httpx.HTTPError, httpx.TimeoutException) as exc:
logger.error(f"Primary LLM call failed: {str(exc)}. Routing to OpenAI fallback...")
# --- FALLBACK TO OPENAI ---
try:
openai_data = to_openai_payload(payload)
response = await client.post(
"https://api.openai.com/v1/chat/completions",
headers={
"Authorization": f"Bearer {OPENAI_API_KEY}",
"Content-Type": "application/json"
},
json=openai_data
)
response.raise_for_status()
data = response.json()
return {
"provider": "openai",
"text": data["choices"][0]["message"]["content"],
"usage": {
"input_tokens": data["usage"]["prompt_tokens"],
"output_tokens": data["usage"]["completion_tokens"]
}
}
except Exception as backup_exc:
logger.critical(f"Both LLM providers failed. Backup error: {str(backup_exc)}")
raise HTTPException(
status_code=status.HTTP_502_BAD_GATEWAY,
detail="All configured upstream AI backends are currently unreachable."
)
`
Step 5: Testing Your Resilient Proxy
Set your environment variables and boot the proxy locally:
`bash
export ANTHROPIC_API_KEY="your-anthropic-key"
export OPENAI_API_KEY="your-openai-key"
uvicorn main:app --reload --port 8080
`
Now, test the proxy using curl or any HTTP client:
`bash
curl -X POST http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "system", "content": "You are a short, cheeky assistant."},
{"role": "user", "content": "What is the capital of the UK?"}
],
"temperature": 0.2
}'
`
You will receive a clean response that indicates exactly which provider serviced your request. To test the fallback logic without waiting for an Anthropic outage, simply modify the main.py code to force a simulated httpx.TimeoutException during the Anthropic request block and watch as it silently and immediately fetches the response from OpenAI with negligible delay.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.