Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build an Automated LLM Fallback Gate in Python with Claude 3.5 Sonnet and Gemini 1.5 Flash

When your primary model hits a rate limit or service disruption, your app shouldn't fail. Build a resilient fallback gate using Claude and Gemini.

Updated 10/5/2026

No matter how much money you throw at high-tier APIs, the reality of building in production is simple: APIs fail. You will encounter 429 Too Many Requests when your user base spikes, or simple 503 Service Unavailable errors during platform-wide brownouts.

If your application relies on a single model like Claude 3.5 Sonnet, a service disruption means your product is dead in the water.

To build software that actually survives production workloads, you need an automated fallback strategy. In this tutorial, we will write a clean Python wrapper class that routes requests to Claude 3.5 Sonnet by default, catches specific API or rate-limit errors, and instantly routes the call to Gemini 1.5 Flash with minimal latency penalty.

Why Gemini 1.5 Flash is the Ideal Fallback

When building a fallback gate, your alternative model needs to meet three strict requirements: 1. Near-Zero Cold Start: It must respond fast enough that the user barely notices the transition. 2. Low Cost: You shouldn't be penalised financially because your primary provider went down. 3. Broad Context Support: It needs to accept similar structures, including system prompts and multi-modal elements.

Gemini 1.5 Flash checks every single box. It is incredibly cheap, fast, and has a massive context window. If you want to dive deeper into how these models compare for major tasks, you can review our technical breakdown on /platforms/claude and /platforms/gemini capabilities.

The Architecture of a Fallback Gate

Our system will implement a simple decorator pattern. We will define an abstract LLM interface, implement it for both Anthropic and Google, and wrap them in a routing manager that automatically catches rate-limit and status exceptions.

Let’s start by installing our dependencies:

`bash pip install anthropic google-generativeai pydantic `

Ensure you have your environment variables set:

`bash export ANTHROPIC_API_KEY="your-anthropic-key" export GEMINI_API_KEY="your-gemini-key" `

Step 1: Defining a Unified Message Schema

Because Anthropic and Google use slightly different schemas for their message history, we need to create a unified schema using Pydantic. This ensures that regardless of which LLM gets hit, the data interface looks exactly the same to your application.

Create a file named gateway.py:

`python from pydantic import BaseModel from typing import List, Literal, Optional

class ChatMessage(BaseModel): role: Literal["user", "assistant", "system"] content: str

class LLMResponse(BaseModel): content: str provider_used: Literal["anthropic", "google"] latency_ms: float `

Step 2: Writing the Client Wrappers

Next, we need helper functions to transform our unified ChatMessage payload into client-specific API calls.

Add the client handlers to your gateway.py file:

`python import time import os import anthropic import google.generativeai as genai

class LLMGateway: def __init__(self): self.anthropic_client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY")) genai.configure(api_key=os.environ.get("GEMINI_API_KEY")) self.gemini_model = genai.GenerativeModel('gemini-1.5-flash')

def _call_claude(self, messages: List[ChatMessage], system_prompt: Optional[str] = None) -> str: # Format messages for Anthropic anthropic_msgs = [] for m in messages: if m.role == "system": # System prompt is passed as a top-level parameter in Claude 3.5 system_prompt = m.content continue anthropic_msgs.append({"role": m.role, "content": m.content})

response = self.anthropic_client.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=1024, system=system_prompt or "", messages=anthropic_msgs ) return response.content[0].text

def _call_gemini(self, messages: List[ChatMessage], system_prompt: Optional[str] = None) -> str: # Construct contents for Gemini contents = [] for m in messages: # Gemini handles role labels slightly differently role = "user" if m.role == "user" else "model" contents.append({"role": role, "parts": [m.content]})

config = genai.types.GenerationConfig(max_output_tokens=1024) # Pass system prompt within the GenerativeModel constructor if required model = self.gemini_model if system_prompt: model = genai.GenerativeModel( 'gemini-1.5-flash', system_instruction=system_prompt ) response = model.generate_content(contents, generation_config=config) return response.text `

Step 3: Implementing the Fallback Logic

Now, let's write our gatekeeper method. It will attempt the Anthropic call. If Anthropic raises a RateLimitError or an InternalServerError, it will catch the error, log a warning, and route the message to Gemini instead.

We will also track latency to see if the fallback degrades our user experience.

`python def generate(self, messages: List[ChatMessage], system_prompt: Optional[str] = None) -> LLMResponse: start_time = time.time() try: # Try our primary model first print("[GATEWAY] Attempting call with Claude 3.5 Sonnet...") output = self._call_claude(messages, system_prompt) latency = (time.time() - start_time) * 1000 return LLMResponse(content=output, provider_used="anthropic", latency_ms=latency) except (anthropic.RateLimitError, anthropic.InternalServerError, anthropic.APIConnectionError) as e: print(f"[WARNING] Primary model failed with: {str(e)}. Swapping to fallback...") fallback_start = time.time() try: # Gracefully route to Gemini output = self._call_gemini(messages, system_prompt) latency = (time.time() - fallback_start) * 1000 return LLMResponse(content=output, provider_used="google", latency_ms=latency) except Exception as gemini_err: # If both are down, it is time to alert DevOps raise RuntimeError(f"Critical failure: Both LLM pipelines failed. Gemini error: {str(gemini_err)}") `

Let's Test It

Let's run a test file to verify our fallback works under simulated duress. Create a file called test_gate.py:

`python from gateway import LLMGateway, ChatMessage

Initialize the resilient gateway gateway = LLMGateway()

messages = [ ChatMessage(role="user", content="Identify the core architectural difference between REST and gRPC.") ]

Standard run try: response = gateway.generate(messages, system_prompt="Be concise and technical.") print(f"\nSuccess! Response provider: {response.provider_used}") print(f"Latency: {response.latency_ms:.2f}ms") print(f"Content:\n{response.content}\n") except Exception as e: print(f"Failed: {e}") ```

If you want to test the fallback behaviour manually without actual downtime, you can simply comment out the Claude invocation inside _call_claude and raise a mock anthropic.RateLimitError to see your gateway switch to Gemini in milliseconds.

Optimising Your Prompt Strategies

While Gemini 1.5 Flash is highly versatile, switching models on the fly means your prompts must be structured to perform well on both models. System parameters or formatting tags that Claude understands perfectly can occasionally confuse Gemini.

To ensure your prompts remain model-agnostic, run them through our /prompts to construct cross-compatible schema models. For ongoing issues with Gemini API credentials or routing exceptions, consult our troubleshooting workflows on /platforms/gemini/articles to stay ahead of production bugs.

pythonclaudegeminireliabilitytutorials

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.