Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build a Resilient JSON Parser for LLM Outputs That Fall Back Safely When Pydantic Fails

Structured outputs aren't bulletproof. When your LLM truncates a response or hallucinates a trailing comma, your app shouldn't crash. Here is how to build a robust, self-healing parser in Python that handles invalid JSON gracefully.

Updated 10/5/2026

The Myth of the Perfect Structured Output

We were promised a world of pristine, guaranteed JSON. When OpenAI and Anthropic introduced structured outputs and tool use, developers everywhere breathed a sigh of relief. We thought we could finally bin our fragile regular expressions and custom validation loops.

But if you have run any high-throughput AI application in production, you already know the painful truth: structured outputs still fail.

An LLM might run out of output tokens mid-sentence, leaving you with an unclosed bracket. It might get confused by a nested schema and inject an unexpected markdown code block inside a JSON string. Or, in moments of sheer panic, it might simply hallucinate a trailing comma that throws off your standard parser.

If your application relies on a strict pydantic.BaseModel to parse these payloads directly, a single missing bracket will trigger a crash. To keep your system robust, you need a defensive parsing strategy. Here is how to build a multi-layered, self-healing JSON parser in Python that catches errors, repairs malformed strings, and falls back gracefully before your user ever notices a hiccup.

Step 1: The Ideal Path (Strict Pydantic)

We always want to try the cleanest method first. If you are using /platforms/openai or tool calling with Claude, the model should ideally return a valid JSON string that maps perfectly to your Pydantic model.

Let's define a target schema for our tutorial. Imagine we are building an agent that extracts key action items and sentiment from customer emails.

`python from pydantic import BaseModel, Field from typing import List, Optional

class ActionItem(BaseModel): task: str = Field(description="The specific task to be completed.") assignee: Optional[str] = Field(None, description="The person or role assigned to the task.") priority: str = Field(description="High, Medium, or Low.")

class EmailAnalysis(BaseModel): sentiment: str = Field(description="Overall tone of the email.") urgency_score: int = Field(description="Scale from 1 to 10.") actions: List[ActionItem] `

In a perfect world, parsing this is a one-liner. But when a network hiccup or token limit cuts off the response, standard parsing fails. It’s the silent engine that keeps your production system ticking along without a hitch, so let's prepare for when things go wrong.

Step 2: Extracting JSON from Markdown Wrappers

Even when instructed to return raw JSON, LLMs love wrapping their output in triple backticks: `json ... ` . Before passing the raw string to any parser, we need to strip these wrappers away.

Here is a simple, robust utility function to clean up the raw text:

`python import re

def clean_json_string(raw_text: str) -> str: # Strip leading/trailing whitespace text = raw_text.strip() # Use regex to find content inside markdown code blocks markdown_pattern = r"^`(?:json)?\s([\s\S]?)\s*`$" match = re.match(markdown_pattern, text) if match: text = match.group(1).strip() return text `

Step 3: Implement Partial Parsing and Self-Healing

If the JSON is still malformed—perhaps due to a missing closing bracket because the model hit its context limit—we need to attempt a soft repair.

While you can write custom regex to balance braces, a brilliant open-source library called json-repair does this work beautifully. It uses a lightweight state machine to insert missing quotes, brackets, and commas without resorting to another slow LLM API call.

First, install it: `bash pip install json-repair `

Now, let's write a robust loader that leverages json_repair as a primary fallback before we hit the Pydantic validation stage.

`python import json from json_repair import repair_json

def robust_json_loads(raw_text: str) -> dict: cleaned = clean_json_string(raw_text) try: # Try standard loading first return json.loads(cleaned) except json.JSONDecodeError: # If standard loading fails, attempt a silent repair try: repaired = repair_json(cleaned) return json.loads(repaired) except Exception as e: raise ValueError(f"Failed to parse and repair JSON: {str(e)}") `

Step 4: The Graceful Pydantic Fallback

Once we have a valid Python dictionary (even a partially repaired one), we must map it to our Pydantic model.

However, if the LLM completely missed a mandatory field, or generated an invalid type (e.g., a string instead of an integer), Pydantic will throw a ValidationError.

Instead of bubbling this error up to the client, we can write a fallback wrapper that instantiates the model with default or placeholder values for missing data, ensuring your application pipeline doesn't break. You can read more about debugging these validation issues in our dedicated guide on /platforms/claude/articles.

Here is the complete implementation of our resilient parser:

`python from pydantic import ValidationError

def force_to_model(schema: type[BaseModel], data: dict) -> BaseModel: try: # Try direct parsing return schema(**data) except ValidationError as e: # Log the error in production print(f"Validation failed: {e}. Attempting smart fallback construction.") fallback_data = {} # Iterate through the schema fields to build a partial model for field_name, field_info in schema.model_fields.items(): val = data.get(field_name) # If the value exists, we try to use it if val is not None: fallback_data[field_name] = val else: # If the field has a default value, use it if field_info.default is not None and field_info.default != ...: fallback_data[field_name] = field_info.default # If the field is an optional type, set to None elif type(None) in getattr(field_info.annotation, "__args__", []): fallback_data[field_name] = None # Otherwise, insert a safe placeholder based on the type else: if field_info.annotation == str: fallback_data[field_name] = "[Missing Data]" elif field_info.annotation == int: fallback_data[field_name] = 0 elif getattr(field_info.annotation, "__origin__", None) == list: fallback_data[field_name] = [] else: fallback_data[field_name] = None return schema(**fallback_data) `

Step 5: Tying It All Together

Let’s test our pipeline with a notoriously difficult output: a truncated payload that cut off right in the middle of a nested list.

`python raw_llm_output = """ `json { "sentiment": "Frustrated", "urgency_score": 9, "actions": [ { "task": "Refund client for damaged items", "assignee": "Finance team", "priority": "High" }, { "task": "Send apology gift card", "assignee": "Support """

Execution try: # 1. Parse and repair the raw string to a dictionary parsed_dict = robust_json_loads(raw_llm_output) print("Repaired Dict:", parsed_dict) # 2. Map safely to our Pydantic model final_model = force_to_model(EmailAnalysis, parsed_dict) print("Final Validated Model:", final_model.model_dump_json(indent=2)) except Exception as e: print(f"Pipeline failed entirely: {e}") ```

When you run this code, the json-repair engine automatically detects the unclosed string, the unclosed dictionary, and the unclosed array, outputting a syntactically correct dictionary. Our force_to_model function then steps in to ensure any fields corrupted by the abrupt termination are assigned safe, type-correct placeholders instead of crashing.

Build Defensively

Relying on LLMs to behave perfectly is a recipe for high error rates and midnight paging alerts. By placing a resilient parsing layer between your raw LLM outputs and your core application logic, you guarantee that downstream functions can always rely on predictable, schema-compliant data—no matter how chaotic the model's response behaves. For more architectural tips on structuring reliable agent systems, dive into our /glossary.

pythonpydanticstructured-outputsllm-opstutorials

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.