Tutorials & Guides
How to Parse Messy PDF Receipts into Structured JSON Using OpenAI Structured Outputs and Pydantic
Extracting line items from unstructured invoice PDFs with regex is a nightmare. Learn how to get guaranteed, validated JSON every time using OpenAI's Structured Outputs and Pydantic.
Updated 10/5/2026
Why Receipt Parsing Always Breaks
If you have ever tried to write a parser for digital receipts, invoices, or utility bills, you have probably wanted to throw your laptop out of a window.
No two vendors use the same layout. One merchant puts the tax in the top-right corner; another buries it in a tiny subtotal line at the bottom. Writing regular expressions to handle this variance is a fool's errand.
Historically, using Large Language Models (LLMs) to parse these files solved the layout problem but introduced a new one: schema drift. You'd ask for JSON, and the LLM would occasionally return markdown, missing fields, or invalid syntax that crashed your downstream database insert.
That problem is officially solved. By leveraging /platforms/openai feature alongside Python's pydantic library, we can guarantee that the model's output strictly adheres to our database schema. If the model fails to match our structure, the API call fails at the source. No more manual validation. No more parsing errors.
In this guide, we will build a robust script that processes receipt images or PDFs and outputs a perfectly typed JSON object.
Step 1: Set Up Your Project
For this tutorial, we will use openai to handle the LLM interaction and pydantic to define our strict data schema. We will also use pypdf to extract text from digital PDFs. If you are dealing with flat scanned images instead, you can send them directly to gpt-4o-mini as an image input.
First, install the required packages:
`bash
pip install openai pydantic pypdf python-dotenv
`
Create a .env file containing your API token:
`env
OPENAI_API_KEY=your_openai_api_key_here
`
(If you experience network issues or API authentication errors, consult the official guide at [https://www.openai-support.com](https://www.openai-support.com) to verify your account tier status.)
Step 2: Define Your Data Schema with Pydantic
We need to tell the OpenAI API exactly what format we expect. To do this, we write standard Pydantic models. We want to extract the merchant's name, the date, individual line items, tax, and the total amount.
`python
from typing import List, Optional
from pydantic import BaseModel, Field
class ReceiptItem(BaseModel): description: str = Field(description="The name or description of the individual item purchased.") quantity: int = Field(description="The quantity purchased. Default to 1 if not specified.") price_per_unit: float = Field(description="The cost of a single unit of this item.") total_price: float = Field(description="The total price for this line item (quantity * price_per_unit).")
class StructuredReceipt(BaseModel):
merchant_name: str = Field(description="The name of the store, restaurant, or business.")
date: Optional[str] = Field(description="The date of purchase in YYYY-MM-DD format if available.")
currency: str = Field(description="The three-letter ISO currency code (e.g. USD, GBP, EUR).")
line_items: List[ReceiptItem] = Field(description="List of individual items on the receipt.")
subtotal: float = Field(description="The calculated subtotal before tax and tips.")
tax_amount: float = Field(description="The total tax charged.")
tip_amount: Optional[float] = Field(description="The tip or gratuity amount, if applicable.")
total_amount: float = Field(description="The absolute final amount paid, including taxes and tips.")
`
Step 3: Extract Text and Call the Parsing Engine
Next, we will write a function that extracts raw text from our target PDF and feeds it to gpt-4o-mini. We will use the beta.chat.completions.parse endpoint to enforce our Pydantic schema.
`python
import os
from openai import OpenAI
from pypdf import PdfReader
from dotenv import load_dotenv
load_dotenv() client = OpenAI()
def extract_text_from_pdf(pdf_path: str) -> str: """Extract raw text from a target PDF receipt.""" reader = PdfReader(pdf_path) text = "" for page in reader.pages: text += page.extract_text() or "" return text
def parse_receipt(receipt_text: str) -> StructuredReceipt:
"""Sends unstructured receipt text to OpenAI and returns a validated Pydantic object."""
response = client.beta.chat.completions.parse(
model="gpt-4o-mini",
messages=[
{
"role": "system",
"content": "You are a precise data extraction agent. Extract invoice data from the provided text into the structured format required. If you cannot find a value, leave it null or use logical defaults."
},
{
"role": "user",
"content": receipt_text
}
],
response_format=StructuredReceipt,
)
# The parsed object is guaranteed to match the StructuredReceipt schema
return response.choices[0].message.parsed
`
Step 4: Putting It All Together
Now, let's create a quick execution script to test our pipeline. Imagine you have a sample digital invoice PDF called receipt.pdf in your directory:
`python
if __name__ == "__main__":
pdf_file_path = "receipt.pdf"
if not os.path.exists(pdf_file_path):
print(f"Please place a sample receipt.pdf in this directory.")
else:
print("Reading PDF text...")
raw_text = extract_text_from_pdf(pdf_file_path)
print("Sending unstructured text to OpenAI for parsing...")
parsed_data = parse_receipt(raw_text)
# Output the parsed data as formatted JSON
print("\n--- Parse Successful! ---\n")
print(parsed_data.model_dump_json(indent=4))
`
Why This Method is Bulletproof
Before Structured Outputs, LLMs were guided by raw prompt formatting tricks. If you want to see how developers used to hack this, check out our collection of legacy prompt templates in our /prompts library.
However, prompting alone was never 100% reliable. Under the hood, OpenAI's new parsing mechanism actually constrains the model's vocabulary during inference. The neural network is physically blocked from generating tokens that violate the JSON schema you defined in your Pydantic model.
This means:
Zero Schema Deviations*: You will never get a syntax error.
Type Coercion*: Fields defined as float or int will always be numeric—never returned as strings (like "$14.99").
Less Validation Boilerplate*: You can pipe this data straight to your database with absolute peace of mind.
By moving this complexity from custom parsing middleware directly to the model's output gate, you save hours of debugging and eliminate the brittle regexes of the past. Your accounting automation is now production-ready.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.