Comparisons
Claude 3.5 Sonnet vs GPT-4o for Writing Technical API Documentation: Which Engine Actually Reads Source Code Without Hallucinating?
We put Claude 3.5 Sonnet and GPT-4o head-to-head on documenting a messy, real-world TypeScript API. Discover which model actually reads your source code and which one resorts to lazy hallucinations.
Updated 10/10/2026
The Documentation Dilemma
We have all been there. You have a sprawling, semi-documented Express or Fastify API, three days until release, and a frontend team screaming for accurate endpoint specs. Writing OpenAPI schemas by hand is a form of digital torture, so you turn to an LLM.
But here is the catch: API documentation demands absolute precision. If an LLM hallucinates a single query parameter, misses a nested middleware validation rule, or invents a HTTP status code, your integration breaks.
In this comparison, we pit Anthropic's Claude 3.5 Sonnet against OpenAI's flagship GPT-4o in a raw, fluff-free battle. We fed both engines a messy, real-world TypeScript codebase with undocumented nested routes, custom validation middleware, and implicit error handling. We wanted to see which LLM actually parses the code logic, and which one gets lazy and starts guessing.
The Test: A Messy Express & TypeScript Route
To make this a fair fight, we did not use a clean, textbook code snippet. We used a real-world route file containing:
1. A custom authorization middleware that mutates the request object.
2. A Zod schema validation step that transforms input data.
3. A database call with implicit try/catch blocks throwing custom database errors.
4. An undocumented 500 fallback handler.
We instructed both models to generate a production-ready OpenAPI 3.0 specification in YAML, alongside a markdown developer guide. No hand-holding, no pre-digested summaries—just raw code analysis.
Round 1: Tracking Data Flow and Parameter Extraction
To write accurate API documentation, an LLM must track how a variable moves from the HTTP request body through your validation schemas and into the database query.
GPT-4o's Performance:
GPT-4o is fast—unbelievably so. It spat out a beautifully formatted YAML file in seconds. However, when we looked closely at the query parameters, we noticed a classic GPT shortcut. It correctly identified the parameters validated by our Zod schema, but completely missed a query parameter extracted directly from req.query further down the controller logic. Instead of reading the controller logic step-by-step, GPT-4o assumed the validation schema contained the entire truth. It also missed the custom headers injected by our authorization middleware.
Claude 3.5 Sonnet's Performance:
Claude took a few seconds longer, but the output was vastly superior. It did not just read the schema; it traced the execution path. It correctly noted that while the Zod schema validated the request body, a separate X-Tenant-ID header was required by the auth middleware to scope the database query. Claude even called out that this header was mandatory, despite it being defined in an imported middleware file we had context-packed into the prompt.
Round 2: Diagnosing Errors and Edge Cases
Nothing is worse than API documentation that pretends errors do not exist. We purposefully left an unhandled custom database error (RecordConflictError) in the code to see if the models would identify it and document the corresponding 409 Conflict response.
- GPT-4o defaulted to standard boilerplate. It generated a generic
400 Bad Requestand a500 Internal Server Errorblock. It completely ignored the custom error class because it was defined in an imported file. It assumed standard behaviour rather than verifying it. - Claude 3.5 Sonnet actually caught the custom error throw. It documented a
409 Conflictresponse, extracted the exact error payload structure from our error-handling class, and wrote a clean, helpful description explaining why this error would trigger (e.g., trying to register an email address that already exists in the database).
If you want to keep your developer pipeline ticking over without a hitch, Claude's attention to these edge cases is a game-changer.
Round 3: Code Block Quality and Formatting
Writing the YAML is only half the battle; developers need clean, copy-pasteable curl commands and response payloads in their markdown guides.
GPT-4o excels at producing visually striking, highly structured markdown. Its formatting is clean, using bold tables and tidy code blocks. However, the actual mock JSON payloads it generated for the responses were slightly off—it included camelCase keys for some fields that our TypeScript interfaces clearly defined as snake_case.
Claude 3.5 Sonnet’s markdown was slightly less flashy, but the data was 100% accurate. The mock payloads perfectly matched our Zod transformation rules (including date-to-string serialisation). Claude’s curl commands also included the necessary auth headers we discovered it had mapped in Round 1.
Pricing, Limits, and Caching: The Cold Hard Math
If you are running these documentation jobs programmatically across a large codebase containing hundreds of files, API costs and rate limits will hit you quickly.
| Metric | Claude 3.5 Sonnet | OpenAI GPT-4o | | :--- | :--- | :--- | | Input Cost (per 1M tokens) | $3.00 | $2.50 | | Output Cost (per 1M tokens) | $15.00 | $10.00 | | Prompt Caching | Yes (Up to 90% discount on read tokens) | Yes (Automatic, up to 50% discount) | | Context Window | 200,000 tokens | 128,000 tokens |
While GPT-4o is cheaper on raw list pricing, Claude’s manual Prompt Caching is a massive financial lifesaver when processing codebases. Because code documentation requires sending the same base files (middleware, database schemas, utility helpers) with different route files over and over, caching those system files can slash your API bill by up to 90%.
The Verdict
For simple, CRUD-based APIs where your routing is straightforward, GPT-4o is a fast, cost-effective tool that will get the job done quickly.
However, for complex, real-world codebases with nested logic, custom middleware, and strict TypeScript types, Claude 3.5 Sonnet is the undisputed winner. It behaves like a senior engineer doing a code review, whereas GPT-4o behaves like an eager intern skimming the surface. Claude actually reads your code, tracks variable lifecycles, and flags custom errors without making things up.
If you run into issues when parsing large-scale repositories, check out our Claude troubleshooting guides for tips on handling large token payloads and optimising your context packing scripts.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.