Tutorials & Guides
How to Build a Local CLI Tool to Auto-Document Legacy ExpressJS APIs Using Claude 3.5 Sonnet and Swagger
Inheriting legacy Express.js backends with zero documentation is a rite of passage for software engineers. This step-by-step tutorial shows you how to programmatically parse messy route directories and use Claude 3.5 Sonnet to draft valid OpenAPI/Swagger schemas.
Updated 10/5/2026
Nothing says "welcome to your new job" quite like being handed a 15,000-line Express.js backend codebase with absolutely zero documentation. If you want to know how the /api/v1/billing/subscribe endpoint behaves, your only choice is to wade through nested middleware functions, loosely typed request bodies, and database queries that lack clear structure.
Writing OpenAPI (Swagger) specifications for existing legacy systems is tedious, error-prone, and a terrible use of engineering hours. This is where AI models excel. Claude 3.5 Sonnet stands out as the absolute standard for parsing abstract code patterns, mapping controller logic, and outputting pristine OpenAPI/Swagger specs.
In this tutorial, we will build a local CLI tool using Python to recursively read Express.js route files, parse controller code, and use /platforms/claude to construct a comprehensive, single-file Swagger configuration.
The Architecture of the Spec-Generator
To build a reliable documentation generator, we must avoid the temptation to dump our entire repository into a massive Claude chat box. This is unstructured, expensive, and risks hitting prompt length limits with minimal formatting control.
Instead, our CLI will follow a structured process:
1. Index and Scan: Locate route configurations (e.g., files matching *.routes.js or containing patterns like router.get or app.post).
2. Isolate and Chunk: Read the routing file and any immediately imported/referenced controllers locally.
3. Map and Generate: Pass targeted controller snippets to Claude to extract HTTP methods, query parameters, expected JSON bodies, and headers.
4. Assemble: Combine the individual YAML blocks into a final, valid openapi.yaml configuration.
If you run into issues during token construction or hit prompt size constraints, you can read more on the Claude support site to resolve system error messages.
Setting Up Your Environment
Ensure you have Python 3.10+ and set up your project dependencies:
`bash
mkdir express-swagger-gen && cd express-swagger-gen
python3 -m venv venv
source venv/bin/activate
pip install anthropic click pyyaml
`
Make sure your Anthropic API key is correctly exposed to your system path:
`bash
export ANTHROPIC_API_KEY="your-api-key-here"
`
Step 1: Writing the Target Finder
First, we need a reliable utility function to recursively walk our node project directory, find route files, and resolve any local relative imports (like controller functions) so we can feed Claude the real business logic. Save this file as file_parser.py:
`python
import os
import re
def scan_route_files(directory): """Scan directory for common Express route patterns.""" route_files = [] for root, _, files in os.walk(directory): if "node_modules" in root or ".git" in root: continue for file in files: if file.endswith(".js") or file.endswith(".ts"): filepath = os.path.join(root, file) with open(filepath, 'r', encoding='utf-8', errors='ignore') as f: content = f.read() # Simple heuristic to identify Express routers if "express.Router" in content or "router." in content or "app.get(" in content: route_files.append(filepath) return route_files
def resolve_associated_controllers(route_filepath): """Analyze a route file, discover local imports, and pack their code.""" base_dir = os.path.dirname(route_filepath) with open(route_filepath, 'r', encoding='utf-8') as f: route_code = f.read()
Simple regex to catch require statements pointing to local controllers imports = re.findall(r"require\(['\"](\.\.?/[^'\"]+)['\"]\)", route_code) combined_context = f"=== ROUTE FILE: {os.path.basename(route_filepath)} ===\n{route_code}\n" for imp in imports: # Resolve paths dynamically candidate_path = os.path.normpath(os.path.join(base_dir, imp)) if not candidate_path.endswith(('.js', '.ts')): candidate_path += '.js' if os.path.exists(candidate_path): with open(candidate_path, 'r', encoding='utf-8') as f: combined_context += f"\n=== IMPORTED FILE: {os.path.basename(candidate_path)} ===\n{f.read()}\n" return combined_context ```
Step 2: Formulating the Extraction System Prompt
We need Claude to parse code strings and extract endpoints in a clean format. Getting the AI to produce valid yaml fragments saves us from parsing broken code blocks later. We will instruct Claude to return exclusively a sub-set of Swagger's paths section. To build highly refined structural outputs like this yourself, take a look at our dedicated guide on drafting structural parameters inside our /prompts.
Create a file named claude_agent.py to hold your API integration:
`python
import os
from anthropic import Anthropic
class ClaudeDocAgent: def __init__(self): self.client = Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
def extract_api_paths(self, codebase_context):
system_prompt = (
"You are an elite software architect specialized in RESTful APIs and OpenAPI 3.0 specs.\n"
"Your task is to analyze the provided ExpressJS router code and any of its associated "
"controller code. Extract and document every API endpoint defined.\n\n"
"CRITICAL requirements:\n"
"1. Generate only the OpenAPI paths schema fragment in valid YAML format.\n"
"2. Detail any request parameters, headers, URL query parameters, and JSON request payloads.\n"
"3. Deduce correct response status codes (e.g., 200, 201, 400, 404, 500) and response structural models based on code paths.\n"
"4. Start your output directly with the paths: element. Do not include markdown code wrappers (e.g., `yaml) "
"and write absolutely no introductory or explanatory text. Go straight to YAML."
)
try:
# Using Claude 3.5 Sonnet for rich, context-aware architectural inference
message = self.client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=4000,
temperature=0.1, # Low temp for accurate parsing mapping
system=system_prompt,
messages=[
{"role": "user", "content": f"Code Context:\n{codebase_context}"}
]
)
return message.content[0].text.strip()
except Exception as e:
print(f"An error occurred while calling Claude: {e}")
return ""
`
Step 3: Compiling Individual Paths Into a Complete Spec
Once Claude returns individual path chunks, we need to stitch them back together under a unified head. Rather than relying on fragile manual string concatenation, we will parse the YAML outputs using the standard pyyaml library. This ensures we produce syntactically valid documents.
Create main.py:
`python
import click
import yaml
import os
from file_parser import scan_route_files, resolve_associated_controllers
from claude_agent import ClaudeDocAgent
BASE_SWAGGER_TEMPLATE = """ openapi: 3.0.0 info: title: Auto-Generated ExpressJS API Documentation description: Generated dynamically using local context analysis and Claude 3.5 Sonnet. version: 1.0.0 paths: {} """
@click.command() @click.option('--dir', '-d', required=True, help="Path to Express.js project folder.") @click.option('--out', '-o', default="openapi_generated.yaml", help="Output YAML document filepath.") def main(dir, out): """Analyze legacy directories and construct beautiful Swagger specifications effortlessly.""" click.echo(f"Scanning directory: {dir}...") route_files = scan_route_files(dir) if not route_files: click.echo("No Express route files detected. Please verify your directories.") return click.echo(f"Found {len(route_files)} target route files. Parsing and processing...") agent = ClaudeDocAgent() master_paths = {}
for index, filepath in enumerate(route_files): click.echo(f"[{index + 1}/{len(route_files)}] Extracting endpoints from: {os.path.basename(filepath)}") context = resolve_associated_controllers(filepath) # Retrieve parsed schema chunk yaml_chunk = agent.extract_api_paths(context) if not yaml_chunk: continue try: # Clean clean up leading paths marker if generated parsed_data = yaml.safe_load(yaml_chunk) if parsed_data and "paths" in parsed_data: parsed_data = parsed_data["paths"] if isinstance(parsed_data, dict): master_paths.update(parsed_data) else: click.echo(f"Warning: Output from {filepath} did not parse into expected dict schema structure.") except Exception as e: click.echo(f"YAML Parsing Error for route {os.path.basename(filepath)}: {e}")
Generate full complete schema document full_swagger = yaml.safe_load(BASE_SWAGGER_TEMPLATE) full_swagger["paths"] = master_paths
Write back clean configuration with open(out, 'w', encoding='utf-8') as f: yaml.dump(full_swagger, f, default_flow_style=False, sort_keys=False)
click.echo(f"\nš Complete specification compiled successfully into: {out}")
if __name__ == '__main__':
main()
`
Running the Code on Your Project
To try out your fresh code extraction system, point your new CLI tool at an legacy Express project. Here's how you might trigger it from your terminal:
`bash
python main.py --dir /Users/dev/projects/legacy-billing-service --out billing-docs.yaml
`
The tool scans your endpoint directories, extracts the underlying database logic and response codes, sends it to Claude, and compiles it into a cleanly structured YAML schema in your current directory.
You can visually preview the output file by pasting the schema straight into the official online Swagger Editor. This allows you to verify that everything has rendered flawlessly without having to manually document a single API endpoint yourself.
To dive deeper into standard structures and explore key API pattern terms, visit our general /glossary for deep definitions of API engineering jargon.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.