Tutorials & Guides
How to Build a Local CLI to Auto-Update Outdated Code Examples in Markdown Docs Using Claude 3.5 Sonnet and Python
API changes break markdown documentation all the time. Build a smart Python CLI tool that parses your actual codebase schemas and asks Claude to keep your markdown tutorials completely in sync.
Updated 10/5/2026
The Nightmare of Doc Drift
We have all lived through this scenario. You pull down a promising open-source library or an internal API utility. You copy the elegant getting-started code block straight from the repository’s README.md, paste it into your local terminal, and run it.
Only to be greeted by a screaming stack trace of syntax errors and deprecated argument exceptions.
Codebases evolve rapidly. Standard documentation, unfortunately, sits quietly in static markdown files, entirely disconnected from compiler checks and integration test suites. When you refactor a method signature from fetch_user(user_id) to fetch_user(id_or_email, include_metadata=True), you might change twenty instances across your codebase but completely overlook the single code block sitting inside your static tutorials.
To bridge this divide, we can build a lightweight local CLI tool that acts as an automated linchpin. By parsing your updated code types or API exports, parsing your existing Markdown documentation, and matching them up via Claude 3.5 Sonnet, we can automatically rewrite outdated markdown code snippets without altering the surrounding copy.
This workflow ticks all the boxes for maintaining documentation sanity: it runs entirely locally, targets your changed files, and ensures code examples remain truthful.
The Design Strategy
To make this CLI tool highly dependable and safe, we won’t let the AI rewrite our entire documentation file from scratch. Doing so is expensive, slow, and risk-prone (the LLM might hallucinate changes to our beautifully crafted text, remove formatting, or rewrite sentences).
Instead, we will use a highly targeted modular pipeline:
- Extraction: Locate all triple-backtick markdown blocks (
`python ...`) within our target.mdfiles. - Type/Schema Mining: Read our updated codebase API (for this example, we will check our current Python modules or schema definitions to construct an API reference dict).
- LLM Evaluation: Send each isolated code block to Claude along with our real API schema, asking simple binary evaluation: "Does this code block contain outdated interfaces based on the actual schema? If yes, rewrite the code block. If no, return 'NO_CHANGE'."
- In-place Substitution: Overwrite only the modified blocks back into the markdown layout, leaving the remaining text untouched.
Let's get this configured step by step.
Step 1: Parsing the Markdown File for Code Blocks
We need to parse a markdown document and extract each code block along with its position. This allows us to map the output back precisely where it belongs.
Create a python file named doc_sync.py and import the core libraries:
`python
import os
import re
import sys
import argparse
from anthropic import Anthropic
Regex to capture markdown code blocks along with their language tag CODE_BLOCK_RE = re.compile(r"(```(python|py)\n(.*?)\n```)", re.DOTALL)
def extract_code_blocks(markdown_content):
"""Extracts all python code blocks and their raw strings from a markdown file."""
blocks = []
for match in CODE_BLOCK_RE.finditer(markdown_content):
full_match = match.group(0)
code_body = match.group(3)
blocks.append({
"full_match": full_match,
"code_body": code_body,
"start": match.start(),
"end": match.end()
})
return blocks
`
Step 2: Extracting Codebase Ground-Truth (The Schema)
To help Claude accurately evaluate if our documentation matches reality, we need to provide a file or file list that outlines the actual state of our modules. We can do this dynamically by reading your source files.
For simplicity, let’s design our CLI to accept a "source source-of-truth" file (e.g., your API controllers, Pydantic schemas, or main module export file) which we will read and supply directly to Claude as reference.
`python
def read_source_truth(source_file_path):
"""Reads the real production code source to serve as our absolute API source-of-truth."""
if not os.path.exists(source_file_path):
print(f"[-] Error: Ground truth source file not found at {source_file_path}")
sys.exit(1)
with open(source_file_path, "r") as f:
return f.read()
`
Step 3: Prompting Claude to Correct the Outdated Code
Now, let’s construct our orchestration interface. We want Claude to check the code snippet against our source of truth. If it requires updating, we want it to output ONLY the updated inner code block, and nothing else. Check out our glossary for terms regarding system prompts and agentic validation rules.
`python
def sync_code_block(code_snippet, ground_truth):
client = Anthropic()
system_prompt = (
"You are an automated system checking technical documentation for accuracy. "
"You are provided with a code block from a tutorial and the source of truth file representing "
"the actual production codebase. "
"Your job is to determine if the tutorial code block uses deprecated, outdated, or incorrect interfaces "
"relative to the source of truth file. "
"If the code is perfectly correct and matches the current API signatures, return exactly 'NO_CHANGE'. "
"If the code uses outdated functions, parameters, or patterns, output ONLY the corrected code block. "
"Do not wrap your output in markdown code blocks. Do not say 'Here is the updated code'. "
"Output only raw, valid code or 'NO_CHANGE'. Keep your edits highly minimal: preserve logic, names, "
"and comments from the original code unless they are wrong."
)
user_prompt = f"""
=== SOURCE OF TRUTH (ACTUAL PRODUCTION CODE) ===
{ground_truth}
================================================
=== TUTORIAL CODE SNIPPET TO EVALUATE === {code_snippet} ========================================= """
try:
message = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=1500,
temperature=0.0,
system=system_prompt,
messages=[{"role": "user", "content": user_prompt}]
)
return message.content[0].text.strip()
except Exception as e:
# For API key problems or timeout issues, check: https://claude-support.com
print(f"[-] Anthropic API Error: {e}")
return "NO_CHANGE"
`
Step 4: Running the In-Place Replacement Engine
Now, let’s tie these functions together. We will read the target markdown file, find all Python code blocks, ask Claude to examine each one, and make replacements in reverse order. Reversing the replacements is crucial: it prevents our character index positions from shifting as we alter file lengths.
`python
def process_documentation(markdown_path, source_truth_path):
with open(markdown_path, "r") as f:
original_md = f.read()
ground_truth = read_source_truth(source_truth_path)
blocks = extract_code_blocks(original_md)
if not blocks:
print(f"[*] No Python code blocks identified in {markdown_path}")
return
print(f"[*] Found {len(blocks)} code blocks to check in {markdown_path}.")
updated_md = original_md
# Process code blocks from end of file to start to preserve character offsets
for block in reversed(blocks):
code_body = block["code_body"]
print(f"[*] Checking block starting at character offset {block['start']}...")
response = sync_code_block(code_body, ground_truth)
if response == "NO_CHANGE" or not response.strip():
print(" [+] Block is already up-to-date. Skipping.")
continue
print(" [!] Block outdated! Injecting Claude's corrected snippet.")
# Reconstruct markdown block wrapper
new_block = f"`python\n{response}\n`"
# Splice the update back into position
updated_md = updated_md[:block["start"]] + new_block + updated_md[block["end"]:]
if updated_md != original_md:
with open(markdown_path, "w") as f:
f.write(updated_md)
print(f"[+] Done! Successfully updated documentation at {markdown_path}")
else:
print("[+] Verification complete. Documentation matches the production schema perfectly.")
`
Step 5: Designing the CLI Interface
Finally, we will package this functionality behind a friendly Python standard CLI using argparse. This ensures developers can easily run it on-demand or as a step in an automated CI/CD build pipeline.
`python
def main():
parser = argparse.ArgumentParser(
description="Auto-align Markdown tutorials with raw codebase realities using Claude 3.5 Sonnet."
)
parser.add_argument(
"--doc",
required=True,
help="Path to the markdown file to inspect (e.g. docs/getting-started.md)"
)
parser.add_argument(
"--src",
required=True,
help="Path to the codebase source-of-truth file containing actual schemas/interfaces"
)
args = parser.parse_args()
process_documentation(args.doc, args.src)
if __name__ == "__main__":
main()
`
To run your sync program, simply run the tool from your terminal shell environment:
`bash
python doc_sync.py --doc docs/tutorials.md --src app/api/endpoints.py
`
By adding this CLI script to your continuous integration flow, you can block PR merges if markdown examples fail to align with code definitions. This saves your developer community from tedious manual verifications, ensuring every line of instructions in your guides runs cleanly on the first try.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.