Tutorials & Guides
How to Build an Automated Markdown Link Rot Repairer Using Python and Claude 3.5 Sonnet
Broken documentation links are a silent drain on developer productivity. Stop hunting for 404 targets manually—build a smart script that scans markdown files, identifies dead links, and uses Claude 3.5 Sonnet to predict and patch the correct destinations.
Updated 10/2/2026
Every engineering manager and technical writer knows the pain of maintaining a documentation hub. You reorganise a folder structure, an external SaaS platform deprecates an API version, or a blog post gets renamed. Suddenly, your carefully curated markdown docs are riddled with broken links.
Standard CI/CD link checkers are great at highlighting the problem, but they are incredibly tedious to fix. They spit out a list of fifty 404 errors, leaving you to manually search through external sites, crawl wayback machines, or dig through git logs to locate where those pages went.
We can automate the entire repair process. In this tutorial, we will write a Python tool that parses markdown files, verifies all local and external hyperlinks asynchronously, and leverages Claude 3.5 Sonnet to intelligently identify or predict the correct replacements, rewriting the files on disk automatically.
The Logic Behind Intelligent Repair
Why use an LLM instead of a simple lookup table? Because broken links usually fall into three categories that require context to solve:
- Changed Paths: A local file moved from
/docs/setup/install.mdto/docs/getting-started/installation.md. - Platform Reorganisation: An external vendor updated their documentation hierarchy (e.g., shifting
/v1/apito/docs/reference/api). - Domain Migrations: The target site changed domains entirely, but kept a similar path structure.
By feeding the broken link, its surrounding paragraph context, and the structure of the target domain or local repository into /platforms/claude, we can let the model use its reasoning capabilities to suggest the new, correct path. If you are comparing model suites for content processing pipelines, you might also look at /platforms/openai, but Sonnet's balance of structured output and spatial understanding of file trees makes it an exceptional choice for repository archaeology.
Step 1: Scanning and Extracting Markdown Links
First, we need to extract all markdown links from our documents. A standard regex will handle basic markdown links ([Text](URL)), but we also want to record the surrounding text block to give our LLM the necessary context.
Create a python file named repair_links.py and set up the extraction layer:
`python
import re
from pathlib import Path
from typing import NamedTuple, List
class MarkdownLink(NamedTuple): filepath: Path link_text: str url: str context: str # The paragraph containing the link
LINK_REGEX = re.compile(r'\[([^\]]+)\]\(([^\)]+)\)')
def extract_links_from_file(file_path: Path) -> List[MarkdownLink]:
links = []
content = file_path.read_text(encoding='utf-8')
paragraphs = content.split('\n\n')
for paragraph in paragraphs:
matches = LINK_REGEX.findall(paragraph)
for text, url in matches:
# Ignore internal anchors within the same page
if url.startswith('#'):
continue
links.append(MarkdownLink(
filepath=file_path,
link_text=text,
url=url,
context=paragraph.strip().replace('\n', ' ')
))
return links
`
Step 2: Testing Links Asynchronously
To keep our execution times reasonable, we should check URLs in parallel using httpx. We want to identify any link returning a non-200 status code or failing to resolve.
`python
import asyncio
import httpx
async def check_url(client: httpx.AsyncClient, link: MarkdownLink) -> tuple[MarkdownLink, bool]: # If it is a relative local link, check if the file exists on disk if not link.url.startswith(('http://', 'https://')): target_path = link.filepath.parent / link.url.split('#')[0] return link, target_path.exists()
For external links, perform a rapid HEAD or GET request try: response = await client.get(link.url, follow_redirects=True, timeout=10.0) return link, response.status_code == 200 except httpx.RequestError: return link, False
async def scan_for_broken_links(links: List[MarkdownLink]) -> List[MarkdownLink]:
async with httpx.AsyncClient() as client:
tasks = [check_url(client, link) for link in links]
results = await asyncio.gather(*tasks)
return [link for link, is_valid in results if not is_valid]
`
Step 3: Feeding the Broken Links to Claude
Now we design the prompt that directs Claude to find the replacement URLs. We will pass Claude the context of the link, the current broken URL, and a map of our local file tree (if it is a local link) to help it find where the resource might have relocated.
We will structure our prompt dynamically. You can learn more about formatting complex JSON request structures at our guide on /prompts.
`python
import os
from anthropic import Anthropic
SYSTEM_PROMPT = """ You are a precise technical writer and site reliability assistant. Your task is to fix a broken hyperlink in a markdown document.
Analyze the broken URL, the text anchor, and the surrounding paragraph context. Provide the best possible correction.
Rules:
1. If you can confidently identify the updated URL, return it.
2. If it is a local reference (e.g., relative path) and you are given the workspace list, select the file path that closest matches the intent.
3. Return your response strictly as a JSON object with two keys:
- "fixed_url": "the new URL or file path"
- "explanation": "a short reason why you chose this correction"
4. Do not include markdown wraps like `json in your final output, return only raw JSON.
"""
def resolve_broken_link(broken_link: MarkdownLink, local_files: List[str]) -> str:
client = Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
prompt_payload = f"""
File being edited: {broken_link.filepath}
Broken URL: {broken_link.url}
Anchor Text: {broken_link.link_text}
Paragraph Context: "{broken_link.context}"
Available local files (if target is local): {', '.join(local_files[:200])}
"""
try:
message = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=300,
temperature=0.2,
system=SYSTEM_PROMPT,
messages=[{"role": "user", "content": prompt_payload}]
)
import json
data = json.loads(message.content[0].text.strip())
return data.get("fixed_url", broken_link.url)
except Exception as e:
print(f"Could not auto-correct {broken_link.url}: {e}")
return broken_link.url
`
Step 4: Putting It All Together
Finally, we will write a script that orchestrates the scanning, targets the broken links, gathers Claude's corrections, and updates the markdown files in place.
`python
def update_file_with_correction(filepath: Path, old_url: str, new_url: str):
if old_url == new_url:
return
content = filepath.read_text(encoding='utf-8')
# Replace exact instances of the old target URL
updated_content = content.replace(f"]({old_url})", f"]({new_url})")
filepath.write_text(updated_content, encoding='utf-8')
print(f"Updated {filepath.name}: {old_url} -> {new_url}")
def main(): # Walk through your target directory (e.g., a docs/ folder) docs_dir = Path("./docs") all_md_files = list(docs_dir.glob("*/.md")) local_file_paths = [str(f.relative_to(docs_dir)) for f in all_md_files] print(f"Scanning {len(all_md_files)} markdown files for links...") all_links = [] for f in all_md_files: all_links.extend(extract_links_from_file(f)) print(f"Checking {len(all_links)} total links...") broken_links = asyncio.run(scan_for_broken_links(all_links)) print(f"Found {len(broken_links)} broken links. Querying Claude for corrections...") for link in broken_links: new_url = resolve_broken_link(link, local_file_paths) update_file_with_correction(link.filepath, link.url, new_url)
if __name__ == "__main__":
main()
`
Ensuring Quality and Safety
Auto-patching your codebase is incredibly empowering, but doing it blindly invites unexpected side effects. To maintain high-quality documentation, always run this workflow on a clean Git branch and inspect the output using git diff before committing the updates. If you run into API timeout issues or rate limits, consult our debugging guide over at /platforms/claude/articles for handling transient network faults seamlessly.
With this automated loop running on a weekly Cron, you can effectively eliminate stale paths and keep your developer documentation highly reliable without manual effort.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.