Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build a Local Codebase Context Packer for Claude 3.5 Sonnet Using Python

Stop wasting your token budget on copy-pasting individual files. Build a lightweight Python tool that respects your .gitignore and bundles your codebase into Claude-friendly XML.

Updated 10/5/2026

Why Your Copy-Paste Workflow is Slowing You Down

We have all done it. You are deep in a debugging session, and you find yourself frantically tab-switching between your editor and your browser, copying and pasting five different files into Claude just so the LLM has enough context to fix a minor database integration bug.

Not only is this incredibly tedious, but it also leads to bloated, messy prompts. Claude loves structure, and Anthropic’s models are famously optimised to parse XML tags better than almost any other format. If you dump raw, unlabelled code blocks into the chat window, the model has to work twice as hard to figure out where auth.py ends and models.py begins.

To make your development workflow actually tick, you need a local, lightweight context packer. This tutorial will guide you through building a command-line interface (CLI) tool in Python that automatically crawls your project directory, respects your .gitignore rules, and bundles your entire codebase into a clean, XML-structured file ready to drop straight into your prompts.

Designing the Context Packer

Before we write a single line of code, let us establish what a robust context packer must do:

  1. Respect `.gitignore`: We do not want to upload node_modules, compiled .pyc files, or binary assets. Sending 40,000 lines of minified JavaScript to an LLM is a spectacular way to burn through your context window.
  2. Use XML Formatting: We will wrap each file's contents inside explicit <file> tags, complete with a path attribute. This aligns perfectly with Anthropic's system prompt guidelines.
  3. Token Warning: We want a rough estimate of how many tokens we are packing before we copy it, so we do not accidentally break the input threshold.

If you want to read more about how Claude interprets structured templates, check out our prompt generator tool for tips on setting up system contexts.

Step 1: Setting Up the Project

Create a new folder and set up a basic Python script. You do not need any heavy external dependencies; we will use standard libraries alongside pathspec to handle gitignore pattern matching properly.

`bash mkdir context-packer cd context-packer pip install pathspec touch packer.py `

Step 2: Parsing `.gitignore` Dynamically

We need to make sure we do not manually hardcode ignored folders. The cleanest way is to load your project’s existing .gitignore and use pathspec to match files against those rules.

Create the boilerplate and the ignore-loading logic in packer.py:

`python import os import sys from pathlib import Path import pathspec

def load_gitignore_spec(project_dir: Path) -> pathspec.PathSpec: gitignore_path = project_dir / ".gitignore" fallback_patterns = [".git/", "__pycache__/", "*.pyc", ".DS_Store", "node_modules/"] if gitignore_path.exists(): with open(gitignore_path, "r", encoding="utf-8") as f: lines = f.read().splitlines() # Combine user's gitignore rules with our sensible fallbacks patterns = lines + fallback_patterns else: patterns = fallback_patterns return pathspec.PathSpec.from_lines("gitwildmatch", patterns) `

Step 3: Traversal and XML Generation

Now, we need to recursively traverse our directories, check each file against our loaded ignore patterns, and format the output into clean XML blocks.

We will skip binary files entirely, as Claude cannot read raw binary bytes anyway.

`python def is_binary(file_path: Path) -> bool: try: with open(file_path, "tr", encoding="utf-8") as f: f.read(1024) return False except UnicodeDecodeError: return True

def pack_codebase(project_dir: Path, spec: pathspec.PathSpec) -> str: output = [] output.append("<codebase>") for root, dirs, files in os.walk(project_dir): # Modify dirs in-place to prevent os.walk from entering ignored directories dirs[:] = [d for d in dirs if not spec.match_file(str(Path(root) / d))] for file in files: full_path = Path(root) / file relative_path = full_path.relative_to(project_dir) if spec.match_file(str(relative_path)): continue if is_binary(full_path): continue try: with open(full_path, "r", encoding="utf-8", errors="ignore") as f: content = f.read() output.append(f' <file path="{relative_path}">') output.append("<![CDATA[") output.append(content) output.append("]]>") output.append(" </file>") except Exception as e: print(f"Skipping {relative_path} due to error: {e}", file=sys.stderr) output.append("</codebase>") return "\n".join(output) `

Using <![CDATA[ ... ]]> tags is an essential safety measure. It ensures that if your code contains characters that look like XML tags (like Python comparison operators or HTML templates in a frontend framework), Claude's parser will not get confused.

Step 4: Estimating Tokens and Writing the CLI

While a precise token count requires calling the official API tokenizer, a quick-and-dirty rule of thumb is that 1 token is roughly equivalent to 4 characters of English text or code. Let us build a clean CLI wrapper that prints the packed XML to standard output, saves it to a file, and alerts us if the output is suspiciously massive.

`python def main(): if len(sys.argv) < 2: print("Usage: python packer.py <path-to-project>") sys.exit(1) project_path = Path(sys.argv[1]).resolve() if not project_path.is_dir(): print(f"Error: {project_path} is not a valid directory.") sys.exit(1) print(f"Scanning {project_path}...", file=sys.stderr) spec = load_gitignore_spec(project_path) packed_xml = pack_codebase(project_path, spec) char_count = len(packed_xml) estimated_tokens = char_count // 4 print(f"Packed size: {char_count} characters (approx. {estimated_tokens} tokens)", file=sys.stderr) if estimated_tokens > 150000: print("WARNING: This payload is massive. Consider ignoring large folders manually.", file=sys.stderr) # Output the result to stdout print(packed_xml)

if __name__ == "__main__": main() `

How to Use It in Your Workflow

You can pipe the output of this script straight to your clipboard. If you are on macOS:

`bash python packer.py /path/to/your/project | pbcopy `

Or on Linux:

`bash python packer.py /path/to/your/project | xclip -sel clip `

When you head over to Claude, you can structure your prompt like this:

`text I am working on a Python application. Below is the current state of my codebase packed in XML tags.

[PASTE YOUR PACKED XML HERE]

Please inspect the files in <codebase> and find why the authentication callback in auth.py is returning a 403 error on local environments. `

Keeping Things Running Smoothly

If you find yourself hitting unexpected errors or empty outputs while running this script, you may have run into some common parsing quirks. For detailed debugging steps on handling weird file encodings or system-specific directory paths, head over to our dedicated troubleshooting guides in our Claude Hub.

By spending ten minutes setting up this local pipeline, you completely eliminate copy-paste fatigue and present Claude with a highly readable, deterministic representation of your workspace. No bloat, no clutter—just clean code contexts ready to build.

claudepythontutorialsdev-workflowprompting

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.