Tutorials & Guides
How to Build a Local Dev Harness to Test Prompt Degradation Across Model Updates Using Python and Claude 3.5 Sonnet
Stop relying on manual 'vibe checks' for your AI features. Learn how to build a lightweight, local testing script in Python to catch prompt drift and API regressions before your users do.
Updated 10/5/2026
Why Prompt Testing is Broken (And How We Fix It)
We have all been there. You spend three days crafting the perfect system prompt. Your structured JSON output is flawless, your tone is exactly on-brand, and your app behaves beautifully. Then, a provider drops a silent model update, or you decide to switch models to save on API costs, and suddenly your parsing pipeline is throwing 500 errors because the model decided to wrap its response in markdown backticks again.
Most developers test their prompts using what we affectionately call "the vibe check." You write a prompt, run it three times in the web playground, say "looks good to me," and push it to production. That is a ticking time bomb for any serious application.
To build robust AI features, you need a local development harness that treats prompts like code. This tutorial will walk you through building a simple, lightweight Python CLI tool to run structured, reproducible assertions against your LLM outputs using Claude 3.5 Sonnet.
If you are new to managing system instructions programmatically, you might want to look at our guide on writing clean prompt templates first. If you run into API issues with Anthropic during this build, you can check their developer portal or reach out via Anthropic Support.
The Anatomy of a Local Prompt Test
Our local testing harness will run against a list of test cases defined in a local JSON file. Each test case will contain: 1. Input variables (e.g., user queries or raw text to process). 2. Assertions (e.g., checking if the output contains specific keys, passes a schema validation, or stays under a target character limit).
Let us organise our directory like this:
`text
prompt-tester/
├── prompts/
│ └── summariser_prompt.txt
├── tests/
│ └── summariser_cases.json
├── test_runner.py
└── requirements.txt
`
First, make sure you have the required dependencies installed. Create your requirements.txt file:
`text
anthropic>=0.30.0
pydantic>=2.0.0
pytest>=8.0.0
`
Run pip install -r requirements.txt to get started.
Step 1: Defining Your Prompt and Test Cases
Let us create a system prompt designed to clean up messy transcription data into a structured summary. Save this inside prompts/summariser_prompt.txt:
`text
You are a precise backend parsing assistant. Your task is to extract actionable items from a raw, messy meeting transcript.
Your output must be raw JSON with the following structure:
{
"action_items": [
{
"assignee": "Name or Unknown",
"task": "The specific task to complete",
"priority": "High", "Medium", or "Low"
}
]
}
Do not include any introductory or concluding conversational text. Do not wrap the output in markdown code blocks.
`
Next, we need structured test cases to verify that our prompt behaves consistently across model changes. Create tests/summariser_cases.json:
`json
[
{
"id": "case_1_standard_transcript",
"variables": {
"transcript": "Dave: I will fix the login bug by Friday. Sarah, can you update the Figma mockups? Sarah: Yeah, I will get to that tomorrow."
},
"assertions": {
"contains_keys": ["action_items"],
"minimum_items": 2,
"expected_strings": ["login bug", "Figma mockups"]
}
},
{
"id": "case_2_empty_transcript",
"variables": {
"transcript": "Um, hello? Is this thing on? Can everyone hear me? Okay, goodbye."
},
"assertions": {
"contains_keys": ["action_items"],
"minimum_items": 0,
"expected_strings": []
}
}
]
`
Step 2: Coding the Testing Harness
Now, let us build test_runner.py. This script will load our prompt, parse the test cases, call the Claude 3.5 Sonnet API, and run programmatic assertions against the returned payload.
`python
import os
import json
import sys
from anthropic import Anthropic
Ensure you have your ANTHROPIC_API_KEY environment variable set. client = Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
def load_file(path): with open(path, "r", encoding="utf-8") as f: return f.read()
def run_assertions(response_text, assertions): failures = [] # Attempt to parse the response as JSON try: data = json.loads(response_text.strip()) except json.JSONDecodeError: failures.append("Response is not valid JSON") return failures
Assertion 1: Check for required top-level JSON keys for key in assertions.get("contains_keys", []): if key not in data: failures.append(f"Missing expected key: '{key}'")
Assertion 2: Verify minimum number of list items if "minimum_items" in assertions and isinstance(data.get("action_items"), list): item_count = len(data["action_items"]) min_items = assertions["minimum_items"] if item_count < min_items: failures.append(f"Expected at least {min_items} action items, got {item_count}")
Assertion 3: Substring matching in raw response for substring in assertions.get("expected_strings", []): if substring.lower() not in response_text.lower(): failures.append(f"Expected substring '{substring}' not found in model output") return failures
def execute_tests(): system_prompt = load_file("prompts/summariser_prompt.txt") test_cases = json.loads(load_file("tests/summariser_cases.json")) all_passed = True print(f"\n🚀 Starting Prompt Evaluation for model: claude-3-5-sonnet-latest\n") for case in test_cases: print(f"🧪 Running {case['id']}...") # Prepare the user prompt user_content = f"Please process this transcript:\n{case['variables']['transcript']}" try: message = client.messages.create( model="claude-3-5-sonnet-latest", max_tokens=1000, temperature=0.0, # Keep it deterministic for testing system=system_prompt, messages=[ {"role": "user", "content": user_content} ] ) response_text = message.content[0].text failures = run_assertions(response_text, case["assertions"]) if failures: print(f"❌ Failed! Errors:") for err in failures: print(f" - {err}") print(f"Raw Output was: {response_text}\n") all_passed = False else: print(f"✅ Passed!\n") except Exception as e: print(f"💥 API Error occurred: {e}") all_passed = False if not all_passed: sys.exit(1) else: print("🎉 All prompt assertions passed successfully!") sys.exit(0)
if __name__ == "__main__":
execute_tests()
`
Step 3: Running Your Harness
Ensure your API key is exported in your environment:
`bash
export ANTHROPIC_API_KEY="your-api-key-here"
python test_runner.py
`
If everything is working as intended, you will see a clean output showing that your assertions passed. Now, whenever Anthropic updates its models or you decide to switch system prompts, you simply rerun this local script.
If you want to view examples of highly structured prompting patterns to improve your reliability, you can browse through the official Anthropic developer cookbook or explore live prompting architectures in our prompts generator.
Taking It Further
This basic script is highly extensible. If your system relies on more complex outputs, you can integrate Pydantic for full runtime validation of your JSON responses instead of basic string assertions. You can also integrate this into your CI/CD pipeline, forcing your build to fail if a developer pushes a prompt change that degrades output performance.
Now you can stop guessing and start measuring. Happy building!
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.