Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build a Local Prompt Evaluation Runner with Pytest and Claude 3.5 Sonnet

Stop guessing if your prompt updates broke your production outputs. Here is a step-by-step guide to building a lightweight, automated prompt evaluation runner using Python and Pytest.

Updated 10/7/2026

Stop Guessing, Start Asserting

We have all been there. You spend three hours tweaking a system prompt for Claude to make sure it formats structured data perfectly. You change one word to fix an edge case, ship it to production, and immediately discover you have broken three other downstream features.

Prompt engineering without automated tests is just vibes-based development. To build reliable AI features, you need to treat prompts like code. This means writing assertions, maintaining a local dataset of test cases, and running evaluations every single time you modify your system instructions.

In this tutorial, we will build a lightweight, local prompt evaluation runner using Pytest and Claude 3.5 Sonnet. By the end of this guide, you will have a pipeline that runs assertions against your LLM outputs to verify JSON schemas, check semantic alignment, and flag regressions before they reach your users.

The Architecture of a Prompt Eval

Our evaluation framework needs three distinct components to work effectively:

  1. The Target Function: The wrapper function that sends our prompt and system instructions to the Anthropic API.
  2. The Test Cases: A local dataset (we will use a simple JSON file) containing various test inputs and their expected output characteristics.
  3. The Assertion Runner: A Pytest file that runs our test cases, calls the target function, and evaluates the output using standard Python checks and LLM-as-a-judge criteria.

Let’s set up our project directory first:

`bash mkdir prompt-evaluator cd prompt-evaluator python3 -m venv venv source venv/bin/activate pip install anthropic pytest python-dotenv `

Create a .env file in your root folder and add your API key:

`env ANTHROPIC_API_KEY=your_actual_api_key_here `

Step 1: Write the Target Function

Let’s build the application function we want to evaluate. Let’s assume we are building a feature that takes unstructured customer support feedback and extracts key metadata: the sentiment, the core issue, and whether the user requires urgent follow-up. We want this returned as strict JSON.

Create a file named app.py and write the parsing logic:

`python import os from anthropic import Anthropic from dotenv import load_dotenv

load_dotenv() client = Anthropic()

SYSTEM_PROMPT = """ You are a precise support log analyst. Extract data from user feedback. You must return ONLY a JSON block. No markdown, no conversational text. Schema: { "sentiment": "positive" | "neutral" | "negative", "primary_issue": string, "requires_urgency": boolean } """

def analyze_feedback(feedback_text: str) -> str: message = client.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=1000, temperature=0.0, system=SYSTEM_PROMPT, messages=[ {"role": "user", "content": feedback_text} ] ) return message.content[0].text.strip() `

Note that we are setting the temperature to 0.0. Determinism is critical for stable prompt evaluations; it ensures we are checking the actual instruction-following capabilities of the prompt under normal conditions rather than fighting random creative variance.

Step 2: Create Your Test Dataset

Next, we need a set of realistic test cases that represent the typical inputs our app will receive, alongside the criteria we want to assert. Create a file called test_cases.json:

`json [ { "id": "case_01", "input": "My order arrived shattered in the box! This was supposed to be a birthday gift for tomorrow.", "expected_sentiment": "negative", "expected_urgency": true }, { "id": "case_02", "input": "The setup took longer than expected, but the product works incredibly well. Very pleased with the build quality.", "expected_sentiment": "positive", "expected_urgency": false } ] `

Step 3: Writing the Pytest Runner with LLM-as-a-Judge

Now we write the test suite. We will check both hard assertions (like verifying that the output is indeed valid JSON and contains the correct keys) and soft semantic assertions (such as asking Claude to evaluate its own output's accuracy as a judge).

If you find yourself stuck on parsing issues while building these pipelines, check out our Claude troubleshooting hub for common parsing fixes.

Create a file named test_prompt_eval.py:

`python import json import pytest from app import analyze_feedback from anthropic import Anthropic

Helper to load test cases def load_tests(): with open("test_cases.json", "r") as f: return json.load(f)

LLM-as-a-judge helper to check semantic alignment def evaluate_with_judge(input_text: str, output_text: str, expected_sentiment: str) -> bool: client = Anthropic() judge_prompt = f""" You are an impartial test evaluator. Original input feedback: "{input_text}" LLM Generated Output: "{output_text}" Does the generated output accurately identify the customer's sentiment as '{expected_sentiment}'? Answer with exactly 'YES' or 'NO' only. """ response = client.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=10, temperature=0.0, messages=[{"role": "user", "content": judge_prompt}] ) return response.content[0].text.strip().upper() == "YES"

@pytest.mark.parametrize("test_case", load_tests()) def test_prompt_outputs(test_case): raw_output = analyze_feedback(test_case["input"]) # Test 1: Verify valid JSON formatting try: parsed_data = json.loads(raw_output) except json.JSONDecodeError: pytest.fail("Output was not valid JSON") # Test 2: Verify required keys exist required_keys = {"sentiment", "primary_issue", "requires_urgency"} assert required_keys.issubset(parsed_data.keys()), f"Missing key fields in {parsed_data}" # Test 3: Verify Boolean values are correct assert parsed_data["requires_urgency"] == test_case["expected_urgency"], \ f"Urgency misidentified. Expected {test_case['expected_urgency']}, got {parsed_data['requires_urgency']}" # Test 4: Run LLM-as-a-judge for semantic evaluation judge_result = evaluate_with_judge( test_case["input"], raw_output, test_case["expected_sentiment"] ) assert judge_result, f"Judge failed the output sentiment accuracy for input: {test_case['input']}" `

Running the Evaluation Suite

Executing your brand-new prompt evaluation runner is as simple as running your normal pytest command. Open your terminal and run:

`bash pytest test_prompt_eval.py -v `

If all assertions pass, you will see green passes. If you modify your system prompt in app.py in a way that alters the JSON keys or misclassifies sentiments, your assertions will catch it instantly.

Using this setup allows you to find exactly what makes your prompts tick. If you want to take your prompt design a step further, try generating tailored test templates using our prompt generator to easily test edge cases.

tutorialsclaudepythontestingevals

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.