Tutorials & Guides
How to Build a Local PII Sanitiser CLI for Your Cloud LLM API Calls Using Python
Sending sensitive user data to cloud LLMs is a compliance nightmare. Here is how to build a local, reversible PII redaction pipeline using Python and Microsoft Presidio.
Updated 10/4/2026
We have all been there. You are building an internal tool or processing user feedback with a cloud LLM, and suddenly you spot it: a raw SQL dump, a plain-text password, or a customer’s home address sitting in the raw prompt. Sending that straight to a third-party API is not just bad form; it is a rapid way to get your security team breathing down your neck.
While platforms like /platforms/openai and Anthropic have robust privacy policies, the safest data is the data you never send in the first place.
In this guide, we are going to build a lightweight, local Command Line Interface (CLI) in Python that intercepts your prompts, redacts Personally Identifiable Information (PII) using Microsoft Presidio, sends the anonymised text to your LLM of choice, and then seamlessly reconstructs the response locally. Your cloud provider gets clean, sanitised text. Your user gets their data back intact. Everybody wins.
Why Existing Cloud Redaction is a Trap
You might be tempted to use a cloud-based PII redaction service. Don't. If you are shipping data to a third-party API to strip out the PII before shipping it to another API, you have simply doubled your attack surface and added extra network latency to your loop.
Running your sanitiser locally on your own container or development machine ensures your data never leaves your boundary until it is scrubbed clean.
Our tool will handle three specific tasks:
1. Analyze: Scan incoming text for entities like emails, phone numbers, names, and IP addresses.
2. Anonymise: Replace those entities with reversible placeholders (e.g., [EMAIL_1]).
3. De-anonymise: Take the LLM's response and swap the real values back in before presenting it to the user.
Setting Up Your Workspace
First, let's get our environment set up. We will use Microsoft's presidio-analyzer and presidio-anonymizer, along with the official OpenAI SDK.
Run the following in your terminal:
`bash
pip install presidio-analyzer presidio-anonymizer openai
python -m spacy download en_core_web_lg
`
Note: We are downloading the large spaCy English model (`en_core_web_lg`) because it provides significantly better accuracy for entity recognition than the small model. Accuracy is what makes this entire workflow tick.
Step 1: Building the Reversible Map
To reconstruct our LLM's output later, we need to keep track of what we replaced. We will create a class called PIISanitiser that manages this state.
Create a file named sanitiser.py and add the following template:
`python
import re
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
from presidio_anonymizer.entities import OperatorConfig
class PIISanitiser: def __init__(self): self.analyzer = AnalyzerEngine() self.anonymizer = AnonymizerEngine() self.mapping = {}
def sanitise(self, text: str) -> str:
# Step 1: Analyze
results = self.analyzer.analyze(
text=text,
language="en",
entities=["EMAIL_ADDRESS", "PHONE_NUMBER", "PERSON", "IP_ADDRESS"]
)
# Step 2: Anonymise with custom operators to track replacements
anonymized_result = self.anonymizer.anonymize(
text=text,
analyzer_results=results,
operators={
"DEFAULT": OperatorConfig("replace", {"new_value": "[REDACTED]"})
}
)
return anonymized_result.text
`
This basic setup replaces everything with [REDACTED]. However, this is useless for reconstruction because we cannot tell [REDACTED] apart from another [REDACTED]. We need unique tokens.
Step 2: Implementing Unique Token Mapping
To make this reversible, we will write a custom routine that dynamically populates our local mapping dictionary with unique indices.
Replace your sanitise method with the following implementation:
`python
def sanitise(self, text: str) -> str:
results = self.analyzer.analyze(
text=text,
language="en",
entities=["EMAIL_ADDRESS", "PHONE_NUMBER", "PERSON", "IP_ADDRESS"]
)
# Sort results backward to avoid offset drift while modifying text
sorted_results = sorted(results, key=lambda x: x.start, reverse=True)
sanitised_text = text
for result in sorted_results:
original_value = text[result.start:result.end]
entity_type = result.entity_type
# Generate a unique key
if original_value not in self.mapping.values():
key_idx = len(self.mapping) + 1
token = f"[{entity_type}_{key_idx}]"
self.mapping[token] = original_value
else:
# Retrieve existing token if entity appears multiple times
token = [k for k, v in self.mapping.items() if v == original_value][0]
# Replace in string
sanitised_text = sanitised_text[:result.start] + token + sanitised_text[result.end:]
return sanitised_text
`
Now, let's write the inverse function to restore those original values once the LLM is finished processing.
`python
def restore(self, text: str) -> str:
restored_text = text
for token, original_value in self.mapping.items():
# Use re.escape to handle any weird characters safely
restored_text = re.sub(re.escape(token), original_value, restored_text)
return restored_text
`
Step 3: Integrating the LLM Pipeline
Now we will create our runner loop. This script takes user input, sanitises it, shoots it over to OpenAI, and reconstructs the response.
Create main.py:
`python
import os
from openai import OpenAI
from sanitiser import PIISanitiser
Initialize your clients client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY")) sanitiser = PIISanitiser()
def process_prompt_safely(prompt: str) -> str: # 1. Redact locally clean_prompt = sanitiser.sanitise(prompt) print(f"\n[DEBUG] Clean Prompt sent to API: {clean_prompt}\n") # 2. Call your LLM response = client.chat.completions.create( model="gpt-4o-mini", messages=[ {"role": "system", "content": "You are a helpful assistant. You must preserve any tokens format like [PERSON_1] or [EMAIL_2] exactly as written in the prompt. Do not translate or change them."}, {"role": "user", "content": clean_prompt} ], temperature=0.3 ) raw_output = response.choices[0].message.content print(f"[DEBUG] Raw LLM Output: {raw_output}\n") # 3. Restore PII locally final_output = sanitiser.restore(raw_output) return final_output
if __name__ == "__main__":
user_input = "Can you write a short draft email to Alice Vance (alice.vance@example.com) asking her to update the server at 192.168.1.55?"
print(f"Original Prompt: {user_input}")
result = process_prompt_safely(user_input)
print(f"Final Output: {result}")
`
How it Looks in Action
When you run python main.py, the CLI will output the following steps:
- Original Prompt:
Can you write a short draft email to Alice Vance (alice.vance@example.com)... - Clean Prompt sent to API:
Can you write a short draft email to [PERSON_1] ([EMAIL_ADDRESS_2])... - Raw LLM Output:
Subject: Action Required... Hi [PERSON_1], please update the server at [IP_ADDRESS_3]... - Final Output:
Subject: Action Required... Hi Alice Vance, please update the server at 192.168.1.55...
Because we instructed the model to leave the tags intact, we can parse them out easily with simple string replacement.
Troubleshooting Edge Cases
Occasionally, the model might try to "correct" your placeholders (e.g., turning [EMAIL_ADDRESS_1] into [EMAIL_1]). If you run into issues with structured outputs or format drift, swing by our guide on /platforms/openai/articles for strategies on handling schema validation under constraint.
Additionally, if you want to experiment with alternative prompting configurations to prevent the model from touching your bracketed tokens, our automated /prompts can help you construct rock-solid system instructions that enforce absolute token preservation.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.