Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build a Local OpenAI API Cost Monitor in Python with Slack Alerts

OpenAI's dashboard billing alerts can take up to 24 hours to trigger. Build an asynchronous local cost-tracking wrapper in Python to alert your Slack channel the instant a runaway agent starts burning money.

Updated 10/5/2026

The Runaway Agent Problem

We have all heard the horror stories: an engineer builds an agentic pipeline, sets it running before heading out for lunch, and returns to find the script got caught in an infinite self-correction loop, generating thousands of requests. By the time OpenAI’s native email alert arrives up to 24 hours later, the damage is done, and your credit card is crying.

While OpenAI allows you to set hard limits on your developer account, these are account-wide ceilings. If you are running multiple projects, testing new prompts, or hosting experimental tools, you need a granular way to keep your development infrastructure ticking over without risking a surprise bill.

In this tutorial, we will build a local, non-blocking cost-tracking wrapper in Python using SQLite. It will intercept your calls to OpenAI, log the token usage, calculate real-time pricing, and immediately post a warning message to your Slack channel if your project’s hourly spending crosses a threshold.

How the Architecture Works

To keep our wrapper as clean as possible, we will use a Python decorator pattern. This lets you monitor your LLM calls by simply adding @monitor_cost above your API client functions.

Every time the function runs, the decorator will: 1. Intercept the returned OpenAI object to read the exact usage metadata. 2. Query SQLite to calculate the running total spent over the last hour. 3. Send an asynchronous Slack alert via Webhooks if you have crossed your pre-configured budget.

To better understand how LLM providers calculate input and output tokens, take a look at our comprehensive AI Glossary.

Step 1: SQLite Schema & Cost Metrics

First, let us set up a basic SQLite database to log every API call. We need to store the timestamp, model name, prompt tokens, and completion tokens.

Create a new project directory and initialise monitor.py:

`python import sqlite3 import time from pathlib import Path

DB_PATH = Path("api_costs.db")

def init_db(): with sqlite3.connect(DB_PATH) as conn: conn.execute(""" CREATE TABLE IF NOT EXISTS api_calls ( id INTEGER PRIMARY KEY AUTOINCREMENT, timestamp REAL, model TEXT, prompt_tokens INTEGER, completion_tokens INTEGER, estimated_cost REAL ) """) conn.commit()

init_db() `

Step 2: Mapping Model Pricing Dynamically

LLM pricing changes frequently, and different models have different rates for input versus output tokens. Let us map out a pricing dictionary for the most common OpenAI models. We will calculate all costs per million tokens to make the numbers easier to work with.

`python MODEL_PRICING = { "gpt-4o": { "input": 2.50, # Cost per 1,000,000 tokens "output": 10.00 }, "gpt-4o-mini": { "input": 0.150, "output": 0.600 } }

def calculate_cost(model: str, prompt_tokens: int, completion_tokens: int) -> float: # Default to gpt-4o pricing if the model isn't mapped explicitly rates = MODEL_PRICING.get(model, MODEL_PRICING["gpt-4o"]) input_cost = (prompt_tokens / 1_000_000) * rates["input"] output_cost = (completion_tokens / 1_000_000) * rates["output"] return input_cost + output_cost `

Step 3: Setting Up Slack Webhook Alerts

Go to your Slack workspace, create a new App, enable Incoming Webhooks, and copy your Webhook URL. We will write a lightweight helper to post warning cards directly to your Slack channel using the standard urllib library to avoid dragging in another dependency like requests.

`python import json import urllib.request

SLACK_WEBHOOK_URL = "https://hooks.slack.com/services/YOUR/WEBHOOK/HERE"

def send_slack_alert(current_hourly_spend: float, threshold: float): payload = { "text": f"⚠️ API Cost Alert: High spending detected on local environment!", "attachments": [ { "color": "#ff0000", "fields": [ { "title": "Last 60 Min Spend", "value": f"${current_hourly_spend:.4f}", "short": True }, { "title": "Configured Threshold", "value": f"${threshold:.4f}", "short": True } ], "footer": "Action required: Verify that your loops or agents are executing as expected." } ] } try: req = urllib.request.Request( SLACK_WEBHOOK_URL, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json"} ) with urllib.request.urlopen(req) as response: response.read() except Exception as e: print(f"Failed to send Slack alert: {e}") `

Step 4: The Decorator Code

Now we tie everything together inside our Python decorator. The decorator expects a function that returns an OpenAI API response object, parses the returned usage dictionary, saves the data to our SQLite instance, and performs the budget check.

`python import functools

BUDGET_THRESHOLD_HOURLY = 1.50 # Set your budget ceiling in USD

def get_recent_spend_total(window_seconds: int = 3600) -> float: cutoff = time.time() - window_seconds with sqlite3.connect(DB_PATH) as conn: cursor = conn.cursor() cursor.execute( "SELECT SUM(estimated_cost) FROM api_calls WHERE timestamp > ?", (cutoff,) ) row = cursor.fetchone() return row[0] if row[0] is not None else 0.0

def monitor_cost(func): @functools.wraps(func) def wrapper(args, *kwargs): response = func(args, *kwargs) try: # Safely extract usage metadata model = response.model usage = response.usage prompt_tokens = usage.prompt_tokens completion_tokens = usage.completion_tokens cost = calculate_cost(model, prompt_tokens, completion_tokens) # Save to SQLite log with sqlite3.connect(DB_PATH) as conn: conn.execute( "INSERT INTO api_calls (timestamp, model, prompt_tokens, completion_tokens, estimated_cost) VALUES (?, ?, ?, ?, ?)", (time.time(), model, prompt_tokens, completion_tokens, cost) ) conn.commit() # Evaluate total cost in the last hour hourly_spend = get_recent_spend_total() if hourly_spend > BUDGET_THRESHOLD_HOURLY: send_slack_alert(hourly_spend, BUDGET_THRESHOLD_HOURLY) except AttributeError: print("Cost Monitor Error: Response did not contain standard token usage parameters.") return response return wrapper `

Putting It to Work

To use your brand new cost monitor, simply decorate the function where you invoke your OpenAI client call:

`python from openai import OpenAI

client = OpenAI(api_key="your-openai-api-key")

@monitor_cost def get_chat_completion(prompt: str): return client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": prompt}] )

Run a test call to see it log response = get_chat_completion("Explain quantum mechanics in three simple sentences.") print("Response received. Local db logged usage successfully.") ```

Getting Stale Data Alerts?

If your local Python environment runs into weird network latency or fails to connect to the Slack endpoint properly when handling async requests, you might need to adjust your threading model. Read through our detailed guides in our OpenAI Hub to see how to run this tracking pipeline completely asynchronously using asyncio.

Now, you can confidently run local experimental agents knowing that if your script gets caught in a wild logic loop, you will get a ping on your phone in seconds rather than discovering a three-figure invoice on your credit card the next morning.

openaipythontutorialsautomationsdev-workflow

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.