Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build a Lightweight Semantic Router in Python Without Heavy Frameworks

Stop importing massive AI orchestration frameworks just to route user queries. Learn how to build a fast, dependency-free semantic router using local sentence embeddings and cosine similarity.

Updated 9/23/2026

Why You Should Bin the Heavy Orchestration Frameworks

If you want to route a user's prompt to a specific LLM agent, database query, or customer support flow, you don’t need to drag a massive orchestration framework into your codebase.

Many tutorials will tell you to pull in complex libraries with hundreds of deep dependencies just to perform simple classification. This introduces massive dependency trees, slows down boot-up times, and makes debugging a nightmare.

Instead, you can achieve incredibly fast, accurate, and completely deterministic query routing by building your own lightweight semantic router. By converting user input into an embedding vector and comparing it to a small set of static reference vectors using simple cosine similarity, you can route queries in milliseconds. Here is how to build your own clean, framework-free semantic router using Python.

What is Semantic Routing?

Before we write the code, let’s define the mechanism. Traditional routers rely on keyword matching (e.g., looking for the word "refund" to trigger a billing flow). This fails when a user says, "I want my money back," which contains no overlapping keywords but carries identical intent.

Semantic routing solves this by mapping the meaning of the input query to a predefined route. To learn more about how vector spaces handle this, check out our /glossary.

Our lightweight architecture will operate as follows: 1. We define our application routes and write a few representative "utterances" (phrases) for each. 2. We convert these utterances into vector embeddings using a fast, low-cost API like /platforms/gemini or OpenAI. 3. When a user input arrives, we embed it. 4. We calculate the cosine similarity between the user input vector and all our reference vectors. 5. We route the user to the category with the highest similarity score, provided it crosses a minimum threshold.

Step 1: Setting Up the Route Schema

Let’s write clean, standard Python to define our routes. We will use dataclasses to avoid bloated schemas.

`python from dataclasses import dataclass from typing import List, Callable, Any

@dataclass class Route: name: str utterances: List[str] handler: Callable[[str], Any] threshold: float = 0.70 # Minimum similarity to trigger this route `

Now, let's create some dummy handlers and define our routes. Suppose we are building an assistant for a SaaS application. We have three main paths: billing questions, technical support, and general chit-chat.

`python def handle_billing(prompt: str): return f"Routing to billing engine for: '{prompt}'"

def handle_tech_support(prompt: str): return f"Spinning up technical diagnostics for: '{prompt}'"

def handle_general(prompt: str): return f"Passing to standard conversational LLM for: '{prompt}'"

Define the routes and their exemplary training utterances ROUTES = [ Route( name="billing", utterances=[ "How do I update my credit card?", "Can I get a refund for last month?", "Where can I download my invoice?", "Why was I charged twice?" ], handler=handle_billing, threshold=0.72 ), Route( name="tech_support", utterances=[ "The server is throwing a 500 error.", "I cannot log into my dashboard.", "My API integration is failing with a timeout.", "How do I reset my authentication token?" ], handler=handle_tech_support, threshold=0.72 ) ] ```

Step 2: Generating Vector Embeddings

To compare prompts, we need to convert them into floats. For this tutorial, we will use a simple, robust wrapper around OpenAI's embedding API. Make sure you have the official SDK installed (pip install openai). If you hit API keys or rate limits issues, our troubleshooting section at /platforms/openai/articles has your back.

`python import os from openai import OpenAI

Initialize client client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))

def get_embedding(text: str, model: str = "text-embedding-3-small") -> List[float]: # Replace newlines with spaces as recommended by OpenAI text = text.replace("\n", " ") response = client.embeddings.create(input=[text], model=model) return response.data[0].embedding `

Step 3: Vector Math Without the Overhead

Normally, tutorials will tell you to pull in PyTorch, scikit-learn, or a dedicated vector database. We don’t need any of that. Cosine similarity is a straightforward mathematical formula that we can write in five lines of pure Python using standard library functions.

Cosine similarity measures the cosine of the angle between two vectors. It returns a value between -1 and 1, where 1 means the vectors are pointing in the exact same direction (semantically identical).

`python import math

def dot_product(v1: List[float], v2: List[float]) -> float: return sum(a * b for a, b in zip(v1, v2))

def magnitude(v: List[float]) -> float: return math.sqrt(sum(a * a for a in v))

def cosine_similarity(v1: List[float], v2: List[float]) -> float: mag_v1 = magnitude(v1) mag_v2 = magnitude(v2) if not mag_v1 or not mag_v2: return 0.0 return dot_product(v1, v2) / (mag_v1 * mag_v2) `

Step 4: Compiling the Router

Now, let's build the engine. When our application boots up, we want to pre-calculate (or "compile") the embeddings for all our static route utterances. This ensures we aren't generating reference embeddings on every incoming request, saving both time and API costs.

`python class SemanticRouter: def __init__(self, routes: List[Route]): self.routes = routes self.compiled_routes = {} self.compile_routes() def compile_routes(self): print("Compiling route embeddings...") for route in self.routes: embeddings = [] for utterance in route.utterances: # Fetch embedding for each reference sample emb = get_embedding(utterance) embeddings.append(emb) self.compiled_routes[route.name] = embeddings print("Compilation complete.")

def route(self, user_query: str) -> str: # 1. Embed the incoming user query query_vector = get_embedding(user_query) best_route = None highest_similarity = -1.0 # 2. Compare against all compiled routes for route in self.routes: reference_vectors = self.compiled_routes[route.name] # Find the highest similarity score for this specific route for ref_vector in reference_vectors: similarity = cosine_similarity(query_vector, ref_vector) if similarity > highest_similarity: highest_similarity = similarity best_route = route # 3. Apply safety threshold check if best_route and highest_similarity >= best_route.threshold: print(f"Route matched: '{best_route.name}' (Confidence: {highest_similarity:.4f})") return best_route.handler(user_query) # 4. Fallback to default conversational handler print(f"No strong route match. Highest similarity was {highest_similarity:.4f}.") return handle_general(user_query) `

Step 5: Testing Our Lean Architecture

Let’s instantiate our clean router and see how it handles natural variations of our target queries.

`python if __name__ == "__main__": # Initialize our compiled router router = SemanticRouter(ROUTES) # Test Case 1: Billing Intent (No matching keywords) test_prompt_1 = "I'm trying to look at my bills, where are they?" response_1 = router.route(test_prompt_1) print(response_1) print("-" * 40) # Test Case 2: Tech Support Intent test_prompt_2 = "My site is down and showing some sort of gateway error" response_2 = router.route(test_prompt_2) print(response_2) print("-" * 40) # Test Case 3: Out-of-bounds Query (Should trigger general fallback) test_prompt_3 = "What is the capital of France?" response_3 = router.route(test_prompt_3) print(response_3) `

Keep Your Stack Fast and Maintainable

By avoiding heavy pipelines, your system boots up instantly, operates with minimal latency, and contains zero dependencies beyond a basic API client. You can easily tweak confidence thresholds for each individual route to prevent false positives, or dynamically load routes from a simple JSON config file.

If you want to design complex prompts to feed into your downstream route handlers, check out our interactive tool at /prompts to craft cleaner, more reliable instructions.

pythonembeddingssemantic-routingai-agentstutorials

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.