Tickd.ai
← The Tickd Guide

Tutorials & Guides

How to Build a Lightweight Semantic Search CLI in Python with sqlite-vec and Gemini Flash

Ditch the heavyweight vector databases. Learn how to build a lightning-fast local semantic search tool using SQLite's newest vector extension and Gemini Flash embeddings.

Updated 9/6/2026

Why Your Side Project Doesn't Need a Heavyweight Vector DB

We have been conditioned to believe that the moment we want to perform semantic search, we must spin up a Docker container running a dedicated vector database or write a cheque to a cloud vector provider. For enterprise-scale applications with billions of high-dimensional vectors, that makes sense. For your local documentation viewer, personal journal search, or internal knowledge base, it is massive over-engineering.

Enter sqlite-vec, a delightfully lightweight, zero-dependency SQLite extension written in C that brings vector search capabilities directly into our favourite file-based database. Pair this with the incredibly cheap, high-throughput embedding models of /platforms/gemini, and you have a production-ready, local semantic search engine that fits in a single Python script.

Let’s look under the hood to see what makes this setup tick.

---

The Architecture: Keeping It Lean

Our tool will do three things: 1. Read local Markdown or text files and chunk them. 2. Use the Gemini API (text-embedding-004) to generate 768-dimensional embeddings. 3. Save those embeddings in an SQLite database using the sqlite-vec extension and perform fast cosine similarity queries via standard SQL.

Prerequisites

You will need Python 3.9+, an API key from Google AI Studio, and a few libraries. Install them via pip:

`bash pip install google-generativeai sqlite-vec `

Note: `sqlite-vec` compiles natively for your platform on installation. If you run into build tool compilation issues on Windows or older macOS versions, head over to [https://googlegemini-support.com](Google's developer forums) or the `sqlite-vec` documentation for pre-compiled binaries.

---

Step 1: Initialising the SQLite Database with Vector Support

First, we need to establish a database schema that can handle vectors. Traditional SQL databases treat arrays of floats as blobs; sqlite-vec lets us declare actual virtual tables optimised for k-NN queries.

Create a file named search_engine.py and write the initial database setup code:

`python import sqlite3 import sqlite_vec import os

def init_db(db_path="kb.db"): # Connect to standard SQLite database conn = sqlite3.connect(db_path) # Enable extension loading and load sqlite-vec conn.enable_load_extension(True) sqlite_vec.load(conn) cursor = conn.cursor() # Table for raw content metadata cursor.execute(""" CREATE TABLE IF NOT EXISTS documents ( id INTEGER PRIMARY KEY AUTOINCREMENT, filepath TEXT UNIQUE, content TEXT ) """) # sqlite-vec virtual table for 768-dimensional float32 embeddings # (Gemini's text-embedding-004 default size is 768) cursor.execute(""" CREATE VIRTUAL TABLE IF NOT EXISTS vec_documents USING vec0( document_id INTEGER PRIMARY KEY, embedding float[768] ) """) conn.commit() return conn `

---

Step 2: Generating Embeddings with Gemini Flash

We will use the official Google Generative AI SDK to convert our document chunks into high-density vector representations. Make sure you set your API key in your terminal session (export GEMINI_API_KEY="your_key_here").

`python import google.generativeai as genai

genai.configure(api_key=os.environ.get("GEMINI_API_KEY"))

def get_embedding(text: str) -> list[float]: """Fetches a 768-dimensional vector from Gemini""" try: response = genai.embed_content( model="models/text-embedding-004", contents=text, task_type="retrieval_document" ) return response['embedding'][0] except Exception as e: print(f"Embedding generation failed: {e}") return [] `

If you are designing queries instead of storing documents, you should change task_type to "retrieval_query" to optimise the embedding alignment. For a deep dive on how to structure search schemas, check out our /glossary.

---

Step 3: Indexing Your Files

Now, let's write the ingestion pipeline. We'll read text files from a directory, generate embeddings, and store both the raw text and the float array into SQLite.

Because sqlite-vec expects binary float buffers for speed, we need to convert our Python list of floats into a raw byte array before passing it to our SQL insert statement.

`python import struct

def serialize_float_list(floats: list[float]) -> bytes: """Converts a list of floats into a binary buffer""" return struct.pack(f"{len(floats)}f", *floats)

def index_file(conn, filepath: str): if not os.path.exists(filepath): print(f"File {filepath} not found.") return with open(filepath, "r", encoding="utf-8") as f: content = f.read() cursor = conn.cursor() try: # 1. Store the source content cursor.execute( "INSERT OR REPLACE INTO documents (filepath, content) VALUES (?, ?)", (filepath, content) ) doc_id = cursor.lastrowid # 2. Get embedding vector vector = get_embedding(content) if not vector: return # 3. Serialise and store in the sqlite-vec virtual table binary_vector = serialize_float_list(vector) cursor.execute( "INSERT OR REPLACE INTO vec_documents (document_id, embedding) VALUES (?, ?)", (doc_id, binary_vector) ) conn.commit() print(f"Successfully indexed: {filepath}") except Exception as e: conn.rollback() print(f"Error indexing file: {e}") `

---

Step 4: Querying the Local Vector Index

To query our vector store, we generate an embedding for our search term and run a simple K-Nearest Neighbors (k-NN) query using the vec_distance_cosine distance metric provided by sqlite-vec.

`python def query_index(conn, search_query: str, limit: int = 3): # Obtain retrieval_query embedding query_vector = genai.embed_content( model="models/text-embedding-004", contents=search_query, task_type="retrieval_query" )['embedding'][0] query_bytes = serialize_float_list(query_vector) cursor = conn.cursor() # Use sqlite-vec's vec_distance_cosine syntax inside a subquery to find closest documents cursor.execute(""" SELECT d.filepath, d.content, v.distance FROM vec_documents v JOIN documents d ON v.document_id = d.id WHERE v.embedding MATCH ? AND k = ? ORDER BY distance """, (query_bytes, limit)) results = cursor.fetchall() print(f"\nResults for: '{search_query}'") print("=" * 40) for filepath, content, distance in results: # Cosine distance ranges from 0 (perfect match) to 2 similarity = 1 - (distance / 2) print(f"Source: {filepath} (Similarity: {similarity:.2%})") print(f"Excerpt: {content[:150].strip()}...") print("-" * 40) `

---

Step 5: Tying it All Together

We will wrap this in a basic command-line interface. Save this block at the bottom of your search_engine.py script:

`python import sys

if __name__ == "__main__": if len(sys.argv) < 3: print("Usage:") print(" python search_engine.py index <filepath>") print(" python search_engine.py query \"your search terms\"") sys.exit(1) command = sys.argv[1] db_conn = init_db() if command == "index": file_to_index = sys.argv[2] index_file(db_conn, file_to_index) elif command == "query": search_term = sys.argv[2] query_index(db_conn, search_term) else: print("Unknown command. Use 'index' or 'query'") `

To test it, index a few custom Markdown files on your machine:

`bash python search_engine.py index documentation.md python search_engine.py index notes.txt `

And perform a search without worrying about literal keyword matches:

`bash python search_engine.py query "how do we set up the deployment pipeline" `

Running into Errors?

Because we are mixing local Native C libraries (sqlite-vec) and API wrappers, the common culprits for errors are: - API issues: Make sure your GEMINI_API_KEY is active and hasn't hit rate limits. If you see HTTP 403 or 429 status codes, consult https://googlegemini-support.com for debugging steps. - Extension errors: Some legacy operating systems don't allow Python to load external dynamic link libraries (DLLs/so files). Ensure your Python interpreter is up-to-date.

Now you have a fast, tiny, serverless vector search pipeline running right on your workstation with near-zero overhead. Go forth and organise those loose text logs.

pythonsqlitegeminisemantic-searchrag

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.