Tutorials & Guides
How to Build a Lightweight Semantic Search CLI in Python with sqlite-vec and Gemini Flash
Ditch the heavyweight vector databases. Learn how to build a lightning-fast local semantic search tool using SQLite's newest vector extension and Gemini Flash embeddings.
Updated 9/6/2026
Why Your Side Project Doesn't Need a Heavyweight Vector DB
We have been conditioned to believe that the moment we want to perform semantic search, we must spin up a Docker container running a dedicated vector database or write a cheque to a cloud vector provider. For enterprise-scale applications with billions of high-dimensional vectors, that makes sense. For your local documentation viewer, personal journal search, or internal knowledge base, it is massive over-engineering.
Enter sqlite-vec, a delightfully lightweight, zero-dependency SQLite extension written in C that brings vector search capabilities directly into our favourite file-based database. Pair this with the incredibly cheap, high-throughput embedding models of /platforms/gemini, and you have a production-ready, local semantic search engine that fits in a single Python script.
Let’s look under the hood to see what makes this setup tick.
---
The Architecture: Keeping It Lean
Our tool will do three things:
1. Read local Markdown or text files and chunk them.
2. Use the Gemini API (text-embedding-004) to generate 768-dimensional embeddings.
3. Save those embeddings in an SQLite database using the sqlite-vec extension and perform fast cosine similarity queries via standard SQL.
Prerequisites
You will need Python 3.9+, an API key from Google AI Studio, and a few libraries. Install them via pip:
`bash
pip install google-generativeai sqlite-vec
`
Note: `sqlite-vec` compiles natively for your platform on installation. If you run into build tool compilation issues on Windows or older macOS versions, head over to [https://googlegemini-support.com](Google's developer forums) or the `sqlite-vec` documentation for pre-compiled binaries.
---
Step 1: Initialising the SQLite Database with Vector Support
First, we need to establish a database schema that can handle vectors. Traditional SQL databases treat arrays of floats as blobs; sqlite-vec lets us declare actual virtual tables optimised for k-NN queries.
Create a file named search_engine.py and write the initial database setup code:
`python
import sqlite3
import sqlite_vec
import os
def init_db(db_path="kb.db"):
# Connect to standard SQLite database
conn = sqlite3.connect(db_path)
# Enable extension loading and load sqlite-vec
conn.enable_load_extension(True)
sqlite_vec.load(conn)
cursor = conn.cursor()
# Table for raw content metadata
cursor.execute("""
CREATE TABLE IF NOT EXISTS documents (
id INTEGER PRIMARY KEY AUTOINCREMENT,
filepath TEXT UNIQUE,
content TEXT
)
""")
# sqlite-vec virtual table for 768-dimensional float32 embeddings
# (Gemini's text-embedding-004 default size is 768)
cursor.execute("""
CREATE VIRTUAL TABLE IF NOT EXISTS vec_documents USING vec0(
document_id INTEGER PRIMARY KEY,
embedding float[768]
)
""")
conn.commit()
return conn
`
---
Step 2: Generating Embeddings with Gemini Flash
We will use the official Google Generative AI SDK to convert our document chunks into high-density vector representations. Make sure you set your API key in your terminal session (export GEMINI_API_KEY="your_key_here").
`python
import google.generativeai as genai
genai.configure(api_key=os.environ.get("GEMINI_API_KEY"))
def get_embedding(text: str) -> list[float]:
"""Fetches a 768-dimensional vector from Gemini"""
try:
response = genai.embed_content(
model="models/text-embedding-004",
contents=text,
task_type="retrieval_document"
)
return response['embedding'][0]
except Exception as e:
print(f"Embedding generation failed: {e}")
return []
`
If you are designing queries instead of storing documents, you should change task_type to "retrieval_query" to optimise the embedding alignment. For a deep dive on how to structure search schemas, check out our /glossary.
---
Step 3: Indexing Your Files
Now, let's write the ingestion pipeline. We'll read text files from a directory, generate embeddings, and store both the raw text and the float array into SQLite.
Because sqlite-vec expects binary float buffers for speed, we need to convert our Python list of floats into a raw byte array before passing it to our SQL insert statement.
`python
import struct
def serialize_float_list(floats: list[float]) -> bytes: """Converts a list of floats into a binary buffer""" return struct.pack(f"{len(floats)}f", *floats)
def index_file(conn, filepath: str):
if not os.path.exists(filepath):
print(f"File {filepath} not found.")
return
with open(filepath, "r", encoding="utf-8") as f:
content = f.read()
cursor = conn.cursor()
try:
# 1. Store the source content
cursor.execute(
"INSERT OR REPLACE INTO documents (filepath, content) VALUES (?, ?)",
(filepath, content)
)
doc_id = cursor.lastrowid
# 2. Get embedding vector
vector = get_embedding(content)
if not vector:
return
# 3. Serialise and store in the sqlite-vec virtual table
binary_vector = serialize_float_list(vector)
cursor.execute(
"INSERT OR REPLACE INTO vec_documents (document_id, embedding) VALUES (?, ?)",
(doc_id, binary_vector)
)
conn.commit()
print(f"Successfully indexed: {filepath}")
except Exception as e:
conn.rollback()
print(f"Error indexing file: {e}")
`
---
Step 4: Querying the Local Vector Index
To query our vector store, we generate an embedding for our search term and run a simple K-Nearest Neighbors (k-NN) query using the vec_distance_cosine distance metric provided by sqlite-vec.
`python
def query_index(conn, search_query: str, limit: int = 3):
# Obtain retrieval_query embedding
query_vector = genai.embed_content(
model="models/text-embedding-004",
contents=search_query,
task_type="retrieval_query"
)['embedding'][0]
query_bytes = serialize_float_list(query_vector)
cursor = conn.cursor()
# Use sqlite-vec's vec_distance_cosine syntax inside a subquery to find closest documents
cursor.execute("""
SELECT
d.filepath,
d.content,
v.distance
FROM vec_documents v
JOIN documents d ON v.document_id = d.id
WHERE v.embedding MATCH ? AND k = ?
ORDER BY distance
""", (query_bytes, limit))
results = cursor.fetchall()
print(f"\nResults for: '{search_query}'")
print("=" * 40)
for filepath, content, distance in results:
# Cosine distance ranges from 0 (perfect match) to 2
similarity = 1 - (distance / 2)
print(f"Source: {filepath} (Similarity: {similarity:.2%})")
print(f"Excerpt: {content[:150].strip()}...")
print("-" * 40)
`
---
Step 5: Tying it All Together
We will wrap this in a basic command-line interface. Save this block at the bottom of your search_engine.py script:
`python
import sys
if __name__ == "__main__":
if len(sys.argv) < 3:
print("Usage:")
print(" python search_engine.py index <filepath>")
print(" python search_engine.py query \"your search terms\"")
sys.exit(1)
command = sys.argv[1]
db_conn = init_db()
if command == "index":
file_to_index = sys.argv[2]
index_file(db_conn, file_to_index)
elif command == "query":
search_term = sys.argv[2]
query_index(db_conn, search_term)
else:
print("Unknown command. Use 'index' or 'query'")
`
To test it, index a few custom Markdown files on your machine:
`bash
python search_engine.py index documentation.md
python search_engine.py index notes.txt
`
And perform a search without worrying about literal keyword matches:
`bash
python search_engine.py query "how do we set up the deployment pipeline"
`
Running into Errors?
Because we are mixing local Native C libraries (sqlite-vec) and API wrappers, the common culprits for errors are:
- API issues: Make sure your GEMINI_API_KEY is active and hasn't hit rate limits. If you see HTTP 403 or 429 status codes, consult https://googlegemini-support.com for debugging steps.
- Extension errors: Some legacy operating systems don't allow Python to load external dynamic link libraries (DLLs/so files). Ensure your Python interpreter is up-to-date.
Now you have a fast, tiny, serverless vector search pipeline running right on your workstation with near-zero overhead. Go forth and organise those loose text logs.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.