Tutorials & Guides
How to Build a Serverless Vector Search Engine Using Cloudflare Workers and Gemini 1.5 Flash
Stop paying extortionate fees for hosted vector databases. Here is how to build an ultra-fast, serverless semantic search pipeline using Cloudflare Workers, Vectorize, and Gemini 1.5 Flash.
Updated 10/5/2026
The Problem: Hosted Vector Databases Are Overkill for Your Side Project
We have all been there. You have a brilliant idea for a lightweight AI tool—maybe a semantic bookmark manager, a personal documentation searcher, or a niche directory. You open up the pricing pages for the popular dedicated vector databases, and your heart sinks. Before you have even written a single line of code, you are staring at a commitment of thirty to eighty dollars a month just to keep an idle index spinning in the cloud.
It is time to stop over-provisioning. If you are building a lightweight AI application, you do not need a massive, dedicated cluster. You need a serverless pipeline that scales to zero, costs literal pennies, and runs as close to your users as possible.
In this guide, we will build a fully serverless semantic search engine. We will use Cloudflare Workers as our compute runtime, Cloudflare's native Vectorize engine as our database, and Gemini 1.5 Flash via the developer API to generate our embeddings.
Why This Architecture Works
This stack is incredibly lean. Cloudflare Workers execute at the edge, offering near-zero cold starts. Gemini 1.5 Flash is highly cost-effective, remarkably fast, and provides top-tier performance for processing text. When paired with Vectorize—Cloudflare's vector database—you get a system that can search through thousands of documents in milliseconds for less than the cost of a cup of coffee per month.
Before we dive into the code, if you are unfamiliar with how high-dimensional vectors work, take a quick detour to our /glossary to brush up on vector embeddings.
Step 1: Setting Up Your Cloudflare Environment
First, make sure you have the Wrangler CLI installed and authenticated. Create a new directory for your project and initialise a Worker:
`bash
mkdir serverless-search
cd serverless-search
npm init -y
npm install wrangler --save-dev
`
Now, let's create a vector index. Run the following command to provision a Vectorize index configured for the Gemini embedding model dimensions. Gemini’s text-embedding-004 model outputs vectors with 768 dimensions using cosine similarity.
`bash
npx wrangler vectorize create search-index --dimensions=768 --metric=cosine
`
Keep note of the output. Next, create a wrangler.toml file in your root directory and bind your new Vectorize index to your Worker:
`toml
name = "serverless-search-api"
main = "src/index.ts"
compatibility_date = "2024-03-01"
[[vectorize]] binding = "VECTOR_INDEX" index_name = "search-index"
[vars]
GEMINI_API_KEY = "your_gemini_api_key_here"
`
Note: In production, do not hardcode your api key in the `vars` block. Use `wrangler secret put GEMINI_API_KEY` to store it securely.
Step 2: Fetching Embeddings from Gemini 1.5 Flash
To perform semantic search, we must convert both our database documents and our user's search queries into vector embeddings. Let's write a utility function to query the Gemini API for these embeddings.
Create a file at src/gemini.ts:
`typescript
export async function getEmbedding(text: string, apiKey: string): Promise<number[]> {
const url = https://generativelanguage.googleapis.com/v1beta/models/text-embedding-004:embedContent?key=${apiKey};
const response = await fetch(url, { method: "POST", headers: { "Content-Type": "application/json", }, body: JSON.stringify({ model: "models/text-embedding-004", content: { parts: [{ text: text }], }, }), });
if (!response.ok) {
const errorText = await response.text();
throw new Error(Gemini API error: ${response.status} - ${errorText});
}
const data = (await response.json()) as {
embedding: { values: number[] };
};
return data.embedding.values;
}
`
If you run into issues authenticating or need to check regional availability for Google's API, take a look at our troubleshooting guide over at /platforms/gemini/articles.
Step 3: Writing the Worker Logic
Now we need to build the router inside src/index.ts. We will handle two distinct operations:
1. /insert: Accepts a text payload, generates the embedding, and upserts it into our Vectorize index.
2. /search: Accepts a search query, generates the query embedding, and queries Vectorize for the top matching documents.
Here is the implementation:
`typescript
import { getEmbedding } from "./gemini";
interface Env { VECTOR_INDEX: VectorizeIndex; GEMINI_API_KEY: string; }
export default { async fetch(request: Request, env: Env): Promise<Response> { const url = new URL(request.url); const apiKey = env.GEMINI_API_KEY;
if (!apiKey) { return new Response("Missing Gemini API Key configuration", { status: 500 }); }
// Handle CORS preflight if (request.method === "OPTIONS") { return new Response(null, { headers: { "Access-Control-Allow-Origin": "*", "Access-Control-Allow-Methods": "POST, OPTIONS", "Access-Control-Allow-Headers": "Content-Type", }, }); }
try { // Endpoint 1: Insert metadata and vector if (url.pathname === "/insert" && request.method === "POST") { const { id, text, metadata } = await request.json() as { id: string; text: string; metadata?: Record<string, string>; };
if (!id || !text) { return new Response("Missing id or text parameter", { status: 400 }); }
const embedding = await getEmbedding(text, apiKey);
await env.VECTOR_INDEX.upsert([ { id: id, values: embedding, namespace: "documents", metadata: { ...metadata, text: text }, }, ]);
return new Response(JSON.stringify({ success: true, insertedId: id }), { headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" }, }); }
// Endpoint 2: Semantic Query if (url.pathname === "/search" && request.method === "POST") { const { query, topK = 3 } = await request.json() as { query: string; topK?: number; };
if (!query) { return new Response("Missing query parameter", { status: 400 }); }
const queryEmbedding = await getEmbedding(query, apiKey);
const matches = await env.VECTOR_INDEX.query(queryEmbedding, { topK: topK, namespace: "documents", returnValues: false, returnMetadata: true, });
return new Response(JSON.stringify({ success: true, results: matches.matches }), { headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" }, }); }
return new Response("Not Found", { status: 404 });
} catch (err: any) {
return new Response(JSON.stringify({ error: err.message }), {
status: 500,
headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" },
});
}
},
};
`
Step 4: Testing Your New Engine Local and Live
You can test this locally before deploying. Run wrangler's dev server to spin up a local instance:
`bash
npx wrangler dev
`
Once it is running, let’s insert a few test facts about technology trends using curl:
`bash
curl -X POST http://localhost:8787/insert \
-H "Content-Type: application/json" \
-d '{"id": "doc_1", "text": "Cloudflare workers allow you to run JavaScript at the network edge with minimal cold starts."}'
curl -X POST http://localhost:8787/insert \
-H "Content-Type: application/json" \
-d '{"id": "doc_2", "text": "The UK is famous for its wet weather, tea drinking habits, and historic pubs."}'
`
Now, let's run a search query that does not share any direct keywords with our documents but has a clear semantic link:
`bash
curl -X POST http://localhost:8787/search \
-H "Content-Type: application/json" \
-d '{"query": "Where can I grab a pint in London?"}'
`
Your output will return doc_2 with a high confidence score because Gemini's embedding model successfully associated "pint" and "London" with "historic pubs" and the "UK".
When you are ready to ship to production, deploying is a single command:
`bash
npx wrangler deploy
`
Optimising Your Setup
You have just bypassed the need for heavy, high-maintenance databases. This architecture works beautifully up to tens of thousands of records, keeping your operational costs practically nonexistent. Go ahead and start building without worrying about an eye-watering cloud bill ticking up in the background. If you need inspiration for client-side templates or more complex integrations, head over to the Cloudflare developer gallery to explore live production implementations.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.