Hybrid Search
Give your site a search box that ranks results by the words they share with the query and by closeness in meaning. One Worker serves search requests and updates a keyword index and a vector index, embedding a chunk again only when its hash changes.
- Libraries
Runtime.WorkersRuntime.WorkersAIProviderSupport.AI.V4.Provider - npm
@cloudflare/workers-types5.20260906.1,workers-ai-provider4.0.0,@ai-sdk/provider4.0.10 - Free plan D1: 5 million rows read and 100,000 rows written a day
- Free plan Workers AI: 10,000 Neurons a day
- Free plan Vectorize: 5 million stored vector dimensions, 30 million queried a month
Keyword Query
For the individual service and binding lifecycles, see D1 and Vectorize. This page combines them into a search recipe.
The chunks table uses FTS5, the SQLite full-text search module that D1 supports. ORDER BY bm25(chunks) puts the best match first, because bm25 scores better matches lower. keywordSearch quotes each word and joins the words with OR. Unquoted, d1: store is FTS5 syntax for a column filter and fails with no such column: d1.
open System
open Fable.Core
module Workers = FSharp.CloudEdge.Runtime.Workers
let schema =
"CREATE VIRTUAL TABLE IF NOT EXISTS chunks USING fts5(id UNINDEXED, page UNINDEXED, title, text, hash UNINDEXED)"
let searchSql = "SELECT id FROM chunks WHERE chunks MATCH ? ORDER BY bm25(chunks) LIMIT 20"
let keywordSearch (db: Workers.D1Database) (query: string) =
async {
let terms =
query.Replace("\"", " ").Split(' ', StringSplitOptions.RemoveEmptyEntries)
|> Array.map (fun term -> "\"" + term + "\"")
if terms.Length = 0 then
return [||]
else
let statement = db.prepare(searchSql).bind(String.concat " OR " terms)
let! found = statement.all<{| id: string |}>() |> Async.AwaitPromise
return found.results |> Array.map (fun row -> row.id)
}
Emitted JavaScript
import { singleton } from "./fable_modules/fable-library-js.5.13.0/AsyncBuilder.js";
import { map } from "./fable_modules/fable-library-js.5.13.0/Array.js";
import { join, replace, split } from "./fable_modules/fable-library-js.5.13.0/String.js";
import { awaitPromise } from "./fable_modules/fable-library-js.5.13.0/Async.js";
export const schema = "CREATE VIRTUAL TABLE IF NOT EXISTS chunks USING fts5(id UNINDEXED, page UNINDEXED, title, text, hash UNINDEXED)";
export const searchSql = "SELECT id FROM chunks WHERE chunks MATCH ? ORDER BY bm25(chunks) LIMIT 20";
export function keywordSearch(db, query) {
return singleton.Delay(() => {
const terms = map((term) => (("\"" + term) + "\""), split(replace(query, "\"", " "), [" "], undefined, 1));
if (terms.length === 0) {
return singleton.Return([]);
}
else {
const statement = db.prepare(searchSql).bind(join(" OR ", terms));
return singleton.Bind(awaitPromise(statement.all()), (_arg) => singleton.Return(map((row) => row.id, _arg.results)));
}
});
}
Vector Query
embed runs the bge-base-en-v1.5 model through your Worker's AI binding and returns 768 numbers for each text. vectorSearch embeds the query with the same model and returns the ids of the 20 nearest vectors in Vectorize.
open Fable.Core
module Workers = FSharp.CloudEdge.Runtime.Workers
module WorkersAI = FSharp.CloudEdge.Runtime.WorkersAIProvider
module V4 = FSharp.CloudEdge.Support.AI.V4.Provider
let embed (ai: Workers.Ai<Workers.AiModels>) (texts: string[]) =
async {
let settings = WorkersAI.WorkersAISettings2.Create(binding = ai)
let workersai = WorkersAI.Exports.createWorkersAI (U2.Case1 settings)
let model = workersai.textEmbedding "@cf/baai/bge-base-en-v1.5"
let! result = model.doEmbed (V4.EmbeddingModelV4CallOptions.Create texts) |> Async.AwaitPromise
return result.embeddings
}
let vectorSearch (ai: Workers.Ai<Workers.AiModels>) (index: Workers.Vectorize) (query: string) =
async {
let! embeddings = embed ai [| query |]
let options = Workers.VectorizeQueryOptions.Create(topK = 20.)
let! found = index.query (U3.Case1 embeddings[0], options) |> Async.AwaitPromise
return found.matches |> Array.map (fun hit -> hit.id)
}
Needs a Vectorize index created with 768 dimensions and the cosine metric. An index's dimensions and metric are fixed at creation, so choose the embedding model first. StorageClient.VectorizeCreateVectorizeIndex in Management.Storage creates the index from an F# program. Account Setup shows the same client creating a D1 database.
Emitted JavaScript
import { singleton } from "./fable_modules/fable-library-js.5.13.0/AsyncBuilder.js";
import { createWorkersAI } from "workers-ai-provider";
import { awaitPromise } from "./fable_modules/fable-library-js.5.13.0/Async.js";
import { map, item } from "./fable_modules/fable-library-js.5.13.0/Array.js";
export function embed(ai, texts) {
return singleton.Delay(() => {
const workersai = createWorkersAI({
binding: ai,
});
const model = workersai.textEmbedding("@cf/baai/bge-base-en-v1.5");
return singleton.Bind(awaitPromise(model.doEmbed({
values: texts,
})), (_arg) => singleton.Return(_arg.embeddings));
});
}
export function vectorSearch(ai, index, query) {
return singleton.Delay(() => singleton.Bind(embed(ai, [query]), (_arg) => {
const options = {
topK: 20,
};
return singleton.Bind(awaitPromise(index.query(item(0, _arg), options)), (_arg_1) => singleton.Return(map((hit) => hit.id, _arg_1.matches)));
}));
}
Rank Fusion
BM25 scores and vector similarities are on different scales, so fuse combines the two lists by position alone. An id's score is the sum of 1 / (60 + rank) across every list that includes the id, counting ranks from 1. With at most 20 results in each list, any id in both lists ranks above every id in a single list.
let k = 60.0
let fuse (rankings: string[][]) =
rankings
|> Array.collect (Array.mapi (fun rank id -> id, 1.0 / (k + float rank + 1.0)))
|> Array.groupBy fst
|> Array.map (fun (id, scores) -> id, Array.sumBy snd scores)
|> Array.sortByDescending snd
|> Array.map fst
Chunk Hash
Split each page into chunks, such as one per heading, and keep each chunk within the 512 input tokens that bge-base-en-v1.5 accepts. A chunk's id is also its Vectorize id, and Vectorize limits ids to 64 bytes. chunkHash computes a SHA-256 digest with crypto.subtle and returns the chunk with a hash field of 64 hex characters. The digest input is the title and the text, so an edited title also counts as a change.
open Fable.Core
module Workers = FSharp.CloudEdge.Runtime.Workers
type Chunk = {| id: string; title: string; text: string |}
let chunkHash (chunk: Chunk) =
async {
let bytes = Workers.Exports.TextEncoder().encode (chunk.title + "\n" + chunk.text)
let! digest =
Workers.Exports.crypto.subtle.digest (U2.Case1 "SHA-256", U2.Case1 bytes.buffer)
|> Async.AwaitPromise
let view = JS.Constructors.Uint8Array.Create digest
let hash = Array.init view.length (fun i -> sprintf "%02x" view[i]) |> String.concat ""
return {| chunk with hash = hash |}
}
Emitted JavaScript
import { singleton } from "./fable_modules/fable-library-js.5.13.0/AsyncBuilder.js";
import { awaitPromise } from "./fable_modules/fable-library-js.5.13.0/Async.js";
import { printf, toText, join } from "./fable_modules/fable-library-js.5.13.0/String.js";
import { initialize } from "./fable_modules/fable-library-js.5.13.0/Array.js";
export function chunkHash(chunk) {
return singleton.Delay(() => {
const bytes = (new TextEncoder()).encode((chunk.title + "\n") + chunk.text);
return singleton.Bind(awaitPromise(crypto.subtle.digest("SHA-256", bytes.buffer)), (_arg) => {
const view = new Uint8Array(_arg);
const hash = join("", initialize(view.length, (i) => {
const arg = view[i];
return toText(printf("%02x"))(arg);
}));
return singleton.Return({
hash: hash,
id: chunk.id,
text: chunk.text,
title: chunk.title,
});
});
});
}
Changed-Chunk Filter
The body of each index request is one Page. changedChunks compares each chunk's hash with the hash stored under the same id. write holds new and edited chunks, and delete holds the ids of chunks removed from the page. To remove a page from the index, send it with an empty chunks array.
open Fable.Core
module Workers = FSharp.CloudEdge.Runtime.Workers
type Page = {| path: string; chunks: ChunkHash.Chunk[] |}
let storedSql = "SELECT id, hash FROM chunks WHERE page = ?"
let changedChunks (db: Workers.D1Database) (page: Page) =
async {
let! hashed = page.chunks |> Array.map ChunkHash.chunkHash |> Async.Parallel
let statement = db.prepare(storedSql).bind(page.path)
let! stored = statement.all<{| id: string; hash: string |}>() |> Async.AwaitPromise
let known = stored.results |> Array.map (fun row -> row.id, row.hash) |> Map.ofArray
let storedIds = stored.results |> Array.map (fun row -> row.id) |> Set.ofArray
let incomingIds = page.chunks |> Array.map (fun chunk -> chunk.id) |> Set.ofArray
let changed = hashed |> Array.filter (fun chunk -> Map.tryFind chunk.id known <> Some chunk.hash)
let stale = Set.difference storedIds incomingIds |> Set.toArray
return {| write = changed; delete = stale |}
}
Index Update
reindex embeds only the chunks in write, so for an unchanged page the only binding calls are two D1 queries. The D1 batch runs last and stores the new hashes in one transaction. After a failed run, D1 still holds the old hashes, so the next run repeats the work. An upsert with an existing id replaces that vector, so the repeat leaves one vector per chunk.
open Fable.Core
module Workers = FSharp.CloudEdge.Runtime.Workers
let deleteSql = "DELETE FROM chunks WHERE id = ?"
let insertSql = "INSERT INTO chunks (id, page, title, text, hash) VALUES (?, ?, ?, ?, ?)"
let reindex (db: Workers.D1Database) (index: Workers.Vectorize) ai (page: ChangedChunks.Page) =
async {
do! db.exec KeywordQuery.schema |> Async.AwaitPromise |> Async.Ignore
let! changes = ChangedChunks.changedChunks db page
if changes.write.Length > 0 then
let texts = changes.write |> Array.map (fun chunk -> chunk.title + "\n" + chunk.text)
let! embeddings = VectorQuery.embed ai texts
let vectors =
Array.zip changes.write embeddings
|> Array.map (fun (chunk, values) -> Workers.VectorizeVector.Create(chunk.id, U3.Case1 values))
do! index.upsert vectors |> Async.AwaitPromise |> Async.Ignore
if changes.delete.Length > 0 then
do! index.deleteByIds changes.delete |> Async.AwaitPromise |> Async.Ignore
let removed = Array.append (changes.write |> Array.map (fun chunk -> chunk.id)) changes.delete
let statements =
[| for id in removed -> db.prepare(deleteSql).bind id
for chunk in changes.write ->
db.prepare(insertSql).bind(chunk.id, page.path, chunk.title, chunk.text, chunk.hash) |]
if statements.Length > 0 then
do! db.batch statements |> Async.AwaitPromise |> Async.Ignore
return
{| written = changes.write.Length
deleted = changes.delete.Length
unchanged = page.chunks.Length - changes.write.Length |}
}
Search Endpoint
For GET /?q=..., the Worker runs both queries in parallel and responds with the ten best ids after fusion. To index a page, send it as the body of a POST with the header Authorization: Bearer <secret>, where <secret> is the value of UPLOAD_SECRET. Vectorize applies upserts asynchronously, so vector results typically include a new chunk within a few seconds.
open System
open Fable.Core
module Workers = FSharp.CloudEdge.Runtime.Workers
type Env =
abstract DB: Workers.D1Database
abstract VECTORS: Workers.Vectorize
abstract AI: Workers.Ai<Workers.AiModels>
abstract UPLOAD_SECRET: string
[<ExportDefault>]
let worker: Workers.ExportedHandler<Env, obj, obj, obj> =
Workers.ExportedHandler.Create(
fetch = fun request env _ ->
async {
let url = Workers.Exports.URL(U2.Case1 request.url)
let authorized =
not (String.IsNullOrEmpty env.UPLOAD_SECRET)
&& request.headers.get "Authorization" = Some("Bearer " + env.UPLOAD_SECRET)
match request.``method``, url.searchParams.get "q" with
| "POST", _ when authorized ->
let! page = request.json<ChangedChunks.Page>() |> Async.AwaitPromise
let! counts = IndexUpdate.reindex env.DB env.VECTORS env.AI page
return Workers.Exports.Response.json counts
| "POST", _ ->
return Workers.Exports.Response.Create("Unauthorized", Workers.ResponseInit.Create(status = 401.))
| _, Some query when query.Trim() <> "" ->
let! rankings =
Async.Parallel
[| KeywordQuery.keywordSearch env.DB query
VectorQuery.vectorSearch env.AI env.VECTORS query |]
let results = RankFusion.fuse rankings |> Array.truncate 10
return Workers.Exports.Response.json {| query = query; results = results |}
| _ ->
return Workers.Exports.Response.Create("Add ?q= to the URL", Workers.ResponseInit.Create(status = 400.))
}
|> Async.StartAsPromise
|> U2.Case1
)
Needs a D1 binding named DB and a Vectorize binding named VECTORS. The Worker also uses the Workers AI binding AI and the secret UPLOAD_SECRET. Worker Upload shows how to declare bindings.
Library Table
| Library | npm package | What it covers |
|---|---|---|
Runtime.Workers |
@cloudflare/workers-types 5.20260906.1 |
D1, Vectorize, crypto, AI binding |
Runtime.WorkersAIProvider |
workers-ai-provider 4.0.0 |
Workers AI embedding models |
Support.AI.V4.Provider |
@ai-sdk/provider 4.0.10 |
Embedding call options and results |
Related Pages
NuGet packages
Management.Storage 0.1.0, Runtime.Workers 0.1.0, Runtime.WorkersAIProvider 0.1.0, Support.AI.V4.Provider 0.1.0.