TRENDING
Close-up of the Rosetta Stone showing the Demotic script above and the Greek script below, the same text written in two different scripts
October 6, 2026
How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
A row of green and grey fibre broadband street cabinets on a pavement beside a fence in Iver, England
October 6, 2026
BT’s TalkTalk Rescue Turns Telecom Continuity Into a New Merger-Control Ground
An ornate cast-iron wall mailbox with its door hanging open, stuffed with colorful flyers and a yellow flyer bulging out of the top slot
October 6, 2026
Google Stops Accepting Product Bug Reports for Its Open-Source Bounty, Citing Automated Submissions
Chronophotograph by Étienne-Jules Marey of a man riding a bicycle, showing five snapshots of the same ride taken at regular intervals
October 6, 2026
How to Find Slow Python Code With the Python 3.15 Tachyon Sampling Profiler
Close-up of an airport baggage tag reading Stockholm Arlanda and ARN
October 6, 2026
Cloudflare Traces Turns Distributed Tracing Into a Trust Decision at the Edge
06 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Shelves of old books fastened by iron chains in the Francis Trigge Chained Library in Grantham, England, a picture of data that can be read but not changed
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
October 5, 2026
Row of capsule hotel pods with white pillows and folded blankets, each capsule an idle sleeper packed into a shared rack
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
October 5, 2026
Denmark’s oldest church book, from Holmens parish, open on a stack of books; its handwritten pages record births between 1617 and 1639
Denmark Says 8.8 Million Population Register Records Were Pulled Through One Company’s Lawful Access
October 5, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 226 Posts
News 228 Posts
Learning Hub 198 Posts
Home/Learning Hub/How to Build Hybrid Search in Python With BM25, Local Embeddings, and Reciprocal Rank Fusion
Learning Hub

How to Build Hybrid Search in Python With BM25, Local Embeddings, and Reciprocal Rank Fusion

Build hybrid search in Python from scratch: BM25 verified against SQLite FTS5, local Ollama embeddings, and rank fusion, then measure where each method wins and where fusion fails.

October 1, 2026 40 Min Read
51

Search that matches keywords and search that matches meaning fail in opposite ways. A keyword engine finds the page that contains HG-4021, but it has nothing to offer for my calls to the backend hang for ages and then die, because no document uses those words. An embedding model understands that sentence, yet it can lose a bare error code among a dozen look-alike codes. Hybrid search runs both methods and merges the two ranked lists, so each one covers the other’s blind spot.

Table Of Content

  • The ideas in plain language
  • Keyword search and BM25
  • Semantic search and embeddings
  • Rank fusion
  • A golden set, hit@k and MRR
  • Prerequisites
  • Step 1: Create the project and check that Ollama answers
  • Check before moving on
  • Step 2: Write the knowledge base and the golden query set
  • What is in the knowledge base
  • What the queries test
  • An honest caveat before you trust any number
  • Step 3: Build BM25 from scratch
  • How the code works
  • Step 4: Look inside the index
  • Reading the output
  • Common mistakes at this step
  • Step 5: Prove your BM25 is right by checking it against SQLite FTS5
  • Reading the output
  • Common mistakes at this step
  • Step 6: Measure BM25 on the golden set
  • Reading the output
  • Step 7: Add semantic search with a local embedding model
  • What to know about this file
  • Reading the output
  • Step 8: Fuse the two rankings
  • Reciprocal rank fusion, by hand
  • Min-max score fusion
  • Reading the output
  • Step 9: Break it on purpose
  • Pitfall 1: adding raw scores from different scales
  • Pitfall 2: every retriever gets an equal vote
  • Pitfall 3: ranking documents that did not match
  • Pitfall 4: the knobs
  • Step 10: Package it as a searcher and add tests
  • The test suite
  • Step 11: Verify everything end to end
  • Common mistakes and how to spot them
  • Where to go next
  • References

In this tutorial you will build a small hybrid search engine in plain Python and, more importantly, learn how to measure it. You will write:

  • a BM25 keyword index, checked number for number against the BM25 function built into SQLite;
  • a semantic index that uses a local embedding model served by Ollama;
  • two ways of fusing the results, reciprocal rank fusion (RRF) and min-max score fusion;
  • a golden query set with hit@k and MRR, so every claim in this article is a number you can reproduce;
  • a pytest suite that guards the parts most likely to break.

Every output block below was copied from a real run on Windows 11 with Python 3.13.14, SQLite 3.50.4, Ollama 0.32.9 and the nomic-embed-text embedding model. That includes the results that did not go the way a tidy tutorial would like. On the simplest query in the test set, a bare error code, plain reciprocal rank fusion ranks the right answer fifth. That is not a bug in the code. It is the most useful thing the experiment shows, and Step 9 explains why it happens and what you can do about it.

The ideas in plain language

Four ideas carry the whole tutorial. If you already know them, skim this section and move on to the prerequisites.

Keyword search and BM25

Keyword (also called lexical) search finds documents that share words with the query. To rank them, a classic scoring recipe is BM25 (Okapi BM25), which Apache Lucene and SQLite FTS5 both implement, as you will see below. Three vocabulary words first:

  • A token is one normalized piece of text. Here a token is a lowercase run of letters and digits, so HG-4021 becomes the two tokens hg and 4021.
  • Term frequency (tf) is how many times a token appears in one document.
  • Document frequency (df) is how many documents contain the token at least once. A token found in every document tells you nothing; a token found in one document is a strong clue. BM25 turns that idea into inverse document frequency (IDF): rare tokens earn big weights.

BM25 adds two refinements. Repeating a word helps less and less (the term frequency saturates, controlled by a constant called k1), and a long document gets a small penalty so it cannot win just by containing more words (length normalization, controlled by b). Wikipedia calls both constants “free parameters, usually chosen, in absence of an advanced optimization”, and gives the usual values as a k1 between 1.2 and 2.0 and a b of 0.75. This tutorial uses 1.2 and 0.75.

Semantic search and embeddings

An embedding model turns a piece of text into a vector, a list of numbers (768 of them for the model used here). Texts with similar meaning end up with similar vectors, even when they share no words. To search, you embed every document once, embed the query, and rank documents by cosine similarity, a score where 1.0 means the two vectors point the same way. When every vector is scaled to length 1, cosine similarity is just a dot product, which is how the code below computes it.

Rank fusion

Fusion combines the ranked lists from two retrievers into one. The two methods produce numbers that mean different things: a BM25 score has no upper limit, while the cosine scores in this tutorial stay in a narrow band (0.41 to 0.66 for the query examined in Step 9). Reciprocal rank fusion (RRF) sidesteps the mismatch by ignoring scores entirely. A document earns 1 / (k + rank) points from every list it appears in, and the points are added up. The RRF paper (Cormack, Clarke and Büttcher, SIGIR 2009) says “k = 60 was fixed during a pilot investigation and not altered during subsequent validation”. The alternative is score fusion: rescale each retriever’s scores to the range 0 to 1 and take a weighted average.

A golden set, hit@k and MRR

To compare methods you need a golden set: queries paired with the documents a person decided are correct answers. Three numbers summarize how a retriever does:

  • hit@1: the share of queries whose first result is correct.
  • hit@3: the share of queries with a correct result in the top three.
  • MRR@10 (mean reciprocal rank): the average of 1 / rank for the first correct result, counting 0 when it is not in the top ten. Rank 1 scores 1.0, rank 2 scores 0.5, rank 4 scores 0.25.

The same metrics are explained in the retrieval evaluation tutorial on this site.

Prerequisites

  • Python 3.10 or newer (tested with 3.13.14). The code uses only the standard library plus requests, and pytest runs the tests.
  • Windows, macOS or Linux. The commands below are for Windows PowerShell; macOS and Linux differ only in how the virtual environment is activated.
  • Ollama, installed and running (tested with 0.32.9). If you have never used it, start with the Ollama install tutorial.
  • SQLite with the FTS5 extension, which Step 5 uses. If your Python’s SQLite lacks FTS5, that one script stops with an error and everything else still works.
  • Comfort reading Python: functions, lists, dictionaries and classes. No machine learning background is assumed.

Step 1: Create the project and check that Ollama answers

Create a folder and a virtual environment, install the two packages, and download the embedding model:

mkdir hybrid-search
cd hybrid-search
python -m venv .venv
.venv\Scripts\activate
pip install requests pytest
ollama pull nomic-embed-text

On macOS or Linux use source .venv/bin/activate in place of the activation line. The model download is about 274 MB. Ollama’s library page lists that size for both the latest and the v1.5 tag, so Nomic’s v1.5 model card is the reference for how to use the model.

Save the next file as versions.py. It records exactly what you are running and proves that Ollama is reachable:

# versions.py
import platform
import sqlite3

import pytest
import requests

print("Python", platform.python_version(), "|", platform.system(), platform.release())
print("requests", requests.__version__, "| pytest", pytest.__version__, "| SQLite", sqlite3.sqlite_version)
print("Ollama", requests.get("http://127.0.0.1:11434/api/version", timeout=10).json()["version"])
info = requests.post("http://127.0.0.1:11434/api/show", json={"model": "nomic-embed-text"}, timeout=10).json()
details = info.get("details", {})
model_info = info.get("model_info", {})
print("model nomic-embed-text:", details.get("family"), details.get("parameter_size"), details.get("quantization_level"))
for key, value in model_info.items():
    if key.endswith("embedding_length") or key.endswith("context_length"):
        print("  %s = %s" % (key, value))

Run it with python versions.py:

Python 3.13.14 | Windows 11
requests 2.34.2 | pytest 9.1.1 | SQLite 3.50.4
Ollama 0.32.9
model nomic-embed-text: nomic-bert 137M F16
  nomic-bert.context_length = 2048
  nomic-bert.embedding_length = 768

Check before moving on

You should see your own versions on the first three lines. The model line proves Ollama answered, and two numbers matter: an embedding length of 768 (every text becomes 768 numbers) and a context length of 2048 tokens, far more than any document here needs. A connection error means Ollama is not running; start it and try again.

Step 2: Write the knowledge base and the golden query set

Save this as corpus.py. It holds a small, fictional knowledge base for an imaginary API gateway called the Halyard Gateway, plus the queries you will use to test every retriever.

# corpus.py
"""A small, entirely fictional knowledge base for the Halyard Gateway, plus a golden query set."""

DOCS = [
    ("err-4012", "HG-4012 token expired. The bearer token presented to the gateway has passed its exp claim. Clients should request a fresh token and retry once. The gateway tolerates small clock differences only when token_clock_skew_s is set."),
    ("err-4021", "HG-4021 upstream timeout. The upstream service did not answer within upstream_timeout_ms. The gateway cancels the call and returns HTTP 504 to the client. Raise the limit for slow endpoints, or fix the slow upstream."),
    ("err-4029", "HG-4029 quota exceeded. The caller used up its allowance of requests for the current window. The gateway returns HTTP 429 and a Retry-After header. Allowances are defined per API key in rate_limits.yaml."),
    ("err-4033", "HG-4033 denied by policy. A policy rule rejected the request. The decision log names the rule that matched, which tells you which team owns the restriction."),
    ("err-4102", "HG-4102 body too large. The request body is larger than max_body_bytes, which defaults to 1048576. The gateway answers HTTP 413 without contacting any upstream."),
    ("err-4201", "HG-4201 upstream refused connection. The gateway could not open a TCP connection to the upstream. Check that the address resolves and the port is listening. Retries are limited by retry_budget_percent."),
    ("err-4210", "HG-4210 connection pool exhausted. Every connection to the upstream is busy. Raise max_conns_per_upstream or reduce concurrency. Clients see HTTP 503 with the header X-HG-Reason set to pool."),
    ("err-5013", "HG-5013 TLS handshake failure. The gateway could not finish a TLS handshake with the upstream. Typical causes are an expired upstream certificate, or a tls_min_version that is higher than the upstream supports."),
    ("err-5031", "HG-5031 circuit open. The circuit breaker for this upstream tripped after repeated failures. Calls fail immediately with HTTP 503 until a probe succeeds. Tune breaker_failure_threshold and breaker_cooldown_s."),
    ("err-5040", "HG-5040 reload rejected. A configuration reload failed validation, so the gateway kept running on the previous configuration. Run hgctl validate to find the offending line."),
    ("cfg-upstream-timeout", "upstream_timeout_ms is the longest time, in milliseconds, that the gateway waits for an upstream to deliver a complete response. Default 30000. For streaming responses the timer restarts on every chunk."),
    ("cfg-connect-timeout", "upstream_connect_timeout_ms is the longest time, in milliseconds, allowed to establish the TCP and TLS connection to an upstream. Default 3000. It does not limit how long the response itself may take."),
    ("cfg-max-upstream", "max_conns_per_upstream caps how many connections the gateway keeps open to a single upstream at the same time. Default 256."),
    ("cfg-max-client", "max_conns_per_client caps how many simultaneous connections the gateway accepts from one client address. Default 64. Extra connections are closed at once."),
    ("cfg-retry-budget", "retry_budget_percent limits what share of requests may be retried, so that retries cannot multiply the load during an outage. Default 20."),
    ("cfg-tls-min", "tls_min_version sets the oldest TLS protocol the gateway will speak, both to clients and to upstreams. Accepted values are 1.2 and 1.3. Default 1.2."),
    ("cfg-clock-skew", "token_clock_skew_s is the number of seconds of clock drift tolerated when checking the exp and nbf claims of a token. Default 30."),
    ("proc-session", "Browser sessions use a cookie and end after session_idle_timeout_min minutes without activity. The default is 15. Someone who leaves the keyboard for a while is signed out and has to log in again."),
    ("proc-keys", "To rotate signing keys without downtime, publish the new public key in the JWKS document first, wait for downstream caches to expire, switch the signer to the new key, and remove the old public key once the longest token lifetime has passed."),
    ("proc-bluegreen", "To release a new gateway version, start the green fleet beside the blue one, move traffic across in 10 percent steps with weighted routing, and keep blue running until error rates stay flat."),
    ("proc-drain", "Before maintenance on a node, mark it as draining so it stops taking new connections while work in progress finishes. Run hgctl drain and wait until the active connection count reaches zero."),
    ("proc-certs", "Replace the serving certificate with hgctl rotate-certs. The command loads the new key pair, swaps it in atomically, and logs the fingerprint of the certificate it replaced."),
    ("proc-dump", "hgctl dump-config prints the effective configuration after defaults and environment overrides are applied. Secret values are masked."),
    ("proc-logs", "Access logs are written as JSON lines that any file-tailing agent can ship. Each line carries a request id, which the gateway also forwards upstream in the X-Request-Id header."),
    ("proc-allowlist", "To limit an API to office networks, add CIDR ranges to the route's allow list. Requests from any other address receive HTTP 403."),
    ("proc-cors", "Browsers refuse cross-origin calls unless the response carries the right headers. Configure the permitted origins separately for every route."),
    ("proc-health", "Active health checks probe each upstream every few seconds. An upstream that fails its probes leaves the rotation until the probes succeed again."),
    ("proc-upgrade", "Upgrading from 3.x to 4.x requires migrating routes.yaml to the new schema. Run hgctl migrate on a copy first and review the diff."),
    ("adv-2026-007", "Advisory HAL-SA-2026-007: request smuggling through duplicate Content-Length headers affects versions before 4.8.2. Upgrade to 4.8.2, or set reject_duplicate_headers to true as a stopgap."),
    ("adv-2026-011", "Advisory HAL-SA-2026-011: the admin API served the /metrics page without authentication when bound to 0.0.0.0. Fixed in 4.9.0. On older versions, bind the admin listener to loopback."),
    ("adv-2025-019", "Advisory HAL-SA-2025-019: unescaped newlines in the User-Agent header allowed forged lines in access logs. Fixed in 4.5.1."),
    ("rel-482", "Release notes 4.8.2: fixes HAL-SA-2026-007, improves reload validation messages, and raises the default max_conns_per_upstream from 128 to 256."),
]

QUERIES = [
    # exact: the user types an identifier that appears verbatim in one document
    {"id": "q01", "kind": "exact", "text": "HG-4021", "relevant": {"err-4021"}},
    {"id": "q02", "kind": "exact", "text": "HG-5013", "relevant": {"err-5013"}},
    {"id": "q03", "kind": "exact", "text": "hgctl rotate-certs", "relevant": {"proc-certs"}},
    {"id": "q04", "kind": "exact", "text": "HAL-SA-2026-007", "relevant": {"adv-2026-007"}},
    {"id": "q05", "kind": "exact", "text": "max_conns_per_client", "relevant": {"cfg-max-client"}},
    {"id": "q06", "kind": "exact", "text": "upstream_connect_timeout_ms", "relevant": {"cfg-connect-timeout"}},
    {"id": "q07", "kind": "exact", "text": "retry_budget_percent", "relevant": {"cfg-retry-budget"}},
    # paraphrase: the user describes the problem in words the documents never use
    {"id": "q08", "kind": "paraphrase", "text": "why do people get kicked out when they step away from their desk", "relevant": {"proc-session"}},
    {"id": "q09", "kind": "paraphrase", "text": "my calls to the backend hang for ages and then die", "relevant": {"err-4021", "cfg-upstream-timeout"}},
    {"id": "q10", "kind": "paraphrase", "text": "how do I replace the server's TLS identity without restarting it", "relevant": {"proc-certs"}},
    {"id": "q11", "kind": "paraphrase", "text": "I edited the settings file but the old behaviour carried on", "relevant": {"err-5040"}},
    {"id": "q12", "kind": "paraphrase", "text": "how do I stop one customer from hammering the service", "relevant": {"err-4029", "cfg-max-client"}},
    {"id": "q13", "kind": "paraphrase", "text": "releasing a new version to a small slice of users first", "relevant": {"proc-bluegreen"}},
    {"id": "q14", "kind": "paraphrase", "text": "taking a machine out for maintenance without cutting off requests in progress", "relevant": {"proc-drain"}},
    {"id": "q15", "kind": "paraphrase", "text": "only let people on the corporate network reach this API", "relevant": {"proc-allowlist"}},
    {"id": "q16", "kind": "paraphrase", "text": "the service keeps turning away big uploads", "relevant": {"err-4102"}},
    # mixed: an identifier or a distinctive term plus ordinary words
    {"id": "q17", "kind": "mixed", "text": "what does HG-4210 mean and how do I fix it", "relevant": {"err-4210"}},
    {"id": "q18", "kind": "mixed", "text": "raise the timeout behind HG-4021 errors", "relevant": {"err-4021", "cfg-upstream-timeout"}},
    {"id": "q19", "kind": "mixed", "text": "clients on TLS 1.0 broke after the upgrade, which setting is the floor", "relevant": {"cfg-tls-min"}},
    {"id": "q20", "kind": "mixed", "text": "tokens get rejected right after expiry even though clocks differ by seconds", "relevant": {"cfg-clock-skew", "err-4012"}},
    {"id": "q21", "kind": "mixed", "text": "is 4.8.1 exposed to smuggling with doubled Content-Length headers", "relevant": {"adv-2026-007"}},
    {"id": "q22", "kind": "mixed", "text": "the metrics page is readable without a password on every network interface", "relevant": {"adv-2026-011"}},
]

_ids = {doc_id for doc_id, _ in DOCS}
for _q in QUERIES:
    assert _q["relevant"] <= _ids, _q["id"]

What is in the knowledge base

DOCS has 32 short articles in four families: 10 error-code pages (for example HG-4021 upstream timeout), 7 configuration-key references, 11 how-to procedures, and 4 security advisories or release notes. I made the error codes and configuration keys confusable on purpose, because real documentation is full of near-duplicates: HG-4021 and HG-4210 use the same four digits in a different order, and upstream_timeout_ms sits next to upstream_connect_timeout_ms.

What the queries test

There are 22 queries of three kinds, each with the set of documents that count as correct:

  • exact (7): the user types an identifier that appears verbatim in one document, such as HG-5013 or max_conns_per_client.
  • paraphrase (9): the user describes the problem in words the documents never use, such as the service keeps turning away big uploads for the HG-4102 body too large page.
  • mixed (6): an identifier or a distinctive term plus ordinary words, such as what does HG-4210 mean and how do I fix it.

An honest caveat before you trust any number

I wrote both the documents and the queries before running any retriever, and I did not edit them after seeing results. Even so, 22 queries is tiny: moving one query’s rank by a few places shifts the averages by several points. Treat everything below as a demonstration of how to measure, not as a benchmark of any method. The point of a golden set is that you replace mine with real questions from your own users.

Step 3: Build BM25 from scratch

BM25 scores a document against a query by adding up one contribution for every query token that appears in the document:

score(doc, query) = sum, over query tokens t found in doc, of
    idf(t) * tf * (k1 + 1) / (tf + k1 * (1 - b + b * len(doc) / avg_len))

Read it from the outside in. idf(t) makes rare tokens count more. The fraction makes repeated tokens count more, with diminishing returns, and shrinks the score of documents longer than average (avg_len is the average document length in tokens). Save this as bm25.py:

# bm25.py
import math
import re
from collections import Counter

TOKEN_RE = re.compile(r"[a-z0-9]+")


def tokenize(text):
    """Lowercase the text and keep runs of letters and digits."""
    return TOKEN_RE.findall(text.lower())


class BM25Index:
    def __init__(self, docs, k1=1.2, b=0.75, idf_mode="lucene"):
        self.k1 = k1
        self.b = b
        self.idf_mode = idf_mode
        self.doc_ids = []
        self.term_freqs = []       # one Counter per document: term -> count
        self.doc_lens = []         # tokens per document
        self.doc_freq = Counter()  # term -> number of documents containing it
        for doc_id, text in docs:
            tokens = tokenize(text)
            self.doc_ids.append(doc_id)
            self.term_freqs.append(Counter(tokens))
            self.doc_lens.append(len(tokens))
            self.doc_freq.update(set(tokens))
        self.n_docs = len(self.doc_ids)
        self.avg_len = sum(self.doc_lens) / self.n_docs

    def idf(self, term):
        n = self.doc_freq.get(term, 0)
        ratio = (self.n_docs - n + 0.5) / (n + 0.5)
        if self.idf_mode == "classic":
            return math.log(ratio)  # negative when a term is in more than half the documents
        if self.idf_mode == "fts5":
            value = math.log(ratio)
            return value if value > 0 else 1e-6  # SQLite FTS5 clamps non-positive values
        return math.log(1.0 + ratio)  # "lucene": never negative

    def score(self, terms, index):
        counts = self.term_freqs[index]
        length_norm = 1 - self.b + self.b * self.doc_lens[index] / self.avg_len
        total = 0.0
        for term in terms:
            f = counts.get(term, 0)
            if f:
                total += self.idf(term) * f * (self.k1 + 1) / (f + self.k1 * length_norm)
        return total

    def search(self, query, top_n=10):
        """Return [(doc_id, score)], best first, for documents that share at least one term."""
        terms = list(dict.fromkeys(tokenize(query)))
        scored = []
        for i, doc_id in enumerate(self.doc_ids):
            if any(t in self.term_freqs[i] for t in terms):
                scored.append((doc_id, self.score(terms, i)))
        scored.sort(key=lambda pair: (-pair[1], pair[0]))
        return scored[:top_n]

How the code works

  • tokenize lowercases the text and keeps runs of letters and digits. Documents and queries go through the same function, which matters more than it sounds (see the mistakes list at the end).
  • The constructor makes one pass over the documents, recording each document’s token counts and length, and how many documents contain each token.
  • idf supports three variants chosen by idf_mode. All three start from the ratio (N - n + 0.5) / (n + 0.5), where N is the number of documents and n the number that contain the token. "classic" takes the logarithm of that ratio. "fts5" does the same but replaces any result that is not positive with 0.000001. "lucene", the default here, takes the logarithm of 1 plus the ratio, so it can never be negative. Apache Lucene’s BM25Similarity computes its IDF that way, which is why I named the variant after it. Step 4 shows why the difference matters.
  • search returns only documents that share at least one token with the query, best first, with ties broken by document id so the output is repeatable.

Step 4: Look inside the index

Before trusting any ranking function, look at the numbers it produces. Save this as s1_inspect.py:

# s1_inspect.py
from bm25 import BM25Index, tokenize
from corpus import DOCS, QUERIES

classic = BM25Index(DOCS, idf_mode="classic")
lucene = BM25Index(DOCS, idf_mode="lucene")

print("documents: %d | average length: %.1f tokens | golden queries: %d" % (lucene.n_docs, lucene.avg_len, len(QUERIES)))
print()
for text in ("HG-4021", "max_conns_per_client", "hgctl rotate-certs", "HAL-SA-2026-007"):
    print("%-22s -> %s" % (text, tokenize(text)))

print()
print("term        in docs   classic idf   lucene idf")
for term in ("the", "gateway", "upstream", "timeout", "4021", "007"):
    print("%-10s %3d/%d %13.3f %12.3f" % (term, lucene.doc_freq[term], lucene.n_docs, classic.idf(term), lucene.idf(term)))

print()
print("Query 'the', best and worst matching document under each variant:")
for name, index in (("classic", classic), ("lucene", lucene)):
    ranked = index.search("the", top_n=len(DOCS))
    print("  %-8s best %-14s %7.3f | worst %-14s %7.3f" % (name, ranked[0][0], ranked[0][1], ranked[-1][0], ranked[-1][1]))

print()
print("What k1 and b do (idf left out): the multiplier a term earns in one document")
k1, b = 1.2, 0.75
for f in (1, 2, 3, 5, 10, 50):
    print("  term appears %2d times, average-length document: %.3f" % (f, f * (k1 + 1) / (f + k1)))
for label, ratio in (("twice", 2.0), ("half", 0.5)):
    norm = 1 - b + b * ratio
    print("  term appears once, document %s the average length: %.3f" % (label, (k1 + 1) / (1 + k1 * norm)))

Run it with python s1_inspect.py:

documents: 32 | average length: 29.5 tokens | golden queries: 22

HG-4021                -> ['hg', '4021']
max_conns_per_client   -> ['max', 'conns', 'per', 'client']
hgctl rotate-certs     -> ['hgctl', 'rotate', 'certs']
HAL-SA-2026-007        -> ['hal', 'sa', '2026', '007']

term        in docs   classic idf   lucene idf
the         31/32        -3.045        0.047
gateway     13/32         0.368        0.894
upstream    12/32         0.495        0.971
timeout      4/32         1.846        1.992
4021         1/32         3.045        3.091
007          2/32         2.501        2.580

Query 'the', best and worst matching document under each variant:
  classic  best err-5031        -2.944 | worst proc-certs      -5.443
  lucene   best proc-certs       0.083 | worst proc-drain       0.045

What k1 and b do (idf left out): the multiplier a term earns in one document
  term appears  1 times, average-length document: 1.000
  term appears  2 times, average-length document: 1.375
  term appears  3 times, average-length document: 1.571
  term appears  5 times, average-length document: 1.774
  term appears 10 times, average-length document: 1.964
  term appears 50 times, average-length document: 2.148
  term appears once, document twice the average length: 0.710
  term appears once, document half the average length: 1.257

Reading the output

Tokens. The tokenizer splits identifiers at hyphens and underscores, so HG-4021 becomes hg and 4021. That is a trade-off: a user who types HG 4021 still matches, but hg is now a token shared by every error page. The part that discriminates is 4021, and BM25 notices because its IDF is high.

IDF. The word the appears in 31 of 32 documents. The classic formula gives it an IDF of -3.045, which is negative: matching the subtracts from a document’s score, and a document is rewarded for using the word less. The last lines of that block show the absurd result for the one-word query the: every score is negative, with err-5031 best at -2.944 and proc-certs worst at -5.443. The plus-one variant keeps every IDF positive. the gets a tiny but sensible 0.047, and 4021, found in a single document, gets 3.091. Wikipedia gives the plus-one form as the usual way to compute IDF, and SQLite’s own source code notes that the classic formula goes negative for terms found in more than half the rows and clamps the value at 0.000001.

The k1 and b table. The final block shows what the two constants do, with IDF set aside. With k1 at 1.2, a token that appears once earns a multiplier of 1.000, twice 1.375, five times 1.774 and 50 times 2.148. The multiplier approaches k1 + 1 = 2.2 but never reaches it: that is saturation. For length, a token that appears once in a document twice the average length earns 0.710, and in a document half the average length 1.257.

Common mistakes at this step

  • Using the classic IDF without noticing that it can go negative, which quietly reverses the ranking for common words.
  • Forgetting that the average length (29.5 tokens here) describes short articles. Length normalization matters more when your documents vary a lot in size.

Step 5: Prove your BM25 is right by checking it against SQLite FTS5

You now have a BM25 that looks plausible, but plausible is not a test. SQLite ships a full-text search module called FTS5 with a built-in bm25() function, so you can compare your numbers with an independent implementation. Two details of the SQLite documentation shape the script:

  • FTS5 returns negated scores. The documentation says “The better the match, the numerically smaller the value returned”, so the script flips the sign back.
  • FTS5’s unicode61 tokenizer splits text much like ours. Its documentation says “By default all space and punctuation characters, as defined by Unicode 6.1, are considered separators, and all other characters as token characters”. For plain ASCII text such as ours, that gives the same tokens.

The SQLite source file also fixes the constants k1 = 1.2 and b = 0.75 and applies the clamp from Step 3, which is why the script builds our index with idf_mode="fts5". Save it as s2_fts5_check.py:

# s2_fts5_check.py
import sqlite3

from bm25 import BM25Index, tokenize
from corpus import DOCS, QUERIES

db = sqlite3.connect(":memory:")
db.execute("CREATE VIRTUAL TABLE kb USING fts5(doc_id UNINDEXED, body, tokenize = 'unicode61')")
db.executemany("INSERT INTO kb (doc_id, body) VALUES (?, ?)", DOCS)
print("SQLite version:", sqlite3.sqlite_version)


def fts5_scores(query):
    terms = list(dict.fromkeys(tokenize(query)))
    match = " OR ".join('"%s"' % t for t in terms)
    rows = db.execute("SELECT doc_id, -bm25(kb) FROM kb WHERE kb MATCH ?", (match,)).fetchall()
    return dict(rows)  # FTS5 returns negated scores so that ORDER BY works; we flip the sign back


mine = BM25Index(DOCS, idf_mode="fts5")

demo = "connection pool upstream timeout"
theirs, ours = fts5_scores(demo), dict(mine.search(demo, top_n=len(DOCS)))
print()
print("query: %r" % demo)
print("%-22s %12s %12s" % ("document", "FTS5 bm25()", "my BM25"))
for doc_id, score in sorted(theirs.items(), key=lambda kv: -kv[1])[:4]:
    print("%-22s %12.6f %12.6f" % (doc_id, score, ours[doc_id]))

largest, pairs = 0.0, 0
for q in QUERIES:
    theirs, ours = fts5_scores(q["text"]), dict(mine.search(q["text"], top_n=len(DOCS)))
    assert set(theirs) == set(ours), q["id"]
    for doc_id in theirs:
        largest = max(largest, abs(theirs[doc_id] - ours[doc_id]))
        pairs += 1
print()
print("fts5-style idf: compared %d (query, document) scores over %d queries" % (pairs, len(QUERIES)))
print("largest absolute difference: %.2e" % largest)

lucene = BM25Index(DOCS)  # the default variant used in the rest of this tutorial
agree = 0
for q in QUERIES:
    best_fts5 = max(fts5_scores(q["text"]).items(), key=lambda kv: (kv[1], kv[0]))[0]
    agree += best_fts5 == lucene.search(q["text"], top_n=1)[0][0]
print("lucene-style idf: same top-1 document as FTS5 on %d of %d queries" % (agree, len(QUERIES)))

Run it with python s2_fts5_check.py:

SQLite version: 3.50.4

query: 'connection pool upstream timeout'
document                FTS5 bm25()      my BM25
err-4210                   7.167649     7.167649
cfg-connect-timeout        4.128885     4.128885
err-4021                   3.171538     3.171538
err-4201                   3.115303     3.115303

fts5-style idf: compared 456 (query, document) scores over 22 queries
largest absolute difference: 3.55e-15
lucene-style idf: same top-1 document as FTS5 on 21 of 22 queries

Reading the output

The first table shows the top four documents for one query, and the two columns agree to six decimal places. Across all 22 golden queries the script compared 456 (query, document) scores. The largest absolute difference is 3.55e-15, which is floating-point rounding noise. Your BM25 arithmetic is the same as SQLite’s.

The last line is the useful warning. With the plus-one IDF that the rest of this tutorial uses, the top result matches FTS5’s on 21 of 22 queries. Both variants are legitimate, and they disagree only where very common words decide a close call. A cross-check proves the arithmetic is right. It does not prove the search is good, which is what the next step measures.

Common mistakes at this step

  • Forgetting that bm25() is negative and sorting FTS5 results in the wrong direction.
  • Comparing against FTS5 while your tokenizer differs. If the scores disagree by more than rounding noise, compare the tokens first.

Step 6: Measure BM25 on the golden set

Now the evaluation helpers. Save this as evalkit.py:

# evalkit.py
KINDS = ("exact", "paraphrase", "mixed")


def first_relevant_rank(ranked_ids, relevant):
    """1-based position of the first relevant document, or None if it is not in the list."""
    for position, doc_id in enumerate(ranked_ids, start=1):
        if doc_id in relevant:
            return position
    return None


def evaluate(search_ids, queries):
    """search_ids(query_text) must return document ids, best first. Returns {query_id: rank}."""
    return {q["id"]: first_relevant_rank(search_ids(q["text"]), q["relevant"]) for q in queries}


def summarize(ranks):
    n = len(ranks)
    return {
        "n": n,
        "hit@1": sum(1 for r in ranks if r == 1) / n,
        "hit@3": sum(1 for r in ranks if r is not None and r <= 3) / n,
        "mrr@10": sum(1.0 / r for r in ranks if r is not None and r <= 10) / n,
    }


def print_summary(results, queries, kinds=KINDS + ("all",)):
    """results maps a method name to the {query_id: rank} dictionary that evaluate() returns."""
    kind_of = {q["id"]: q["kind"] for q in queries}
    print("%-12s %-11s %3s %6s %6s %7s" % ("method", "queries", "n", "hit@1", "hit@3", "mrr@10"))
    for method, ranks in results.items():
        for kind in kinds:
            subset = [r for qid, r in ranks.items() if kind == "all" or kind_of[qid] == kind]
            s = summarize(subset)
            print("%-12s %-11s %3d %6.2f %6.2f %7.3f" % (method, kind, s["n"], s["hit@1"], s["hit@3"], s["mrr@10"]))
  • first_relevant_rank finds the 1-based position of the first correct document, or None if the retriever never returned one.
  • evaluate runs a search function over every query and returns one rank per query.
  • summarize turns ranks into hit@1, hit@3 and MRR@10. A missing result counts as zero in all three.
  • print_summary prints the numbers overall and for each kind of query.

Save s3_eval_bm25.py and run it with python s3_eval_bm25.py:

# s3_eval_bm25.py
from bm25 import BM25Index
from corpus import DOCS, QUERIES
from evalkit import evaluate, print_summary

index = BM25Index(DOCS)


def bm25_ids(query):
    return [doc_id for doc_id, _ in index.search(query, top_n=len(DOCS))]


ranks = evaluate(bm25_ids, QUERIES)
print_summary({"bm25": ranks}, QUERIES)
print()
print("queries where BM25 does not put a relevant document in the top 3:")
for q in QUERIES:
    r = ranks[q["id"]]
    if r is None or r > 3:
        print("  %s %-10s rank %-4s %s" % (q["id"], q["kind"], r or "-", q["text"][:62]))
method       queries       n  hit@1  hit@3  mrr@10
bm25         exact         7   0.86   1.00   0.929
bm25         paraphrase    9   0.33   0.67   0.513
bm25         mixed         6   0.67   0.83   0.792
bm25         all          22   0.59   0.82   0.721

queries where BM25 does not put a relevant document in the top 3:
  q09 paraphrase rank 5    my calls to the backend hang for ages and then die
  q15 paraphrase rank 4    only let people on the corporate network reach this API
  q16 paraphrase rank 18   the service keeps turning away big uploads
  q20 mixed      rank 4    tokens get rejected right after expiry even though clocks diff

Reading the output

BM25 puts a correct document first for 13 of 22 queries (hit@1 = 0.59) and in the top three for 18 (hit@3 = 0.82). The split by kind is the story: MRR is 0.929 on exact queries and only 0.513 on paraphrases. The four misses show why. For the service keeps turning away big uploads, the correct page (HG-4102 body too large) shares exactly one word with the query, the, so BM25 ranks it 18th. A method that counts shared words has no way to connect turning away big uploads with body too large.

Step 7: Add semantic search with a local embedding model

Save this as dense.py. It sends text to Ollama, scales each returned vector to length 1, and ranks documents by dot product.

# dense.py
import hashlib
import json
import math
import os

import requests

OLLAMA_URL = "http://127.0.0.1:11434/api/embed"
MODEL = "nomic-embed-text"
CACHE_PATH = os.path.join(os.path.dirname(os.path.abspath(__file__)), "cache", "embeddings.json")


def _key(text):
    return hashlib.sha256((MODEL + "\n" + text).encode("utf-8")).hexdigest()


def embed(texts):
    """Return one unit-length vector per text. Vectors are cached in a JSON file."""
    cache = {}
    if os.path.exists(CACHE_PATH):
        with open(CACHE_PATH, encoding="utf-8") as fh:
            cache = json.load(fh)
    missing = [t for t in dict.fromkeys(texts) if _key(t) not in cache]
    if missing:
        reply = requests.post(OLLAMA_URL, json={"model": MODEL, "input": missing}, timeout=120)
        reply.raise_for_status()
        for text, vector in zip(missing, reply.json()["embeddings"]):
            norm = math.sqrt(sum(x * x for x in vector))
            cache[_key(text)] = [x / norm for x in vector]
        os.makedirs(os.path.dirname(CACHE_PATH), exist_ok=True)
        with open(CACHE_PATH, "w", encoding="utf-8") as fh:
            json.dump(cache, fh)
    return [cache[_key(t)] for t in texts]


class DenseIndex:
    def __init__(self, docs):
        self.doc_ids = [doc_id for doc_id, _ in docs]
        self.vectors = embed(["search_document: " + text for _, text in docs])

    def search(self, query, top_n=10):
        """Return [(doc_id, cosine)], best first. Vectors are unit length, so a dot product is the cosine."""
        q = embed(["search_query: " + query])[0]
        scored = [(doc_id, sum(a * b for a, b in zip(q, vec))) for doc_id, vec in zip(self.doc_ids, self.vectors)]
        scored.sort(key=lambda pair: (-pair[1], pair[0]))
        return scored[:top_n]

What to know about this file

  • Ollama’s embedding endpoint. POST /api/embed takes a model and an input, which the API documentation describes as “text or list of text to generate embeddings for”. The code sends every missing text as one list, so embedding the whole corpus is a single request. It calls the loopback address 127.0.0.1 directly.
  • Task prefixes. Nomic’s model card explains that the model expects a prefix “instructing the model which task is being performed”. Documents are embedded as search_document: followed by the text, and queries as search_query: followed by the text. Skipping the prefixes never raises an error, so a mistake there would go unnoticed. I did not measure the effect on this corpus, so I make no claim about its size.
  • The cache. Vectors are stored in cache/embeddings.json, keyed by a hash of the model name and the exact text. Re-running any step is instant, works offline and gives identical numbers. Because the model name is part of the key, switching models cannot mix incompatible vectors.

Save s4_eval_dense.py and run it with python s4_eval_dense.py. The first run embeds the corpus and takes a few seconds:

# s4_eval_dense.py
from corpus import DOCS, QUERIES
from dense import DenseIndex
from evalkit import evaluate, print_summary

index = DenseIndex(DOCS)
print("embedding size:", len(index.vectors[0]), "| documents embedded:", len(index.vectors))


def dense_ids(query):
    return [doc_id for doc_id, _ in index.search(query, top_n=len(DOCS))]


ranks = evaluate(dense_ids, QUERIES)
print()
print_summary({"dense": ranks}, QUERIES)
print()
print("queries where dense retrieval does not put a relevant document in the top 3:")
for q in QUERIES:
    r = ranks[q["id"]]
    if r is None or r > 3:
        top = dense_ids(q["text"])[0]
        print("  %s %-10s rank %-3s top hit %-18s %s" % (q["id"], q["kind"], r or "-", top, q["text"][:46]))
embedding size: 768 | documents embedded: 32

method       queries       n  hit@1  hit@3  mrr@10
dense        exact         7   0.71   0.86   0.804
dense        paraphrase    9   0.44   0.56   0.572
dense        mixed         6   0.83   1.00   0.917
dense        all          22   0.64   0.77   0.740

queries where dense retrieval does not put a relevant document in the top 3:
  q01 exact      rank 8   top hit err-4033           HG-4021
  q09 paraphrase rank 7   top hit proc-drain         my calls to the backend hang for ages and then
  q11 paraphrase rank 6   top hit adv-2025-019       I edited the settings file but the old behavio
  q12 paraphrase rank 7   top hit proc-allowlist     how do I stop one customer from hammering the 
  q16 paraphrase rank 5   top hit adv-2026-007       the service keeps turning away big uploads

Reading the output

Overall, semantic search lands close to BM25: hit@1 of 0.64 (14 of 22), hit@3 of 0.77 and MRR@10 of 0.740, against 0.721 for BM25. The pattern underneath is different. It is strongest on mixed queries (MRR 0.917) and weaker on exact ones (0.804). The worst miss is the simplest query in the set: for the bare code HG-4021 the right page comes 8th and err-4033 comes first. My reading is that a bare identifier gives an embedding model almost no meaning to hold on to, so it returns pages that merely look alike, the other HG-40xx codes. I did not test that explanation further. Semantic search also misses four paraphrases (q09, q11, q12 and q16), so it is not the cure-all that some descriptions suggest.

Step 8: Fuse the two rankings

Now the step this tutorial is named for. Save this as fusion.py:

# fusion.py


def rrf(rankings, k=60, weights=None):
    """Reciprocal Rank Fusion. Each ranking is a list of document ids, best first."""
    weights = weights or [1.0] * len(rankings)
    fused = {}
    for ranking, weight in zip(rankings, weights):
        for rank, doc_id in enumerate(ranking, start=1):
            fused[doc_id] = fused.get(doc_id, 0.0) + weight / (k + rank)
    return sorted(fused.items(), key=lambda pair: (-pair[1], pair[0]))


def min_max(scored):
    """Rescale a list of (doc_id, score) pairs to the range 0..1."""
    if not scored:
        return {}
    values = [score for _, score in scored]
    low, span = min(values), max(values) - min(values)
    return {doc_id: ((score - low) / span if span else 0.0) for doc_id, score in scored}


def weighted_sum(lexical, semantic, alpha=0.5, normalize=True):
    """Blend two lists of (doc_id, score) pairs. Missing documents count as 0."""
    if normalize:
        lex, sem = min_max(lexical), min_max(semantic)
    else:
        lex, sem = dict(lexical), dict(semantic)
    fused = {doc_id: alpha * lex.get(doc_id, 0.0) + (1 - alpha) * sem.get(doc_id, 0.0)
             for doc_id in set(lex) | set(sem)}
    return sorted(fused.items(), key=lambda pair: (-pair[1], pair[0]))

Reciprocal rank fusion, by hand

Elastic’s documentation for its rrf retriever gives the formula as pseudocode: for every ranked list, if the document is in that list, add 1.0 / ( k + rank( result(q), d ) ), where ranks start at 1. The rank_constant setting, the documentation says, “determines how much influence documents in individual result sets per query have over the final ranked result set”, and it defaults to 60. A document missing from a list earns nothing from it, which is why the code simply skips it.

Try it with numbers. A document that is first in the BM25 list and eighth in the semantic list earns 1/61 + 1/68 = 0.01639 + 0.01471 = 0.03110. A document that is fourth in one list and first in the other earns 1/64 + 1/61 = 0.01563 + 0.01639 = 0.03202, and wins. Keep that example in mind, because it is exactly what happens in Step 9.

Min-max score fusion

weighted_sum is the alternative. min_max rescales one retriever’s scores so its lowest returned score becomes 0 and its highest becomes 1. weighted_sum then blends the two lists with a weight alpha on BM25, where 0.5 is an even blend. A document missing from a list counts as 0 for that list. If every score in a list is equal, min_max returns zeros instead of dividing by zero.

Save s5_eval_fusion.py and run it with python s5_eval_fusion.py:

# s5_eval_fusion.py
from bm25 import BM25Index
from corpus import DOCS, QUERIES
from dense import DenseIndex
from evalkit import evaluate, print_summary
from fusion import rrf, weighted_sum

bm25 = BM25Index(DOCS)
dense = DenseIndex(DOCS)
N = len(DOCS)


def bm25_ids(query):
    return [doc_id for doc_id, _ in bm25.search(query, top_n=N)]


def dense_ids(query):
    return [doc_id for doc_id, _ in dense.search(query, top_n=N)]


def rrf_ids(query):
    return [doc_id for doc_id, _ in rrf([bm25_ids(query), dense_ids(query)], k=60)]


def minmax_ids(query):
    return [doc_id for doc_id, _ in weighted_sum(bm25.search(query, top_n=N), dense.search(query, top_n=N), alpha=0.5)]


methods = {"bm25": bm25_ids, "dense": dense_ids, "rrf": rrf_ids, "minmax": minmax_ids}
results = {name: evaluate(fn, QUERIES) for name, fn in methods.items()}
print_summary(results, QUERIES)

print()
print("rank of the first relevant document (- means not retrieved)")
print("%-4s %-10s %5s %6s %5s %7s  %s" % ("id", "kind", "bm25", "dense", "rrf", "minmax", "query"))
worse = {"rrf": [], "minmax": []}
for q in QUERIES:
    r = {m: results[m][q["id"]] for m in methods}
    best_single = min(x for x in (r["bm25"], r["dense"]) if x) if (r["bm25"] or r["dense"]) else None
    for m in worse:
        if r[m] and best_single and r[m] > best_single:
            worse[m].append(q["id"])
    show = {m: (str(x) if x else "-") for m, x in r.items()}
    print("%-4s %-10s %5s %6s %5s %7s  %s" % (q["id"], q["kind"], show["bm25"], show["dense"], show["rrf"], show["minmax"], q["text"][:46]))

print()
for m in worse:
    print("%-6s ranks the answer below the better single retriever on: %s" % (m, ", ".join(worse[m]) or "none"))


def rank_or_big(rank):
    return rank if rank else 10 ** 6


better = sum(1 for q in QUERIES if rank_or_big(results["minmax"][q["id"]]) < rank_or_big(results["rrf"][q["id"]]))
poorer = sum(1 for q in QUERIES if rank_or_big(results["minmax"][q["id"]]) > rank_or_big(results["rrf"][q["id"]]))
print("min-max vs rrf over %d queries: better on %d, worse on %d, same on %d" % (len(QUERIES), better, poorer, len(QUERIES) - better - poorer))
method       queries       n  hit@1  hit@3  mrr@10
bm25         exact         7   0.86   1.00   0.929
bm25         paraphrase    9   0.33   0.67   0.513
bm25         mixed         6   0.67   0.83   0.792
bm25         all          22   0.59   0.82   0.721
dense        exact         7   0.71   0.86   0.804
dense        paraphrase    9   0.44   0.56   0.572
dense        mixed         6   0.83   1.00   0.917
dense        all          22   0.64   0.77   0.740
rrf          exact         7   0.71   0.86   0.814
rrf          paraphrase    9   0.33   0.78   0.597
rrf          mixed         6   1.00   1.00   1.000
rrf          all          22   0.64   0.86   0.776
minmax       exact         7   0.86   1.00   0.929
minmax       paraphrase    9   0.56   0.78   0.694
minmax       mixed         6   0.83   1.00   0.917
minmax       all          22   0.73   0.91   0.830

rank of the first relevant document (- means not retrieved)
id   kind        bm25  dense   rrf  minmax  query
q01  exact          1      8     5       1  HG-4021
q02  exact          1      1     1       1  HG-5013
q03  exact          1      1     1       1  hgctl rotate-certs
q04  exact          2      2     2       2  HAL-SA-2026-007
q05  exact          1      1     1       1  max_conns_per_client
q06  exact          1      1     1       1  upstream_connect_timeout_ms
q07  exact          1      1     1       1  retry_budget_percent
q08  paraphrase     1      1     1       1  why do people get kicked out when they step aw
q09  paraphrase     5      7     4       4  my calls to the backend hang for ages and then
q10  paraphrase     2      2     2       1  how do I replace the server's TLS identity wit
q11  paraphrase     3      6     2       3  I edited the settings file but the old behavio
q12  paraphrase     1      7     2       1  how do I stop one customer from hammering the 
q13  paraphrase     3      1     1       1  releasing a new version to a small slice of us
q14  paraphrase     1      1     1       1  taking a machine out for maintenance without c
q15  paraphrase     4      1     2       2  only let people on the corporate network reach
q16  paraphrase    18      5     8       6  the service keeps turning away big uploads
q17  mixed          2      1     1       1  what does HG-4210 mean and how do I fix it
q18  mixed          1      1     1       1  raise the timeout behind HG-4021 errors
q19  mixed          1      2     1       1  clients on TLS 1.0 broke after the upgrade, wh
q20  mixed          4      1     1       2  tokens get rejected right after expiry even th
q21  mixed          1      1     1       1  is 4.8.1 exposed to smuggling with doubled Con
q22  mixed          1      1     1       1  the metrics page is readable without a passwor

rrf    ranks the answer below the better single retriever on: q01, q12, q15, q16
minmax ranks the answer below the better single retriever on: q15, q16, q20
min-max vs rrf over 22 queries: better on 4, worse on 2, same on 16

Reading the output

Both fusion methods beat both single retrievers on hit@3 and MRR@10. RRF reaches 0.86 and 0.776, min-max reaches 0.91 and 0.830, against 0.82 and 0.721 for BM25 and 0.77 and 0.740 for semantic search. On the mixed queries RRF puts the right page first every time (hit@1 = 1.00). That is the case for hybrid search, and it is real.

The per-query table adds the caveats that a summary hides:

  • Neither method dominates. RRF ranks the answer below the better single retriever on four queries (q01, q12, q15 and q16), and min-max does so on three (q15, q16 and q20). Query q16 is the extreme case: BM25 says 18th, semantic search says 5th, and fusion lands at 8th with RRF and 6th with min-max, worse than the better input.
  • Min-max beats RRF on four queries and loses on two, with 16 unchanged. A four-to-two split over six differing queries is far too thin to crown a winner.

Step 9: Break it on purpose

The fastest way to learn what fusion can and cannot do is to run its failure modes yourself. The next script has four parts. Add each part to the same file, s6_pitfalls.py, one after another, and run the finished file once with python s6_pitfalls.py. The output below is shown in four slices that match the four parts.

Pitfall 1: adding raw scores from different scales

The most tempting fusion is the simplest: add the two scores. The first part sets up the helpers, prints the score ranges for one query, and compares the raw sum with min-max fusion and RRF.

# s6_pitfalls.py
from bm25 import BM25Index, tokenize
from corpus import DOCS, QUERIES
from dense import DenseIndex
from evalkit import evaluate, print_summary, summarize
from fusion import rrf, weighted_sum

bm25 = BM25Index(DOCS)
dense = DenseIndex(DOCS)
N = len(DOCS)


def ids(pairs):
    return [doc_id for doc_id, _ in pairs]


def rank_of(ranked, doc_id):
    return ranked.index(doc_id) + 1 if doc_id in ranked else None


def bm25_hits(query):
    return bm25.search(query, top_n=N)


def dense_hits(query):
    return dense.search(query, top_n=N)


print("=== Pitfall 1: adding raw scores from different scales ===")
q17 = next(q for q in QUERIES if q["id"] == "q17")["text"]
lex, sem = bm25_hits(q17), dense_hits(q17)
print("query: %r" % q17)
print("BM25 scores run from %.2f to %.2f; cosine scores run from %.2f to %.2f" % (lex[-1][1], lex[0][1], sem[-1][1], sem[0][1]))
results = {
    "raw-sum": evaluate(lambda q: ids(weighted_sum(bm25_hits(q), dense_hits(q), normalize=False)), QUERIES),
    "min-max": evaluate(lambda q: ids(weighted_sum(bm25_hits(q), dense_hits(q))), QUERIES),
    "rrf": evaluate(lambda q: ids(rrf([ids(bm25_hits(q)), ids(dense_hits(q))])), QUERIES),
}
print_summary(results, QUERIES, kinds=("all",))
=== Pitfall 1: adding raw scores from different scales ===
query: 'what does HG-4210 mean and how do I fix it'
BM25 scores run from 0.59 to 7.79; cosine scores run from 0.41 to 0.66
method       queries       n  hit@1  hit@3  mrr@10
raw-sum      all          22   0.59   0.86   0.736
min-max      all          22   0.73   0.91   0.830
rrf          all          22   0.64   0.86   0.776

For the query shown, BM25 scores run from 0.59 to 7.79 while cosine scores run from 0.41 to 0.66, so BM25 swamps any raw sum. The raw sum is the weakest of the three on hit@1 and MRR@10 (0.59 and 0.736), barely different from either single retriever. Never add raw scores from different systems; normalize them or use ranks.

Pitfall 2: every retriever gets an equal vote

Now the failure from the introduction. This part inspects the bare query HG-4021 and then tries weights.

# s6_pitfalls.py (continued)
print()
print("=== Pitfall 2: every retriever gets an equal vote ===")
query = "HG-4021"
lex_ids, sem_ids = ids(bm25_hits(query)), ids(dense_hits(query))
print("query %r: err-4021 is rank %s for BM25 and rank %s for dense retrieval" % (query, rank_of(lex_ids, "err-4021"), rank_of(sem_ids, "err-4021")))
print("BM25 scores:", [(d, round(s, 2)) for d, s in bm25_hits(query)[:3]])
print("cosine scores:", [(d, round(s, 3)) for d, s in dense_hits(query)[:3]], "| err-4021:", round(dict(dense_hits(query))["err-4021"], 3))
for doc_id, score in rrf([lex_ids, sem_ids])[:3]:
    print("  rrf %-10s %.5f  (bm25 rank %s, dense rank %s)" % (doc_id, score, rank_of(lex_ids, doc_id) or "-", rank_of(sem_ids, doc_id) or "-"))
print("weights (bm25, dense)   hit@1  hit@3  mrr@10")
for weights in ((1, 1), (2, 1), (3, 1), (1, 2)):
    ranks = evaluate(lambda q: ids(rrf([ids(bm25_hits(q)), ids(dense_hits(q))], weights=weights)), QUERIES)
    s = summarize(list(ranks.values()))
    print("%-22s %5.2f  %5.2f  %6.3f" % (weights, s["hit@1"], s["hit@3"], s["mrr@10"]))
=== Pitfall 2: every retriever gets an equal vote ===
query 'HG-4021': err-4021 is rank 1 for BM25 and rank 8 for dense retrieval
BM25 scores: [('err-4021', 3.84), ('err-4210', 1.52), ('err-5040', 1.2)]
cosine scores: [('err-4033', 0.572), ('err-4102', 0.558), ('err-5040', 0.557)] | err-4021: 0.528
  rrf err-4033   0.03202  (bm25 rank 4, dense rank 1)
  rrf err-5040   0.03175  (bm25 rank 3, dense rank 3)
  rrf err-4102   0.03151  (bm25 rank 5, dense rank 2)
weights (bm25, dense)   hit@1  hit@3  mrr@10
(1, 1)                  0.64   0.86   0.776
(2, 1)                  0.64   0.91   0.777
(3, 1)                  0.73   0.91   0.822
(1, 2)                  0.59   0.86   0.745

BM25 is certain: err-4021 scores 3.84 and the runner-up 1.52. Semantic search is not: its top cosine scores are 0.572, 0.558 and 0.557, with err-4021 at 0.528, a spread of only 0.044. RRF sees neither fact. It sees ranks, so BM25’s runaway first place earns the same single vote as semantic search’s hair’s-breadth first place, and err-4021 finishes fifth, which is the arithmetic from Step 8. Elastic’s documentation states the design openly (“Each child retriever carries an equal weight as part of the RRF formula”) and presents it as a benefit, because RRF removes “the need to figure out what the appropriate weighting is using linear combination”. The cost is that RRF throws away confidence.

The weights table shows one remedy. Giving BM25 three times the vote lifts MRR@10 to 0.822 and hit@1 to 0.73, while favoring semantic search drops MRR@10 to 0.745. Hold off on treating that as a recipe; the last pitfall shows why.

Pitfall 3: ranking documents that did not match

A BM25 function that scores every document and sorts will happily rank documents that share no token with the query. Their scores are all 0, so the sort leaves them in corpus order, and RRF then rewards whichever documents happen to sit early in your data.

# s6_pitfalls.py (continued)
print()
print("=== Pitfall 3: ranking documents that did not match ===")


def padded_bm25_ids(query):
    """The tempting version: score every document and sort. Zero-score documents keep corpus order."""
    terms = tokenize(query)
    scores = {doc_id: bm25.score(terms, i) for i, doc_id in enumerate(bm25.doc_ids)}
    return sorted(scores, key=lambda d: -scores[d])


lonely = "kicked away desk"
semantic = ids(dense_hits(lonely))
print("BM25 matches for %r: %s" % (lonely, bm25.search(lonely)))
for label, fused in (("dense alone", semantic),
                     ("rrf, honest BM25 list", ids(rrf([ids(bm25_hits(lonely)), semantic]))),
                     ("rrf, padded BM25 list", ids(rrf([padded_bm25_ids(lonely), semantic])))):
    print("%-22s proc-session is rank %2d | top 3: %s" % (label, rank_of(fused, "proc-session"), fused[:3]))
results = {
    "rrf": evaluate(lambda q: ids(rrf([ids(bm25_hits(q)), ids(dense_hits(q))])), QUERIES),
    "rrf-padded": evaluate(lambda q: ids(rrf([padded_bm25_ids(q), ids(dense_hits(q))])), QUERIES),
}
print_summary(results, QUERIES, kinds=("all",))
=== Pitfall 3: ranking documents that did not match ===
BM25 matches for 'kicked away desk': []
dense alone            proc-session is rank  1 | top 3: ['proc-session', 'proc-keys', 'adv-2025-019']
rrf, honest BM25 list  proc-session is rank  1 | top 3: ['proc-session', 'proc-keys', 'adv-2025-019']
rrf, padded BM25 list  proc-session is rank  8 | top 3: ['err-4021', 'err-4033', 'err-5040']
method       queries       n  hit@1  hit@3  mrr@10
rrf          all          22   0.64   0.86   0.776
rrf-padded   all          22   0.64   0.86   0.776

The query kicked away desk has no BM25 matches at all. With an honest, empty BM25 list, fusion returns the semantic ranking and proc-session, the right page, stays first. With the padded list it falls to eighth, and the top three are all error pages from the front of the corpus. On the golden set the two versions score identically. That is expected: ordinary words such as the make nearly every query match some documents, so the padding only reorders the far tail of the BM25 list. It is also how a bug like this hides, since the average looks fine and the damage appears only on rare queries.

Pitfall 4: the knobs

Last, the settings every fusion tutorial invites you to tune: the RRF constant k, the depth (how many results each retriever contributes), and the min-max weight alpha.

# s6_pitfalls.py (continued)
print()
print("=== Pitfall 4: the knobs ===")
print("rrf k      hit@1  hit@3  mrr@10")
for k in (1, 10, 60, 200):
    ranks = evaluate(lambda q: ids(rrf([ids(bm25_hits(q)), ids(dense_hits(q))], k=k)), QUERIES)
    s = summarize(list(ranks.values()))
    print("%-10d %5.2f  %5.2f  %6.3f" % (k, s["hit@1"], s["hit@3"], s["mrr@10"]))
print("rrf depth  hit@1  hit@3  mrr@10")
for depth in (1, 3, 5, 10, N):
    ranks = evaluate(lambda q: ids(rrf([ids(bm25_hits(q))[:depth], ids(dense_hits(q))[:depth]])), QUERIES)
    s = summarize(list(ranks.values()))
    print("%-10d %5.2f  %5.2f  %6.3f" % (depth, s["hit@1"], s["hit@3"], s["mrr@10"]))
print("min-max alpha (weight on BM25)  hit@1  hit@3  mrr@10")
for alpha in (0.3, 0.5, 0.7):
    ranks = evaluate(lambda q: ids(weighted_sum(bm25_hits(q), dense_hits(q), alpha=alpha)), QUERIES)
    s = summarize(list(ranks.values()))
    print("%-30.1f %5.2f  %5.2f  %6.3f" % (alpha, s["hit@1"], s["hit@3"], s["mrr@10"]))
=== Pitfall 4: the knobs ===
rrf k      hit@1  hit@3  mrr@10
1           0.59   0.86   0.749
10          0.64   0.86   0.778
60          0.64   0.86   0.776
200         0.64   0.91   0.780
rrf depth  hit@1  hit@3  mrr@10
1           0.68   0.77   0.727
3           0.68   0.86   0.780
5           0.64   0.82   0.756
10          0.64   0.91   0.774
32          0.64   0.86   0.776
min-max alpha (weight on BM25)  hit@1  hit@3  mrr@10
0.3                             0.77   0.91   0.854
0.5                             0.73   0.91   0.830
0.7                             0.64   0.91   0.777

On this set, any k from 10 to 200 changes MRR@10 by 0.004 or less, and only the extreme k = 1 stands out (0.749). Depth is not even monotonic (0.727, 0.780, 0.756, 0.774, 0.776). The weights tell the real story. Weighted RRF improved when BM25 got more weight, while min-max improved when BM25 got less (alpha of 0.3 scores 0.854). Two fusion methods, opposite advice: with 22 queries you are fitting noise. Tune only on a golden set large enough that a one-query change does not move the answer, and re-check it whenever the corpus changes.

Step 10: Package it as a searcher and add tests

Save this as search.py. It wraps both indexes and either fusion method behind one class, and its command line prints each result with the rank each retriever gave it, so you can see why a document won.

# search.py
import sys

from bm25 import BM25Index
from corpus import DOCS
from dense import DenseIndex
from fusion import rrf, weighted_sum


class HybridSearcher:
    def __init__(self, docs, method="rrf", k=60, alpha=0.5, depth=50):
        self.texts = dict(docs)
        self.bm25 = BM25Index(docs)
        self.dense = DenseIndex(docs)
        self.method, self.k, self.alpha, self.depth = method, k, alpha, depth

    def search(self, query, top_n=3):
        lexical = self.bm25.search(query, top_n=self.depth)
        semantic = self.dense.search(query, top_n=self.depth)
        lexical_ids = [doc_id for doc_id, _ in lexical]
        semantic_ids = [doc_id for doc_id, _ in semantic]
        if self.method == "minmax":
            fused = weighted_sum(lexical, semantic, alpha=self.alpha)
        else:
            fused = rrf([lexical_ids, semantic_ids], k=self.k)
        return [{
            "doc_id": doc_id,
            "score": score,
            "bm25_rank": lexical_ids.index(doc_id) + 1 if doc_id in lexical_ids else None,
            "dense_rank": semantic_ids.index(doc_id) + 1 if doc_id in semantic_ids else None,
        } for doc_id, score in fused[:top_n]]


if __name__ == "__main__":
    args = sys.argv[1:]
    method = args.pop(0) if args and args[0] in ("rrf", "minmax") else "rrf"
    query = " ".join(args) or "HG-4021"
    searcher = HybridSearcher(DOCS, method=method)
    print("method: %s | query: %s" % (method, query))
    print("%-3s %-22s %-8s %-9s %-10s %s" % ("#", "document", "score", "bm25 rank", "dense rank", "text"))
    for position, hit in enumerate(searcher.search(query), start=1):
        print("%-3d %-22s %-8.4f %-9s %-10s %s" % (
            position, hit["doc_id"], hit["score"], hit["bm25_rank"] or "-", hit["dense_rank"] or "-",
            searcher.texts[hit["doc_id"]][:48] + "..."))

Run the bare code through both fusion methods:

python search.py rrf HG-4021
python search.py minmax HG-4021
method: rrf | query: HG-4021
#   document               score    bm25 rank dense rank text
1   err-4033               0.0320   4         1          HG-4033 denied by policy. A policy rule rejected...
2   err-5040               0.0317   3         3          HG-5040 reload rejected. A configuration reload ...
3   err-4102               0.0315   5         2          HG-4102 body too large. The request body is larg...
method: minmax | query: HG-4021
#   document               score    bm25 rank dense rank text
1   err-4021               0.9128   1         8          HG-4021 upstream timeout. The upstream service d...
2   err-4210               0.5312   2         6          HG-4210 connection pool exhausted. Every connect...
3   err-4033               0.5308   4         1          HG-4033 denied by policy. A policy rule rejected...

RRF puts err-4033 first and leaves the right page out of the top three. Min-max puts err-4021 first with a score of 0.9128, because it keeps the size of BM25’s lead. Now a paraphrase and a mixed query:

python search.py rrf "why do people get kicked out when they step away from their desk"
python search.py minmax "clients on TLS 1.0 broke after the upgrade, which setting is the floor"
method: rrf | query: why do people get kicked out when they step away from their desk
#   document               score    bm25 rank dense rank text
1   proc-session           0.0328   1         1          Browser sessions use a cookie and end after sess...
2   adv-2026-011           0.0303   7         5          Advisory HAL-SA-2026-011: the admin API served t...
3   err-4012               0.0294   8         8          HG-4012 token expired. The bearer token presente...
method: minmax | query: clients on TLS 1.0 broke after the upgrade, which setting is the floor
#   document               score    bm25 rank dense rank text
1   cfg-tls-min            0.9784   1         2          tls_min_version sets the oldest TLS protocol the...
2   err-5013               0.7419   3         1          HG-5013 TLS handshake failure. The gateway could...
3   adv-2026-011           0.7166   2         6          Advisory HAL-SA-2026-011: the admin API served t...

For the paraphrase, both retrievers rank proc-session first, so the fused result is confident. For the TLS question, BM25 had the right page first and semantic search second, and fusion keeps it on top.

The test suite

Save this as test_hybrid.py. The first group of tests needs nothing but Python, and the last one uses the cached embeddings and acts as a regression gate: if a later change makes hybrid search clearly worse than either retriever alone, it fails.

# test_hybrid.py
import os
import sqlite3

import pytest

from bm25 import BM25Index, tokenize
from corpus import DOCS, QUERIES
from dense import CACHE_PATH, DenseIndex
from evalkit import evaluate, first_relevant_rank, summarize
from fusion import min_max, rrf, weighted_sum

TINY = [
    ("a", "red apple on the table"),
    ("b", "green apple and green pear"),
    ("c", "the table is made of oak"),
    ("d", "pear tree in the garden"),
]


def test_tokenize_splits_identifiers():
    assert tokenize("HG-4021 max_conns Per-Client") == ["hg", "4021", "max", "conns", "per", "client"]


def test_bm25_prefers_the_rarer_term():
    index = BM25Index(TINY)
    assert index.search("oak table")[0][0] == "c"


def test_bm25_returns_nothing_when_no_term_matches():
    assert BM25Index(TINY).search("zebra") == []


def test_classic_idf_goes_negative_but_lucene_does_not():
    docs = [("x%d" % i, "the cat %d" % i) for i in range(6)]
    assert BM25Index(docs, idf_mode="classic").idf("the") < 0
    assert BM25Index(docs, idf_mode="lucene").idf("the") > 0


def test_bm25_matches_sqlite_fts5():
    db = sqlite3.connect(":memory:")
    try:
        db.execute("CREATE VIRTUAL TABLE kb USING fts5(doc_id UNINDEXED, body)")
    except sqlite3.OperationalError:
        pytest.skip("this SQLite build has no FTS5")
    db.executemany("INSERT INTO kb VALUES (?, ?)", TINY)
    rows = dict(db.execute("SELECT doc_id, -bm25(kb) FROM kb WHERE kb MATCH ?", ('"green" OR "table"',)).fetchall())
    mine = dict(BM25Index(TINY, idf_mode="fts5").search("green table", top_n=10))
    assert set(rows) == set(mine)
    for doc_id in rows:
        assert rows[doc_id] == pytest.approx(mine[doc_id], abs=1e-9)


def test_rrf_uses_ranks_not_scores():
    fused = dict(rrf([["a", "b"], ["b", "a"]], k=60))
    assert fused["a"] == pytest.approx(1 / 61 + 1 / 62)
    assert fused["a"] == pytest.approx(fused["b"])


def test_rrf_rewards_agreement_over_one_lonely_first_place():
    fused = rrf([["x", "a", "b"], ["y", "a", "c"]], k=60)
    assert fused[0][0] == "a"


def test_rrf_weights_shift_the_balance():
    assert rrf([["a", "b"], ["b", "a"]], weights=[2, 1])[0][0] == "a"
    assert rrf([["a", "b"], ["b", "a"]], weights=[1, 2])[0][0] == "b"


def test_rrf_ignores_an_empty_ranking():
    assert rrf([[], ["a", "b"]]) == rrf([["a", "b"]])


def test_min_max_survives_equal_scores():
    assert min_max([("a", 2.0), ("b", 2.0)]) == {"a": 0.0, "b": 0.0}
    assert min_max([]) == {}


def test_weighted_sum_counts_missing_documents_as_zero():
    fused = dict(weighted_sum([("a", 4.0), ("b", 2.0)], [("b", 0.9), ("c", 0.1)], alpha=0.5))
    assert fused["a"] == pytest.approx(0.5)
    assert fused["c"] == pytest.approx(0.0)


def test_metrics():
    assert first_relevant_rank(["a", "b", "c"], {"c"}) == 3
    assert first_relevant_rank(["a", "b"], {"z"}) is None
    s = summarize([1, 2, None, 4])
    assert s["hit@1"] == 0.25 and s["hit@3"] == 0.5
    assert s["mrr@10"] == pytest.approx((1 + 0.5 + 0 + 0.25) / 4)


@pytest.mark.skipif(not os.path.exists(CACHE_PATH), reason="run the embedding steps first to fill the cache")
def test_golden_set_hybrid_is_not_worse_than_either_retriever():
    bm25, dense = BM25Index(DOCS), DenseIndex(DOCS)

    def bm25_ids(q):
        return [d for d, _ in bm25.search(q, top_n=len(DOCS))]

    def dense_ids(q):
        return [d for d, _ in dense.search(q, top_n=len(DOCS))]

    def hybrid_ids(q):
        return [d for d, _ in rrf([bm25_ids(q), dense_ids(q)])]

    mrr = {name: summarize(list(evaluate(fn, QUERIES).values()))["mrr@10"]
           for name, fn in (("bm25", bm25_ids), ("dense", dense_ids), ("hybrid", hybrid_ids))}
    assert mrr["hybrid"] >= max(mrr["bm25"], mrr["dense"]) - 0.02

Run it with python -m pytest -q test_hybrid.py:

.............                                                                                                                                                      [100%]
13 passed in 0.46s

Each test protects one lesson from this tutorial: identifiers split into tokens, BM25 prefers the rarer term, the classic IDF really does go negative, your BM25 equals SQLite’s, RRF depends on ranks and not scores, an empty list adds nothing, min_max survives equal scores, and the metrics count what they should.

Step 11: Verify everything end to end

From a clean folder with the files above, run every step in order:

python versions.py
python s1_inspect.py
python s2_fts5_check.py
python s3_eval_bm25.py
python s4_eval_dense.py
python s5_eval_fusion.py
python s6_pitfalls.py
python search.py rrf HG-4021
python search.py minmax HG-4021
python -m pytest -q test_hybrid.py

You are done when you see these things:

  • The FTS5 check reports a largest difference around 1e-15, and the BM25 numbers in Steps 4 and 6 match mine exactly. They involve no model, so they must.
  • The semantic and fusion numbers are close to mine. Vectors from a different Ollama version or different hardware can differ in the last digits, which can move a result by a place or two on a near-tie, so a query or two of difference is normal.
  • Hybrid search beats each single retriever on hit@3 and MRR@10, and the per-query table still shows queries where fusion loses.
  • All 13 tests pass.

Common mistakes and how to spot them

  • Different tokenizers for documents and queries. The index and the query must go through the same function, or tokens silently fail to match and results look empty.
  • A negative IDF. Check the IDF of your most common word. If it is below zero, your ranking rewards using it less.
  • Adding raw scores. Pitfall 1: use ranks or normalize.
  • Padding a ranked list with non-matches. Pitfall 3: a retriever that matched nothing should contribute an empty list.
  • Skipping the task prefixes. The embedding model does not complain, so you would not notice; follow the model card.
  • Mixing vectors from different models. Documents and queries must be embedded by the same model; keep the model name in your cache key.
  • Tuning on a tiny golden set. Pitfall 4: the best-looking setting flipped direction between two fusion methods.
  • Judging by averages alone. The summary hid the HG-4021 failure; read the per-query table.
  • Ollama not running or the model not pulled. The request fails with a connection error in the first case and an HTTP error in the second.

Where to go next

  • Add a second stage that rescores the top results with a cross-encoder, and measure whether it helps, with the cross-encoder reranking tutorial.
  • Tune how documents are split before they reach any retriever with the chunk size and top-k tuning tutorial.
  • Grow the golden set with real user questions and track precision and recall alongside MRR, as in the retrieval evaluation tutorial.
  • Move the keyword half into SQLite FTS5 itself once the corpus is large. Step 5 showed that its scores match your Python version’s when both use the same IDF variant.
  • Feed the fused results to a model in a complete pipeline, as in the local RAG agent tutorial.

References

  • Okapi BM25 (Wikipedia), for the formula, the usual values of k1 and b, and the plus-one IDF.
  • SQLite FTS5 documentation and its fts5_aux.c source, for the bm25() function and the IDF clamp.
  • Elasticsearch reciprocal rank fusion documentation, for the formula, the rank_constant default of 60 and the equal-weight design.
  • Cormack, Clarke and Büttcher, Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods (SIGIR 2009).
  • Nomic Embed Text v1.5 model card, Ollama’s nomic-embed-text page and the Ollama API documentation.

Tags:

EmbeddingsHybrid SearchOllamaPythonRAGSQLite

Share

A red emergency stop button on a factory control panel, a literal kill switch
Previous Post

Paragon’s CEO Turns Spyware Ethics Into a Promise the Vendor Cannot Audit

Wooden mousetrap with five spring-loaded traps set over drilled holes, with scattered grain beside a sack in a mill at Malbork
Next Post

Attackers Were Probing a Zimbra Mail Server Flaw Eight Days After Its Patch, Microsoft Says

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
05 Oct
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
05 Oct
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
Trending
October 5, 2026
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
October 5, 2026
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
October 5, 2026
Denmark Says 8.8 Million Population Register Records Were Pulled Through One Company’s Lawful Access
October 5, 2026
How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
October 5, 2026
BT’s TalkTalk Rescue Turns Telecom Continuity Into a New Merger-Control Ground
October 5, 2026
Google Stops Accepting Product Bug Reports for Its Open-Source Bounty, Citing Automated Submissions

Related Posts

Technician working with a laptop beside server racks, representing enterprise AI retrieval infrastructure
Articles

Google’s Agentic RAG Push Makes Enterprise AI Less of a One-Shot Guess

June 7, 2026
A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026