TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Evaluate RAG Retrieval Quality With Precision, Recall, and MRR in Python
Learning Hub

How to Evaluate RAG Retrieval Quality With Precision, Recall, and MRR in Python

A hands-on Python tutorial that builds a retriever, a golden dataset, and precision, recall, and MRR metrics from scratch, then uses them to catch a real retrieval regression before it reaches...

August 11, 2026 21 Min Read
53

If you have built a RAG (retrieval-augmented generation) system, you already know the uncomfortable feeling of shipping a change and just hoping it did not break anything. You try three or four queries by hand, the answers look fine, you merge. Then a week later someone asks a question the old version answered correctly and the new version gets wrong, and nobody notices until a user complains. This tutorial fixes that by teaching you to measure retrieval quality with real numbers instead of a gut feeling.

Table Of Content

  • What “Evaluating a RAG Retriever” Actually Means
  • Why Trying a Few Queries by Hand Is Not Enough
  • The Retrieval Half vs. the Generation Half
  • Prerequisites
  • Step 1: Set Up an Isolated Project
  • Step 2: Build a Tiny Knowledge Base and a TF-IDF Retriever
  • Step 3: Build a Golden Dataset
  • Step 4: Implement Retrieval Metrics From Scratch
  • Precision@k
  • Recall@k
  • Reciprocal Rank and Mean Reciprocal Rank (MRR)
  • Hit Rate
  • Step 5: Run the Evaluation and Read the Results
  • Step 6: Reproduce a Real Retrieval Regression
  • The Change: an “Improved” Tokenizer
  • What the Metrics Reveal That Eyeballing Would Miss
  • Step 7: Turn the Evaluation Into a pytest Gate That Blocks Bad Deploys
  • Watching the Gate Catch the Regression
  • Step 8: Wire the Gate Into CI (Reference)
  • Common Mistakes and Gotchas
  • How to Confirm It All Works End to End
  • Next Steps

By the end you will have built, from scratch and in plain Python, a tiny document retriever, a hand-labeled test set called a golden dataset, four standard information retrieval metrics (precision, recall, mean reciprocal rank, and hit rate), and a pytest based eval gate that automatically fails when retrieval quality drops. Along the way you will reproduce a real, realistic retrieval bug and watch the metrics catch it, including one result that will probably surprise you: the metric most people reach for first, precision, actually goes up when the system gets worse. Every command below was run for real while writing this post, and every output shown is copied from that real run, not invented.

What “Evaluating a RAG Retriever” Actually Means

A RAG system answers questions in two stages. First a retriever searches a collection of documents (a knowledge base, a support library, a codebase, whatever you have indexed) and pulls out the passages that look most relevant to the user’s question. Second, a language model reads those passages and writes an answer grounded in them. This tutorial is entirely about the first stage: did the retriever find the right passages at all. Whether the language model then wrote a good answer from that material is a separate question, covered briefly in Next Steps.

Retrieval quality matters even if you never look at it directly, because a language model cannot answer correctly from information it was never given. A confident, well-written answer built on the wrong source passage is still wrong. If your retriever is quietly missing the right document half the time, no amount of prompt engineering on the generation side will fix that; it is treating a broken foundation as if the plumbing were the problem.

Why Trying a Few Queries by Hand Is Not Enough

Manually testing a handful of queries has three specific problems. You only test the queries you happen to think of, so you systematically miss whatever you did not anticipate. You have no number to compare between yesterday’s system and today’s, so “it looks about the same” is doing a lot of unverifiable work. And a change that fixes the one query you tested can silently break five others you did not, because search systems rarely fail or improve uniformly across every kind of question.

The fix, borrowed directly from how automated testing works for regular code, is a golden dataset: a fixed list of test questions, each paired with the answer you already know is correct, that you can run against your system in seconds and get back an actual score. That score is what makes retrieval quality something you can track over time instead of something you vaguely remember.

The Retrieval Half vs. the Generation Half

Real evaluation platforms for RAG systems usually score two separate failure surfaces. Retrieval metrics, the subject of this tutorial, check whether the right source documents came back at all. Generation metrics, often scored with a second language model acting as a judge, check whether the final written answer is actually faithful to those documents and actually answers the question. sxz.io already has a tutorial on the second half, How to Build a Local LLM-as-a-Judge Evaluation Harness for AI Agents, which is worth reading once you are comfortable with the retrieval side built here. Keeping the two separate is not just tidy organization: retrieval metrics are cheap, deterministic, and fast because they only compare lists of document IDs, while generation metrics need a language model to judge nuance in written text, which is slower, costs more to run, and can itself be noisy. You want the cheap, deterministic check running constantly, and the expensive one running less often.

Prerequisites

  • A terminal you are comfortable typing commands into. Examples below use Windows PowerShell paths; the same commands work on macOS or Linux with the standard venv path differences noted inline.
  • Python 3 installed, with pip available. This tutorial was built and tested against Python 3.13.14, but nothing here depends on version-specific behavior; any actively supported Python 3 release will work the same way.
  • Basic Python familiarity: functions, dictionaries, and list comprehensions. If you can read a for loop, you can follow this.
  • No prior machine learning or search-engine background required. TF-IDF, cosine similarity, precision, recall, and mean reciprocal rank are all explained here before they are used.
  • No API keys and no GPU. Everything in this tutorial runs on a CPU with packages installed from the Python Package Index.

Step 1: Set Up an Isolated Project

Create a fresh project folder and a virtual environment, so the packages this tutorial installs do not collide with anything else on your machine:

mkdir rag-eval-tutorial
cd rag-eval-tutorial
python -m venv venv

Activate it. On Windows PowerShell:

.\venv\Scripts\Activate.ps1

On macOS or Linux:

source venv/bin/activate

Now install the three packages this tutorial needs:

pip install numpy scikit-learn pytest

Confirm the install worked:

python -c "import numpy, sklearn, pytest; print('numpy', numpy.__version__); print('scikit-learn', sklearn.__version__); print('pytest', pytest.__version__)"

Expected output (your exact version numbers may differ slightly):

numpy 2.5.2
scikit-learn 1.9.0
pytest 9.1.1

What each package is for: numpy gives us fast numeric arrays, scikit-learn provides a ready-made TF-IDF vectorizer and a cosine similarity function so we are not hand-rolling linear algebra, and pytest is the testing framework that becomes our automated eval gate in Step 7.

Step 2: Build a Tiny Knowledge Base and a TF-IDF Retriever

Every RAG system needs a document store: the collection of text chunks it searches through. For this tutorial, create a file called corpus.py with ten short fictional support articles for an imaginary SaaS product. A small, readable corpus makes it possible to reason about every single result by hand, which is exactly what you want while learning what the metrics mean.

"""Tiny support-article knowledge base used as the RAG "system under test"."""

DOCUMENTS = [
    {
        "doc_id": "pw-reset",
        "text": (
            "To reset your password, open the login page and click Forgot password. "
            "Enter the email address on your account and we will send a reset link. "
            "The link expires after 30 minutes. If you do not receive the email, check "
            "your spam folder before requesting a new link."
        ),
    },
    {
        "doc_id": "2fa-setup",
        "text": (
            "Two factor authentication (2FA) adds a second login step using an "
            "authenticator app. Open Account Settings, select Security, then Enable "
            "2FA. Scan the QR code with an authenticator app and enter the six digit "
            "code to confirm setup. Store your backup codes somewhere safe."
        ),
    },
    {
        "doc_id": "sso-setup",
        "text": (
            "Single sign-on (SSO) lets your team log in with an existing identity "
            "provider such as Okta or Azure AD. An account admin configures SSO under "
            "Organization Settings by entering the identity provider metadata URL and "
            "assigning which email domains are required to use SSO."
        ),
    },
    {
        "doc_id": "webhook-config",
        "text": (
            "Webhooks send an HTTP POST request to your server whenever an event "
            "happens, such as a new order or a status change. Add a webhook endpoint "
            "under Developer Settings, choose which events to subscribe to, and verify "
            "each request using the signing secret before trusting the payload."
        ),
    },
    {
        "doc_id": "api-rate-limits",
        "text": (
            "The API allows 600 requests per minute per account on the standard plan. "
            "If you exceed the limit you will receive an HTTP 429 response with a "
            "Retry-After header. Enterprise plans can request a higher rate limit "
            "through their account manager."
        ),
    },
    {
        "doc_id": "billing-cycle",
        "text": (
            "Subscriptions renew monthly on the date you first subscribed. You can "
            "view your next billing date and download past invoices from the Billing "
            "tab. Switching plans mid cycle prorates the difference on your next "
            "invoice."
        ),
    },
    {
        "doc_id": "refund-policy",
        "text": (
            "Refunds are available within 14 days of a charge for annual plans. "
            "Monthly plans are not refundable but you can cancel at any time to stop "
            "future charges. Contact billing support with your invoice number to "
            "request a refund."
        ),
    },
    {
        "doc_id": "data-export",
        "text": (
            "You can export all of your account data as a CSV or JSON archive from "
            "Account Settings, Privacy, Export Data. Large exports are emailed as a "
            "download link within 24 hours instead of downloading directly in the "
            "browser."
        ),
    },
    {
        "doc_id": "team-roles",
        "text": (
            "Teams support three roles: Owner, Admin, and Member. Owners manage "
            "billing and can delete the workspace. Admins manage members and "
            "integrations. Members can use the product but cannot change workspace "
            "settings. Roles are assigned from the Members page."
        ),
    },
    {
        "doc_id": "delete-account",
        "text": (
            "To permanently delete your account, go to Account Settings, Danger Zone, "
            "Delete Account. This removes all of your data after a 7 day grace period "
            "during which you can cancel the deletion by logging back in."
        ),
    },
]

Next, the retriever itself. This is the part of a RAG system that turns a question into a ranked list of document IDs. Create retriever.py:

"""Minimal TF-IDF retriever: the RAG "system under test"."""

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np


class TfidfRetriever:
    def __init__(self, documents, token_pattern=None):
        self.doc_ids = [d["doc_id"] for d in documents]
        self.texts = [d["text"] for d in documents]
        kwargs = {"lowercase": True}
        if token_pattern is not None:
            kwargs["token_pattern"] = token_pattern
        self.vectorizer = TfidfVectorizer(**kwargs)
        self.doc_matrix = self.vectorizer.fit_transform(self.texts)

    def search(self, query, k=3):
        query_vec = self.vectorizer.transform([query])
        scores = cosine_similarity(query_vec, self.doc_matrix)[0]
        ranked_indices = np.argsort(scores)[::-1][:k]
        return [self.doc_ids[i] for i in ranked_indices if scores[i] > 0]

What TF-IDF means, in plain terms: TF-IDF stands for term frequency times inverse document frequency. It scores how important a word is to one specific document relative to the whole collection. A word that appears in almost every document, like “the” or “account,” gets a low score because it does not help distinguish anything. A word that appears in only one or two documents, like “webhook” or “refund,” gets a high score because its presence is a strong signal about what that document is actually about. Every document becomes a long list of these word-importance scores, called a vector.

What cosine similarity means: once a query and every document are turned into these vectors, cosine similarity measures the angle between two vectors rather than their raw distance. That matters because it makes the comparison insensitive to document length: a short document and a long document that both emphasize the same words in the same proportions still score as similar. The retriever ranks all documents by this similarity score against the query and returns the top k.

Real production RAG systems increasingly use dense embedding models instead of, or alongside, TF-IDF. This tutorial deliberately uses TF-IDF anyway, for three practical reasons: it needs no model download, it runs instantly on a laptop CPU, and every score it produces is fully inspectable, which makes it easier to see exactly what the metrics you are about to build are actually measuring. That choice does not limit what you learn: every metric in this tutorial only looks at which doc_id values came back and in what order, never at how the retriever chose them, so the same metric code works completely unchanged if you swap in an embedding-based retriever later.

Verify the retriever works before moving on:

python -c "from corpus import DOCUMENTS; from retriever import TfidfRetriever; r = TfidfRetriever(DOCUMENTS); print(r.search('How do I reset my password?'))"

Real output from this exact command:

['pw-reset']

That is the correct document, and only that document came back, because no other document in this tiny corpus shares enough vocabulary with the query to score above zero. If you see an empty list or a different doc_id here, double check that corpus.py was saved correctly before continuing.

Step 3: Build a Golden Dataset

A golden dataset is a fixed list of test queries, each labeled by a human with the document or documents that correctly answer it. It plays the same role in a RAG system that a suite of unit tests plays in regular software: a fixed, trusted reference you can run your system against and get back a pass or fail signal that means something. Create golden_dataset.py:

"""Golden dataset: queries paired with human-labeled relevant doc_ids."""

GOLDEN_DATASET = [
    {"query": "How do I reset my password?", "relevant_doc_ids": ["pw-reset"]},
    {
        "query": "How do I protect my account from someone else logging in?",
        "relevant_doc_ids": ["pw-reset", "2fa-setup"],
    },
    {"query": "How do I enable 2FA?", "relevant_doc_ids": ["2fa-setup"]},
    {"query": "How do I set up SSO for my company?", "relevant_doc_ids": ["sso-setup"]},
    {
        "query": "How do I get an HTTP callback when a new order is created?",
        "relevant_doc_ids": ["webhook-config"],
    },
    {"query": "What is the API rate limit?", "relevant_doc_ids": ["api-rate-limits"]},
    {"query": "When does my subscription renew?", "relevant_doc_ids": ["billing-cycle"]},
    {"query": "Can I cancel and get a refund?", "relevant_doc_ids": ["refund-policy"]},
    {
        "query": "How do I export my data before deleting my account?",
        "relevant_doc_ids": ["data-export", "delete-account"],
    },
    {"query": "What can a Member do versus an Admin?", "relevant_doc_ids": ["team-roles"]},
]

Notice most entries have exactly one relevant document, but two of them (the “protect my account” and “export my data” queries) list two. That is intentional: a query with only one correct answer cannot tell you anything about whether your system is missing a second correct answer alongside the one it found. Real usage includes both single-answer and multi-answer questions, and the golden dataset needs to include both kinds too, or your metrics will only ever measure the easy case.

Notice also this dataset labels which document IDs are correct, not what the ideal generated answer text should say. That is a deliberate scope decision: comparing document IDs is fast, exact, and needs no judgment call, which is what keeps retrieval evaluation cheap enough to run on every single change. Scoring whether generated prose is a good answer is a separate, harder problem, and it is exactly what an LLM-as-judge is for.

In a real product, the source material this tutorial draws from puts this bluntly: the highest-quality eval cases come from production failures, not from your imagination, because production surfaces real user phrasing, real failure modes, and real retrieved context that nobody would think to invent by hand. Ten hand-written queries here are a stand-in for that same process at a scale small enough to reason about by hand.

Step 4: Implement Retrieval Metrics From Scratch

With a retriever and a golden dataset in place, the next step is scoring how well they match. Create metrics.py with four standard information retrieval metrics, each explained below before the code.

Precision@k

Precision@k answers: of the k documents the retriever actually returned, what fraction were genuinely relevant? A low precision score means the retriever is noisy, burying good context inside irrelevant clutter. That matters even when the right document does get retrieved, because real RAG pipelines have a limited context window, and irrelevant passages compete for that space and can distract the language model during generation.

Recall@k

Recall@k answers the opposite question: of all the documents that were actually relevant, what fraction did the retriever manage to find? A low recall score means the system is missing information the model needed to answer correctly, and no amount of clever prompting on the generation side can recover a fact that was never retrieved in the first place.

Reciprocal Rank and Mean Reciprocal Rank (MRR)

Precision and recall both treat every position in the results list equally, but position matters in practice: a relevant document buried at rank 10 is far less useful than one at rank 1, especially since most RAG pipelines only pass the first few results to the language model at all. Reciprocal rank captures this directly: it is 1 / rank of the first relevant result found (so rank 1 scores 1.0, rank 2 scores 0.5, rank 3 scores 0.33), or 0 if no relevant result appears at all. MRR is simply the average of this score across every query in the golden dataset.

Hit Rate

Hit rate is the simplest metric here: the fraction of queries where at least one relevant document showed up anywhere in the top k, full stop. It does not care about precision or rank, only about whether retrieval failed completely. Think of it as the first triage question, worth checking before the more nuanced metrics.

Here is the full implementation:

"""Retrieval metrics implemented from scratch, no eval library involved."""


def precision_at_k(retrieved_ids, relevant_ids, k):
    top_k = retrieved_ids[:k]
    if not top_k:
        return 0.0
    hits = sum(1 for doc_id in top_k if doc_id in relevant_ids)
    return hits / len(top_k)


def recall_at_k(retrieved_ids, relevant_ids, k):
    if not relevant_ids:
        return 0.0
    top_k = retrieved_ids[:k]
    hits = sum(1 for doc_id in top_k if doc_id in relevant_ids)
    return hits / len(relevant_ids)


def reciprocal_rank(retrieved_ids, relevant_ids):
    for rank, doc_id in enumerate(retrieved_ids, start=1):
        if doc_id in relevant_ids:
            return 1.0 / rank
    return 0.0


def evaluate(retriever, golden_dataset, k=3):
    rows = []
    for case in golden_dataset:
        retrieved = retriever.search(case["query"], k=k)
        relevant = case["relevant_doc_ids"]
        rows.append(
            {
                "query": case["query"],
                "retrieved": retrieved,
                "relevant": relevant,
                "precision": precision_at_k(retrieved, relevant, k),
                "recall": recall_at_k(retrieved, relevant, k),
                "reciprocal_rank": reciprocal_rank(retrieved, relevant),
                "hit": reciprocal_rank(retrieved, relevant) > 0,
            }
        )

    n = len(rows)
    summary = {
        "k": k,
        "mean_precision": sum(r["precision"] for r in rows) / n,
        "mean_recall": sum(r["recall"] for r in rows) / n,
        "mrr": sum(r["reciprocal_rank"] for r in rows) / n,
        "hit_rate": sum(1 for r in rows if r["hit"]) / n,
    }
    return rows, summary

Notice precision_at_k divides by len(top_k), the number of documents actually returned, not by k itself. That detail matters once a query returns fewer than k results (this retriever only returns documents that score above zero similarity), and it becomes important again in Step 6.

Step 5: Run the Evaluation and Read the Results

Tie it together with a small report script. Create eval_config.py, which will act as the single place that decides which retriever configuration is currently under test:

"""Swap TOKEN_PATTERN to simulate shipping the min-4-character tokenizer regression."""

TOKEN_PATTERN = None

Then create evaluate.py:

"""Run the golden dataset through the retriever and print a metrics report."""

from corpus import DOCUMENTS
from eval_config import TOKEN_PATTERN
from golden_dataset import GOLDEN_DATASET
from metrics import evaluate
from retriever import TfidfRetriever

retriever = TfidfRetriever(DOCUMENTS, token_pattern=TOKEN_PATTERN)
print("Vocabulary size:", len(retriever.vectorizer.vocabulary_))
for term in ["2fa", "sso", "api"]:
    print(f"  {term!r} in vocabulary: {term in retriever.vectorizer.vocabulary_}")

rows, summary = evaluate(retriever, GOLDEN_DATASET, k=3)
print()
for row in rows:
    mark = "HIT " if row["hit"] else "MISS"
    print(f"[{mark}] {row['query']}")
    print(f"       retrieved={row['retrieved']}")
    print(f"       relevant ={row['relevant']}")
    print(
        f"       precision@3={row['precision']:.2f}  recall@3={row['recall']:.2f}  "
        f"reciprocal_rank={row['reciprocal_rank']:.2f}"
    )
print()
print("SUMMARY:", summary)

Run it:

python evaluate.py

Real captured output:

Vocabulary size: 226
  '2fa' in vocabulary: True
  'sso' in vocabulary: True
  'api' in vocabulary: True

[HIT ] How do I reset my password?
       retrieved=['pw-reset']
       relevant =['pw-reset']
       precision@3=1.00  recall@3=1.00  reciprocal_rank=1.00
[HIT ] How do I protect my account from someone else logging in?
       retrieved=['delete-account', 'data-export', 'pw-reset']
       relevant =['pw-reset', '2fa-setup']
       precision@3=0.33  recall@3=0.50  reciprocal_rank=0.33
[HIT ] How do I enable 2FA?
       retrieved=['2fa-setup', 'pw-reset']
       relevant =['2fa-setup']
       precision@3=0.50  recall@3=1.00  reciprocal_rank=1.00
[HIT ] How do I set up SSO for my company?
       retrieved=['sso-setup', 'refund-policy', 'pw-reset']
       relevant =['sso-setup']
       precision@3=0.33  recall@3=1.00  reciprocal_rank=1.00
[HIT ] How do I get an HTTP callback when a new order is created?
       retrieved=['webhook-config', 'pw-reset', 'api-rate-limits']
       relevant =['webhook-config']
       precision@3=0.33  recall@3=1.00  reciprocal_rank=1.00
[HIT ] What is the API rate limit?
       retrieved=['api-rate-limits', 'pw-reset', 'billing-cycle']
       relevant =['api-rate-limits']
       precision@3=0.33  recall@3=1.00  reciprocal_rank=1.00
[HIT ] When does my subscription renew?
       retrieved=['billing-cycle']
       relevant =['billing-cycle']
       precision@3=1.00  recall@3=1.00  reciprocal_rank=1.00
[HIT ] Can I cancel and get a refund?
       retrieved=['refund-policy', 'team-roles', 'delete-account']
       relevant =['refund-policy']
       precision@3=0.33  recall@3=1.00  reciprocal_rank=1.00
[HIT ] How do I export my data before deleting my account?
       retrieved=['data-export', 'delete-account', 'pw-reset']
       relevant =['data-export', 'delete-account']
       precision@3=0.67  recall@3=1.00  reciprocal_rank=1.00
[HIT ] What can a Member do versus an Admin?
       retrieved=['team-roles', 'sso-setup', 'pw-reset']
       relevant =['team-roles']
       precision@3=0.33  recall@3=1.00  reciprocal_rank=1.00

SUMMARY: {'k': 3, 'mean_precision': 0.5166666666666666, 'mean_recall': 0.95, 'mrr': 0.9333333333333333, 'hit_rate': 1.0}

Reading this result: hit_rate of 1.0 means every single query found at least one relevant document somewhere in its top 3 results, a clean baseline. mean_recall of 0.95 is high but not perfect, because of exactly one row: the “protect my account” query found pw-reset but not 2fa-setup, scoring recall 0.50 on that query alone and pulling the average down. mean_precision of 0.52 looks unremarkable at first, but look closer at why: queries like “reset my password” and “subscription renew” score a perfect 1.00 precision because only one document scored above zero similarity at all, while queries like “cancel and get a refund” score only 0.33 because two irrelevant documents filled out the remaining two of three slots. That is expected behavior given a golden dataset built mostly from single-answer queries and a fixed k=3, not itself a bug. The number worth remembering going forward is MRR at 0.93: that is the baseline this tutorial protects in Step 7.

Step 6: Reproduce a Real Retrieval Regression

The Change: an “Improved” Tokenizer

Here is a realistic scenario. An engineer inspects the retriever’s vocabulary, notices a lot of very short tokens, and assumes they are noise: things like stray punctuation fragments or meaningless short words. To “clean it up,” they add a rule requiring every token to be at least 4 characters long, believing this only removes stopword-like clutter. Make that exact change by editing eval_config.py:

"""Swap TOKEN_PATTERN to simulate shipping the min-4-character tokenizer regression."""

TOKEN_PATTERN = r"(?u)\b\w{4,}\b"

In plain terms, this regular expression means: a word boundary (\b), followed by four or more word characters (\w{4,}), followed by another word boundary. Anything shorter than 4 characters, including whole words, is now invisible to the retriever.

Run python evaluate.py again with this one-line change in place. Real captured output:

Vocabulary size: 180
  '2fa' in vocabulary: False
  'sso' in vocabulary: False
  'api' in vocabulary: False

[HIT ] How do I reset my password?
       retrieved=['pw-reset']
       relevant =['pw-reset']
       precision@3=1.00  recall@3=1.00  reciprocal_rank=1.00
[MISS] How do I protect my account from someone else logging in?
       retrieved=['delete-account', 'data-export', 'api-rate-limits']
       relevant =['pw-reset', '2fa-setup']
       precision@3=0.00  recall@3=0.00  reciprocal_rank=0.00
[HIT ] How do I enable 2FA?
       retrieved=['2fa-setup']
       relevant =['2fa-setup']
       precision@3=1.00  recall@3=1.00  reciprocal_rank=1.00
[MISS] How do I set up SSO for my company?
       retrieved=[]
       relevant =['sso-setup']
       precision@3=0.00  recall@3=0.00  reciprocal_rank=0.00
[HIT ] How do I get an HTTP callback when a new order is created?
       retrieved=['webhook-config', 'api-rate-limits']
       relevant =['webhook-config']
       precision@3=0.50  recall@3=1.00  reciprocal_rank=1.00
[HIT ] What is the API rate limit?
       retrieved=['api-rate-limits']
       relevant =['api-rate-limits']
       precision@3=1.00  recall@3=1.00  reciprocal_rank=1.00
[HIT ] When does my subscription renew?
       retrieved=['billing-cycle']
       relevant =['billing-cycle']
       precision@3=1.00  recall@3=1.00  reciprocal_rank=1.00
[HIT ] Can I cancel and get a refund?
       retrieved=['refund-policy', 'delete-account']
       relevant =['refund-policy']
       precision@3=0.50  recall@3=1.00  reciprocal_rank=1.00
[HIT ] How do I export my data before deleting my account?
       retrieved=['data-export', 'delete-account', 'pw-reset']
       relevant =['data-export', 'delete-account']
       precision@3=0.67  recall@3=1.00  reciprocal_rank=1.00
[HIT ] What can a Member do versus an Admin?
       retrieved=['team-roles', 'sso-setup']
       relevant =['team-roles']
       precision@3=0.50  recall@3=1.00  reciprocal_rank=1.00

SUMMARY: {'k': 3, 'mean_precision': 0.6166666666666667, 'mean_recall': 0.8, 'mrr': 0.8, 'hit_rate': 0.8}

The vocabulary shrank from 226 tokens to 180, and the three acronyms printed above, “2fa,” “sso,” and “api,” all vanished, because every one of them is exactly 3 characters long. The damage shows up immediately in two queries. “How do I set up SSO for my company?” now retrieves nothing at all: once “sso” and every other word in that query shorter than 4 characters are stripped away, “company” is the only content word left, and it does not appear anywhere in the SSO document. “How do I protect my account from someone else logging in?” also drops to a complete miss, retrieving three wrong documents instead of the two correct ones it found before.

What the Metrics Reveal That Eyeballing Would Miss

Here is the counterintuitive part, and the real reason this tutorial exists. Compare the two SUMMARY lines directly:

baseline: mean_precision=0.52  mean_recall=0.95  mrr=0.93  hit_rate=1.00
broken:   mean_precision=0.62  mean_recall=0.80  mrr=0.80  hit_rate=0.80

mean_precision went up, from 0.52 to 0.62, while hit_rate, recall, and MRR all dropped. If this engineer had checked precision alone, believing “fewer noisy short words” would obviously mean cleaner, more precise results, they would have concluded their change was a genuine improvement and shipped it. The reason precision rose is mechanical, not meaningful: recall that precision_at_k divides by however many documents were actually returned, and the broken tokenizer now returns fewer documents per query overall (some queries return only 1, one returns 0), which shrinks the denominator right along with the numerator. A retriever that gives up and returns almost nothing can look artificially precise on the documents it does still return.

This is exactly why real eval platforms build what the source material behind this tutorial calls a diagnostic matrix: pairing each metric against the specific failure mode it actually catches, instead of trusting any single number in isolation. Precision alone would have shipped this regression. Hit rate, recall, and MRR all caught it immediately.

Step 7: Turn the Evaluation Into a pytest Gate That Blocks Bad Deploys

Reading a metrics dictionary and deciding by eye whether it looks acceptable does not scale, and it is exactly the judgment call that just went wrong in Step 6. An eval gate replaces that judgment call with automated assertions: fixed thresholds that either pass or fail, the same way a unit test does, so a regression like the one above fails a build automatically instead of reaching production.

Notice evaluate.py already imports its retriever configuration from eval_config.py rather than hardcoding it. That is deliberate: the automated gate you are about to write will import from that exact same file. If the gate tested a different configuration than what a developer actually runs locally or ships, passing the gate would prove nothing about the real system.

First, make sure eval_config.py is back to the good configuration before continuing:

"""Swap TOKEN_PATTERN to simulate shipping the min-4-character tokenizer regression."""

TOKEN_PATTERN = None

Now create test_eval_gate.py:

import pytest

from corpus import DOCUMENTS
from eval_config import TOKEN_PATTERN
from golden_dataset import GOLDEN_DATASET
from metrics import evaluate
from retriever import TfidfRetriever

THRESHOLDS = {"mean_recall": 0.90, "mrr": 0.90, "hit_rate": 0.95}


@pytest.fixture(scope="module")
def summary():
    retriever = TfidfRetriever(DOCUMENTS, token_pattern=TOKEN_PATTERN)
    _, run_summary = evaluate(retriever, GOLDEN_DATASET, k=3)
    return run_summary


def test_recall_meets_threshold(summary):
    assert summary["mean_recall"] >= THRESHOLDS["mean_recall"], (
        f"mean_recall {summary['mean_recall']:.2f} fell below the "
        f"{THRESHOLDS['mean_recall']} gate"
    )


def test_mrr_meets_threshold(summary):
    assert summary["mrr"] >= THRESHOLDS["mrr"], (
        f"MRR {summary['mrr']:.2f} fell below the {THRESHOLDS['mrr']} gate"
    )


def test_hit_rate_meets_threshold(summary):
    assert summary["hit_rate"] >= THRESHOLDS["hit_rate"], (
        f"hit_rate {summary['hit_rate']:.2f} fell below the {THRESHOLDS['hit_rate']} gate"
    )

The thresholds are set just under the real baseline measured in Step 5 (recall 0.95, MRR 0.93, hit_rate 1.00), not at a perfect 1.00. That gap is intentional: it gives ordinary, harmless variation (a golden dataset query rewritten slightly, a new document added to the corpus) a small amount of room, so the gate fails on genuine regressions instead of on every unrelated change.

Watching the Gate Catch the Regression

With eval_config.py set back to TOKEN_PATTERN = None, run:

python -m pytest test_eval_gate.py -v

Real captured output:

test_eval_gate.py::test_recall_meets_threshold PASSED                    [ 33%]
test_eval_gate.py::test_mrr_meets_threshold PASSED                       [ 66%]
test_eval_gate.py::test_hit_rate_meets_threshold PASSED                  [100%]

============================== 3 passed in 0.62s ==============================

Now make the exact same edit as Step 6, changing TOKEN_PATTERN to r"(?u)\b\w{4,}\b", and run the same command again. Real captured output:

test_eval_gate.py::test_recall_meets_threshold FAILED                    [ 33%]
test_eval_gate.py::test_mrr_meets_threshold FAILED                       [ 66%]
test_eval_gate.py::test_hit_rate_meets_threshold FAILED                  [100%]

================================== FAILURES ===================================
_________________________ test_recall_meets_threshold _________________________

E       AssertionError: mean_recall 0.80 fell below the 0.9 gate
E       assert 0.8 >= 0.9

_________________________ test_mrr_meets_threshold ____________________________

E       AssertionError: MRR 0.80 fell below the 0.9 gate
E       assert 0.8 >= 0.9

________________________ test_hit_rate_meets_threshold ________________________

E       AssertionError: hit_rate 0.80 fell below the 0.95 gate
E       assert 0.8 >= 0.95

=========================== short test summary info ===========================
FAILED test_eval_gate.py::test_recall_meets_threshold
FAILED test_eval_gate.py::test_mrr_meets_threshold
FAILED test_eval_gate.py::test_hit_rate_meets_threshold
============================== 3 failed in 0.69s ==============================

All three tests fail, each with the exact metric name and exact numbers that caused the failure. That specificity is what makes this useful in a real workflow: a teammate looking at a failed build does not have to reproduce anything locally to know what broke and by how much. pytest also exits with a non-zero status code on failure (confirm this yourself with echo $LASTEXITCODE on PowerShell or echo $? on macOS/Linux immediately after the run), which is the signal a CI system uses to mark a build as failed.

Revert eval_config.py back to TOKEN_PATTERN = None before continuing, so your project is left in its known-good state.

Step 8: Wire the Gate Into CI (Reference)

Everything so far runs entirely on your own machine. The natural next step is running that same pytest command automatically on every pull request, so a regression like the one in Step 6 never gets merged in the first place. This reference workflow was verified against GitHub’s own actions/checkout and actions/setup-python documentation, but it was not executed as part of this tutorial, since doing so requires pushing to a real GitHub repository:

# .github/workflows/rag-eval-gate.yml
name: RAG Eval Gate
on:
  pull_request:

jobs:
  eval-gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: actions/setup-python@v7
        with:
          python-version: "3.13"
      - run: pip install numpy scikit-learn pytest
      - run: pytest test_eval_gate.py -v

When a pull request introduces a regression like Step 6’s tokenizer change, this job fails the same way it just failed on your machine, and the pull request shows a failing check. To actually stop a bad pull request from being merged, mark this job as a required status check under the repository’s branch protection settings, a standard GitHub feature for exactly this purpose.

Common Mistakes and Gotchas

  • Trusting precision alone. Step 6 is the whole lesson: a system that returns fewer results can look more precise while actually retrieving less of what matters. Always read precision next to recall and hit_rate, never by itself.
  • Letting the golden dataset go stale. If nobody updates it as the product changes, it quietly stops reflecting reality, and a system can drift badly while still passing an outdated gate. Treat it as living test data, not a one-time setup task.
  • Over-trusting a small golden dataset’s exact numbers. With only 10 queries here, one flipped result swings every average by 10 percentage points. That is fine for learning the mechanics, but a real threshold-setting exercise needs enough golden cases, realistically dozens to hundreds, that a single query’s outcome cannot dominate the aggregate.
  • Ignoring empty results. A query that retrieves zero documents and a query that retrieves three weak ones both produce a defined precision and recall score (thanks to the length guard in precision_at_k), but they are different failures. That is exactly why hit_rate exists as its own separate check: a total retrieval failure deserves its own signal, not just a low number buried in an average.
  • Confusing a perfect retrieval score with a good final answer. A perfect score here only proves the language model had the right source material available. It says nothing about whether the model actually used that material correctly when writing its answer, which is a separate problem generation-side metrics are built to catch.

How to Confirm It All Works End to End

  1. Fresh checkout: recreate the virtual environment and reinstall with pip install numpy scikit-learn pytest, confirming no step above depended on hidden local state.
  2. With eval_config.py set to TOKEN_PATTERN = None, run python evaluate.py and confirm hit_rate reads 1.0 in the printed summary.
  3. Run python -m pytest test_eval_gate.py -v and confirm all 3 tests show PASSED.
  4. Edit eval_config.py to TOKEN_PATTERN = r"(?u)\b\w{4,}\b", rerun the same pytest command, and confirm all 3 tests now show FAILED with the exact assertion messages shown in Step 7.
  5. Revert eval_config.py to TOKEN_PATTERN = None and rerun once more to confirm the gate passes cleanly again.

If every one of those five checks matches what is described, your eval gate is working correctly end to end.

Next Steps

  • Swap the TF-IDF retriever for a real embedding-based one when you are ready to move past a laptop demo. Nothing in metrics.py or test_eval_gate.py needs to change, since both only ever look at the list of doc_id values a retriever returns.
  • Once retrieval is solid, move on to scoring the generation half with How to Build a Local LLM-as-a-Judge Evaluation Harness for AI Agents, which covers whether the final written answer is actually faithful to whatever the retriever found.
  • If you want a full working RAG system to point this evaluation approach at instead of the toy corpus used here, see How to Build a Local RAG Q&A Agent With LangChain v1 and Ollama.
  • Grow the golden dataset from real user questions once you have a live system, and revisit the thresholds in test_eval_gate.py periodically as the corpus and the product both change.

Tags:

information-retrievalpytestPythonRAGscikit-learn

Share

Mark Zuckerberg, Meta's CEO, in a 2025 official portrait
Previous Post

Zuckerberg’s AI Manifesto Turns Philosophy Into a Regulatory Wish List

Front view of an NVIDIA DGX Spark AI computer, one of the hardware platforms that runs Nemotron 3.5 Lightning
Next Post

NVIDIA’s Nemotron 3.5 Lightning Ships With Day-One Ubuntu Support From Canonical

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Technician working with a laptop beside server racks, representing enterprise AI retrieval infrastructure
Articles

Google’s Agentic RAG Push Makes Enterprise AI Less of a One-Shot Guess

June 7, 2026
A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026