TRENDING
Two orange safety relief valves on grey pressure vessels in an industrial plant
October 9, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
October 9, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
A lugworm lying on wet sand and mud at low tide
October 9, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
An 1840 Mulready postal envelope with a red Leicester postmark dated 4 May 1840 and a handwritten address
How to Audit SPF, DKIM, and DMARC in Python to Stop Spoofed Email From Using Your Domain
October 9, 2026
Wooden two-dial chess clock with brass-rimmed white faces showing different times
A CNCF Post on NIS2 and DORA Turns Compliance Into a Backlog and Leaves the Classification Call Unowned
October 9, 2026
A hand holding an egg against a bright light in a dark room, with the light shining through the shell to show what is inside
Anthropic Launches OSS Scanner to Email Open-Source Maintainers AI Bug Reports No Human Has Reviewed
October 9, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 233 Posts
News 235 Posts
Learning Hub 205 Posts
Home/Learning Hub/How to Audit SPF, DKIM, and DMARC in Python to Stop Spoofed Email From Using Your Domain
Learning Hub

How to Audit SPF, DKIM, and DMARC in Python to Stop Spoofed Email From Using Your Domain

Build a Python toolkit that evaluates SPF, verifies DKIM signatures, and applies DMARC under RFC 9989, then audit real domains for the gaps spoofers use.

October 9, 2026 46 Min Read
5

Every email carries two senders. One is the name and address you see in the From: line. The other is the envelope sender, the address mail servers use while they hand the message along (the SMTP MAIL FROM command, also called the return path). Plain email never checks that the first one is true, so anyone can type your domain into it. SPF, DKIM, and DMARC are the three DNS-based checks that close that gap, and they only help when the records behind them are correct. A record that exists but is broken, such as two SPF records or an SPF policy that needs eleven DNS lookups, fails silently.

Table Of Content

  • What each check proves
  • Prerequisites
  • Set up the lab
  • Step 1: Look at real SPF, DKIM, and DMARC records
  • A small DNS layer
  • Run it against real domains
  • A pretend Internet for the experiments
  • Step 2: Teach Python to evaluate SPF
  • Parse a record into terms
  • Evaluate the terms for one sender
  • Run it
  • Step 3: Reproduce the SPF traps
  • What each row teaches
  • Count the worst case with analyze()
  • Step 4: Sign and verify a message with DKIM
  • Tamper with a signed message
  • What the results mean
  • The extra From header
  • Step 5: Add DMARC alignment and the DNS tree walk
  • Why the tree walk replaced the Public Suffix List
  • Parse the record and walk the tree
  • Choose the policy and decide pass or fail
  • Test the discovery rules
  • Nine delivery scenarios
  • Step 6: Turn it into an auditor
  • Audit four invented domains
  • Audit real domains
  • Use it as a CI gate
  • Step 7: Test it
  • Common mistakes and how to avoid them
  • Treating “SPF pass” as proof of the sender
  • Publishing a second SPF record
  • Checking the lookup count only once
  • Going straight to p=reject
  • Forgetting domains that send nothing
  • Weak or unmanaged DKIM
  • Trusting old advice
  • Forgetting that DNS itself can be forged
  • Confirm it works end to end
  • Where to go next

In this tutorial you will build a small Python toolkit that does what a receiving mail server does: it evaluates SPF, verifies DKIM signatures, and applies DMARC under the current rules. Then you will turn it into an auditor command that you can point at a real domain or run in CI. You will reproduce the mistakes that let spoofed mail through, watch each one fail on purpose, and then fix it. Nothing in the lab sends a single email.

The timing is good for learning this. Google’s email sender guidelines say that since February 1, 2024, anyone sending more than 5,000 messages a day to Gmail accounts must “Set up SPF and DKIM email authentication for your domain” and set up DMARC (“Your DMARC enforcement policy can be set to none”), and that “the domain in the sender’s From: header must be aligned with either the SPF domain or the DKIM domain.” And in May 2026 the IETF published RFC 9989, which states “This document obsoletes RFCs 7489 and 9091.” DMARC now has a new specification, so many older tutorials describe rules that have changed. This one follows RFC 9989.

What each check proves

Before any code, it helps to see what question each check answers. A DNS TXT record is a small piece of text published under a domain name, and all three checks keep their rules in TXT records.

Check The question it answers What it looks at Where the record lives
SPF Is this server allowed to send mail for the envelope domain? The connecting IP address and the MAIL FROM domain A TXT record at the domain itself
DKIM Was this message signed by a key the domain published, and is it unchanged? The signed headers and the body A TXT record at selector._domainkey.domain
DMARC Does the domain in the visible From: line match what SPF or DKIM proved, and what should happen if not? The From: domain plus the SPF and DKIM results A TXT record at _dmarc.domain

Neither SPF nor DKIM requires the domain it authenticates to match the visible From: line. That is DMARC’s job, and the idea it adds is alignment: the domain that SPF or DKIM authenticated has to match the domain in From:. RFC 9989 calls the From: domain the Author Domain. Keep that word in mind, because most of the surprises in this tutorial come from the gap between “authenticated” and “aligned”.

Prerequisites

You need Python 3.10 or newer (I ran everything on Python 3.13.14 on Windows 11), a terminal, and internet access for the two steps that read real DNS records. Those lookups are read-only. No mail server, no domain of your own, and no DNS changes are needed, because the lab invents its own domains. The invented names end in .test, a top-level domain reserved for testing by RFC 2606, and the IP addresses come from the documentation ranges in RFC 5737 (192.0.2.0/24, 198.51.100.0/24, and 203.0.113.0/24). No real person or company owns them.

You will install four packages: dnspython to read real DNS, dkimpy to sign and verify DKIM, cryptography to make RSA keys, and pytest for the tests.

Set up the lab

Make an empty folder, open a terminal in it, and create a virtual environment so the packages stay out of your system Python. On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install dnspython==2.9.0 dkimpy==1.1.8 cryptography==50.0.2 pytest==9.1.1

On macOS or Linux:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install dnspython==2.9.0 dkimpy==1.1.8 cryptography==50.0.2 pytest==9.1.1

Every file in this tutorial goes into that one folder. The first line of each code block is a comment holding the file name, so you can see where it belongs. When a block starts with # spf.py (continued), append it to the file you already have. Check: python -c "import dns, dkim, cryptography; print('ok')" should print ok.

Step 1: Look at real SPF, DKIM, and DMARC records

Start by looking at the real thing. Each check publishes its rules at a predictable DNS name, so a few lookups show how real domains configure them.

A small DNS layer

The rest of the lab asks DNS questions in one place, so the same code can run against a pretend Internet in tests and against the real one in the audit. Two details matter later. First, a DNS answer can be ok, nodata (the name exists but has no record of that type), or nxdomain (the name does not exist). SPF counts the last two as “void lookups”. Second, a timeout or SERVFAIL is not an answer at all, so it raises DnsTemporaryError, which SPF reports as temperror.

# dnsio.py
"""A small DNS layer: an in-memory zone for repeatable tests and a live resolver for real domains."""
from __future__ import annotations

from dataclasses import dataclass


class DnsTemporaryError(Exception):
    """SERVFAIL, a timeout or no reachable name server (SPF calls this result temperror)."""


@dataclass(frozen=True)
class Answer:
    status: str          # "ok", "nodata" (the name exists, the record type does not) or "nxdomain"
    records: tuple = ()  # TXT: str, A and AAAA: str, MX: (preference, host)

    @property
    def void(self) -> bool:
        return self.status != "ok"


class ZoneResolver:
    """Answers from a dict such as {"example.test": {"TXT": ["v=spf1 -all"], "A": ["192.0.2.1"]}}."""

    def __init__(self, zone, broken=()):
        self.zone = {name.lower().rstrip("."): data for name, data in zone.items()}
        self.broken = {name.lower().rstrip(".") for name in broken}
        self.log = []

    def query(self, name, rtype):
        name = name.lower().rstrip(".")
        self.log.append((name, rtype))
        if name in self.broken:
            raise DnsTemporaryError(f"SERVFAIL for {name}")
        node = self.zone.get(name)
        if node is None:
            if any(other.endswith("." + name) for other in self.zone):
                return Answer("nodata")  # an empty non-terminal: names exist below it
            return Answer("nxdomain")
        records = node.get(rtype, [])
        return Answer("ok", tuple(records)) if records else Answer("nodata")


class LiveResolver:
    """The same interface backed by dnspython, for auditing real domains."""

    def __init__(self, nameserver=None, lifetime=6.0):
        import dns.resolver

        self._resolver = dns.resolver.Resolver()
        if nameserver:
            self._resolver.nameservers = [nameserver]
        self._resolver.lifetime = lifetime
        self.log = []

    def query(self, name, rtype):
        import dns.exception
        import dns.resolver

        name = name.lower().rstrip(".")
        self.log.append((name, rtype))
        try:
            answer = self._resolver.resolve(name, rtype)
        except dns.resolver.NXDOMAIN:
            return Answer("nxdomain")
        except dns.resolver.NoAnswer:
            return Answer("nodata")
        except (dns.resolver.NoNameservers, dns.exception.Timeout) as exc:
            raise DnsTemporaryError(f"{type(exc).__name__} for {name}") from exc
        records = []
        for rdata in answer:
            if rtype == "TXT":
                records.append(b"".join(rdata.strings).decode("ascii", "replace"))
            elif rtype == "MX":
                records.append((rdata.preference, str(rdata.exchange).rstrip(".").lower()))
            else:
                records.append(rdata.address)
        return Answer("ok", tuple(records))

ZoneResolver answers from a Python dictionary and remembers every question in log, which later lets you count queries. LiveResolver wraps dnspython and has the same query method.

Run it against real domains

# step01_live_records.py
from dnsio import LiveResolver

dns = LiveResolver()

for domain in ["example.com", "google.com", "gmail.com"]:
    print(f"== {domain}")
    for name in (domain, f"_dmarc.{domain}"):
        for text in dns.query(name, "TXT").records:
            if text.lower().startswith(("v=spf1", "v=dmarc1")):
                print(f"  {name:22} {text}")

print("== DKIM keys live under <selector>._domainkey.<domain>")
for name in ("s1._domainkey.github.com", "20230601._domainkey.gmail.com"):
    for text in dns.query(name, "TXT").records:
        shown = text if len(text) < 60 else text[:60] + f"... ({len(text)} characters)"
        print(f"  {name:32} {shown}")

Run python step01_live_records.py. These are the records I got on October 9, 2026. Yours may differ, because domain owners change their records.

== example.com
  example.com            v=spf1 -all
  _dmarc.example.com     v=DMARC1;p=reject;sp=reject;adkim=s;aspf=s
== google.com
  google.com             v=spf1 include:_spf.google.com ~all
  _dmarc.google.com      v=DMARC1; p=reject; rua=mailto:[email protected]
== gmail.com
  gmail.com              v=spf1 redirect=_spf.google.com
  _dmarc.gmail.com       v=DMARC1; p=none; sp=quarantine; rua=mailto:[email protected]
== DKIM keys live under <selector>._domainkey.<domain>
  s1._domainkey.github.com         k=rsa; t=s; p=MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAyn... (406 characters)
  20230601._domainkey.gmail.com    v=DKIM1; k=rsa; p=

Read the output line by line. example.com publishes v=spf1 -all, a “null” SPF record that says no server is allowed to send as this domain, and a DMARC record with p=reject and strict alignment (adkim=s; aspf=s). That is a sensible posture for a domain that never sends mail. google.com authorizes the servers listed in another domain’s record (include:_spf.google.com) and ends with ~all. gmail.com uses redirect=_spf.google.com, which hands the whole policy to another name, and its DMARC record has p=none for the domain and sp=quarantine for subdomains.

The last two lines are DKIM keys. A DKIM record sits at selector._domainkey.domain, and an ordinary DNS lookup cannot list a domain’s selectors. The usual way to learn one is to read the s= tag of a signed message, which is why the auditor you build later takes selectors as input. The GitHub key is a 2048-bit RSA public key (the long p= value). The Gmail selector has an empty p=, and RFC 6376 says of that: “An empty value means that this public key has been revoked.” Nobody can verify a signature that claims that selector any more.

Check: you should see all three domains with an SPF line and a DMARC line. If a lookup raises DnsTemporaryError, your network blocks the resolver; try LiveResolver("8.8.8.8") in the script.

A pretend Internet for the experiments

For repeatable experiments the lab builds its own zone. The cast is a shop (example.test) that sends from its own servers and through an email service provider (mailer.test), an attacker (evil.test) who owns valid records for their own domain, a mailing-list server (lists.test), and a few domains used later to test the DMARC lookup rules. keypair() creates an RSA key once and saves it under keys/, so every run publishes the same DKIM record. These are throwaway lab keys. Never publish or share a real private key.

# fixtures.py
"""Synthetic DNS data for the lab. Names end in .test and addresses come from the RFC 5737 documentation ranges."""
from __future__ import annotations

import base64
from pathlib import Path

from cryptography.hazmat.primitives import serialization
from cryptography.hazmat.primitives.asymmetric import rsa

from dnsio import ZoneResolver

KEYS = Path(__file__).parent / "keys"
PEM, DER = serialization.Encoding.PEM, serialization.Encoding.DER


def keypair(label, bits=2048):
    """Create a private key once and reuse it, so every run publishes the same DNS record."""
    KEYS.mkdir(exist_ok=True)
    path = KEYS / f"{label}.pem"
    if path.exists():
        key = serialization.load_pem_private_key(path.read_bytes(), password=None)
    else:
        key = rsa.generate_private_key(public_exponent=65537, key_size=bits)
        path.write_bytes(key.private_bytes(PEM, serialization.PrivateFormat.TraditionalOpenSSL,
                                           serialization.NoEncryption()))
    pem = key.private_bytes(PEM, serialization.PrivateFormat.TraditionalOpenSSL, serialization.NoEncryption())
    public = key.public_key().public_bytes(DER, serialization.PublicFormat.SubjectPublicKeyInfo)
    return pem, "v=DKIM1; k=rsa; p=" + base64.b64encode(public).decode()


def mail_zone():
    """example.test (the shop), mailer.test (its email provider), evil.test, lists.test and friends."""
    _, shop_txt = keypair("shop")
    _, esp_txt = keypair("esp")
    _, evil_txt = keypair("evil")
    return {
        "example.test": {"TXT": ["v=spf1 ip4:192.0.2.0/24 include:mailer.test -all"], "A": ["192.0.2.80"]},
        "_dmarc.example.test": {"TXT": ["v=DMARC1; p=reject; sp=quarantine; np=reject; rua=mailto:[email protected]"]},
        "mail1._domainkey.example.test": {"TXT": [shop_txt]},
        "esp1._domainkey.example.test": {"TXT": [esp_txt]},  # the shop hands one selector to its provider
        "news.example.test": {"A": ["192.0.2.81"]},
        "mailer.test": {"TXT": ["v=spf1 ip4:198.51.100.0/24 -all"]},
        "mailer1._domainkey.mailer.test": {"TXT": [esp_txt]},
        "evil.test": {"TXT": ["v=spf1 ip4:203.0.113.0/26 -all"]},
        "k1._domainkey.evil.test": {"TXT": [evil_txt]},
        "lists.test": {"TXT": ["v=spf1 ip4:203.0.113.64/26 -all"]},
        # a decentralised organisation: corp.test and eu.corp.test each publish a record
        "corp.test": {"A": ["192.0.2.90"]},
        "_dmarc.corp.test": {"TXT": ["v=DMARC1; p=reject"]},
        "eu.corp.test": {"A": ["192.0.2.91"]},
        "_dmarc.eu.corp.test": {"TXT": ["v=DMARC1; p=quarantine; psd=n"]},
        "news.eu.corp.test": {"A": ["192.0.2.92"]},
        # a public suffix operator that publishes a policy for everything below bank.test
        "_dmarc.bank.test": {"TXT": ["v=DMARC1; p=reject; psd=y"]},
        "shop.bank.test": {"A": ["192.0.2.95"]},
        # a domain that is testing a reject policy
        "testing.test": {"A": ["192.0.2.96"]},
        "_dmarc.testing.test": {"TXT": ["v=DMARC1; p=reject; t=y; rua=mailto:[email protected]"]},
    }


def audit_zone():
    """Four domains for the auditor: good, sloppy, legacy and parked."""
    _, good_txt = keypair("shop")
    _, weak_txt = keypair("weak1024", 1024)
    zone = {
        "good.test": {"TXT": ["v=spf1 ip4:192.0.2.0/24 -all"]},
        "_dmarc.good.test": {"TXT": ["v=DMARC1; p=reject; rua=mailto:[email protected]"]},
        "s1._domainkey.good.test": {"TXT": [good_txt]},
        "sloppy.test": {"TXT": ["v=spf1 " + " ".join(f"include:{c}.sloppy.test" for c in "abcdefghi") + " a mx ?all"],
                        "A": ["192.0.2.50"], "MX": [(10, "mx.sloppy.test")]},
        "mx.sloppy.test": {"A": ["192.0.2.51"]},
        "_dmarc.sloppy.test": {"TXT": ["v=DMARC1; p=none"]},
        "s1._domainkey.sloppy.test": {"TXT": [weak_txt + "; t=y"]},
        "legacy.test": {"TXT": ["v=spf1 ip4:192.0.2.0/24 -all", "v=spf1 include:mailer.test -all"]},
        "_dmarc.legacy.test": {"TXT": ["v=DMARC1; p=quarantine; sp=none; pct=50; rua=mailto:[email protected]"]},
        "parked.test": {"TXT": ["v=spf1 -all"]},
        "_dmarc.parked.test": {"TXT": ["v=DMARC1; p=reject; sp=reject; adkim=s; aspf=s"]},
        "mailer.test": {"TXT": ["v=spf1 ip4:198.51.100.0/24 -all"]},
    }
    for letter in "abcdefghi":
        zone[f"{letter}.sloppy.test"] = {"TXT": [f"v=spf1 ip4:203.0.113.{ord(letter) - 96} -all"]}
    return zone


def mail_resolver(**kwargs):
    return ZoneResolver(mail_zone(), **kwargs)

Step 2: Teach Python to evaluate SPF

An SPF record is a list of terms that a receiver checks from left to right, using the IP address that connected and the domain from the envelope sender. The first term that matches decides the result. Each term has an optional qualifier: + pass (the default), - fail, ~ softfail, or ? neutral. RFC 7208 defines these mechanisms: ip4 and ip6 match an address range, a and mx match the addresses of a host, include asks another domain’s policy, exists tests whether a name resolves, all matches everything, and ptr uses reverse DNS. Two modifiers appear too: redirect swaps in another domain’s policy, and exp supplies an explanation string.

The checker can end in seven results: pass, fail, softfail, neutral, none (no SPF record), temperror (a DNS failure that may clear), and permerror (the record itself is broken). Your version covers every mechanism except ptr, and it does not expand macros or evaluate exp. The RFC says of ptr, “This mechanism SHOULD NOT be published,” so leaving it out costs little. The auditor still flags it.

Parse a record into terms

Parsing comes first. parse() splits the record, peels off the qualifier, and turns each piece into a Term. Anything it cannot read raises SpfError, because RFC 7208 says that “if there are any syntax errors anywhere in the record, check_host() returns immediately with the result ‘permerror’”.

# spf.py
"""A small SPF checker built from RFC 7208: check_host() for one sender, analyze() for the worst case."""
from __future__ import annotations

import ipaddress
import re
from dataclasses import dataclass, field

from dnsio import DnsTemporaryError

RESULT_FOR_QUALIFIER = {"+": "pass", "-": "fail", "~": "softfail", "?": "neutral"}
MECHANISMS = {"all", "include", "a", "mx", "ptr", "ip4", "ip6", "exists"}
LOOKUP_TERMS = {"include", "a", "mx", "ptr", "exists", "redirect"}  # RFC 7208 section 4.6.4
SPF_TEXT = re.compile(r"v=spf1(\s.*)?", re.IGNORECASE)


class SpfError(Exception):
    """A permanent error. RFC 7208 calls the result permerror."""


@dataclass(frozen=True)
class Term:
    name: str
    qualifier: str = "+"
    value: str = ""
    cidr4: int | None = None
    cidr6: int | None = None
    modifier: bool = False

    def __str__(self):
        if self.modifier:
            return f"{self.name}={self.value}"
        text = ("" if self.qualifier == "+" else self.qualifier) + self.name
        if self.value:
            text += ":" + self.value
        if self.cidr4 is not None:
            text += f"/{self.cidr4}"
        if self.cidr6 is not None:
            text += f"//{self.cidr6}"
        return text


def parse(record: str) -> list[Term]:
    parts = record.split()
    if not parts or parts[0].lower() != "v=spf1":
        raise SpfError("the record does not start with v=spf1")
    terms, seen = [], set()
    for raw in parts[1:]:
        qualifier, body = "+", raw
        if raw[0] in RESULT_FOR_QUALIFIER:
            qualifier, body = raw[0], raw[1:]
        name = re.split(r"[:/=]", body, maxsplit=1)[0].lower()
        if name in MECHANISMS:
            terms.append(_mechanism(name, qualifier, body))
        elif re.match(r"[A-Za-z][A-Za-z0-9._-]*=", raw):
            key, value = raw.split("=", 1)
            key = key.lower()
            if key in ("redirect", "exp"):
                if key in seen:
                    raise SpfError(f"the record has more than one {key} modifier")
                seen.add(key)
            terms.append(Term(key, "", value, modifier=True))  # unknown modifiers are ignored later
        else:
            raise SpfError(f"unknown mechanism {raw!r}")
    return terms


def _mechanism(name, qualifier, body):
    rest = body[len(name):]
    if name in ("a", "mx"):
        match = re.fullmatch(r"(?::([^/]+))?(?:/(\d{1,2}))?(?://(\d{1,3}))?", rest)
        if not match:
            raise SpfError(f"cannot parse {body!r}")
        cidr4 = int(match.group(2)) if match.group(2) else None
        cidr6 = int(match.group(3)) if match.group(3) else None
        return Term(name, qualifier, match.group(1) or "", cidr4, cidr6)
    if name == "all":
        if rest:
            raise SpfError("all takes no argument")
        return Term(name, qualifier)
    if name == "ptr":
        return Term(name, qualifier, rest[1:] if rest.startswith(":") else "")
    if not rest.startswith(":") or len(rest) == 1:
        raise SpfError(f"{name} needs a value")
    return Term(name, qualifier, rest[1:])

Evaluate the terms for one sender

Now the evaluator. The class _Checker keeps two counters. count() charges every term that costs a DNS lookup against a budget of ten, and query() charges every address or MX lookup that comes back empty against a budget of two void lookups. Section 4.6.4 of the RFC is the source of both limits: “SPF implementations MUST limit the total number of those terms to 10 during SPF evaluation, to avoid unreasonable load on the DNS,” and “SPF implementations SHOULD limit ‘void lookups’ to two.” Exceeding either limit produces a permerror, although the void-lookup limit is only a SHOULD in the RFC.

The most important line is in the include branch of matches(). An include does not paste another record into yours. It evaluates that record and then asks one question: did it pass? Anything else, including a fail, simply means not a match and evaluation continues. The RFC admits the name is misleading: “In hindsight, the name ‘include’ was poorly chosen.”

# spf.py (continued)
@dataclass
class SpfResult:
    verdict: str  # pass, fail, softfail, neutral, none, temperror or permerror
    detail: str = ""
    matched: str = ""
    lookups: int = 0
    voids: int = 0
    trace: list = field(default_factory=list)


class _Checker:
    def __init__(self, resolver, ip, max_lookups, max_voids):
        self.dns = resolver
        self.ip = ipaddress.ip_address(ip)
        self.max_lookups, self.max_voids = max_lookups, max_voids
        self.lookups = 0
        self.voids = 0
        self.trace: list[str] = []

    def count(self, what):
        self.lookups += 1
        self.trace.append(f"lookup {self.lookups}: {what}")
        if self.lookups > self.max_lookups:
            raise SpfError(f"more than {self.max_lookups} DNS lookups (stopped at {what})")

    def query(self, name, rtype):
        answer = self.dns.query(name, rtype)
        if answer.void:
            self.voids += 1
            self.trace.append(f"void lookup {self.voids}: {rtype} {name} returned {answer.status}")
            if self.voids > self.max_voids:
                raise SpfError(f"more than {self.max_voids} void lookups")
        return answer

    @staticmethod
    def spec(value):
        if "%" in value:
            raise SpfError("macros are not supported in this tutorial")
        return value.lower().rstrip(".")

    def record(self, domain):
        found = [t for t in self.dns.query(domain, "TXT").records if SPF_TEXT.fullmatch(t)]
        if len(found) > 1:
            raise SpfError(f"{domain} publishes {len(found)} SPF records")
        return found[0] if found else None

    def run(self, domain, depth=0):
        if depth > 10:
            raise SpfError("the include or redirect chain is too deep")
        record = self.record(domain)
        if record is None:
            return "none", None
        terms = parse(record)
        for term in terms:
            if term.modifier:
                continue
            if self.matches(term, domain, depth):
                self.trace.append(f"{domain}: {term} matched")
                return RESULT_FOR_QUALIFIER[term.qualifier], term
        redirect = next((t for t in terms if t.name == "redirect"), None)
        if redirect is None:
            return "neutral", None
        target = self.spec(redirect.value)
        self.count(f"redirect={target}")
        result, term = self.run(target, depth + 1)
        if result == "none":
            raise SpfError(f"the redirect target {target} has no SPF record")
        return result, term

    def matches(self, term, domain, depth):
        name = term.name
        if name == "all":
            return True
        if name in ("ip4", "ip6"):
            try:
                network = ipaddress.ip_network(term.value, strict=False)
            except ValueError as exc:
                raise SpfError(f"bad {name} value {term.value!r}") from exc
            if network.version != (4 if name == "ip4" else 6):
                raise SpfError(f"{name}:{term.value} is the wrong address family")
            return network.version == self.ip.version and self.ip in network
        if name == "include":
            target = self.spec(term.value)
            self.count(f"include:{target}")
            result, _ = self.run(target, depth + 1)
            if result == "none":
                raise SpfError(f"include:{target} points at a domain with no SPF record")
            return result == "pass"  # fail, softfail and neutral inside an include are "not a match"
        if name in ("a", "mx"):
            target = self.spec(term.value) if term.value else domain
            self.count(f"{name}:{target}")
            hosts = [target]
            if name == "mx":
                hosts = [host for _, host in sorted(self.query(target, "MX").records)]
                if len(hosts) > 10:
                    raise SpfError(f"mx:{target} names more than 10 hosts")
            return any(self.in_network(addr, term) for host in hosts for addr in self.addresses(host))
        if name == "exists":
            target = self.spec(term.value)
            self.count(f"exists:{target}")
            return self.query(target, "A").status == "ok"
        raise SpfError(f"{name} is not evaluated by this tutorial")

    def addresses(self, host):
        return list(self.query(host, "A" if self.ip.version == 4 else "AAAA").records)

    def in_network(self, address, term):
        prefix = term.cidr4 if self.ip.version == 4 else term.cidr6
        if prefix is None:
            prefix = 32 if self.ip.version == 4 else 128
        return self.ip in ipaddress.ip_network(f"{address}/{prefix}", strict=False)


def check_host(ip, domain, resolver, *, max_lookups=10, max_voids=2) -> SpfResult:
    """Evaluate the SPF policy of `domain` for a connection from `ip` (RFC 7208 section 4)."""
    checker = _Checker(resolver, ip, max_lookups, max_voids)
    try:
        verdict, term = checker.run(domain.lower().rstrip("."))
        detail, matched = "", str(term) if term else ""
    except SpfError as exc:
        verdict, detail, matched = "permerror", str(exc), ""
    except DnsTemporaryError as exc:
        verdict, detail, matched = "temperror", str(exc), ""
    return SpfResult(verdict, detail, matched, checker.lookups, checker.voids, checker.trace)

Run it

# step02_spf_basic.py
from fixtures import mail_resolver
from spf import check_host

dns = mail_resolver()
print("example.test publishes:", dns.query("example.test", "TXT").records[0])
print()
for who, ip in [("the shop's own server", "192.0.2.10"), ("the email provider", "198.51.100.7"),
                ("a stranger", "203.0.113.9")]:
    result = check_host(ip, "example.test", dns)
    print(f"{who:22} {ip:14} -> {result.verdict:5} matched={result.matched or '-':20} DNS lookups={result.lookups}")

print()
print("How the provider's address was decided:")
for line in check_host("198.51.100.7", "example.test", dns).trace:
    print("  " + line)

Run python step02_spf_basic.py.

example.test publishes: v=spf1 ip4:192.0.2.0/24 include:mailer.test -all

the shop's own server  192.0.2.10     -> pass  matched=ip4:192.0.2.0/24     DNS lookups=0
the email provider     198.51.100.7   -> pass  matched=include:mailer.test  DNS lookups=1
a stranger             203.0.113.9    -> fail  matched=-all                 DNS lookups=1

How the provider's address was decided:
  lookup 1: include:mailer.test
  mailer.test: ip4:198.51.100.0/24 matched
  example.test: include:mailer.test matched

The shop’s own server matches ip4:192.0.2.0/24 without any DNS lookup. The provider’s address is found only after the checker follows include:mailer.test, which costs one lookup, and the trace shows the inner term matching first and then the outer include matching. A stranger falls through to -all and fails. Check: the three verdicts should be pass, pass, fail.

Step 3: Reproduce the SPF traps

Most real SPF damage comes from a handful of mistakes. This script builds a tiny zone for each one and shows the verdict. Each record is written out in the script so you can see exactly what is published.

# step03_spf_traps.py
from dnsio import ZoneResolver
from spf import check_host


def run(label, zone, ip="203.0.113.9", **kwargs):
    result = check_host(ip, "t.test", ZoneResolver(zone, **kwargs))
    print(f"{label:34} {result.verdict:9} lookups={result.lookups:2} voids={result.voids}  {result.detail}")


def includes(count, tail="ip4:203.0.113.0/24 -all"):
    zone = {"t.test": {"TXT": ["v=spf1 " + " ".join(f"include:i{n}.test" for n in range(1, count + 1)) + " " + tail]}}
    zone.update({f"i{n}.test": {"TXT": ["v=spf1 -all"]} for n in range(1, count + 1)})
    return zone


run("include with -all inside", {
    "t.test": {"TXT": ["v=spf1 include:partner.test ip4:203.0.113.0/24 -all"]},
    "partner.test": {"TXT": ["v=spf1 ip4:198.51.100.0/24 -all"]}})
run("two SPF records", {"t.test": {"TXT": ["v=spf1 -all", "v=spf1 ip4:203.0.113.0/24 -all"]}})
run("include of a domain with no SPF", {
    "t.test": {"TXT": ["v=spf1 include:plain.test -all"]}, "plain.test": {"TXT": ["hello"]}})
run("10 includes, then a matching ip4", includes(10))
run("11 includes, then a matching ip4", includes(11))
run("three a: terms on missing names", {
    "t.test": {"TXT": ["v=spf1 a:x1.test a:x2.test a:x3.test ip4:203.0.113.0/24 -all"]}})
run("DNS failure inside an include", {"t.test": {"TXT": ["v=spf1 include:down.test -all"]}}, broken=["down.test"])
run("ends in +all", {"t.test": {"TXT": ["v=spf1 ip4:198.51.100.0/24 +all"]}})
run("no all and no redirect", {"t.test": {"TXT": ["v=spf1 ip4:198.51.100.0/24"]}})
run("redirect instead of all", {
    "t.test": {"TXT": ["v=spf1 redirect=policy.test"]}, "policy.test": {"TXT": ["v=spf1 ip4:203.0.113.0/24 -all"]}})

Run python step03_spf_traps.py.

include with -all inside           pass      lookups= 1 voids=0  
two SPF records                    permerror lookups= 0 voids=0  t.test publishes 2 SPF records
include of a domain with no SPF    permerror lookups= 1 voids=0  include:plain.test points at a domain with no SPF record
10 includes, then a matching ip4   pass      lookups=10 voids=0  
11 includes, then a matching ip4   permerror lookups=11 voids=0  more than 10 DNS lookups (stopped at include:i11.test)
three a: terms on missing names    permerror lookups= 3 voids=3  more than 2 void lookups
DNS failure inside an include      temperror lookups= 1 voids=0  SERVFAIL for down.test
ends in +all                       pass      lookups= 0 voids=0  
no all and no redirect             neutral   lookups= 0 voids=0  
redirect instead of all            pass      lookups= 1 voids=0  

What each row teaches

An include with -all inside. The partner’s record ends in -all, yet the verdict is pass, because the include is only not a match and the later ip4 term matches. The RFC says evaluating a “-all” directive in the referenced record “does not terminate the overall processing and does not necessarily result in an overall ‘fail’.”

Two SPF records. The RFC says “A domain name MUST NOT have multiple records that would cause an authorization check to select more than one record,” and the result is permerror, which means SPF fails for every sender, including your real ones. One way to end up here is to add a new email vendor by pasting a second record instead of editing the first.

An include of a domain with no SPF record. Another permerror. The RFC’s table maps an include whose target has no record (none) to permerror, so a typo in an include name breaks the whole policy.

The limit of ten. Ten includes followed by a matching ip4 pass. Eleven includes are a permerror even though the matching ip4 would have been reached next. Receivers stop at the limit, so only the senders whose terms sit near the end of a long record fail, which makes the problem hard to notice.

Void lookups. Three a: terms pointing at names that do not exist give permerror on the third, because the budget is two. A vendor name that no longer exists in DNS is one way to run into this.

A DNS failure. A SERVFAIL for an included name is temperror, not fail. Receivers may retry or defer, so it is a different outcome from a broken record.

Ends in +all, or has no all. +all authorizes the entire Internet. With no all and no redirect, an address that matches nothing gets neutral, which is not a refusal. The RFC describes all as the term that “always matches” and says it is used “as the rightmost mechanism in a record to provide an explicit default.”

A redirect. Instead of ending in all, the record hands its whole policy to policy.test. That costs one lookup and is the tidy way to share one policy between several domains of your own.

Count the worst case with analyze()

The checker above counts lookups only on the path it took. A given sender may match early and use few lookups, while another sender uses all of them. To know whether any sender could hit the limit, you need the count for the whole tree. analyze() walks every include and redirect, counts every term that costs a lookup, counts void answers, and records the qualifier on the final all. The auditor uses it.

# spf.py (continued)
@dataclass
class SpfReport:
    domain: str
    record: str | None = None
    problem: str = ""  # "none", "multiple", "syntax" or ""
    lookups: int = 0  # every DNS-querying term in the whole tree: the worst case
    voids: int = 0
    all_qualifier: str | None = None
    uses_ptr: bool = False
    includes: list = field(default_factory=list)
    error: str = ""


def analyze(domain, resolver, max_depth=10) -> SpfReport:
    """Walk the whole include and redirect tree and count every term that costs a DNS lookup."""
    report = SpfReport(domain.lower().rstrip("."))

    def walk(name, depth, effective, stack):
        if depth > max_depth or name in stack:
            report.error = f"{name} is too deep or loops back"
            return
        try:
            texts = [t for t in resolver.query(name, "TXT").records if SPF_TEXT.fullmatch(t)]
        except DnsTemporaryError as exc:
            report.error = f"DNS error for {name}: {exc}"
            return
        if len(texts) != 1:
            if depth == 0:
                report.problem = "none" if not texts else "multiple"
            else:
                report.error = f"{name} has {len(texts)} SPF records"
            return
        if depth == 0:
            report.record = texts[0]
        try:
            terms = parse(texts[0])
        except SpfError as exc:
            report.problem, report.error = "syntax", str(exc)
            return
        for term in terms:
            if term.name in LOOKUP_TERMS:
                report.lookups += 1
            if term.name == "ptr":
                report.uses_ptr = True
            if term.name == "all" and effective:
                report.all_qualifier = term.qualifier
            if term.name in ("a", "mx", "exists") and not term.modifier:
                target = term.value or name
                rtype = "MX" if term.name == "mx" else "A"
                try:
                    report.voids += resolver.query(target, rtype).void
                except DnsTemporaryError:
                    pass
            if term.name == "include":
                report.includes.append(term.value)
                walk(term.value.lower(), depth + 1, False, stack | {name})
            elif term.name == "redirect":
                walk(term.value.lower(), depth + 1, effective, stack | {name})

    walk(report.domain, 0, True, frozenset())
    return report

I cross-checked it on a real record. For github.com on October 9, 2026, analyze() counted 10 terms that cost a lookup, and a separate short script I wrote with a plain regular expression over the same tree also counted 10. That is exactly the maximum, so one more vendor would break the record for some senders. The count is a worst case: no single sender necessarily needs all ten.

Step 4: Sign and verify a message with DKIM

DKIM works with a key pair. The sending system hashes the message body, then signs the selected headers together with the DKIM-Signature header itself (which holds the body hash) using its private key, and adds the result to that header. The matching public key is published in DNS under the selector named in s=. A receiver finds the key at selector._domainkey.domain, recomputes both hashes, and checks the signature. The tags you will see in a signature are d= (the signing domain), s= (the selector), h= (which headers are signed), bh= (the body hash), b= (the signature), and c= (the canonicalization, which is how much whitespace change the hashes tolerate).

You will not write the cryptography yourself. dkimpy does the signing and verifying, and dkim_lab.py wires it to the lab’s DNS layer through its dnsfunc hook. Two helpers are worth a look. signature_tags() reads the tags back out of a signed message. verify() collects dkimpy’s log messages and boils them down to a short reason, because dkimpy’s raw text includes long hash values.

# dkim_lab.py
"""DKIM signing and verification with dkimpy, wired to the lab's DNS layer."""
from __future__ import annotations

import logging
import re

import dkim

CRLF = "\r\n"
DEFAULT_BODY = "Please pay invoice 1042 to account 1234.\nThanks,\nAlice\n"


def build_message(*, sender="Alice <[email protected]>", to="[email protected]", subject="Invoice 1042",
                  reply_to=None, body=DEFAULT_BODY) -> bytes:
    headers = [("From", sender), ("To", to), ("Subject", subject),
               ("Date", "Fri, 09 Oct 2026 10:00:00 +0000"), ("Message-ID", "<[email protected]>")]
    if reply_to:
        headers.append(("Reply-To", reply_to))
    head = CRLF.join(f"{name}: {value}" for name, value in headers)
    return (head + CRLF + CRLF + body.replace("\n", CRLF)).encode()


def sign(message: bytes, domain: str, selector: str, pem: bytes, *, canon=(b"relaxed", b"relaxed"),
         headers=None, length=False) -> bytes:
    """Return the message with a DKIM-Signature header added at the top."""
    signer = dkim.DKIM(message)
    header = signer.sign(selector.encode(), domain.encode(), pem, canonicalize=canon,
                         include_headers=headers, length=length)
    return header + message


def _dns_func(resolver):
    def lookup(name, timeout=5):
        if isinstance(name, bytes):
            name = name.decode()
        answer = resolver.query(name.rstrip("."), "TXT")
        return "".join(answer.records).encode() if answer.records else None
    return lookup


def signature_tags(message: bytes) -> dict:
    """The tags of the first DKIM-Signature header (unfolded)."""
    head = message.split(b"\r\n\r\n", 1)[0].decode(errors="replace")
    match = re.search(r"(?im)^DKIM-Signature:([^\r\n]*(?:\r\n[ \t][^\r\n]*)*)", head)
    if not match:
        return {}
    text = re.sub(r"\s+", "", match.group(1))
    return dict(part.split("=", 1) for part in text.split(";") if "=" in part)


def verify(message: bytes, resolver):
    """Return (passed, why). `why` is the last complaint dkimpy logged, or an empty string."""
    complaints: list[str] = []

    class Collect(logging.Handler):
        def emit(self, record):
            complaints.append(record.getMessage())

    logger = logging.getLogger(f"dkim-lab-{id(complaints)}")
    logger.setLevel(logging.DEBUG)
    logger.propagate = False
    logger.addHandler(Collect())
    try:
        passed = bool(dkim.verify(message, logger=logger, dnsfunc=_dns_func(resolver)))
    except Exception as exc:  # dkimpy raises on malformed input in some paths
        return False, f"{type(exc).__name__}: {exc}"
    if passed:
        return True, ""
    if any(c.startswith("body hash mismatch") for c in complaints):
        return False, "body hash mismatch"
    if any("valid: False" in c for c in complaints):
        return False, "header signature does not verify"
    return False, complaints[-1] if complaints else "failed"

Tamper with a signed message

The experiment signs one message, then changes it in several ways that real mail systems and real attackers change messages, and checks the signature each time.

# step04_dkim_tamper.py
from dkim_lab import build_message, sign, signature_tags, verify
from fixtures import keypair, mail_resolver

pem, _ = keypair("shop")
dns = mail_resolver()
original = build_message(reply_to="[email protected]")
signed = sign(original, "example.test", "mail1", pem)

tags = signature_tags(signed)
print("signature tags:", {key: tags[key] for key in ("v", "a", "c", "d", "s", "h")})
print()


def edit(message, old, new):
    assert old.encode() in message, old
    return message.replace(old.encode(), new.encode(), 1)


def show(label, message):
    passed, why = verify(message, dns)
    print(f"{label:46} {'pass' if passed else 'FAIL'}  {why}")


show("untouched", signed)
show("body edited (account 1234 -> 9999)", edit(signed, "account 1234", "account 9999"))
show("Subject edited", edit(signed, "Subject: Invoice 1042", "Subject: Invoice 1042 URGENT"))
show("Reply-To edited (signed by default)", edit(signed, "Reply-To: [email protected]", "Reply-To: [email protected]"))
show("extra header added (X-Scanned-By)", b"X-Scanned-By: gateway\r\n" + signed)

print()
narrow = sign(original, "example.test", "mail1", pem, headers=[b"from", b"subject"])
print("signed headers:", signature_tags(narrow)["h"])
show("Reply-To edited, Reply-To not signed", edit(narrow, "Reply-To: [email protected]", "Reply-To: [email protected]"))

print()
for canon in (b"simple", b"relaxed"):
    message = sign(original, "example.test", "mail1", pem, canon=(canon, canon))
    show(f"{canon.decode()}/{canon.decode()}, a relay doubles a space", edit(message, "Please pay", "Please  pay"))

print()
limited = sign(original, "example.test", "mail1", pem, length=True)
print("with length=True the signature carries l=" + signature_tags(limited)["l"])
show("text appended to the body, l= signature", limited + b"Send the money to account 9999 instead.\r\n")
show("text appended to the body, no l=", signed + b"Send the money to account 9999 instead.\r\n")

Run python step04_dkim_tamper.py.

signature tags: {'v': '1', 'a': 'rsa-sha256', 'c': 'relaxed/relaxed', 'd': 'example.test', 's': 'mail1', 'h': 'from:to:subject:date:message-id:reply-to:from'}

untouched                                      pass  
body edited (account 1234 -> 9999)             FAIL  body hash mismatch
Subject edited                                 FAIL  header signature does not verify
Reply-To edited (signed by default)            FAIL  header signature does not verify
extra header added (X-Scanned-By)              pass  

signed headers: from:subject
Reply-To edited, Reply-To not signed           pass  

simple/simple, a relay doubles a space         FAIL  body hash mismatch
relaxed/relaxed, a relay doubles a space       pass  

with length=True the signature carries l=58
text appended to the body, l= signature        pass  
text appended to the body, no l=               FAIL  body hash mismatch

What the results mean

Body and signed headers are protected. Changing one number in the body fails the body hash, and editing the Subject or the signed Reply-To breaks the header signature. Look at the first line of output: dkimpy’s default h= list includes reply-to because the message has one, and it lists from at both ends. More on that in a moment.

Unsigned headers are not protected. Adding an X-Scanned-By header passes, which is right, because gateways add such headers all the time. The flip side is the next row. When you sign only from and subject, anyone can swap the Reply-To address and the signature still passes, so replies go to the attacker. Sign every header whose meaning matters: From, To, Subject, Date, Message-ID, and Reply-To. RFC 6376 insists on the first: “The From header field MUST be signed.”

Canonicalization decides what survives a relay. A relay that doubles a space breaks simple/simple but not relaxed/relaxed, which ignores changes in whitespace. Relaxed tolerates that kind of change, which makes it the safer choice for mail that may cross gateways.

The l= tag is a loophole. With length=True the signature carries l=58, the number of body bytes covered, and extra text appended after those 58 bytes passes. Without it, the same append fails. RFC 6376 describes l= as telling the verifier “the number of octets in the body of the email after canonicalization included in the cryptographic hash,” and its section on misuse of body length limits (8.2) warns that “using the ‘l=’ tag enables attacks in which an intermediary with malicious intent can modify a message to include content that solely benefits the attacker.” Leave it off.

The extra From header

One more attack deserves a look because it targets the same From: line that DMARC cares about. Mail software is often loose about malformed messages, and a message with two From: headers is one of them. DKIM covers only the instance the signature binds to, while a mail program may show a different one. RFC 6376 gives “showing only the first of multiple From: fields” as an example of that looseness. If a verifier does not defend against this, an attacker could take a genuinely signed message, add a second From: at the top, and the signature would still verify while the reader sees the attacker’s name. RFC 6376 section 8.15 warns that “an agent would be incorrect to infer that all instances of a header field are signed just because one is,” and recommends listing the field an extra time in h=, for example h=from:from:... for a message with one From field.

# step05_dkim_duplicate_from.py
import email
import email.policy

from dkim_lab import build_message, sign, signature_tags, verify
from fixtures import keypair, mail_resolver

pem, _ = keypair("shop")
dns = mail_resolver()
original = build_message()
fake_from = b"From: Boss <[email protected]>\r\n"


def inspect(label, message):
    passed, why = verify(message, dns)
    parsed = email.message_from_bytes(message, policy=email.policy.default)
    print(label)
    print(f"  DKIM: {'pass' if passed else 'FAIL'} {why}")
    print(f"  From headers in the message: {[str(v) for v in parsed.get_all('From')]}")
    print(f"  Python's parser reports From: {parsed['From']}")


plain = sign(original, "example.test", "mail1", pem, headers=[b"from", b"to", b"subject", b"date", b"message-id"])
print("h= when you pass a plain list:", signature_tags(plain)["h"])
inspect("plain list, a second From added at the top", fake_from + plain)

print()
default = sign(original, "example.test", "mail1", pem)
print("h= with dkimpy's defaults:", signature_tags(default)["h"])
inspect("dkimpy defaults, a second From added at the top", fake_from + default)

Run python step05_dkim_duplicate_from.py.

h= when you pass a plain list: from:to:subject:date:message-id
plain list, a second From added at the top
  DKIM: FAIL header signature does not verify
  From headers in the message: ['Boss <[email protected]>', 'Alice <[email protected]>']
  Python's parser reports From: Boss <[email protected]>

h= with dkimpy's defaults: from:to:subject:date:message-id:from
dkimpy defaults, a second From added at the top
  DKIM: FAIL header signature does not verify
  From headers in the message: ['Boss <[email protected]>', 'Alice <[email protected]>']
  Python's parser reports From: Boss <[email protected]>

The result is better than the attack. Both messages fail, even the one signed from a plain list with a single from. dkimpy defends on both ends: by default it signs from twice (see the end of its h= list), and its verifier adds an extra from to the headers it hashes. A comment in dkimpy’s source says that line addresses a reported bug about additional From headers. Python’s own parser, shown in the output, would report the attacker’s name, which is exactly why the defense matters. Do not assume every verifier does the same. Over-sign From on your side and the protection travels with the message.

Step 5: Add DMARC alignment and the DNS tree walk

DMARC takes the results of the two checks above and adds two things: a rule for which results count, and a policy for what receivers should do when none does. A message passes DMARC when at least one authenticated identifier is aligned with the Author Domain. RFC 9989 says it plainly: “If one or more of the Authenticated Identifiers align with the Author Domain, the message is considered to pass the DMARC mechanism check.” The authenticated identifiers are the domain SPF passed for and the d= domain of any DKIM signature that passed.

Alignment has two modes, chosen per record with aspf and adkim (both default to relaxed). In relaxed mode, an identifier aligns when it has the same Organizational Domain as the Author Domain: “the two are said to be in relaxed alignment.” In strict mode the two domains must be identical: “the two are said to be in strict alignment.” That raises the question of what an Organizational Domain is.

Why the tree walk replaced the Public Suffix List

The old rules (RFC 7489) answered with the Public Suffix List, a long list of the domain endings under which anyone can register names, such as .com or .co.uk. RFC 9989 replaces it with the DNS tree walk. It notes that RFC 7489 did not require any particular list and gave no guidance on how often to refresh it, so receivers could end up with different answers. In the tree walk, a receiver queries _dmarc at the Author Domain, then at each parent in turn, and stops early when it meets a record carrying psd=n (this domain is its own organization) or psd=y (this is a public suffix operator’s policy). The record at the name with the fewest labels wins, unless a psd flag says otherwise. To stop an attacker from forcing hundreds of queries with a very long domain name, the walk never makes more than eight: “a shortcut is built into the process so that Author Domains with more than eight labels do not result in more than eight DNS queries.”

RFC 9989 also lists what changed in the tags. It added np (policy for non-existent subdomains), psd, and t (a testing flag), and removed pct, rf, and ri. On pct it says “Operational experience showed that the ‘pct’ tag was usually not accurately applied.” With t=y, a receiver is asked to apply one level below the published policy: reject becomes quarantine, and quarantine becomes none. Unknown tags must be ignored, and a record whose first tag is not exactly v=DMARC1 is not a DMARC record at all.

Parse the record and walk the tree

# dmarc.py
"""DMARC as RFC 9989 describes it: record parsing, the DNS tree walk, alignment and the policy to apply."""
from __future__ import annotations

from dataclasses import dataclass, field

POLICIES = ("none", "quarantine", "reject")
LOWER = {"reject": "quarantine", "quarantine": "none", "none": "none"}  # what t=y asks receivers to apply


@dataclass
class DmarcRecord:
    raw: str
    tags: dict

    def get(self, tag, default=None):
        return self.tags.get(tag, default)


def parse_record(text):
    """Return a DmarcRecord, or None when the text is not one (v=DMARC1 must be the first tag)."""
    parts = [part.strip() for part in text.split(";")]
    key, _, value = parts[0].partition("=")
    if key.strip() != "v" or value.strip() != "DMARC1":
        return None
    tags = {}
    for part in parts[1:]:
        key, sep, value = part.partition("=")
        if sep and key.strip():
            tags.setdefault(key.strip().lower(), value.strip())
    return DmarcRecord(text, tags)


def records_at(name, resolver):
    answer = resolver.query("_dmarc." + name, "TXT")
    return [rec for rec in map(parse_record, answer.records) if rec]

records_at() returns every valid DMARC record at a name, because RFC 9989 says several records at one name are all discarded. walk_names() and walk() follow the procedure from section 4.10, and organizational_domain() applies the selection rules: a psd=n record makes that name the organization, a psd=y record on a parent makes the name one label below it the organization, and otherwise the record with the fewest labels wins. discover() then picks the policy record in the order the RFC lists: the Author Domain first, then its Organizational Domain, then the public suffix domain.

# dmarc.py (continued)
def walk_names(domain):
    """The names a DNS tree walk queries, longest first (RFC 9989 section 4.10: never more than 8)."""
    labels = domain.split(".")
    start = 1 if len(labels) < 8 else len(labels) - 7
    return [domain] + [".".join(labels[i:]) for i in range(start, len(labels))]


def walk(domain, resolver, multiple=None):
    """Return [(name, record)] for every name with exactly one valid record, stopping at psd=n or psd=y."""
    found = []
    for name in walk_names(domain):
        records = records_at(name, resolver)
        if len(records) > 1 and multiple is not None:
            multiple.append(name)  # several records at one name are all discarded
        if len(records) != 1:
            continue
        found.append((name, records[0]))
        if records[0].get("psd", "u").lower() in ("n", "y"):
            break
    return found


def organizational_domain(found, start):
    """Pick the Organizational Domain from a walk (RFC 9989 section 4.10.2)."""
    for name, record in found:
        flag = record.get("psd", "u").lower()
        if flag == "n":
            return name
        if flag == "y" and name != start:
            labels = start.split(".")
            return ".".join(labels[-(len(name.split(".")) + 1):])
    return found[-1][0] if found else start


def org_domain(name, resolver):
    return organizational_domain(walk(name, resolver), name)


@dataclass
class Discovery:
    author: str
    found: list
    org_domain: str
    policy_domain: str = ""
    record: DmarcRecord | None = None
    source: str = "none"
    multiple: list = field(default_factory=list)


def discover(author, resolver):
    """Find the policy record for an Author Domain (RFC 9989 section 4.10.1)."""
    author = author.lower().rstrip(".")
    multiple: list[str] = []
    found = walk(author, resolver, multiple)
    org = organizational_domain(found, author)
    by_name = dict(found)
    result = Discovery(author, found, org, multiple=multiple)
    if author in by_name:
        result.policy_domain, result.source = author, "the Author Domain"
    elif org in by_name:
        result.policy_domain, result.source = org, "the Organizational Domain"
    else:
        psd = next((name for name, rec in found if rec.get("psd", "u").lower() == "y"), None)
        if psd:
            result.policy_domain, result.source = psd, "the public suffix domain"
    if result.policy_domain:
        result.record = by_name[result.policy_domain]
    return result

Choose the policy and decide pass or fail

The last part picks which policy tag applies. A record at the Author Domain uses p. A record found higher up uses sp for a subdomain that exists and np for one that does not, and falls back to p when those are absent. evaluate() then checks SPF alignment and DKIM alignment, passes the message if either holds, and otherwise returns the policy, lowered one level if t=y is set. One more honest detail: the RFC says failing messages “are handled in accordance with the Mail Receiver’s local policies,” which may take the published policy into account “at the Mail Receiver’s discretion.” A DMARC policy is a request, and the code reports what the policy asks for.

# dmarc.py (continued)
def policy_for(discovery, author_exists):
    """The policy tag value that applies, or None when the record is unusable (section 4.10.1)."""
    record = discovery.record
    if discovery.policy_domain == discovery.author:
        chosen = record.get("p")
    elif author_exists:
        chosen = record.get("sp") or record.get("p")
    else:
        chosen = record.get("np") or record.get("sp") or record.get("p")
    chosen = (chosen or "").lower()
    if chosen in POLICIES:
        return chosen
    return "none" if "mailto:" in record.get("rua", "").lower() else None


def aligned(identifier, author, mode, resolver):
    """Identifier alignment: identical in strict mode, same Organizational Domain in relaxed mode."""
    if not identifier:
        return False
    identifier, author = identifier.lower().rstrip("."), author.lower().rstrip(".")
    if identifier == author:
        return True
    if mode.lower() == "s":
        return False
    return org_domain(identifier, resolver) == org_domain(author, resolver)


@dataclass
class Outcome:
    result: str  # "pass", "fail" or "none" (no DMARC record applies)
    disposition: str  # "accept", "none", "quarantine" or "reject"
    policy_domain: str = ""
    spf_aligned: bool = False
    dkim_aligned: bool = False
    note: str = ""


def evaluate(author, spf, dkim_results, resolver):
    """spf is (verdict, domain checked); dkim_results is [(d= domain, "pass" or "fail")]."""
    author = author.lower().rstrip(".")
    found = discover(author, resolver)
    if found.record is None:
        return Outcome("none", "accept", note="no DMARC record found by the tree walk")
    adkim, aspf = found.record.get("adkim", "r"), found.record.get("aspf", "r")
    spf_verdict, spf_domain = spf
    spf_ok = spf_verdict == "pass" and aligned(spf_domain, author, aspf, resolver)
    dkim_ok = any(res == "pass" and aligned(d, author, adkim, resolver) for d, res in dkim_results)
    if spf_ok or dkim_ok:
        return Outcome("pass", "accept", found.policy_domain, spf_ok, dkim_ok)
    exists = resolver.query(author, "A").status != "nxdomain"
    policy = policy_for(found, exists)
    if policy is None:
        return Outcome("none", "accept", found.policy_domain, note="the record has no usable policy")
    note = ""
    if found.record.get("t", "n").lower() == "y":
        policy, note = LOWER[policy], "t=y: one level below the published policy"
    return Outcome("fail", policy, found.policy_domain, False, False, note)

Test the discovery rules

This script runs the lookups for six Author Domains and prints where the policy came from, which domain counted as the organization, which policy applies, and how many DNS queries it took.

# step06_dmarc_discovery.py
import dmarc
from fixtures import mail_resolver

print(f"{'Author domain':42} {'policy record from':26} {'org domain':16} {'policy':10} queries")
for author in ["example.test", "news.example.test", "ghost.example.test", "news.eu.corp.test",
               "shop.bank.test", "a.b.c.d.e.f.g.h.i.j.mail.example.test"]:
    dns = mail_resolver()
    found = dmarc.discover(author, dns)
    queries = len(dns.log)
    exists = dns.query(author, "A").status != "nxdomain"
    policy = dmarc.policy_for(found, exists) if found.record else "-"
    print(f"{author:42} {found.policy_domain or '-':26} {found.org_domain:16} {policy:10} {queries}")

print()
dns = mail_resolver()
dmarc.discover("a.b.c.d.e.f.g.h.i.j.mail.example.test", dns)
print("names queried for the 13-label domain:")
for name, rtype in dns.log:
    print("  " + name)

Run python step06_dmarc_discovery.py.

Author domain                              policy record from         org domain       policy     queries
example.test                               example.test               example.test     reject     2
news.example.test                          example.test               example.test     quarantine 3
ghost.example.test                         example.test               example.test     reject     3
news.eu.corp.test                          eu.corp.test               eu.corp.test     quarantine 2
shop.bank.test                             bank.test                  shop.bank.test   reject     2
a.b.c.d.e.f.g.h.i.j.mail.example.test      example.test               example.test     reject     8

names queried for the 13-label domain:
  _dmarc.a.b.c.d.e.f.g.h.i.j.mail.example.test
  _dmarc.g.h.i.j.mail.example.test
  _dmarc.h.i.j.mail.example.test
  _dmarc.i.j.mail.example.test
  _dmarc.j.mail.example.test
  _dmarc.mail.example.test
  _dmarc.example.test
  _dmarc.test

Row by row: example.test has its own record, so p=reject applies. news.example.test exists but has no record, so the walk finds the organization’s record and uses sp=quarantine. ghost.example.test does not exist, so np=reject applies. news.eu.corp.test shows why the tree walk exists: eu.corp.test publishes its own record with psd=n, so it counts as a separate organization with its own p=quarantine, even though corp.test above it publishes p=reject. A Public Suffix List cannot express that. shop.bank.test falls to a public suffix operator’s record (psd=y at bank.test), and its Organizational Domain is the name one label below that record. The long name needs exactly eight queries, and the list printed under the table matches the RFC’s own example of an eight-query walk, with .test in place of .com.

The query count in the table is two for example.test because the walk also asks _dmarc.test. This code always finishes the walk, since it needs the Organizational Domain for alignment even when the Author Domain has its own record.

Nine delivery scenarios

Now put SPF, DKIM, and DMARC together. Each scenario sends a message from an IP address with an envelope domain, signs it or not, and prints all three results.

# step07_dmarc_scenarios.py
import dmarc
import spf
from dkim_lab import build_message, sign, signature_tags, verify
from fixtures import keypair, mail_resolver

dns = mail_resolver()
shop_pem, _ = keypair("shop")
esp_pem, _ = keypair("esp")
evil_pem, _ = keypair("evil")


def judge(label, author, ip, envelope_domain, message):
    check = spf.check_host(ip, envelope_domain, dns)
    passed, _ = verify(message, dns)
    signing_domain = signature_tags(message).get("d")
    dkim_results = [(signing_domain, "pass" if passed else "fail")] if signing_domain else []
    outcome = dmarc.evaluate(author, (check.verdict, envelope_domain), dkim_results, dns)
    dkim_text = f"{'pass' if passed else 'fail'} d={signing_domain}" if signing_domain else "none"
    print(f"{label:30} SPF {check.verdict:8} DKIM {dkim_text:22} -> DMARC {outcome.result:4} "
          f"{outcome.disposition:10} {outcome.note}")


own = sign(build_message(), "example.test", "mail1", shop_pem)
judge("A. the shop's own server", "example.test", "192.0.2.10", "example.test", own)

provider_default = sign(build_message(), "mailer.test", "mailer1", esp_pem)
judge("B1. provider signs as itself", "example.test", "198.51.100.7", "mailer.test", provider_default)

provider_aligned = sign(build_message(), "example.test", "esp1", esp_pem)
judge("B2. provider signs as the shop", "example.test", "198.51.100.7", "mailer.test", provider_aligned)

spoof = sign(build_message(sender="CEO <[email protected]>"), "evil.test", "k1", evil_pem)
judge("C. spoofer with its own SPF+DKIM", "example.test", "203.0.113.9", "evil.test", spoof)

judge("D. forwarded, DKIM intact", "example.test", "203.0.113.50", "example.test", own)

edited = own.replace(b"Subject: Invoice 1042", b"Subject: [list] Invoice 1042")
judge("E. mailing list edits Subject", "example.test", "203.0.113.70", "lists.test", edited)

news = build_message(sender="Shop <[email protected]>")
judge("F1. existing subdomain, unsigned", "news.example.test", "203.0.113.9", "news.example.test", news)
ghost = build_message(sender="Shop <[email protected]>")
judge("F2. non-existent subdomain", "ghost.example.test", "203.0.113.9", "ghost.example.test", ghost)

testing = build_message(sender="Shop <[email protected]>")
judge("G. policy reject with t=y", "testing.test", "203.0.113.9", "testing.test", testing)

Run python step07_dmarc_scenarios.py.

A. the shop's own server       SPF pass     DKIM pass d=example.test    -> DMARC pass accept     
B1. provider signs as itself   SPF pass     DKIM pass d=mailer.test     -> DMARC fail reject     
B2. provider signs as the shop SPF pass     DKIM pass d=example.test    -> DMARC pass accept     
C. spoofer with its own SPF+DKIM SPF pass     DKIM pass d=evil.test       -> DMARC fail reject     
D. forwarded, DKIM intact      SPF fail     DKIM pass d=example.test    -> DMARC pass accept     
E. mailing list edits Subject  SPF pass     DKIM fail d=example.test    -> DMARC fail reject     
F1. existing subdomain, unsigned SPF none     DKIM none                   -> DMARC fail quarantine 
F2. non-existent subdomain     SPF none     DKIM none                   -> DMARC fail reject     
G. policy reject with t=y      SPF none     DKIM none                   -> DMARC fail quarantine t=y: one level below the published policy

A, the shop’s own server. SPF passes for example.test and DKIM passes with d=example.test. Both are aligned, so DMARC passes.

B1, a provider that signs as itself. This is the most common real-world failure. The provider’s default is to use its own envelope domain and its own signing domain, so SPF and DKIM both pass, but for mailer.test, not for the shop. Nothing is aligned, so DMARC fails and the policy says reject. The shop’s own newsletter would be refused.

B2, the fix. The shop publishes a selector under its own name that holds the provider’s key (esp1._domainkey.example.test, which in practice you create with the CNAME or key your provider gives you). The provider now signs as example.test. DKIM aligns, DMARC passes, and SPF is still unaligned, which no longer matters.

C, a spoofer with valid SPF and DKIM of their own. The attacker owns evil.test, so SPF passes and DKIM passes, but both identify evil.test while the visible From: is [email protected]. Nothing aligns, DMARC fails, and the policy is reject. This single row is the reason DMARC exists. Passing SPF or DKIM tells you who authenticated, not whether that is who the reader thinks it is. RFC 9989 is careful about the limit too: the mechanisms “only validate the usage of a DNS domain in an email message” and do not validate the local part or the message’s content.

D, a forwarded message. A forwarder re-sends the message from its own address, so SPF fails for example.test. The message is unchanged, so the DKIM signature still verifies and is aligned. DMARC passes on DKIM alone. This is why you want both checks in place.

E, a mailing list that edits the subject. The list changes the Subject, which breaks DKIM, and its own servers pass SPF only for lists.test. Nothing aligns. A p=reject domain whose people post to mailing lists will see exactly this failure. RFC 9989 has sections titled “Interoperability Issues” and “Interoperability Considerations” for these indirect mail flows.

F1 and F2, subdomains. Unsigned mail from an existing subdomain gets the organization’s sp=quarantine. Mail from a subdomain that does not exist gets np=reject. Non-existent subdomains get their own tag, np, in RFC 9989.

G, a policy in test mode. The record says p=reject; t=y, so the failing message is treated one level lower, as quarantine. That is how t=y lets a domain owner try a stricter policy while watching reports.

Step 6: Turn it into an auditor

The pieces are ready to become a tool. audit.py takes a domain, finds its SPF, DMARC, and DKIM records through the same code, and lists findings sorted by severity. The first half checks SPF and DMARC.

# audit.py
"""Audit the email authentication DNS of a domain: SPF, DMARC and DKIM selectors."""
from __future__ import annotations

import argparse
import base64
import sys
from dataclasses import dataclass

from cryptography.hazmat.primitives import serialization
from cryptography.hazmat.primitives.asymmetric import rsa

import dmarc
import spf
from dnsio import DnsTemporaryError, LiveResolver

RANK = {"high": 0, "medium": 1, "low": 2, "info": 3}
STRENGTH = {"none": 0, "quarantine": 1, "reject": 2}


@dataclass
class Finding:
    severity: str
    code: str
    message: str


def audit_spf(domain, resolver):
    report = spf.analyze(domain, resolver)
    if report.problem == "none":
        return [Finding("medium", "SPF-1", "No SPF record. Receivers get the result none for every sender.")]
    if report.problem == "multiple":
        return [Finding("high", "SPF-2", "More than one SPF record. RFC 7208 makes that a permerror, so SPF fails for everyone.")]
    if report.problem == "syntax":
        return [Finding("high", "SPF-3", f"The record does not parse ({report.error}), which is a permerror.")]
    out = []
    if report.record.lower().split() == ["v=spf1", "-all"]:
        return [Finding("info", "SPF-11", "A null SPF record (v=spf1 -all): the domain says it sends no mail.")]
    if report.error:
        out.append(Finding("high", "SPF-4", f"Part of the include tree is broken: {report.error}."))
    if report.all_qualifier == "+":
        out.append(Finding("high", "SPF-5", "The record ends in +all, so every IP address passes SPF."))
    elif report.all_qualifier == "?":
        out.append(Finding("medium", "SPF-6", "The record ends in ?all (neutral), which authorizes and rejects nothing."))
    elif report.all_qualifier is None:
        out.append(Finding("medium", "SPF-6", "No all mechanism and no redirect, so unlisted senders get neutral."))
    elif report.all_qualifier == "~":
        out.append(Finding("info", "SPF-7", "~all (softfail) is common while a DMARC policy does the enforcing."))
    if report.lookups > 10:
        out.append(Finding("high", "SPF-8", f"{report.lookups} DNS-querying terms in the tree; the limit is 10, so some senders get permerror."))
    elif report.lookups >= 8:
        out.append(Finding("low", "SPF-8", f"{report.lookups} of 10 allowed DNS lookups are already used."))
    if report.voids > 2:
        out.append(Finding("medium", "SPF-9", f"{report.voids} void lookups; RFC 7208 suggests a limit of 2."))
    if report.uses_ptr:
        out.append(Finding("low", "SPF-10", "The ptr mechanism is slow and unreliable, and RFC 7208 says not to use it."))
    return out


def audit_dmarc(domain, resolver):
    found = dmarc.discover(domain, resolver)
    if found.record is None:
        extra = f" (several records at {', '.join(found.multiple)} were discarded)" if found.multiple else ""
        return [Finding("high", "DMARC-1", f"No usable DMARC record{extra}. Anyone can use this domain in From.")]
    record, out = found.record, []
    policy = record.get("p", "").lower()
    if policy not in dmarc.POLICIES:
        out.append(Finding("medium", "DMARC-2", "No valid p tag, so receivers treat the record as p=none."))
    elif policy == "none":
        out.append(Finding("medium", "DMARC-3", "p=none only monitors. Spoofed mail is still delivered."))
    elif policy == "quarantine":
        out.append(Finding("info", "DMARC-4", "p=quarantine asks for the spam folder; reject is stronger once reports are clean."))
    for tag in ("sp", "np"):
        value = record.get(tag, "").lower()
        if value in STRENGTH and policy in STRENGTH and STRENGTH[value] < STRENGTH[policy]:
            out.append(Finding("medium", "DMARC-5", f"{tag}={value} is weaker than p={policy}, so subdomains are the easy way in."))
    if "mailto:" not in record.get("rua", "").lower():
        out.append(Finding("low", "DMARC-6", "No rua tag, so you receive no aggregate reports."))
    if record.get("t", "n").lower() == "y":
        out.append(Finding("info", "DMARC-7", "t=y asks receivers to apply the policy one level lower."))
    legacy = [tag for tag in ("pct", "rf", "ri") if tag in record.tags]
    if legacy:
        out.append(Finding("info", "DMARC-8", f"Legacy tag(s) {', '.join(legacy)} were removed by RFC 9989; receivers may ignore them."))
    return out

Every rule is a plain comparison on the data you already collect. SPF gets high findings for a doubled record, a parse error, a broken include, +all, or more than ten lookups, and medium for a missing record, ?all, a missing all, or too many void lookups. DMARC gets high when no record can be found, medium for p=none and for subdomain policies weaker than p, low for no rua reporting address, and info for t=y and for the removed legacy tags. The second half checks DKIM selectors and adds the command line.

# audit.py (continued)
def tag_values(text):
    pairs = (part.partition("=") for part in text.split(";"))
    return {key.strip().lower(): value.strip() for key, sep, value in pairs if sep}


def audit_dkim(domain, selectors, resolver):
    out = []
    for selector in selectors:
        name = f"{selector}._domainkey.{domain}"
        text = "".join(resolver.query(name, "TXT").records)
        if not text:
            out.append(Finding("medium", "DKIM-1", f"{name} has no key record."))
            continue
        tags = tag_values(text)
        key_b64 = tags.get("p", "").replace(" ", "")
        if not key_b64:
            out.append(Finding("info", "DKIM-2", f"{name} has an empty p= tag, which means the key was revoked."))
            continue
        try:
            key = serialization.load_der_public_key(base64.b64decode(key_b64))
        except ValueError:
            out.append(Finding("high", "DKIM-3", f"{name}: the p= value is not a public key."))
            continue
        if isinstance(key, rsa.RSAPublicKey):
            bits = key.key_size
            if bits < 1024:
                out.append(Finding("high", "DKIM-4", f"{name}: a {bits}-bit RSA key is below the 1024-bit minimum."))
            elif bits < 2048:
                out.append(Finding("medium", "DKIM-4", f"{name}: a {bits}-bit RSA key; 2048 bits is the usual advice."))
            else:
                out.append(Finding("info", "DKIM-4", f"{name}: a {bits}-bit RSA key."))
        if "y" in tags.get("t", "").lower().split(":"):
            out.append(Finding("low", "DKIM-5", f"{name} has t=y, which marks the domain as testing DKIM."))
    return out


def audit_domain(domain, resolver, selectors=()):
    domain = domain.lower().rstrip(".")
    sections = {"SPF": audit_spf(domain, resolver), "DMARC": audit_dmarc(domain, resolver),
                "DKIM": audit_dkim(domain, selectors, resolver)}
    return sorted(((area, f) for area, items in sections.items() for f in items),
                  key=lambda pair: RANK[pair[1].severity])


def main(argv=None):
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("domain", nargs="+")
    parser.add_argument("--selector", action="append", default=[], help="DKIM selector to check (repeatable)")
    parser.add_argument("--fail-on", choices=list(RANK), default="high", help="exit 1 if any finding is this bad or worse")
    args = parser.parse_args(argv)
    resolver, worst = LiveResolver(), 99
    for domain in args.domain:
        print(domain)
        try:
            results = audit_domain(domain, resolver, args.selector)
        except DnsTemporaryError as exc:
            print(f"  DNS error: {exc}")
            worst = 0
            continue
        for area, finding in results:
            print(f"  {finding.severity.upper():6} {area:5} {finding.code:9} {finding.message}")
            worst = min(worst, RANK[finding.severity])
        if not results:
            print("  no findings")
    return 1 if worst <= RANK[args.fail_on] else 0


if __name__ == "__main__":
    sys.exit(main())

The DKIM rules read each key record, decode the p= value, and report the RSA key size: below 1024 bits is high, 1024 up to 2047 is medium, and 2048 or more is info. An empty p= is reported as a revoked key, and t=y as a domain still testing DKIM. The command exits with status 1 when any finding is as bad as --fail-on (default high), so you can drop it into a CI job.

Audit four invented domains

First run the auditor against the pretend zone, where the answers are known in advance. good.test is configured well, sloppy.test has eleven lookups, a neutral ?all, p=none, and a weak test key, legacy.test has two SPF records and a legacy tag, and parked.test is a domain that sends no mail.

# step08_audit_fixtures.py
from audit import audit_domain
from dnsio import ZoneResolver
from fixtures import audit_zone

dns = ZoneResolver(audit_zone())
for domain, selectors in [("good.test", ["s1"]), ("sloppy.test", ["s1"]), ("legacy.test", ["old"]),
                          ("parked.test", [])]:
    print(domain)
    for area, finding in audit_domain(domain, dns, selectors):
        print(f"  {finding.severity.upper():6} {area:5} {finding.code:9} {finding.message}")

Run python step08_audit_fixtures.py.

good.test
  INFO   DKIM  DKIM-4    s1._domainkey.good.test: a 2048-bit RSA key.
sloppy.test
  HIGH   SPF   SPF-8     11 DNS-querying terms in the tree; the limit is 10, so some senders get permerror.
  MEDIUM SPF   SPF-6     The record ends in ?all (neutral), which authorizes and rejects nothing.
  MEDIUM DMARC DMARC-3   p=none only monitors. Spoofed mail is still delivered.
  MEDIUM DKIM  DKIM-4    s1._domainkey.sloppy.test: a 1024-bit RSA key; 2048 bits is the usual advice.
  LOW    DMARC DMARC-6   No rua tag, so you receive no aggregate reports.
  LOW    DKIM  DKIM-5    s1._domainkey.sloppy.test has t=y, which marks the domain as testing DKIM.
legacy.test
  HIGH   SPF   SPF-2     More than one SPF record. RFC 7208 makes that a permerror, so SPF fails for everyone.
  MEDIUM DMARC DMARC-5   sp=none is weaker than p=quarantine, so subdomains are the easy way in.
  MEDIUM DKIM  DKIM-1    old._domainkey.legacy.test has no key record.
  INFO   DMARC DMARC-4   p=quarantine asks for the spam folder; reject is stronger once reports are clean.
  INFO   DMARC DMARC-8   Legacy tag(s) pct were removed by RFC 9989; receivers may ignore them.
parked.test
  LOW    DMARC DMARC-6   No rua tag, so you receive no aggregate reports.
  INFO   SPF   SPF-11    A null SPF record (v=spf1 -all): the domain says it sends no mail.

Check: good.test should show nothing above info, and sloppy.test should show a HIGH SPF-8 finding for eleven lookups. The count comes from nine includes plus a and mx. The parked.test output is the shape you want for a domain that never sends mail: a null SPF record, strict alignment, and p=reject. The only note is the missing reporting address.

Audit real domains

Now point the same code at the real Internet. The auditor reads DNS only, so this is safe to run against any domain.

python audit.py example.com google.com gmail.com
python audit.py github.com --selector s1
python audit.py gmail.com --selector 20230601

On October 9, 2026 the three runs printed this:

example.com
  LOW    DMARC DMARC-6   No rua tag, so you receive no aggregate reports.
  INFO   SPF   SPF-11    A null SPF record (v=spf1 -all): the domain says it sends no mail.
google.com
  INFO   SPF   SPF-7     ~all (softfail) is common while a DMARC policy does the enforcing.
gmail.com
  MEDIUM DMARC DMARC-3   p=none only monitors. Spoofed mail is still delivered.
  INFO   SPF   SPF-7     ~all (softfail) is common while a DMARC policy does the enforcing.
github.com
  LOW    SPF   SPF-8     10 of 10 allowed DNS lookups are already used.
  INFO   SPF   SPF-7     ~all (softfail) is common while a DMARC policy does the enforcing.
  INFO   DMARC DMARC-4   p=quarantine asks for the spam folder; reject is stronger once reports are clean.
  INFO   DMARC DMARC-8   Legacy tag(s) pct were removed by RFC 9989; receivers may ignore them.
  INFO   DKIM  DKIM-4    s1._domainkey.github.com: a 2048-bit RSA key.
gmail.com
  MEDIUM DMARC DMARC-3   p=none only monitors. Spoofed mail is still delivered.
  INFO   SPF   SPF-7     ~all (softfail) is common while a DMARC policy does the enforcing.
  INFO   DKIM  DKIM-2    20230601._domainkey.gmail.com has an empty p= tag, which means the key was revoked.

Treat these as mechanical findings about DNS records, not as grades. example.com shows the null-SPF shape plus a missing reporting address. google.com shows only the informational note about ~all. gmail.com publishes p=none, which the auditor flags as monitoring only; that describes one record, not Google’s whole defense. github.com sits at the lookup limit, which is legal but leaves no room, and its record still carries the old pct tag that RFC 9989 removed. The Gmail selector is reported as revoked, matching the empty p= you saw in Step 1.

Use it as a CI gate

Because the exit code reflects the worst finding, a scheduled job can fail when a record regresses. Asking for --fail-on medium against gmail.com exits with status 1, because of its p=none:

gmail.com
  MEDIUM DMARC DMARC-3   p=none only monitors. Spoofed mail is still delivered.
  INFO   SPF   SPF-7     ~all (softfail) is common while a DMARC policy does the enforcing.

[exit code 1]

On macOS or Linux, echo $? prints the status after a run; in PowerShell it is $LASTEXITCODE.

Step 7: Test it

Everything above ran as demonstrations. Tests keep the rules from drifting when you change a regular expression later. The file below checks the include semantics, the limits, the DKIM tamper cases, the tree walk, alignment, the policy choice, and the auditor, all against the pretend Internet so no test needs a network.

# test_emailauth.py
import dmarc
import spf
from audit import audit_domain
from dkim_lab import build_message, sign, signature_tags, verify
from dnsio import ZoneResolver
from fixtures import audit_zone, keypair, mail_resolver


def verdict(zone, ip="203.0.113.9", **kwargs):
    return spf.check_host(ip, "t.test", ZoneResolver(zone, **kwargs))


def includes(count):
    zone = {"t.test": {"TXT": ["v=spf1 " + " ".join(f"include:i{n}.test" for n in range(1, count + 1))
                               + " ip4:203.0.113.0/24 -all"]}}
    zone.update({f"i{n}.test": {"TXT": ["v=spf1 -all"]} for n in range(1, count + 1)})
    return zone


def test_include_does_not_inherit_the_inner_all():
    zone = {"t.test": {"TXT": ["v=spf1 include:p.test ip4:203.0.113.0/24 -all"]}, "p.test": {"TXT": ["v=spf1 -all"]}}
    assert verdict(zone).verdict == "pass"


def test_two_spf_records_are_a_permerror():
    assert verdict({"t.test": {"TXT": ["v=spf1 -all", "v=spf1 +all"]}}).verdict == "permerror"


def test_lookup_limit_is_ten():
    assert verdict(includes(10)).verdict == "pass"
    assert verdict(includes(11)).verdict == "permerror"


def test_void_lookup_limit_is_two():
    zone = {"t.test": {"TXT": ["v=spf1 a:x1.test a:x2.test a:x3.test +all"]}}
    result = verdict(zone)
    assert (result.verdict, result.voids) == ("permerror", 3)


def test_dns_failure_is_a_temperror():
    zone = {"t.test": {"TXT": ["v=spf1 include:down.test -all"]}}
    assert verdict(zone, broken=["down.test"]).verdict == "temperror"


def test_a_with_cidr_and_ipv6():
    zone = {"t.test": {"TXT": ["v=spf1 a:www.t.test/24 ip6:2001:db8::/32 -all"]}, "www.t.test": {"A": ["203.0.113.9"]}}
    assert verdict(zone, ip="203.0.113.200").verdict == "pass"
    assert verdict(zone, ip="2001:db8::1").verdict == "pass"
    assert verdict(zone, ip="2001:db9::1").verdict == "fail"


def test_redirect_and_macros():
    zone = {"t.test": {"TXT": ["v=spf1 redirect=p.test"]}, "p.test": {"TXT": ["v=spf1 ip4:203.0.113.0/24 -all"]}}
    assert verdict(zone).verdict == "pass"
    assert verdict({"t.test": {"TXT": ["v=spf1 exists:%{i}.t.test -all"]}}).verdict == "permerror"


def test_analyze_counts_the_whole_tree():
    report = spf.analyze("sloppy.test", ZoneResolver(audit_zone()))
    assert (report.lookups, len(report.includes), report.all_qualifier) == (11, 9, "?")


def test_dkim_round_trip_and_body_tamper():
    pem, _ = keypair("shop")
    dns = mail_resolver()
    signed = sign(build_message(), "example.test", "mail1", pem)
    assert verify(signed, dns) == (True, "")
    assert verify(signed.replace(b"1234", b"9999"), dns) == (False, "body hash mismatch")


def test_dkim_ignores_unsigned_headers():
    pem, _ = keypair("shop")
    dns = mail_resolver()
    message = build_message(reply_to="[email protected]")
    narrow = sign(message, "example.test", "mail1", pem, headers=[b"from", b"subject"])
    assert verify(narrow.replace(b"Reply-To: alice@", b"Reply-To: thief@"), dns)[0]
    full = sign(message, "example.test", "mail1", pem)
    assert not verify(full.replace(b"Reply-To: alice@", b"Reply-To: thief@"), dns)[0]


def test_extra_from_header_breaks_the_signature():
    pem, _ = keypair("shop")
    dns = mail_resolver()
    signed = sign(build_message(), "example.test", "mail1", pem)
    assert signature_tags(signed)["h"].endswith(":from")
    assert not verify(b"From: Boss <[email protected]>\r\n" + signed, dns)[0]


def test_tree_walk_never_makes_more_than_eight_queries():
    dns = mail_resolver()
    dmarc.discover("a.b.c.d.e.f.g.h.i.j.mail.example.test", dns)
    assert len(dns.log) == 8


def test_psd_flags_decide_the_organizational_domain():
    dns = mail_resolver()
    assert dmarc.org_domain("news.eu.corp.test", dns) == "eu.corp.test"  # psd=n
    assert dmarc.org_domain("shop.bank.test", dns) == "shop.bank.test"  # psd=y one label up
    assert dmarc.org_domain("news.example.test", dns) == "example.test"  # fewest labels


def test_relaxed_and_strict_alignment():
    dns = mail_resolver()
    assert dmarc.aligned("mail.example.test", "example.test", "r", dns)
    assert not dmarc.aligned("mail.example.test", "example.test", "s", dns)
    assert not dmarc.aligned("evil.test", "example.test", "r", dns)


def test_policy_comes_from_p_sp_or_np():
    dns = mail_resolver()

    def chosen(author):
        found = dmarc.discover(author, dns)
        return dmarc.policy_for(found, dns.query(author, "A").status != "nxdomain")

    assert chosen("example.test") == "reject"
    assert chosen("news.example.test") == "quarantine"
    assert chosen("ghost.example.test") == "reject"


def test_multiple_dmarc_records_are_all_discarded():
    zone = {"d.test": {"A": ["192.0.2.1"]}, "_dmarc.d.test": {"TXT": ["v=DMARC1; p=none", "v=DMARC1; p=reject"]}}
    found = dmarc.discover("d.test", ZoneResolver(zone))
    assert found.record is None and found.multiple == ["d.test"]


def test_attacker_with_valid_spf_and_dkim_still_fails():
    dns = mail_resolver()
    outcome = dmarc.evaluate("example.test", ("pass", "evil.test"), [("evil.test", "pass")], dns)
    assert (outcome.result, outcome.disposition) == ("fail", "reject")


def test_dkim_alone_saves_a_forwarded_message():
    dns = mail_resolver()
    outcome = dmarc.evaluate("example.test", ("fail", "example.test"), [("example.test", "pass")], dns)
    assert (outcome.result, outcome.dkim_aligned) == ("pass", True)


def test_t_y_lowers_the_policy_one_level():
    dns = mail_resolver()
    assert dmarc.evaluate("testing.test", ("none", "testing.test"), [], dns).disposition == "quarantine"


def test_auditor_findings():
    dns = ZoneResolver(audit_zone())
    sloppy = {f.code for _, f in audit_domain("sloppy.test", dns, ["s1"])}
    assert {"SPF-8", "DMARC-3", "DKIM-4"} <= sloppy
    good = {f.severity for _, f in audit_domain("good.test", dns, ["s1"])}
    assert good <= {"info"}

Run python -m pytest -q.

.................... [100%]
20 passed in 0.95s

Check: you should see 20 passed. If a DKIM test fails with a missing file or key error, delete the keys folder and run any step once to regenerate the lab keys.

Common mistakes and how to avoid them

Treating “SPF pass” as proof of the sender

Scenario C is the lesson. SPF and DKIM prove that a domain authenticated the message, and DMARC alignment is what connects that domain to the address the reader sees. Always judge the three together.

Publishing a second SPF record

Adding a vendor is a merge, not an append. Edit the one record, then run the auditor, which reports the doubled record as a permerror.

Checking the lookup count only once

An include you do not control can grow. The count in analyze() is a worst case across the whole tree, so run it on a schedule, and keep some headroom below ten.

Going straight to p=reject

Start at p=none with a rua address, read the reports, fix the senders that fail (scenario B1 and its fix), and only then move to quarantine and reject. Use t=y as the careful middle step, and expect mailing lists and some forwarders to need attention before enforcement (scenario E).

Forgetting domains that send nothing

Domains and subdomains that never send mail still appear in From: lines, so give them a policy too. Publish a null SPF record and a DMARC record with p=reject for them, as example.com does, and set sp and np on the domains that do send.

Weak or unmanaged DKIM

Use 2048-bit keys (Google says sending to personal Gmail accounts “requires a DKIM key of 1024 bits or longer” and recommends 2048 bits if your domain provider supports it), sign the headers that matter (and From twice), avoid l=, and rotate selectors. Revoke an old key by publishing an empty p=, which is what the Gmail selector in Step 1 shows.

Trusting old advice

Much online guidance still describes pct, the Public Suffix List, and RFC 7489. RFC 9989 says an implementation based on the older RFC and a PSL “might arrive at a different answer” in some decentralized setups and notes that the problem “is entirely avoided by the use of strict alignment and publishing explicit DMARC Policy Records for all Author Domains used in an organization’s email.” Check what your receivers actually do by reading your aggregate reports.

Forgetting that DNS itself can be forged

All three checks are only as trustworthy as the DNS answers they read. DNSSEC protects those answers, and our post on the DNSSEC root key rollover explains what a resolver needs for it to work.

Confirm it works end to end

Start in an empty folder with a fresh virtual environment. Create the files from this tutorial, then run the eight step scripts in order, python -m pytest -q, and the audit commands. Success looks like this: step02_spf_basic.py gives pass, pass, fail; step03_spf_traps.py shows the ten verdicts above; step04_dkim_tamper.py fails the tampered messages and passes the untouched one and the one with an extra unsigned header; step05_dkim_duplicate_from.py fails both extra-From messages; step06_dmarc_discovery.py shows eight queries for the long name; step07_dmarc_scenarios.py shows pass for A, B2, and D and fail for B1, C, E, F1, F2, and G; step08_audit_fixtures.py shows the high findings for sloppy.test and legacy.test; and pytest ends with 20 passed. Only the live outputs, from step01_live_records.py and the audit.py runs against real domains, should differ from what you see here.

Where to go next

The toolkit leaves out several things on purpose, and each is a good next exercise: SPF macros and the exp modifier, the HELO identity used when the envelope sender is empty, the ptr mechanism, Ed25519 DKIM keys, and the reports. DMARC’s reporting half now lives in two companion documents that RFC 9989 points to, RFC 9990 for aggregate reports and RFC 9991 for failure reports. Parsing a day of aggregate reports from your rua mailbox is the natural way to find the senders that scenario B1 hides. If you like building verifiers, our tutorial on verifying webhook signatures uses the same habit of checking exactly what was signed, and our write-up of a phishing campaign that hid text from filters shows the kind of mail these checks help you reject.

Sources used: RFC 7208 (SPF), RFC 6376 (DKIM), RFC 9989 (DMARC), and Google’s email sender guidelines. DNS records were read live on October 9, 2026.

Tags:

DKIMDMARCEmail SecurityPhishingPythonSPF

Share

Wooden two-dial chess clock with brass-rimmed white faces showing different times
Previous Post

A CNCF Post on NIS2 and DORA Turns Compliance Into a Backlog and Leaves the Classification Call Unowned

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
09 Oct
How to Audit SPF, DKIM, and DMARC in Python to Stop Spoofed Email From Using Your Domain
09 Oct
A CNCF Post on NIS2 and DORA Turns Compliance Into a Backlog and Leaves the Classification Call Unowned
Trending
October 9, 2026
How to Audit SPF, DKIM, and DMARC in Python to Stop Spoofed Email From Using Your Domain
October 9, 2026
A CNCF Post on NIS2 and DORA Turns Compliance Into a Backlog and Leaves the Classification Call Unowned
October 9, 2026
Anthropic Launches OSS Scanner to Email Open-Source Maintainers AI Bug Reports No Human Has Reviewed
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026