How to Audit Python .pth Startup Hooks and Migrate to Python 3.15 .start Files
Learn how Python .pth import lines run code at every startup, build a standard-library auditor that catches hidden lines and tampering, check wheels before installing them, and migrate to Python 3.15...
Every time Python starts, it does a little housekeeping before your first line of code runs. One piece of that housekeeping is reading small text files that end in .pth from the site-packages folder, the place where installed packages live. Most lines in those files are just folder names that Python adds to its module search path. But a line that starts with import followed by a space or a tab is not a folder name. Python treats it as code and runs it. So a package you install can run code every time any Python program starts in that environment, even if no program ever imports the package.
Table Of Content
- A few terms first
- Prerequisites
- Step 1: Watch a harmless startup hook run
- Write a small helper module
- Create the environment and the hook
- What the output tells you
- Why did everything print twice?
- Check that it worked
- Step 2: Learn the exact rules Python uses to read a .pth file
- The rules, side by side
- Step 3: Build the auditor
- Part 1: parse a startup file the way site.py does
- Part 2: who owns a file, and has it changed?
- Part 3: the decision ladder
- Part 4: the command line
- Common mistake: auditing an environment with its own Python
- Step 4: Audit a realistic environment
- Reading the report
- See tampering get caught
- Step 5: Check a wheel before you install it
- Build a demo wheel
- Write the wheel checker
- Check, stage and install
- The same check on real wheels
- Step 6: See what popular packages ship
- What the survey says, and what it does not
- Step 7: Try Python 3.15 .start files
- Scenario A: an import line only
- Scenario B: the migration
- Scenario C: a .start file alone
- Common mistake: shipping only a .start file
- Scenario D: mistakes and duplicates
- Scenario E: a stale path line
- Scenario F: two lines the versions read differently
- Scenario G: can an audit hook see any of this?
- What .start files change, and what they do not
- Step 8: Turn the audit into a gate and test it
- Test the auditor
- What this auditor does not cover
- Confirm it all works end to end
- Where to go next
That is not a thought experiment. LiteLLM’s own incident report says version 1.82.8 of its package “contained litellm_init.pth and a malicious payload in the LiteLLM AI Gateway proxy_server.py”, and that the compromised releases were “live on March 24, 2026 from 10:39 UTC for about 40 minutes before being quarantined by PyPI”. The researchers at FutureSearch who found the release wrote that the file “executes automatically on every Python process startup when litellm is installed in the environment”. Datadog Security Labs, which traced the incident to a campaign it calls TeamPCP, spelled out the practical difference: a victim of version 1.82.7 ran the payload only after using the compromised package in an application, while 1.82.8 launched it whenever the interpreter started. We covered the group behind that campaign in Google’s TeamPCP Mole Turns Threat Intelligence Into a Spy Operation. This tutorial is the defensive side: learning what your own Python environments run at startup, and how to keep an eye on it.
By the end you will have done seven concrete things. You will watch a harmless startup hook run and see which command-line options switch it off. You will learn the exact rules Python uses to read a .pth file, including a trick that hides an import line behind a comment. You will build a standard-library auditor, pthaudit.py, that reports which package owns each startup file, whether it changed since install, and which lines look risky. You will check a wheel before installing it. You will survey what the 1,000 most downloaded packages on PyPI ship. You will try the new Python 3.15 .start files from PEP 829. And you will turn the auditor into an allowlist gate with tests.
I ran every command in this article on Windows 11 with Python 3.13.14 and, for the 3.15 steps, Python 3.15.0rc3. At the time of writing, release candidate 3 is the newest 3.15 build, and the schedule in PEP 790 lists the final 3.15.0 for Friday, October 9, 2026, so small details could still shift. Where a result depends on the Python version, I say so. All outputs below were captured on October 6, 2026.
A few terms first
Interpreter startup is the short stretch between launching python and running your code. During it, Python imports the site module, which sets up the module search path. site-packages is the folder where pip installs packages; every virtual environment has its own. A .pth file is a text file in that folder, read at startup. An import line is a line in a .pth file that starts with import and a space or a tab; Python executes it. A wheel is the zip-based file format that pip normally installs. Every installed package has a .dist-info folder, and inside it a RECORD file listing every file the installer wrote, with a hash of each. An entry point is a name that points at a function, written package.module:function. Python 3.15 uses these in new .start files, which we meet in Step 7.
Prerequisites
- Python 3.12 or newer for every step except Step 7. I ran 3.13.14, and I read the
site.pysource on the 3.12 and 3.14 branches of CPython, which follow the same rules (3.10 differs, see Step 2). - For Step 7 only, a Python 3.15 build. Release candidate 3 is on the python.org release page; if you read this after October 9, install the final release instead.
- A terminal. The commands are the same on Windows, macOS and Linux unless I say otherwise; I used Windows 11, where the virtual environment folder is
Lib\site-packages. On macOS and Linux it islib/python3.x/site-packages, and the scripts work that out for you. - An internet connection for
pipand for the survey in Step 6, and about 200 MB of disk space. Nothing touches your real Python installation: every experiment lives in a throwaway virtual environment. - Basic comfort running Python scripts and creating a virtual environment with
python -m venv.
Create a folder called pth-lab, open a terminal in it, and install the two packages the lab needs:
python -m pip install pytest requests
Every code block below starts with a comment that names its file. Save each block under that name in pth-lab. A block whose first line ends in “(continued)” is the next part of the same file. The scripts print <lab> in place of the full path to your pth-lab folder so the output reads the same on every machine.
Step 1: Watch a harmless startup hook run
The fastest way to understand a mechanism is to trigger it yourself. We will make a virtual environment, drop a one-line .pth file into it, and start Python in several different ways.
Write a small helper module
All the lab scripts share a few helpers, so put them in one file. The two that matter most are clean_env, which removes every PYTHON* environment variable so your shell settings cannot change what we observe, and site_packages, which finds a virtual environment’s site-packages folder by looking at the file system. That second point is deliberate: asking an environment’s own Python for its paths would start it, and starting it would run the very hooks we want to study.
# labenv.py
"""Small helpers shared by the lab scripts."""
import os
import pathlib
import subprocess
def clean_env():
"""The current environment without PYTHON* variables, so shell settings cannot change what we observe."""
env = {k: v for k, v in os.environ.items() if not k.startswith("PYTHON")}
env["PIP_DISABLE_PIP_VERSION_CHECK"] = "1"
return env
def venv_python(venv):
venv = pathlib.Path(venv)
return venv / ("Scripts/python.exe" if os.name == "nt" else "bin/python")
def site_packages(venv):
"""The venv's site-packages folder, found by looking at the file system (nothing is executed)."""
venv = pathlib.Path(venv)
if os.name == "nt":
return venv / "Lib" / "site-packages"
return sorted(venv.glob("lib/python*/site-packages"))[0]
def run(python, *args, **kwargs):
"""Run an interpreter or any program and capture its text output."""
return subprocess.run([str(python), *args], capture_output=True, text=True,
encoding="utf-8", errors="replace", env=clean_env(), **kwargs)
def show(title, proc, keep=None):
"""Print a command title, then its output lines (all of them, or only those containing a keyword in `keep`).
The lab folder is printed as <lab> so the output reads the same on every machine.
"""
cwd = os.getcwd()
print(f"$ {title}")
for line in (proc.stderr + proc.stdout).splitlines():
line = line.replace(cwd, "<lab>").replace(cwd.replace("\\", "\\\\"), "<lab>")
if line.strip() and (keep is None or any(word in line for word in keep)):
print(" " + line)
if proc.returncode:
print(f" (exit code {proc.returncode})")
Create the environment and the hook
The next script creates demo-env, writes a file called hello_startup.pth into its site-packages, and then starts that environment’s Python seven different ways. The hook is a single line that prints which arguments Python was started with, using sys.orig_argv. A trace_site.py helper runs at the end; we will come back to it.
# step01_demo_hook.py
"""Create a throwaway venv, add a harmless startup hook, and watch it run."""
import pathlib
import shutil
import venv
from labenv import run, show, site_packages, venv_python
ENV = pathlib.Path("demo-env")
shutil.rmtree(ENV, ignore_errors=True)
venv.create(ENV, with_pip=True)
py = venv_python(ENV)
hook = site_packages(ENV) / "hello_startup.pth"
hook.write_text('import sys; print("[hello_startup] ran, args:", sys.orig_argv[1:], file=sys.stderr)\n',
encoding="utf-8", newline="\n")
print("wrote", hook.name, "into", site_packages(ENV).relative_to(ENV))
show("python -c \"print('user code')\"", run(py, "-c", "print('user code')"))
show("python -m pip --version", run(py, "-m", "pip", "--version"))
show("python -S -c \"print('user code')\" (site disabled)", run(py, "-S", "-c", "print('user code')"))
show("python -I -c \"print('user code')\" (isolated mode)", run(py, "-I", "-c", "print('user code')"))
show("python -s -c \"print('user code')\" (no user site)", run(py, "-s", "-c", "print('user code')"))
child = "import subprocess, sys; subprocess.run([sys.executable, '-c', 'print(\"child\")'])"
show("a program that starts a child Python", run(py, "-c", child))
show("python -v -c pass (only the lines about .pth files)", run(py, "-v", "-c", "pass"), keep=("pth",))
show("python -S trace_site.py (who calls addsitedir, and how often)", run(py, "-S", "trace_site.py"), keep=("addsitedir",))
# trace_site.py
"""Show who calls site.addsitedir() during startup. Run it as: python -S trace_site.py (-S keeps site from running by itself)."""
import os
import site
import sys
original = site.addsitedir
def traced(sitedir, known_paths=None):
callers = [sys._getframe(depth).f_code.co_name for depth in (1, 2, 3)]
print("addsitedir", os.path.relpath(sitedir), "<-", " <- ".join(callers))
return original(sitedir, known_paths)
site.addsitedir = traced
site.main()
python step01_demo_hook.py
Expected output:
wrote hello_startup.pth into Lib\site-packages
$ python -c "print('user code')"
[hello_startup] ran, args: ['-c', "print('user code')"]
[hello_startup] ran, args: ['-c', "print('user code')"]
user code
$ python -m pip --version
[hello_startup] ran, args: ['-m', 'pip', '--version']
[hello_startup] ran, args: ['-m', 'pip', '--version']
pip 26.1.2 from <lab>\demo-env\Lib\site-packages\pip (python 3.13)
$ python -S -c "print('user code')" (site disabled)
user code
$ python -I -c "print('user code')" (isolated mode)
[hello_startup] ran, args: ['-I', '-c', "print('user code')"]
[hello_startup] ran, args: ['-I', '-c', "print('user code')"]
user code
$ python -s -c "print('user code')" (no user site)
[hello_startup] ran, args: ['-s', '-c', "print('user code')"]
[hello_startup] ran, args: ['-s', '-c', "print('user code')"]
user code
$ a program that starts a child Python
[hello_startup] ran, args: ['-c', 'import subprocess, sys; subprocess.run([sys.executable, \'-c\', \'print("child")\'])']
[hello_startup] ran, args: ['-c', 'import subprocess, sys; subprocess.run([sys.executable, \'-c\', \'print("child")\'])']
[hello_startup] ran, args: ['-c', 'print("child")']
[hello_startup] ran, args: ['-c', 'print("child")']
child
$ python -v -c pass (only the lines about .pth files)
Processing .pth file: '<lab>\\demo-env\\Lib\\site-packages\\hello_startup.pth'
Processing .pth file: '<lab>\\demo-env\\Lib\\site-packages\\hello_startup.pth'
$ python -S trace_site.py (who calls addsitedir, and how often)
addsitedir demo-env <- addsitepackages <- venv <- main
addsitedir demo-env\Lib\site-packages <- addsitepackages <- venv <- main
addsitedir demo-env <- addsitepackages <- main <- <module>
addsitedir demo-env\Lib\site-packages <- addsitepackages <- main <- <module>
What the output tells you
Look at the first run. The hook printed its line before user code, so it ran before your program started, and nothing in your program asked for it. The second run is python -m pip --version: pip is just a Python program, so the hook ran there too. In the child-process run, the hook fired once for the parent and once more for the child Python it started, which is why the count grew.
Now compare the options. -S (documented here) disables the site module, and the hook disappeared. -I (isolated mode) ignores PYTHON* variables and the user’s own site folder, and -s (no user site) skips only that user folder. Neither touches the environment’s main site-packages, so the hook still ran. Isolated mode makes a script’s behaviour more predictable, but it does not protect the script from a hostile package that is already installed in that environment.
Notice also that nothing imported hello_startup. The name of a .pth file is arbitrary; only the .pth ending makes Python read it.
The last -v run shows a way to see which files Python processed without writing any code: it prints a Processing .pth file line for each one. Try it on your real interpreter later: python -v -c pass 2>&1 | grep pth in bash or zsh, python -v -c pass 2>&1 | findstr pth in cmd.exe, or python -v -c pass 2>&1 | Select-String pth in PowerShell.
Why did everything print twice?
On my machine, every hook line appeared twice per start, and -v listed the file twice. That is not a typo in the script. The trace_site.py helper wraps site.addsitedir and prints who calls it. On Python 3.13.14 in a virtual environment, each of the two site folders is added twice during one startup, once through site.venv() and once through site.main() itself. Python keeps sys.path free of duplicates, but nothing stops an import line from running twice. Python 3.15.0rc3 ran the same kind of file once (scenario F in Step 7 shows it: a hidden line printed twice on 3.13 and once on 3.15). I only measured this on Windows with 3.13.14, so on macOS or Linux you may see one line. The lesson holds either way: write startup code so that running it twice is harmless.
Check that it worked
You should see the [hello_startup] line before user code in every run except the one with -S. If you see nothing at all, the file probably landed in the wrong folder; the script prints the folder it used.
Step 2: Learn the exact rules Python uses to read a .pth file
A scanner is only as good as its parser. If your scanner reads a file differently from the way Python does, someone can write a file that your scanner reads as harmless and Python reads as code. Let us build the obvious scanner first and then break it. The obvious design reads a file line by line and keeps the lines that start with the word import and a space.
The script below also writes a second file, hidden.pth. It contains a comment line, then the character U+2028 (LINE SEPARATOR), then an import line. A text editor may draw nothing for U+2028, so the file can look like a single comment.
# step02_naive_scan.py
"""A first, naive scanner for .pth files, and the file that gets past it."""
import pathlib
import site
from labenv import run, show, site_packages, venv_python
ENV = pathlib.Path("demo-env")
py = venv_python(ENV)
LINE_SEPARATOR = chr(0x2028) # str.splitlines() breaks lines on it, most editors draw nothing
def naive_import_lines(path):
"""Read the file line by line and keep the lines that start with 'import '."""
with open(path, encoding="utf-8") as f:
return [line.rstrip("\n") for line in f if line.startswith("import ")]
hidden = site_packages(ENV) / "hidden.pth"
hidden.write_text("# harmless comment" + LINE_SEPARATOR + 'import sys; print("[hidden] this import line ran", file=sys.stderr)\n',
encoding="utf-8", newline="\n")
print("naive scanner finds import lines in hello_startup.pth:", len(naive_import_lines(site_packages(ENV) / "hello_startup.pth")))
print("naive scanner finds import lines in hidden.pth: ", len(naive_import_lines(hidden)))
text = hidden.read_text(encoding="utf-8")
print("lines seen by file iteration:", len(text.split("\n")) - 1, "| lines seen by str.splitlines():", len(text.splitlines()))
show("python -c pass", run(py, "-c", "pass"), keep=("hidden",))
source = pathlib.Path(site.__file__).read_text(encoding="utf-8").splitlines()
start = next((i for i, line in enumerate(source) if "pth_content.splitlines()" in line), None)
print(f"\nwhat {pathlib.Path(site.__file__).name} does in this Python ({'.'.join(map(str, __import__('sys').version_info[:3]))}):")
if start is None:
print(" (this Python version reads .pth files differently)")
else:
print("\n".join(source[start:start + 9]))
python step02_naive_scan.py
naive scanner finds import lines in hello_startup.pth: 1
naive scanner finds import lines in hidden.pth: 0
lines seen by file iteration: 1 | lines seen by str.splitlines(): 2
$ python -c pass
[hidden] this import line ran
[hidden] this import line ran
what site.py does in this Python (3.13.14):
for n, line in enumerate(pth_content.splitlines(), 1):
if line.startswith("#"):
continue
if line.strip() == "":
continue
try:
if line.startswith(("import ", "import\t")):
exec(line)
continue
The naive scanner found one import line in hello_startup.pth and none in hidden.pth. Python ran the hidden one anyway. The reason is in the last lines of the output, which are Python’s own site.py: it reads the whole file and loops over pth_content.splitlines(). The str.splitlines() method breaks lines on more characters than a newline, among them form feed, U+0085, U+2028 and U+2029, while iterating over an open file breaks only on newline characters. So Python saw two lines where the scanner saw one comment: the second began with the word import and a space, and was executed.
I compared Lib/site.py on the 3.10, 3.12, 3.13, 3.14 and main branches of CPython on October 6. The 3.12, 3.13 and 3.14 branches read the file whole and call splitlines(); the 3.10 branch iterates over the file object instead, so this particular trick does not work there. I did not check 3.11. The lesson is general: copy the interpreter’s rules, not your assumptions.
The rules, side by side
Here is what I measured on 3.13.14 and 3.15.0rc3, from reading site.py on each and from the experiments in Step 7. The 3.15 column matters because the same file can mean different things to different Python versions.
| Question | Python 3.12 to 3.14 (run on 3.13.14) | Python 3.15.0rc3 |
|---|---|---|
| How is the file split into lines? | str.splitlines() |
str.splitlines() |
| What makes a line an import line? | The raw line starts with import or import plus a tab |
The same test, applied after stripping whitespace from both ends |
| What makes a line a comment? | The raw line starts with # |
The stripped line starts with # |
| Do import lines run? | Always | Yes during the deprecation period, except when a matching .start file exists |
| A path line naming a folder that does not exist | Ignored silently | A message on standard error at every start (rc3) |
A .pth file that is not valid UTF-8 |
Decoded with the locale encoding as a fallback | The same fallback, now deprecated; the 3.15 docs say it will be removed in 3.20 |
| An error on one line | Prints an error, then “Remainder of file ignored” | Reports it and continues with the next line |
One consequence: a line such as import sys, with a leading space, is inert on 3.13 and executes on 3.15. A scanner built only from the 3.13 rules would call that file harmless. We will handle this by reporting a line as an import line if either rule says it is one.
Step 3: Build the auditor
Now we write pthaudit.py. Its design goals are simple. It uses only the standard library, so it never imports anything from the environment it inspects. It mirrors Python’s parser, including the quirks from Step 2. It tells you who owns each startup file and whether the file changed since install. It flags risky patterns. And it speaks in exit codes so that a CI job can act on it. Each file gets one of six statuses, checked in this order:
| Status | Meaning | Exit code |
|---|---|---|
TAMPERED |
The file’s hash differs from the hash recorded when the package was installed | 2 |
allowed |
A person reviewed this exact file; its hash is on your allowlist | 0 |
path-only |
Only folders and comments: nothing runs | 0 |
SUSPICIOUS |
Code with risky patterns (such as exec or base64), an unusual line separator, an invalid entry point or a non-UTF-8 encoding |
2 |
UNOWNED |
Code that no installed package claims | 2 |
REVIEW |
Code that belongs to an installed package and shows nothing alarming; a human should still look once | 1 |
Part 1: parse a startup file the way site.py does
decode_startup_file decodes as UTF-8, accepting a byte order mark, and falls back to the locale encoding just as site.py does. We add errors="replace" so a broken file can never crash the audit, and we flag it as non-utf8. parse_pth classifies every line with the union rule from Step 2. risk_flags looks for patterns such as exec(, base64 and network calls. inspect_bytes ties these together for one file, and both the environment audit and the wheel check will call it. It also handles .start files (each entry must look like pkg.mod:callable) and sitecustomize.py or usercustomize.py, which Python imports right after the path setup.
# pthaudit.py
"""Audit Python startup hooks: .pth import lines, .start entry points, sitecustomize files.
Run it with a trusted interpreter and point it at the environment you distrust:
python -I -S pthaudit.py path/to/venv
It uses only the standard library, so nothing from the audited environment is imported or executed.
Exit codes: 0 nothing to review, 1 something needs review, 2 something is tampered, unowned or suspicious.
"""
import argparse
import base64
import csv
import hashlib
import json
import locale
import os
import re
import sys
from dataclasses import dataclass, field
from pathlib import Path
IMPORT_PREFIXES = ("import ", "import\t")
# str.splitlines() breaks lines on these characters; a text editor or a line-by-line file reader does not.
ODD_BREAKS = re.compile("[" + "".join(chr(c) for c in (0x0B, 0x0C, 0x1C, 0x1D, 0x1E, 0x85, 0x2028, 0x2029)) + "]")
ENTRY_POINT = re.compile(r"^[A-Za-z_]\w*(\.[A-Za-z_]\w*)*:[A-Za-z_]\w*(\.[A-Za-z_]\w*)*$")
RISK = [
(r"\bexec\s*\(", "exec"), (r"\beval\s*\(", "eval"), (r"base64|b64decode", "base64"),
(r"\bcompile\s*\(", "compile"), (r"marshal|zlib|bz2|lzma", "packed-code"),
(r"subprocess|os\.system|os\.popen|Popen", "spawns-process"),
(r"socket|urllib|http\.client|requests|urlopen", "network"),
(r"ctypes", "ctypes"), (r"\bopen\s*\(", "file-access"), (r"__import__", "dynamic-import"), (r"pickle", "pickle"),
]
DANGEROUS = {"exec", "eval", "base64", "compile", "packed-code", "spawns-process", "network", "ctypes", "pickle"}
# ---- 1. parse a .pth file the way site.py does -------------------------------------------------
def decode_startup_file(data):
"""Decode like site.py: UTF-8 first (a BOM is accepted), then the locale encoding."""
try:
return data.decode("utf-8-sig"), "utf-8"
except UnicodeDecodeError:
encoding = locale.getencoding()
return data.decode(encoding, errors="replace"), encoding
@dataclass
class Line:
number: int
text: str
kind: str # path, import, comment or blank
note: str = ""
def parse_pth(text):
"""Classify every line. site.py reads the whole file and uses str.splitlines()."""
lines = []
for number, raw in enumerate(text.splitlines(), 1):
if raw.startswith("#"):
lines.append(Line(number, raw, "comment"))
elif not raw.strip():
lines.append(Line(number, raw, "blank"))
elif raw.startswith(IMPORT_PREFIXES):
lines.append(Line(number, raw, "import"))
elif raw.strip().startswith(IMPORT_PREFIXES):
# Python 3.15 strips each line before the check, 3.12 to 3.14 do not: report it either way.
lines.append(Line(number, raw, "import", "leading whitespace"))
else:
lines.append(Line(number, raw, "path"))
return lines
def risk_flags(text):
flags = [name for pattern, name in RISK if re.search(pattern, text)]
if len(text) > 200:
flags.append("long-line")
if any(ord(ch) > 127 for ch in text):
flags.append("non-ascii")
return flags
def inspect_bytes(name, data):
"""Return (kind, code, flags) for one startup file; `code` lists what would run at startup."""
if name.endswith(".pth"):
kind = "pth"
elif name.endswith(".start"):
kind = "start"
elif name in ("sitecustomize.py", "usercustomize.py"):
kind = name[:-3]
else:
return None, [], []
text, encoding = decode_startup_file(data)
flags = []
if encoding != "utf-8":
flags.append("non-utf8")
if ODD_BREAKS.search(text):
flags.append("unusual-line-separator")
code = []
if kind == "pth":
for line in parse_pth(text):
if line.kind == "import":
code.append((line.number, line.text.strip()))
flags.extend(risk_flags(line.text))
if line.note:
flags.append("leading-whitespace-import")
elif kind == "start":
for number, raw in enumerate(text.splitlines(), 1):
entry = raw.strip()
if entry and not entry.startswith("#"):
code.append((number, entry))
if not ENTRY_POINT.match(entry):
flags.append("invalid-entry")
else:
code.append((1, f"<module body, {len(text.splitlines())} lines>"))
flags.extend(risk_flags(text))
return kind, code, sorted(set(flags), key=flags.index)
Part 2: who owns a file, and has it changed?
Every installed package leaves a RECORD file, described in the packaging specification. Each row lists a file, a hash in the form sha256= followed by URL-safe base64 without padding, and a size. If a startup file is listed there, a package owns it. If the file is not listed, nobody installed it through pip: it was written by hand, by another program or by an attacker. If the file is listed but its content no longer matches the recorded hash, someone changed it after the install.
# pthaudit.py (continued)
# ---- 2. who owns a file, and has it changed since install? ------------------------------------
def record_index(sitedir):
"""Map each path listed in a *.dist-info/RECORD file to [(distribution, hash field)]."""
index = {}
for info in sorted(Path(sitedir).glob("*.dist-info")):
record = info / "RECORD"
if record.is_file():
label = " ".join(info.name[: -len(".dist-info")].rsplit("-", 1))
for row in csv.reader(record.read_text(encoding="utf-8").splitlines()):
if row:
index.setdefault(row[0].replace("\\", "/"), []).append((label, row[1] if len(row) > 1 else ""))
return index
def hash_matches(data, hash_field):
algorithm, _, expected = hash_field.partition("=")
if not expected:
return None
digest = base64.urlsafe_b64encode(hashlib.new(algorithm, data).digest()).rstrip(b"=").decode()
return digest == expected
Part 3: the decision ladder
audit_sitedir walks a site-packages folder, inspects each startup file, attaches owner and integrity, and assigns the status using the order in the table above. The allowlist maps a file name to the SHA-256 of the exact bytes a person reviewed, so changing even one character removes the file from the allowlist.
# pthaudit.py (continued)
# ---- 3. audit a site directory ----------------------------------------------------------------
@dataclass
class Finding:
sitedir: str
name: str
kind: str
sha256: str
owner: str = "UNOWNED"
integrity: str = "n/a"
code: list = field(default_factory=list)
flags: list = field(default_factory=list)
status: str = ""
def audit_sitedir(sitedir, allowed=None):
sitedir, allowed = Path(sitedir), allowed or {}
owners = record_index(sitedir)
findings = []
for path in sorted(p for p in sitedir.iterdir() if p.is_file() and not p.name.startswith(".")):
data = path.read_bytes()
kind, code, flags = inspect_bytes(path.name, data)
if kind is None:
continue
found = Finding(str(sitedir), path.name, kind, hashlib.sha256(data).hexdigest(), code=code, flags=flags)
if path.name in owners:
found.owner = ", ".join(sorted({label for label, _ in owners[path.name]}))
checks = [hash_matches(data, field_) for _, field_ in owners[path.name]]
found.integrity = "ok" if any(c for c in checks) else ("MISMATCH" if False in checks else "n/a")
has_code = bool(code) or kind in ("start", "sitecustomize", "usercustomize")
if found.integrity == "MISMATCH":
found.status = "TAMPERED"
elif allowed.get(path.name) == found.sha256:
found.status = "allowed"
elif not has_code:
found.status = "path-only"
elif DANGEROUS & set(flags) or {"unusual-line-separator", "invalid-entry", "non-utf8"} & set(flags):
found.status = "SUSPICIOUS"
elif found.owner == "UNOWNED":
found.status = "UNOWNED"
else:
found.status = "REVIEW"
findings.append(found)
return findings
Part 4: the command line
site_dirs_for accepts either a virtual environment folder (it contains pyvenv.cfg) or a plain site folder. main prints one block per file and returns the worst exit code. The last two lines make the file runnable as a script.
# pthaudit.py (continued)
def site_dirs_for(path):
"""A venv folder (has pyvenv.cfg) maps to its site-packages; anything else is treated as a site folder."""
path = Path(path)
if (path / "pyvenv.cfg").is_file():
found = [path / "Lib" / "site-packages"] if os.name == "nt" else sorted(path.glob("lib/python*/site-packages"))
return [p for p in found if p.is_dir()]
return [path]
EXIT_BY_STATUS = {"TAMPERED": 2, "SUSPICIOUS": 2, "UNOWNED": 2, "REVIEW": 1}
def main(argv=None):
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
parser.add_argument("paths", nargs="+", help="venv folders or site-packages folders")
parser.add_argument("--allowlist", help="JSON file of reviewed files: {name: sha256}")
parser.add_argument("--write-allowlist", help="write every file that needs review to this JSON file")
parser.add_argument("--json", action="store_true", help="print findings as JSON")
args = parser.parse_args(argv)
allowed = json.loads(Path(args.allowlist).read_text(encoding="utf-8")) if args.allowlist else {}
findings = [f for p in args.paths for d in site_dirs_for(p) for f in audit_sitedir(d, allowed)]
if args.json:
print(json.dumps([f.__dict__ for f in findings], indent=2))
else:
for f in findings:
print(f"[{f.status}] {f.name} (owner: {f.owner}; integrity: {f.integrity}; flags: {', '.join(f.flags) or '-'})")
for number, text in f.code:
print(f" line {number}: {text[:100]}{'...' if len(text) > 100 else ''}")
counts = {}
for f in findings:
counts[f.status] = counts.get(f.status, 0) + 1
print("summary:", ", ".join(f"{n} {s}" for s, n in sorted(counts.items())) or "no startup files found")
if args.write_allowlist:
reviewed = {f.name: f.sha256 for f in findings if f.status in ("REVIEW", "UNOWNED")}
Path(args.write_allowlist).write_text(json.dumps(reviewed, indent=2) + "\n", encoding="utf-8")
print(f"wrote {len(reviewed)} entries to {args.write_allowlist}: read each one before you commit this file")
return max((EXIT_BY_STATUS.get(f.status, 0) for f in findings), default=0)
# pthaudit.py (continued)
if __name__ == "__main__":
sys.exit(main())
Common mistake: auditing an environment with its own Python
Run the auditor with a Python you trust and point it at the folder you distrust:
python -I -S pthaudit.py path/to/venv
Do not run path/to/venv/bin/python pthaudit.py. Starting that interpreter would run the environment’s startup hooks, so a hostile hook would execute inside your audit before the first line of the auditor. The -S flag also keeps your trusted Python from running the hooks in its site-packages, and -I ignores PYTHON* variables and the user site folder. The auditor needs nothing outside the standard library, which is what makes this safe.
Step 4: Audit a realistic environment
A fresh virtual environment is almost empty, so let us make demo-env look like a working one. The next script installs setuptools, then creates two tiny projects and installs each with pip install -e (an editable install, the kind a developer makes of their own code). Then it runs the auditor. The two files from Steps 1 and 2, hello_startup.pth and hidden.pth, are still in the environment, so the audit has something to find. Run this after Steps 1 and 2 so those files exist.
# step04_audit_env.py
"""Build a realistic environment, then audit it from the outside with a trusted interpreter."""
import pathlib
import shutil
import sys
from labenv import run, show, site_packages, venv_python
ENV = pathlib.Path("demo-env")
py, SP = venv_python(ENV), site_packages(ENV)
def pip(*args):
result = run(py, "-m", "pip", "install", "-q", *args)
assert result.returncode == 0, result.stderr[-500:]
PYPROJECT = '[build-system]\nrequires = ["setuptools>=64"]\nbuild-backend = "setuptools.build_meta"\n[project]\nname = "{name}"\nversion = "0.1"\n{extra}'
shutil.rmtree("projects", ignore_errors=True)
flat = pathlib.Path("projects/flat_demo")
(flat / "flat_demo").mkdir(parents=True)
(flat / "flat_demo" / "__init__.py").write_text("X = 1\n", encoding="utf-8")
(flat / "pyproject.toml").write_text(PYPROJECT.format(name="flat_demo", extra=""), encoding="utf-8")
src = pathlib.Path("projects/src_demo")
(src / "src" / "src_demo").mkdir(parents=True)
(src / "src" / "src_demo" / "__init__.py").write_text("X = 1\n", encoding="utf-8")
(src / "pyproject.toml").write_text(PYPROJECT.format(name="src_demo", extra='[tool.setuptools.packages.find]\nwhere = ["src"]\n'), encoding="utf-8")
pip("setuptools")
pip("--no-build-isolation", "-e", str(flat))
pip("--no-build-isolation", "-e", str(src))
print("startup files now in site-packages:", sorted(p.name for p in SP.glob("*.pth")))
audit = lambda: run(sys.executable, "-I", "-S", "pthaudit.py", str(ENV))
show("python -I -S pthaudit.py demo-env", audit())
shim = SP / "distutils-precedence.pth"
original = shim.read_bytes()
shim.write_bytes(original + b'import sys; print("[tampered] added after install", file=sys.stderr)\n')
print("\n--- now append one line to distutils-precedence.pth, as malware or a careless edit might ---")
report = audit()
shown, in_block = [], False
for line in report.stdout.splitlines():
in_block = line.startswith("[TAMPERED]") or (in_block and not line.startswith("["))
if in_block or line.startswith("summary"):
shown.append(line)
print("$ python -I -S pthaudit.py demo-env (only the changed file and the summary)")
print("\n".join(" " + line for line in shown))
print(f" (exit code {report.returncode})")
shim.write_bytes(original)
python step04_audit_env.py
startup files now in site-packages: ['__editable__.flat_demo-0.1.pth', '__editable__.src_demo-0.1.pth', 'distutils-precedence.pth', 'hello_startup.pth', 'hidden.pth']
$ python -I -S pthaudit.py demo-env
[REVIEW] __editable__.flat_demo-0.1.pth (owner: flat_demo 0.1; integrity: ok; flags: -)
line 1: import __editable___flat_demo_0_1_finder; __editable___flat_demo_0_1_finder.install()
[path-only] __editable__.src_demo-0.1.pth (owner: src_demo 0.1; integrity: ok; flags: -)
[REVIEW] distutils-precedence.pth (owner: setuptools 84.0.0; integrity: ok; flags: dynamic-import)
line 1: import os; var = 'SETUPTOOLS_USE_DISTUTILS'; enabled = os.environ.get(var, 'local') == 'local'; enab...
[UNOWNED] hello_startup.pth (owner: UNOWNED; integrity: n/a; flags: -)
line 1: import sys; print("[hello_startup] ran, args:", sys.orig_argv[1:], file=sys.stderr)
[SUSPICIOUS] hidden.pth (owner: UNOWNED; integrity: n/a; flags: unusual-line-separator)
line 2: import sys; print("[hidden] this import line ran", file=sys.stderr)
summary: 2 REVIEW, 1 SUSPICIOUS, 1 UNOWNED, 1 path-only
(exit code 2)
Reading the report
Go through the files one at a time, the way you would on a real review.
__editable__.flat_demo-0.1.pth is REVIEW. It belongs to our own editable project, and its import line installs an import hook that finds the project’s code. It is a normal, legitimate use of an import line, but the auditor cannot know that, so it asks for a human look. The other editable project, with a src layout, produced a file with only a folder path, so it is path-only. That difference is real: the same tool wrote two different kinds of file depending on the project layout.
distutils-precedence.pth comes from setuptools. It is also legitimate: the line checks the SETUPTOOLS_USE_DISTUTILS environment variable and installs a shim from a module called _distutils_hack, and the file name, distutils-precedence, tells you the purpose. It carries the flag dynamic-import because it calls __import__. Again REVIEW.
hello_startup.pth is UNOWNED: no installed package lists it. In a real environment built from a lock file, an unowned startup file is the finding you most want to chase. hidden.pth is SUSPICIOUS because it contains an unusual line separator, and the auditor shows the import line behind the comment as line 2. The summary counts the statuses and the exit code is 2, the worst.
See tampering get caught
The second half of the script appends one line to distutils-precedence.pth, the way malware or a careless edit might, runs the audit again, and restores the file.
--- now append one line to distutils-precedence.pth, as malware or a careless edit might ---
$ python -I -S pthaudit.py demo-env (only the changed file and the summary)
[TAMPERED] distutils-precedence.pth (owner: setuptools 84.0.0; integrity: MISMATCH; flags: dynamic-import)
line 1: import os; var = 'SETUPTOOLS_USE_DISTUTILS'; enabled = os.environ.get(var, 'local') == 'local'; enab...
line 2: import sys; print("[tampered] added after install", file=sys.stderr)
summary: 1 REVIEW, 1 SUSPICIOUS, 1 TAMPERED, 1 UNOWNED, 1 path-only
(exit code 2)
The file still belongs to setuptools 84.0.0, but its content no longer matches the hash in that package’s RECORD, so the status jumps from REVIEW to TAMPERED and both lines are listed. This check has a limit that I will repeat at the end: the RECORD file lives in the same folder, so someone who can change the startup file can usually change RECORD too. Treat integrity: ok as “matches what the installer wrote”, not as “safe”.
Step 5: Check a wheel before you install it
Auditing an installed environment tells you what already happened. It is better to look inside a package before it is installed. A wheel is a zip file, so you can open it and list its contents without running anything. The check must look at every folder of the wheel, because the wheel format lets a file ride in a .data/purelib folder and still end up in site-packages.
Build a demo wheel
To practise safely, we build a wheel by hand. It contains a package and a .pth file whose line decodes and runs a base64 string, an obfuscation style that the auditor’s risk flags are written to catch. The payload only prints one sentence.
# make_demo_wheel.py
"""Build a tiny wheel by hand: a harmless package whose .pth file looks like a real dropper."""
import base64
import hashlib
import zipfile
PAYLOAD = base64.b64encode(b'print("[demo-startup] payload ran at interpreter start")').decode()
FILES = {
"demo_startup/__init__.py": "VERSION = '1.0'\n",
"demo_startup_hook.pth": f"import base64; exec(base64.b64decode('{PAYLOAD}'))\n",
"demo_startup-1.0.dist-info/METADATA": "Metadata-Version: 2.1\nName: demo-startup\nVersion: 1.0\n",
"demo_startup-1.0.dist-info/WHEEL": "Wheel-Version: 1.0\nGenerator: by-hand\nRoot-Is-Purelib: true\nTag: py3-none-any\n",
}
def record_line(name, data):
digest = base64.urlsafe_b64encode(hashlib.sha256(data).digest()).rstrip(b"=").decode()
return f"{name},sha256={digest},{len(data)}"
def build(path="demo_startup-1.0-py3-none-any.whl"):
blobs = {name: text.encode("utf-8") for name, text in FILES.items()}
record = "\n".join([record_line(n, d) for n, d in blobs.items()] + ["demo_startup-1.0.dist-info/RECORD,,"]) + "\n"
with zipfile.ZipFile(path, "w", zipfile.ZIP_DEFLATED) as zf:
for name, data in blobs.items():
zf.writestr(name, data)
zf.writestr("demo_startup-1.0.dist-info/RECORD", record)
return path
if __name__ == "__main__":
print("built", build())
Write the wheel checker
# wheel_check.py
"""Look inside wheels for startup files BEFORE installing them.
Exit code 2 if a wheel carries import lines, .start files or sitecustomize modules,
1 if it only carries path-only .pth files, 0 if it carries nothing of the kind.
"""
import glob
import os
import sys
import zipfile
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pthaudit
STARTUP_MODULES = ("sitecustomize.py", "usercustomize.py")
def inspect_wheel(zf):
"""Return [(member, kind, code, flags)] for every startup file anywhere in the wheel."""
found = []
for member in zf.namelist():
base = member.rsplit("/", 1)[-1]
if not base.startswith(".") and (base.endswith((".pth", ".start")) or base in STARTUP_MODULES):
kind, code, flags = pthaudit.inspect_bytes(base, zf.read(member))
found.append((member, kind, code, flags))
return found
def main(patterns):
worst = 0
paths = [path for pattern in patterns for path in (sorted(glob.glob(pattern)) or [pattern])] # expand * ourselves: cmd and PowerShell do not
for path in paths:
with zipfile.ZipFile(path) as zf:
found = inspect_wheel(zf)
print(f"{os.path.basename(path)}: {len(found)} startup file(s)")
for member, kind, code, flags in found:
print(f" {member} [{kind}] flags: {', '.join(flags) or '-'}")
for number, text in code:
print(f" line {number}: {text[:100]}")
worst = max(worst, 2 if code or kind != "pth" else 1)
return worst
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
inspect_wheel reads only the members whose names look like startup files, so checking a large wheel stays fast. The command line expands * itself because neither cmd nor PowerShell does it for programs. The exit code is 2 if the wheel carries code-bearing startup files, 1 if it carries only path-only .pth files, and 0 if it carries none.
Check, stage and install
The script runs the checker on the demo wheel, then installs the wheel with pip install --target staging. The --target option unpacks the files into a plain folder, and because that folder is not a site directory, Python never processes the .pth file inside it. We can audit the staged files with pthaudit.py and nothing runs. Then, to show why this matters, the script installs the same wheel into a new environment with no check at all. Finally it runs the checker on two real wheels downloaded from PyPI.
# step05_wheel.py
"""Inspect a wheel before installing it, stage it without running anything, then install it blindly."""
import pathlib
import shutil
import sys
import venv
import make_demo_wheel
from labenv import run, show, venv_python
wheel = make_demo_wheel.build()
show(f"python wheel_check.py {wheel}", run(sys.executable, "-S", "wheel_check.py", wheel))
shutil.rmtree("staging", ignore_errors=True)
show("pip install --no-deps --target staging <wheel>",
run(sys.executable, "-m", "pip", "install", "-q", "--no-deps", "--target", "staging", wheel))
print("files staged:", sorted(p.name for p in pathlib.Path("staging").iterdir() if p.name != "__pycache__"))
show("python -I -S pthaudit.py staging", run(sys.executable, "-I", "-S", "pthaudit.py", "staging"))
print("\n--- the same check on real wheels from PyPI ---")
shutil.rmtree("real-wheels", ignore_errors=True)
show("pip download --no-deps --only-binary=:all: --dest real-wheels setuptools coverage",
run(sys.executable, "-m", "pip", "download", "-q", "--no-deps", "--only-binary=:all:", "--dest", "real-wheels", "setuptools", "coverage"))
show("python wheel_check.py real-wheels/*.whl", run(sys.executable, "-S", "wheel_check.py", "real-wheels/*.whl"))
shutil.rmtree("wheel-env", ignore_errors=True)
venv.create("wheel-env", with_pip=True)
py = venv_python("wheel-env")
print("\n--- now install the same wheel for real, with no check ---")
show("pip install --no-deps <wheel> (inside wheel-env)", run(py, "-m", "pip", "install", "-q", "--no-deps", wheel))
show("python -c \"print('user code')\"", run(py, "-c", "print('user code')"))
show("python -m pip --version", run(py, "-m", "pip", "--version"))
show("python -m pip uninstall -y demo-startup", run(py, "-m", "pip", "uninstall", "-y", "demo-startup"))
show("python -c \"print('user code')\"", run(py, "-c", "print('user code')"))
python step05_wheel.py
$ python wheel_check.py demo_startup-1.0-py3-none-any.whl
demo_startup-1.0-py3-none-any.whl: 1 startup file(s)
demo_startup_hook.pth [pth] flags: exec, base64
line 1: import base64; exec(base64.b64decode('cHJpbnQoIltkZW1vLXN0YXJ0dXBdIHBheWxvYWQgcmFuIGF0IGludGVycHJldG
(exit code 2)
$ pip install --no-deps --target staging <wheel>
files staged: ['demo_startup', 'demo_startup-1.0.dist-info', 'demo_startup_hook.pth']
$ python -I -S pthaudit.py staging
[SUSPICIOUS] demo_startup_hook.pth (owner: demo_startup 1.0; integrity: ok; flags: exec, base64)
line 1: import base64; exec(base64.b64decode('cHJpbnQoIltkZW1vLXN0YXJ0dXBdIHBheWxvYWQgcmFuIGF0IGludGVycHJldG...
summary: 1 SUSPICIOUS
(exit code 2)
The wheel checker and the auditor both flagged exec and base64 in the demo file, and the staging folder showed the same thing without running it. Then the blind install did what the first section of this article described.
--- now install the same wheel for real, with no check ---
$ pip install --no-deps <wheel> (inside wheel-env)
$ python -c "print('user code')"
[demo-startup] payload ran at interpreter start
[demo-startup] payload ran at interpreter start
user code
$ python -m pip --version
[demo-startup] payload ran at interpreter start
[demo-startup] payload ran at interpreter start
pip 26.1.2 from <lab>\wheel-env\Lib\site-packages\pip (python 3.13)
$ python -m pip uninstall -y demo-startup
[demo-startup] payload ran at interpreter start
[demo-startup] payload ran at interpreter start
Found existing installation: demo-startup 1.0
Uninstalling demo-startup-1.0:
Successfully uninstalled demo-startup-1.0
$ python -c "print('user code')"
user code
Notice three things. The payload ran at the very next start of Python, with no import of demo_startup anywhere. It ran again when we started pip itself, so a hook in an environment gets to run inside the tool you use to manage that environment. And it ran during the pip uninstall that removed it; the last run, after the uninstall, printed only user code. (The doubled lines are the Python 3.13 behaviour from Step 1.)
The same check on real wheels
--- the same check on real wheels from PyPI ---
$ pip download --no-deps --only-binary=:all: --dest real-wheels setuptools coverage
$ python wheel_check.py real-wheels/*.whl
coverage-7.16.2-cp313-cp313-win_amd64.whl: 1 startup file(s)
a1_coverage.pth [pth] flags: exec, long-line
line 1: import sys; exec('import os\n\nif os.getenv("COVERAGE_PROCESS_START") or os.getenv("COVERAGE_PROCESS
setuptools-84.0.0-py3-none-any.whl: 1 startup file(s)
distutils-precedence.pth [pth] flags: dynamic-import
line 1: import os; var = 'SETUPTOOLS_USE_DISTUTILS'; enabled = os.environ.get(var, 'local') == 'local'; enab
(exit code 2)
Both real packages carry startup code, and both are legitimate. The coverage wheel ships a1_coverage.pth, whose import line starts coverage measurement when one of two environment variables is set. The setuptools wheel ships the shim we met in Step 4. The checker cannot tell friend from foe; what it gives you is the list of files that deserve a decision. To download wheels without installing them, use pip download --no-deps --only-binary=:all:, as the script above does. The --only-binary=:all: part matters: it refuses source distributions, which run build code at install time and are outside what this checker can see.
Step 6: See what popular packages ship
How common are startup files in the wild? We can find out without downloading gigabytes. A wheel is a zip file, and a zip file keeps its table of contents at the end. Using HTTP range requests, a small file object can fetch just the last 256 KB of a wheel from PyPI, hand it to Python’s zipfile module, and read the file names and any .pth member. The script below does this for the most downloaded packages, using the public top PyPI packages list and PyPI’s JSON API. It picks one wheel per package, preferring a pure-Python wheel, then a Windows or Linux x86-64 wheel.
# survey_wheels.py
"""Count how many popular PyPI packages ship Python startup files, reading only the end of each wheel."""
import concurrent.futures as cf
import json
import sys
import threading
import zipfile
import requests
import wheel_check
LIST_URL = "https://hugovk.github.io/top-pypi-packages/top-pypi-packages.min.json"
HEADERS = {"User-Agent": "startup-file-survey/1.0 (educational script)", "Accept-Encoding": "identity"}
class HttpFile:
"""A read-only, seekable file object backed by HTTP Range requests, so zipfile can list a remote wheel."""
def __init__(self, url, session, size, tail=1 << 18):
"""`size` is the file size in bytes (PyPI reports it); the last `tail` bytes are fetched up front."""
self.url, self.session, self.pos, self.size, self.requests = url, session, 0, size, 1
self.start = max(0, size - tail)
reply = session.get(url, headers={"Range": f"bytes={self.start}-{size - 1}"}, timeout=30)
reply.raise_for_status()
if reply.status_code == 200: # the server ignored Range and sent the whole file
self.start = 0
self.cache = reply.content
def seekable(self):
return True
def tell(self):
return self.pos
def seek(self, offset, whence=0):
self.pos = offset if whence == 0 else self.pos + offset if whence == 1 else self.size + offset
return self.pos
def read(self, n=-1):
end = self.size if n is None or n < 0 else min(self.pos + n, self.size)
if self.start <= self.pos and end <= self.start + len(self.cache):
data = self.cache[self.pos - self.start:end - self.start]
else:
reply = self.session.get(self.url, headers={"Range": f"bytes={self.pos}-{end - 1}"}, timeout=30)
reply.raise_for_status()
data, self.requests = reply.content, self.requests + 1
self.pos += len(data)
return data
def pick_wheel(files):
"""Prefer a pure-Python wheel, then a Windows or Linux x86-64 wheel; one wheel stands for the package."""
wheels = [f for f in files if f["packagetype"] == "bdist_wheel" and not f.get("yanked")]
def rank(f):
name = f["filename"]
if name.endswith("-none-any.whl"):
return 0
if "-cp313-" in name and "win_amd64" in name:
return 1
if "-abi3-" in name and "win_amd64" in name:
return 2
if "manylinux" in name and "x86_64" in name:
return 3
return 4
return min(wheels, key=rank) if wheels else None
LOCAL = threading.local()
def survey_one(rank, project):
if not hasattr(LOCAL, "session"): # one session per worker thread, reused for every package
LOCAL.session = requests.Session()
LOCAL.session.headers.update(HEADERS)
session = LOCAL.session
try:
reply = session.get(f"https://pypi.org/pypi/{project}/json", timeout=30)
if reply.status_code != 200:
return {"rank": rank, "project": project, "error": f"PyPI answered {reply.status_code}"}
data = reply.json()
wheel = pick_wheel(data["urls"])
result = {"rank": rank, "project": project, "version": data["info"]["version"], "wheel": None, "files": []}
if wheel:
remote = HttpFile(wheel["url"], session, wheel["size"])
with zipfile.ZipFile(remote) as zf:
found = wheel_check.inspect_wheel(zf)
result.update(wheel=wheel["filename"], requests=remote.requests, size=wheel["size"],
files=[{"member": m, "kind": k, "code": [t for _, t in c], "flags": f} for m, k, c, f in found])
return result
except Exception as exc: # one broken package must not stop the survey
return {"rank": rank, "project": project, "error": repr(exc)[:120]}
def main(count):
listing = requests.get(LIST_URL, headers=HEADERS, timeout=60).json()
rows = listing["rows"][:count]
with cf.ThreadPoolExecutor(12) as pool:
results = sorted(pool.map(lambda item: survey_one(item[0] + 1, item[1]["project"]), enumerate(rows)), key=lambda r: r["rank"])
ok = [r for r in results if "error" not in r]
with_files = [r for r in ok if r["files"]]
kinds = lambda kind: [(r, f) for r in ok for f in r["files"] if f["kind"] == kind]
with_imports = [(r, f) for r, f in kinds("pth") if f["code"]]
print(f"download list updated {listing['last_update']}; packages scanned: {len(results)}")
print(f"no wheel on PyPI: {sum(1 for r in ok if not r['wheel'])} | lookup errors: {len(results) - len(ok)}")
print(f"wheels with at least one startup file: {len(with_files)}")
print(f" .pth files: {len(kinds('pth'))}, in {len({r['project'] for r, _ in kinds('pth')})} packages")
print(f" .pth files with import lines: {len(with_imports)}, in {len({r['project'] for r, _ in with_imports})} packages")
print(f" .start files: {len(kinds('start'))} | sitecustomize.py or usercustomize.py: {len(kinds('sitecustomize')) + len(kinds('usercustomize'))}")
print(f" HTTP requests per wheel (median): {sorted(r['requests'] for r in ok if r['wheel'])[len([r for r in ok if r['wheel']]) // 2]}")
print("\nrank package file first import line")
for r, f in with_imports:
print(f"{r['rank']:>4} {r['project'][:18]:<18} {f['member'][:38]:<38} {f['code'][0][:70]}")
with open("survey.json", "w", encoding="utf-8") as out:
json.dump(results, out, indent=1)
if __name__ == "__main__":
main(int(sys.argv[1]) if len(sys.argv) > 1 else 1000)
python survey_wheels.py 1000
download list updated 2026-10-01 12:40:51; packages scanned: 1000
no wheel on PyPI: 5 | lookup errors: 0
wheels with at least one startup file: 10
.pth files: 6, in 6 packages
.pth files with import lines: 6, in 6 packages
.start files: 0 | sitecustomize.py or usercustomize.py: 5
HTTP requests per wheel (median): 1
rank package file first import line
10 setuptools distutils-precedence.pth import os; var = 'SETUPTOOLS_USE_DISTUTILS'; enabled = os.environ.get(
104 coverage a1_coverage.pth import sys; exec('import os\n\nif os.getenv("COVERAGE_PROCESS_START")
288 google-cloud-aipla google_cloud_aiplatform-2.3.0-py3.12-n import sys, types, os;p = os.path.join(sys._getframe(1).f_locals['site
514 coloredlogs coloredlogs.pth import os; exec('try: __import__("coloredlogs").auto_install() if os.e
891 sphinxcontrib-jsma sphinxcontrib_jsmath-1.0.1-py3.7-nspkg import sys, types, os;has_mfs = sys.version_info > (3, 5);p = os.path.
987 pywin32 pywin32.pth import pywin32_bootstrap
What the survey says, and what it does not
Of 995 wheels I could inspect, only 10 contained any startup file at all. Six .pth files appeared in six packages, and every one of them had import lines: setuptools (rank 10 by downloads in the list), coverage (104), google-cloud-aiplatform (288), coloredlogs (514), sphinxcontrib-jsmath (891) and pywin32 (987). The two namespace-package files from Google and Sphinx use a long import line that reads the variable sitedir from the calling frame; I will come back to that detail in Step 7, because Python 3.15’s site.py keeps providing it for compatibility. The other four wheels (opentelemetry-instrumentation, debugpy, ddtrace and gevent) hold five sitecustomize.py files inside subpackages. Those folders are not on the module search path by default, so they run at startup only if something puts them there; I treat them as items to read, not as alarms. No wheel in this sample ships a .start file yet, which is unsurprising a few days before the release.
The caveats matter. The list was last updated on October 1, 2026 and I ran the scan on October 6. I looked at one wheel per package, so a startup file that exists only in a platform-specific wheel I did not pick would be missed. Five packages (pyspark, docopt, swifter, oss2 and ratelimit) had no wheel and were not inspected. And “1,000 most downloaded” is a sample of popularity, not a measure of risk. Still, the result is useful: legitimate import lines are rare but real, so “ban all import lines” would break normal tools today, and a short, reviewed allowlist is a realistic policy.
Step 7: Try Python 3.15 .start files
PEP 829, written by Barry Warsaw and marked Final for Python 3.15, tries to make startup code easier to audit. Its motivation is blunt: import lines are run with exec() during startup, which in the PEP’s words “opens a broad attack surface”. The change has four parts. A new file type, name.start, lists entry points in the form pkg.mod:callable, one per line; Python imports the module and calls the function with no arguments. .pth files keep their path lines unchanged. During a deprecation period, a .start file with the same base name as a .pth file turns off the import lines in that .pth. And the import lines themselves are on a schedule. The 3.15 site documentation says: “import lines in name.pth files are deprecated and will be silently ignored in Python 3.18 and 3.19. In Python 3.20 a warning will be produced for import lines in name.pth files.”
The new files are processed after every path extension has been applied, so an entry point can live in a module reachable only through a .pth path. The documentation’s migration recipe is to move everything after the first semicolon of your import line into a zero-argument function, name it in a .start file, and, if you still support older Pythons, keep an import line of the form import pkg.mod; pkg.mod.callable() in the .pth file. Older Pythons run that line; Python 3.15 ignores it because the .start file wins.
The script below plays the part of a package author migrating a package called hello_pkg. It creates two virtual environments, old-env on your current Python and new-env on Python 3.15, installs the package files into both, and runs seven scenarios, labelled A to G. Note that initialize() guards itself, because Step 1 showed that a hook can be called twice, and because the docs say entry points are not de-duplicated.
# step07_start_files.py
"""Migrate a package from .pth import lines to a .start entry point, and compare Python 3.13 with 3.15.
Usage: python step07_start_files.py C:/path/to/python-3.15/python.exe
"""
import shutil
import subprocess
import sys
import textwrap
from labenv import clean_env, run, show, site_packages, venv_python
PY315 = sys.argv[1]
LINE_SEPARATOR = chr(0x2028)
HOOK = textwrap.dedent('''\
import sys
print("[hello_pkg] hook module imported", file=sys.stderr)
_done = False
def initialize():
"""Startup entry point. It can be called more than once, so it guards itself."""
global _done
if not _done:
_done = True
print("[hello_pkg] initialize() ran", file=sys.stderr)
def ping():
print("[hello_pkg] ping() ran", file=sys.stderr)
def boom():
raise RuntimeError("entry point failed on purpose")
''')
IMPORT = "import hello_pkg.hook; hello_pkg.hook.initialize()\n"
START = "hello_pkg.hook:initialize\n"
for name, interpreter in (("old-env", sys.executable), ("new-env", PY315)):
shutil.rmtree(name, ignore_errors=True)
subprocess.run([interpreter, "-m", "venv", "--without-pip", name], check=True, env=clean_env())
package = site_packages(name) / "hello_pkg"
package.mkdir()
(package / "__init__.py").write_text("", encoding="utf-8")
(package / "hook.py").write_text(HOOK, encoding="utf-8", newline="\n")
OLD, NEW = venv_python("old-env"), venv_python("new-env")
VERSIONS = (("3.13", OLD), ("3.15", NEW))
KEEP_V = ("import lines in", "Executing entry point", "Exec'ing")
def setup(files):
for env in ("old-env", "new-env"):
for stale in [*site_packages(env).glob("*.pth"), *site_packages(env).glob("*.start")]:
stale.unlink()
for name, text in files.items():
(site_packages(env) / name).write_text(text, encoding="utf-8", newline="\n")
def both(*args, keep=None):
for label, py in VERSIONS:
show(f"[{label}] python {' '.join(args)}", run(py, *args), keep=keep)
print("### A. an import line only (what packages ship today)")
setup({"hello_pkg.pth": IMPORT})
both("-c", "pass")
show("[3.15] python -v -c pass (only the lines about import lines)", run(NEW, "-v", "-c", "pass"), keep=KEEP_V)
print("\n### B. the migration: keep the import line and add a .start file")
setup({"hello_pkg.pth": IMPORT, "hello_pkg.start": START})
both("-c", "pass")
show("[3.15] python -v -c pass (only the lines about import lines)", run(NEW, "-v", "-c", "pass"), keep=KEEP_V)
both("-S", "-c", "print('user code')")
print("\n### C. a .start file alone")
setup({"hello_pkg.start": START})
both("-c", "pass")
print("\n### D. mistakes and duplicates inside a .start file (3.15)")
setup({"hello_pkg.start": "# entry points\nhello_pkg.hook:ping\nhello_pkg.hook\nnonexistent.mod:fn\nhello_pkg.hook:boom\n"
"import os; os.system('echo x')\nhello_pkg.hook:ping\n"})
show("[3.15] python -c \"print('user code')\"", run(NEW, "-c", "print('user code')"),
keep=("[hello_pkg]", "Invalid entry", "Error resolving", "Error in entry", "ValueError:", "RuntimeError:", "ModuleNotFoundError:", "user code"))
print("\n### E. a .pth line that names a folder that does not exist")
setup({"stale.pth": "missing-dir\n"})
both("-c", "pass")
print("\n### F. two lines that old and new Python read differently")
setup({"hidden.pth": "# harmless comment" + LINE_SEPARATOR + 'import sys; print("[hidden] ran", file=sys.stderr)\n',
"lead.pth": ' import sys; print("[leading space] ran", file=sys.stderr)\n'})
both("-c", "pass")
show("python -I -S pthaudit.py new-env", run(sys.executable, "-I", "-S", "pthaudit.py", "new-env"))
print("\n### G. can an audit hook installed by sitecustomize see the startup code?")
SITECUSTOMIZE = textwrap.dedent('''\
import sys
def watch(event, args):
if event in ("compile", "import") and "hello_pkg" in str(args[0]):
print(f"[audit] {event} event:", str(args[0])[:60], file=sys.stderr)
sys.addaudithook(watch)
print("[sitecustomize] audit hook installed", file=sys.stderr)
''')
setup({"hello_pkg.pth": IMPORT, "hello_pkg.start": START, "sitecustomize.py": SITECUSTOMIZE})
both("-c", "exec('import hello_pkg.hook')", keep=("[hello_pkg]", "[sitecustomize]", "[audit]"))
Run it, passing the path of your Python 3.15 interpreter (on Windows, something like C:\Python315\python.exe):
python step07_start_files.py /path/to/python3.15
Scenario A: an import line only
### A. an import line only (what packages ship today)
$ [3.13] python -c pass
[hello_pkg] hook module imported
[hello_pkg] initialize() ran
$ [3.15] python -c pass
[hello_pkg] hook module imported
[hello_pkg] initialize() ran
$ [3.15] python -v -c pass (only the lines about import lines)
import lines in <lab>\new-env\Lib\site-packages\hello_pkg.pth are deprecated, use entry points in a <lab>\new-env\Lib\site-packages\hello_pkg.start file instead.
Exec'ing from <lab>\new-env\Lib\site-packages\hello_pkg.pth: import hello_pkg.hook; hello_pkg.hook.initialize()
This is what packages ship today. Both versions run the hook. On 3.15, the -v output carries the only warning the interpreter gives: import lines in <path>.pth are deprecated, use entry points in a <path>.start file instead. The documentation says the diagnostic appears only with -v, so almost nobody will ever see it. The two lines from the hook are worth reading: the first comes from the module’s own top-level code, which runs when the module is imported, and the second is the function call.
Scenario B: the migration
### B. the migration: keep the import line and add a .start file
$ [3.13] python -c pass
[hello_pkg] hook module imported
[hello_pkg] initialize() ran
$ [3.15] python -c pass
[hello_pkg] hook module imported
[hello_pkg] initialize() ran
$ [3.15] python -v -c pass (only the lines about import lines)
import lines in <lab>\new-env\Lib\site-packages\hello_pkg.pth are suppressed due to matching <lab>\new-env\Lib\site-packages\hello_pkg.start file.
Executing entry point: hello_pkg.hook:initialize from <lab>\new-env\Lib\site-packages\hello_pkg.start
$ [3.13] python -S -c print('user code')
user code
$ [3.15] python -S -c print('user code')
user code
With both files present, Python 3.15 says the import lines are suppressed due to matching the .start file and calls the entry point instead. Python 3.13 never reads .start files, so it runs the import line, and the visible result is identical. That is the whole point of the straddle form. The last two commands run with -S, and neither version runs anything.
Scenario C: a .start file alone
### C. a .start file alone
$ [3.13] python -c pass
$ [3.15] python -c pass
[hello_pkg] hook module imported
[hello_pkg] initialize() ran
Common mistake: shipping only a .start file
On Python 3.13 the hook printed nothing. It silently did not run. A package that drops its import line and ships only a .start file works on 3.15 and quietly stops working on every older Python. Keep the straddle line until you drop support for those versions.
Scenario D: mistakes and duplicates
### D. mistakes and duplicates inside a .start file (3.15)
$ [3.15] python -c "print('user code')"
[hello_pkg] hook module imported
[hello_pkg] ping() ran
Invalid entry point syntax in <lab>\new-env\Lib\site-packages\hello_pkg.start: 'hello_pkg.hook'
ValueError: invalid format: 'hello_pkg.hook'
Error resolving entry point nonexistent.mod:fn from <lab>\new-env\Lib\site-packages\hello_pkg.start
ModuleNotFoundError: No module named 'nonexistent'
Error in entry point hello_pkg.hook:boom from <lab>\new-env\Lib\site-packages\hello_pkg.start
RuntimeError: entry point failed on purpose
Invalid entry point syntax in <lab>\new-env\Lib\site-packages\hello_pkg.start: "import os; os.system('echo x')"
ValueError: invalid format: "import os; os.system('echo x')"
[hello_pkg] ping() ran
user code
The .start file in this scenario has a comment and six entries, and the output shows what 3.15 does with each. The first entry, ping, runs. hello_pkg.hook has no colon and no function, so it is rejected with invalid format. nonexistent.mod:fn fails with ModuleNotFoundError. boom raises an exception inside the entry point. The line import os; os.system('echo x') is refused as an invalid entry point, which is the PEP’s narrowing at work: a .start file cannot carry arbitrary statements. Every failure prints a message with a traceback on standard error at every start, without -v, and processing continues. The last ping entry runs a second time because duplicates are not removed.
One more thing is easy to miss. Python imports the entry point’s module to find the function, so any top-level code in that module, and in the package’s __init__.py, runs at startup too. That is why our hook module’s own print appears before the first ping.
Scenario E: a stale path line
### E. a .pth line that names a folder that does not exist
$ [3.13] python -c pass
$ [3.15] python -c pass
In <lab>\new-env\Lib\site-packages\stale.pth: <lab>\new-env\Lib\site-packages\missing-dir does not exist; skipping sys.path append
In 3.15.0rc3, a .pth line that names a folder that does not exist now prints a message at every start; 3.13.14 stays silent. If you have leftover .pth files pointing at deleted folders (for example from an old editable install), expect new noise. That is a good reason to run the auditor and tidy up before upgrading.
Scenario F: two lines the versions read differently
### F. two lines that old and new Python read differently
$ [3.13] python -c pass
[hidden] ran
[hidden] ran
$ [3.15] python -c pass
[hidden] ran
[leading space] ran
$ python -I -S pthaudit.py new-env
[SUSPICIOUS] hidden.pth (owner: UNOWNED; integrity: n/a; flags: unusual-line-separator)
line 2: import sys; print("[hidden] ran", file=sys.stderr)
[UNOWNED] lead.pth (owner: UNOWNED; integrity: n/a; flags: leading-whitespace-import)
line 1: import sys; print("[leading space] ran", file=sys.stderr)
summary: 1 SUSPICIOUS, 1 UNOWNED
(exit code 2)
The comment-then-U+2028 trick from Step 2 works on 3.15 as well, which is no surprise, since 3.15 also uses splitlines(). The leading-space line ran only on 3.15, matching the table in Step 2. The auditor flags both files: hidden.pth as SUSPICIOUS because of the unusual line separator, and lead.pth with the leading-whitespace-import flag.
Scenario G: can an audit hook see any of this?
If you read How to Catch a Malicious Python Package at Import Time With Audit Hooks, you may wonder whether a runtime hook could watch startup code. PEP 829 says that because entry points go through the import system, “the standard audit hooks (PEP 578) can provide monitoring”. Here is what happens when the hook is installed the usual way, from sitecustomize.py:
### G. can an audit hook installed by sitecustomize see the startup code?
$ [3.13] python -c exec('import hello_pkg.hook')
[hello_pkg] hook module imported
[hello_pkg] initialize() ran
[sitecustomize] audit hook installed
[audit] compile event: b"exec('import hello_pkg.hook')\n"
[audit] compile event: b'import hello_pkg.hook'
$ [3.15] python -c exec('import hello_pkg.hook')
[hello_pkg] hook module imported
[hello_pkg] initialize() ran
[sitecustomize] audit hook installed
[audit] compile event: b"exec('import hello_pkg.hook')\n"
[audit] compile event: b'import hello_pkg.hook'
Both versions printed the hook’s two lines before the line saying the audit hook was installed. The documentation explains why: Python imports sitecustomize after the path setup, so everything the .pth and .start files do is already finished. The hook did work for the code that ran later: it printed events for the -c source and for the exec call in it (the source arrives as bytes, hence the b'...' form). So a hook installed from sitecustomize.py cannot watch startup files. Only a hook installed earlier, for example by an application that embeds Python, can, which is why auditing the files on disk matters.
What .start files change, and what they do not
They separate two jobs that import lines mixed together: extending the search path and running code. A directory listing now tells you where code runs: the .start files, plus whatever import lines remain in .pth files. The entry points are named, they are importable functions rather than arbitrary strings, and the PEP leaves room for a future policy that could allow or deny them. It is also honest about its limits: “The overall pre-start code execution attack surface is not eliminated by this PEP.” A malicious package can still name a function that does bad things. The auditor therefore treats a .start file as code, shows each entry point, and flags an entry that does not match the pkg.mod:callable form.
One compatibility detail from the 3.15 source: when it runs import lines, site.py defines the local names sitedir and fullname first, with a comment that setuptools’ -nspkg.pth files read sys._getframe(1).f_locals['sitedir']. That is the pattern in the Google and Sphinx files from the survey, and it is one reason import lines will not disappear overnight.
Step 8: Turn the audit into a gate and test it
An audit you run once is a report. An audit that fails a build is a control. The script below removes our two hand-made demo hooks (as an investigator would after reading them), writes an allowlist from the files that need review, and then checks the environment against it. Finally it installs the demo wheel as if it were a new dependency and checks again.
# step08_gate.py
"""Turn the audit into a gate: an allowlist of reviewed files plus exit codes a CI job can act on."""
import json
import pathlib
import sys
import make_demo_wheel
from labenv import run, show, site_packages, venv_python
ENV = pathlib.Path("demo-env")
SP = site_packages(ENV)
for name in ("hello_startup.pth", "hidden.pth"): # the hand-made hooks, removed after reading them
(SP / name).unlink()
def audit(*extra):
return run(sys.executable, "-I", "-S", "pthaudit.py", str(ENV), *extra)
show("python -I -S pthaudit.py demo-env --write-allowlist allow.json", audit("--write-allowlist", "allow.json"))
print(pathlib.Path("allow.json").read_text(encoding="utf-8"), end="")
show("python -I -S pthaudit.py demo-env --allowlist allow.json", audit("--allowlist", "allow.json"))
print("\n--- a new dependency arrives ---")
wheel = make_demo_wheel.build()
show("pip install --no-deps demo_startup-1.0-py3-none-any.whl", run(venv_python(ENV), "-m", "pip", "install", "-q", "--no-deps", wheel))
show("python -I -S pthaudit.py demo-env --allowlist allow.json", audit("--allowlist", "allow.json"))
python step08_gate.py
$ python -I -S pthaudit.py demo-env --write-allowlist allow.json
[REVIEW] __editable__.flat_demo-0.1.pth (owner: flat_demo 0.1; integrity: ok; flags: -)
line 1: import __editable___flat_demo_0_1_finder; __editable___flat_demo_0_1_finder.install()
[path-only] __editable__.src_demo-0.1.pth (owner: src_demo 0.1; integrity: ok; flags: -)
[REVIEW] distutils-precedence.pth (owner: setuptools 84.0.0; integrity: ok; flags: dynamic-import)
line 1: import os; var = 'SETUPTOOLS_USE_DISTUTILS'; enabled = os.environ.get(var, 'local') == 'local'; enab...
summary: 2 REVIEW, 1 path-only
wrote 2 entries to allow.json: read each one before you commit this file
(exit code 1)
{
"__editable__.flat_demo-0.1.pth": "1a23810d0cc55615e8485bf2f2f41d2978320ed2b27642d7b108fb5c1ac14c80",
"distutils-precedence.pth": "2638ce9e2500e572a5e0de7faed6661eb569d1b696fcba07b0dd223da5f5d224"
}
$ python -I -S pthaudit.py demo-env --allowlist allow.json
[allowed] __editable__.flat_demo-0.1.pth (owner: flat_demo 0.1; integrity: ok; flags: -)
line 1: import __editable___flat_demo_0_1_finder; __editable___flat_demo_0_1_finder.install()
[path-only] __editable__.src_demo-0.1.pth (owner: src_demo 0.1; integrity: ok; flags: -)
[allowed] distutils-precedence.pth (owner: setuptools 84.0.0; integrity: ok; flags: dynamic-import)
line 1: import os; var = 'SETUPTOOLS_USE_DISTUTILS'; enabled = os.environ.get(var, 'local') == 'local'; enab...
summary: 2 allowed, 1 path-only
--- a new dependency arrives ---
$ pip install --no-deps demo_startup-1.0-py3-none-any.whl
$ python -I -S pthaudit.py demo-env --allowlist allow.json
[allowed] __editable__.flat_demo-0.1.pth (owner: flat_demo 0.1; integrity: ok; flags: -)
line 1: import __editable___flat_demo_0_1_finder; __editable___flat_demo_0_1_finder.install()
[path-only] __editable__.src_demo-0.1.pth (owner: src_demo 0.1; integrity: ok; flags: -)
[SUSPICIOUS] demo_startup_hook.pth (owner: demo_startup 1.0; integrity: ok; flags: exec, base64)
line 1: import base64; exec(base64.b64decode('cHJpbnQoIltkZW1vLXN0YXJ0dXBdIHBheWxvYWQgcmFuIGF0IGludGVycHJldG...
[allowed] distutils-precedence.pth (owner: setuptools 84.0.0; integrity: ok; flags: dynamic-import)
line 1: import os; var = 'SETUPTOOLS_USE_DISTUTILS'; enabled = os.environ.get(var, 'local') == 'local'; enab...
summary: 1 SUSPICIOUS, 2 allowed, 1 path-only
(exit code 2)
The first audit exits with code 1 and writes allow.json, which maps two file names to the SHA-256 of the reviewed bytes. Read each entry before you commit a file like this; the tool cannot judge whether a file deserves trust, only record that you decided it does. The second audit passes with no exit code printed, which means 0. Then a “new dependency” arrives with its own startup file, and the gate fails with code 2 and names the file. That is the behaviour you want in CI: any CI system can run python -I -S pthaudit.py .venv --allowlist allow.json and stop the build on a non-zero exit code. Keep allow.json in version control, outside the environment it describes, so that a change to it shows up in code review. Pair it with hash-pinned installs such as the ones in How to Lock Python Dependencies With pip’s pylock.toml so that what you reviewed is what gets installed.
Test the auditor
The auditor parses hostile input, so it deserves tests. This file covers the parsing rules, the hidden line, the owner and integrity checks, the allowlist, the .start validation, wheels with files in .data folders and the exit codes. The last test asks Python 3.15 itself: it feeds the same files to site.StartupState and checks that the interpreter and the auditor agree about which lines are import lines. That test reads private attributes of StartupState, which the class’s own docstring in site.py calls intentionally private, so treat it as a check for this lab rather than an API to build on. It is skipped on older Pythons.
# test_pthaudit.py
import base64
import hashlib
import io
import json
import sys
import zipfile
from pathlib import Path
import pytest
import pthaudit
import wheel_check
LS = chr(0x2028) # LINE SEPARATOR
def record_hash(data):
return "sha256=" + base64.urlsafe_b64encode(hashlib.sha256(data).digest()).rstrip(b"=").decode()
def make_site(tmp_path, files, owner=("pkg", "1.0"), tamper=None):
"""A fake site-packages folder whose RECORD lists `files` with correct hashes (or a changed file if `tamper`)."""
site = tmp_path / "site-packages"
site.mkdir(parents=True)
info = site / f"{owner[0]}-{owner[1]}.dist-info"
info.mkdir()
rows = [f"{name},{record_hash(data)},{len(data)}" for name, data in files.items()]
(info / "RECORD").write_text("\n".join(rows) + "\n", encoding="utf-8")
for name, data in files.items():
(site / name).write_bytes(tamper if tamper and name == "x.pth" else data)
return site
def kinds(text):
return [line.kind for line in pthaudit.parse_pth(text)]
def test_import_needs_a_space_or_a_tab():
assert kinds("import os\nimport\tsys\nimport(os)\nimportos\n") == ["import", "import", "path", "path"]
def test_comments_and_blank_lines():
assert kinds("# note\n\n \n/some/dir\n") == ["comment", "blank", "blank", "path"]
def test_import_line_hidden_behind_a_comment_is_found():
text = "# harmless comment" + LS + "import sys\n"
assert len(list(io.StringIO(text))) == 1 # a line-by-line reader sees one comment line
assert kinds(text) == ["comment", "import"] # site.py splits with str.splitlines()
def test_leading_whitespace_import_is_reported_with_a_note():
line = pthaudit.parse_pth(" import sys\n")[0]
assert (line.kind, line.note) == ("import", "leading whitespace")
def test_bom_is_accepted_and_other_encodings_are_flagged():
text, encoding = pthaudit.decode_startup_file(b"\xef\xbb\xbfimport sys\n")
assert text == "import sys\n" and encoding == "utf-8"
kind, code, flags = pthaudit.inspect_bytes("legacy.pth", "import sys # caf\xe9\n".encode("cp1252"))
assert "non-utf8" in flags and len(code) == 1
def test_risky_import_line_gets_flags():
payload = base64.b64encode(b"print(1)").decode()
kind, code, flags = pthaudit.inspect_bytes("evil.pth", f"import base64; exec(base64.b64decode('{payload}'))\n".encode())
assert kind == "pth" and {"exec", "base64"} <= set(flags)
def test_owner_and_integrity_are_checked_against_record(tmp_path):
data = b"import sys\n"
site = make_site(tmp_path, {"x.pth": data})
(found,) = pthaudit.audit_sitedir(site)
assert (found.owner, found.integrity, found.status) == ("pkg 1.0", "ok", "REVIEW")
changed_site = make_site(tmp_path / "changed", {"x.pth": data}, tamper=b"import sys; print(1)\n")
(changed,) = pthaudit.audit_sitedir(changed_site)
assert changed.status == "TAMPERED"
def test_a_hook_that_no_package_owns_is_flagged(tmp_path):
site = tmp_path / "site-packages"
site.mkdir()
(site / "dropped.pth").write_text("import sys\n", encoding="utf-8")
(found,) = pthaudit.audit_sitedir(site)
assert found.status == "UNOWNED"
def test_path_only_files_are_quiet(tmp_path):
site = make_site(tmp_path, {"x.pth": b"some/dir\n"})
assert [f.status for f in pthaudit.audit_sitedir(site)] == ["path-only"]
def test_allowlist_needs_the_same_hash(tmp_path):
data = b"import sys\n"
site = make_site(tmp_path, {"x.pth": data})
allowed = {"x.pth": hashlib.sha256(data).hexdigest()}
assert pthaudit.audit_sitedir(site, allowed)[0].status == "allowed"
assert pthaudit.audit_sitedir(site, {"x.pth": "0" * 64})[0].status == "REVIEW"
def test_start_file_entries_are_validated():
kind, code, flags = pthaudit.inspect_bytes("p.start", b"# c\npkg.mod:run\npkg.mod\nimport os; os.system('x')\n")
assert kind == "start" and [text for _, text in code] == ["pkg.mod:run", "pkg.mod", "import os; os.system('x')"]
assert "invalid-entry" in flags
assert "invalid-entry" not in pthaudit.inspect_bytes("p.start", b"pkg.mod:run\n")[2]
def test_wheel_check_finds_startup_files_in_data_directories():
buffer = io.BytesIO()
with zipfile.ZipFile(buffer, "w") as zf:
zf.writestr("pkg/__init__.py", "")
zf.writestr("pkg-1.0.data/purelib/deep.pth", "import sys\n")
zf.writestr("sitecustomize.py", "print('hi')\n")
zf.writestr("pkg-1.0.dist-info/METADATA", "x")
with zipfile.ZipFile(buffer) as zf:
found = wheel_check.inspect_wheel(zf)
assert sorted((member, kind) for member, kind, _, _ in found) == [("pkg-1.0.data/purelib/deep.pth", "pth"), ("sitecustomize.py", "sitecustomize")]
def test_exit_codes(tmp_path, capsys):
clean = make_site(tmp_path / "clean", {"x.pth": b"some/dir\n"})
assert pthaudit.main([str(clean)]) == 0
review = make_site(tmp_path / "review", {"x.pth": b"import sys\n"})
assert pthaudit.main([str(review)]) == 1
allow = tmp_path / "allow.json"
assert pthaudit.main([str(review), "--write-allowlist", str(allow)]) == 1
assert pthaudit.main([str(review), "--allowlist", str(allow)]) == 0
bad = make_site(tmp_path / "bad", {"x.pth": b"import sys\n"}, tamper=b"import sys; print(1)\n")
assert pthaudit.main([str(bad), "--allowlist", str(allow)]) == 2
capsys.readouterr()
@pytest.mark.skipif(sys.version_info < (3, 15), reason="needs Python 3.15's site.StartupState")
def test_agrees_with_the_interpreter_on_3_15(tmp_path):
"""Uses private attributes of site.StartupState, so it is a check for this lab, not an API to build on."""
import site
corpus = {
"a.pth": "import os\nimport\tsys\nimport(os)\n/some/path\n",
"b.pth": " import sys\n# note\n",
"c.pth": "# hidden" + LS + "import sys\n",
"d.pth": "import sys; x = 1\n\n\n",
}
for name, text in corpus.items():
(tmp_path / name).write_text(text, encoding="utf-8", newline="\n")
state = site.StartupState(set())
state.addsitedir(str(tmp_path))
theirs = {Path(k).name: v for k, v in state._importexecs.items()}
mine = {}
for name, text in corpus.items():
found = [line.text.strip() for line in pthaudit.parse_pth(text) if line.kind == "import"]
if found:
mine[name] = found
assert mine == theirs
python -m pytest -q test_pthaudit.py
On Python 3.13.14:
13 passed, 1 skipped in 0.09s
On Python 3.15.0rc3 (in a virtual environment with pytest installed):
14 passed in 0.10s
What this auditor does not cover
It reads only .pth, .start, sitecustomize.py and usercustomize.py files in one environment’s own site-packages. It does not look at the system site folders when an environment sets include-system-site-packages, at the user site, at a sitecustomize found elsewhere on sys.path (for instance through PYTHONPATH), or at PYTHONSTARTUP. It does not read the modules that a .start entry point names; you must review those yourself, including their top-level code. The risk flags are heuristics for triage, and a determined author can avoid the keywords, so the allowlist of reviewed hashes is the real control. The integrity check trusts RECORD, which an attacker with write access can edit along with the file. And the wheel checker sees only wheels, not source distributions. Finally, everything here was measured on one Windows 11 machine with Python 3.13.14 and 3.15.0rc3; I did not run it on macOS or Linux.
Confirm it all works end to end
Run through this list once. First, python step01_demo_hook.py prints the hook line before user code in every run except the one with -S. Second, python step02_naive_scan.py reports zero import lines in hidden.pth while Python runs one. Third, python -I -S pthaudit.py demo-env, run before Step 8 deletes the demo hooks, exits with a non-zero code and lists hidden.pth as SUSPICIOUS. Fourth, python wheel_check.py demo_startup-1.0-py3-none-any.whl exits with 2. Fifth, python -m pytest -q test_pthaudit.py passes. Sixth, and most useful: run python -I -S pthaudit.py against a real project environment, read the output, write an allowlist you trust and commit it.
Where to go next
If you want to go deeper on what happens after startup code runs, the audit hooks tutorial shows how to watch a package’s behaviour at import time. To make fresh malicious releases less likely to reach you in the first place, see How to Set Up Dependency Cooldowns in npm and pip. Startup files also cost time on every launch, which connects to How to Cut a Python Command Line Tool’s Startup Time With Python 3.15 Lazy Imports. And because 3.15 deprecates reading .pth files in anything other than UTF-8, the encoding facts in How to Prepare Your Python Code for the Python 3.15 UTF-8 Default are worth a read before you upgrade.








No Comment! Be the first one.