TRENDING
Galvanized steel guardrail bolted to wooden posts along the edge of a bridge approach, with a grassy verge and a gravel road beside it
October 1, 2026
How to Enforce Guardrails on AI-Generated Terraform With Open Policy Agent and Rego
Microscope die shot of an AMD EPYC 7702 engineering sample I/O die, its circuit blocks glowing in teal, gold and violet
October 1, 2026
AMD Agrees to Buy Fei-Fei Li’s World Labs for $8.2 Billion to Steer Its Chip Roadmap
A silver signet ring engraved with a coat of arms between two sticks of red sealing wax on a grey surface
October 1, 2026
How to Build a Merkle Tree Certificate Issuer in Python to Keep Post-Quantum Certificates Small
Brass swing-bar door lock, a secondary latch, mounted on a hotel room door
October 1, 2026
Cloudflare’s Post-Quantum Visibility Turns Quantum Readiness Into a Per-Hop Audit
A seven-spot ladybird with black spots on its orange shell climbs a green plant stem
October 1, 2026
OpenAI Launches Dots, Always-On Agents, and Says It Is Still Fixing Known Vulnerabilities
01 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Faint white watermark of a crown above an oval emblem showing through blue paper, a design that stays invisible until light passes through the sheet
How to Detect and Strip Invisible Unicode in Python to Stop ASCII Smuggling and Trojan Source
September 30, 2026
A small white wooden toll booth with a Pay Point sign and a fare board at Penmaenpool Toll Bridge, with orange traffic cones on the bridge deck
Two Cloudflare Agent Billing Betas Turn Web Monetization Into a Question of Who Holds the Meter
September 30, 2026
Eight silver hex keys of graduated sizes fanned out on a steel ring against a dark green surface
Attackers Exploit a Hex-Encoding Bypass in Cisco SD-WAN Manager, and CISA Sets an October 3 Deadline
September 30, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 216 Posts
News 218 Posts
Learning Hub 188 Posts
Home/Learning Hub/How to Prevent Path Traversal in Python File Downloads and Archive Extraction
Learning Hub

How to Prevent Path Traversal in Python File Downloads and Archive Extraction

This tutorial reproduces path traversal in a FastAPI download endpoint, shows four plausible-looking checks failing before a fifth holds, and builds a tested safe_join() function plus safe tar and...

September 28, 2026 29 Min Read
19

Somewhere in almost every web application there is a route that hands a file back to the user: a report download, a profile picture, a log viewer, a template loader. Somewhere else there is code that unpacks a file the user uploaded. In both places the application takes a piece of text chosen by a stranger and turns it into a location on your disk. If that translation is careless, the stranger can steer it out of the folder you meant to expose and into one you did not, such as the folder that holds your API keys. The bug class is called path traversal (also known as directory traversal), and MITRE catalogs it as CWE-22, titled Improper Limitation of a Pathname to a Restricted Directory (‘Path Traversal’).

Table Of Content

  • Prerequisites
  • A note on operating systems
  • What path traversal is, in plain language
  • Step 1: Build a practice lab
  • Step 2: Write a download endpoint with the classic bug
  • Step 3: Attack it, and see why your first attempt may not work
  • Gotcha: your own test client may be quietly fixing the attack
  • Step 4: Why the bug works, starting with what os.path.join really does
  • Step 5: Removing :path is not a fix
  • Step 6: Five checks that look right, and only one is
  • Check 1: refuse any path containing two dots
  • Check 2: resolve the path, then use startswith
  • Check 3: is_relative_to without resolving
  • Check 4: abspath plus commonpath
  • Check 5: resolve, then is_relative_to
  • One more Windows trap: resolve both sides
  • Step 7: The fix, resolve first and then check containment
  • Step 8: Or let your framework do it
  • Step 9: The same bug in a different costume, archive extraction
  • 9.1 Build some hostile archives
  • 9.2 What tar extraction does
  • What the default does on Python 3.13
  • What the tar and data filters do
  • 9.3 A refused archive can still leave files behind
  • 9.4 The zipfile module plays by different rules
  • Step 10: Lock it in with tests
  • Common mistakes and gotchas
  • How to confirm everything works end to end
  • Next steps

This is not a museum piece. This month sxz.io covered a maximum-severity GitLab flaw that GitLab itself described as a path traversal issue in the repository commits API, and a WordPress Core fix for an unauthenticated local file inclusion bug in page template resolution that had been present since 2016. Different products, closely related bug classes: in both, text an attacker could influence ended up naming a file the software should never have opened.

In this tutorial you will build a small download endpoint on purpose with the classic bug, attack it, and then fix it properly. Along the way you will watch four plausible-looking checks fail before a fifth one holds, learn why a check that works on Linux can quietly fail on Windows, and then apply the same thinking to archives, where the same bug is famous under the name Zip Slip. You will use the extraction filters that the Python tarfile module gained in version 3.12, and you will finish with a small, pytest-verified module you can adapt to your own project.

Prerequisites

  • Python 3.12 or newer. I built and tested everything on Python 3.13.14 on Windows 11. Python 3.12 is the release that added the filter argument to tar extraction, which the archive half of this tutorial relies on.
  • Four packages: fastapi 0.141.1, uvicorn 0.54.0, requests 2.34.2, and pytest 9.1.1. Those are the versions I used. FastAPI installs Starlette (1.7.0 here), and we will look inside it later.
  • Two terminal windows open in the same project folder, one for the server and one for the attacker. One demo also uses the curl command (I used curl 8.21.0), but you can skip it.
  • Basic Python and HTTP knowledge. You should know what a function is and what a URL path looks like. No security background is assumed, and every term is defined when it first appears.

A note on operating systems

Some behavior differs by platform, and I only ran this on Windows 11, so I mark every spot where Linux or macOS would differ. On those systems a backslash is an ordinary filename character rather than a folder separator, and there are no drive letters. Creating a symbolic link on Windows also needs extra permission. The Python documentation for os.symlink says: “On newer versions of Windows 10, unprivileged accounts can create symlinks if Developer Mode is enabled. When Developer Mode is not available/enabled, the SeCreateSymbolicLinkPrivilege privilege is required, or the process must be run as an administrator.” The lab script in Step 1 tries to create one symlink and reports whether it worked, and every later step copes if it could not. It worked on my machine because I ran as an administrator.

Create a project folder, then create and activate a virtual environment and install the packages:

python -m venv .venv
.venv\Scripts\activate           # on macOS or Linux: source .venv/bin/activate
pip install fastapi uvicorn requests pytest

What path traversal is, in plain language

A filesystem path is just text that the operating system interprets when you open a file. Three details of that interpretation cause almost every traversal bug:

  • The segment .. means “the parent folder”. Opening public/../secrets/api_key.txt climbs out of public and into its neighbor secrets. The operating system does this at the moment the file is opened, not when your program builds the string.
  • An absolute path starts from the top of a drive or filesystem instead of from a folder you chose, for example C:\data\file.txt or /var/log/syslog. Many path-joining functions treat an absolute path as an instruction to forget everything that came before it.
  • A symbolic link (symlink) is a special file that acts as an alias for another location. A path can look perfectly innocent, with no dots at all, and still end up somewhere else because one of its folders is a symlink.

What your program wants is a promise, called containment: whatever the user asks for, I will only touch files inside this one base directory. The OWASP description of the attack puts it this way: “A path traversal attack (also known as directory traversal) aims to access files and directories that are stored outside the web root folder. By manipulating variables that reference files with “dot-dot-slash (../)” sequences and its variations or by using absolute file paths, it may be possible to access arbitrary files and directories stored on file system including application source code or configuration and critical system files.”

The plan for the rest of this tutorial is to establish containment correctly. The key idea is to canonicalize first: turn the messy, user-supplied text into the single absolute location the operating system would really open, with every .. removed and every symlink followed. Only then do you ask whether that location is inside your base directory.

Step 1: Build a practice lab

Save every file from this tutorial in the same project folder, starting with the following, which you should name lab.py. Every demo imports it. It builds a tiny tree of folders and files that we can attack safely, and it wipes and rebuilds that tree each time it runs, so you can always get a clean slate by running python lab.py.

"""Builds the small file tree that every demo in this tutorial attacks."""
import os
import shutil
from pathlib import Path

HERE = Path(__file__).resolve().parent
LAB = HERE / "lab"
PUBLIC = LAB / "public"            # the only directory we mean to serve
SECRETS = LAB / "secrets"          # a sibling we must never serve
BACKUP = LAB / "public-backup"     # a sibling whose name STARTS with "public"
SECRET_FILE = SECRETS / "api_key.txt"


def build() -> bool:
    """(Re)create the lab tree. Returns True if the symlink could be created."""
    if LAB.exists():
        shutil.rmtree(LAB)
    (PUBLIC / "notes" / "2026").mkdir(parents=True)
    SECRETS.mkdir()
    BACKUP.mkdir()
    (PUBLIC / "report.txt").write_text("Q3 report: revenue is up\n")
    (PUBLIC / "notes" / "2026" / "q3.txt").write_text("Q3 notes: ship the fix\n")
    (BACKUP / "old.txt").write_text("OLD BACKUP: internal only\n")
    SECRET_FILE.write_text("API_KEY=sk_live_demo_do_not_share\n")
    try:
        # A shortcut inside the public folder that points at the secrets folder.
        os.symlink(SECRETS, PUBLIC / "shortcut", target_is_directory=True)
        return True
    except (OSError, NotImplementedError):
        return False


if __name__ == "__main__":
    linked = build()
    for path in sorted(LAB.rglob("*")):
        print(path.relative_to(HERE).as_posix())
    print("symlink created:", linked)

Run it:

python lab.py

You should see this tree. The last line says False if your account cannot create symlinks, which is fine:

lab/public
lab/public/notes
lab/public/notes/2026
lab/public/notes/2026/q3.txt
lab/public/report.txt
lab/public/shortcut
lab/public-backup
lab/public-backup/old.txt
lab/secrets
lab/secrets/api_key.txt
symlink created: True

Here is what each piece is for. lab/public is the only folder our application is supposed to serve. lab/secrets holds api_key.txt, a fake key, and must never be served. lab/public-backup is a sibling of public whose name happens to start with the same letters, which will matter in Step 6. lab/public/shortcut is a symlink that points at lab/secrets. The recursive listing shows it as a single entry because it does not descend into symlinked folders, but anything that opens a path through it lands in the secrets folder.

Step 2: Write a download endpoint with the classic bug

Now write the vulnerable application. Save this as app_v1.py:

import os

from fastapi import FastAPI, HTTPException
from fastapi.responses import PlainTextResponse

from lab import PUBLIC

app = FastAPI()


@app.get("/files/{name:path}", response_class=PlainTextResponse)
def read_file(name: str) -> str:
    path = os.path.join(PUBLIC, name)  # VULNERABLE: trusts the user's path
    try:
        with open(path, encoding="utf-8") as f:
            return f.read()
    except (FileNotFoundError, IsADirectoryError):
        raise HTTPException(status_code=404, detail="Not found")

Read it slowly, because the bug is one line. The route is /files/{name:path}. The :path part tells the router (Starlette, underneath FastAPI) that name may contain slashes, which is what lets a request like /files/notes/2026/q3.txt reach subfolders. The handler then calls os.path.join(PUBLIC, name), which glues the base folder and the user’s text together, and opens the result. Nothing checks where that path actually leads. The try block only turns a missing file into a clean 404.

Start the server in your first terminal and leave it running:

python -m uvicorn app_v1:app --port 8000

Step 3: Attack it, and see why your first attempt may not work

Now play the attacker. Save the following as attack.py. It sends a list of hostile paths to the server and prints the status code and the first few characters of each response. It uses http.client from the standard library on purpose, for a reason you will see in a moment.

"""Fire a list of hostile paths at a running demo server.

http.client sends the request path exactly as written. Many clients (curl,
requests) collapse ".." before sending, which would hide the bug from you.
"""
import http.client
import sys
from urllib.parse import quote

from lab import SECRET_FILE


def fetch(path: str, port: int) -> tuple[int, str]:
    conn = http.client.HTTPConnection("127.0.0.1", port, timeout=5)
    conn.request("GET", path)
    resp = conn.getresponse()
    body = resp.read().decode("utf-8", "replace").strip().replace("\n", " ")
    conn.close()
    return resp.status, body


PAYLOADS = {
    "legit file": "report.txt",
    "legit subfolder": "notes/2026/q3.txt",
    "dot-dot": "../secrets/api_key.txt",
    "dot-dot, percent-encoded": "..%2Fsecrets%2Fapi_key.txt",
    "backslashes (Windows)": "..%5Csecrets%5Capi_key.txt",
    "absolute path": quote(str(SECRET_FILE), safe=""),
    "sibling directory": "../public-backup/old.txt",
    "through a symlink": "shortcut/api_key.txt",
}

if __name__ == "__main__":
    prefix = sys.argv[1] if len(sys.argv) > 1 else "/files/"
    port = int(sys.argv[2]) if len(sys.argv) > 2 else 8000
    for label, payload in PAYLOADS.items():
        status, body = fetch(prefix + payload, port)
        print(f"{label:26} {status}  {body[:42]!r}")

In your second terminal, run it:

python attack.py

This is what I got:

legit file                 200  'Q3 report: revenue is up'
legit subfolder            200  'Q3 notes: ship the fix'
dot-dot                    200  'API_KEY=sk_live_demo_do_not_share'
dot-dot, percent-encoded   200  'API_KEY=sk_live_demo_do_not_share'
backslashes (Windows)      200  'API_KEY=sk_live_demo_do_not_share'
absolute path              200  'API_KEY=sk_live_demo_do_not_share'
sibling directory          200  'OLD BACKUP: internal only'
through a symlink          200  'API_KEY=sk_live_demo_do_not_share'

The first two rows are honest requests, and they work. Every other row is an attack, and every one of them returned real file contents. Here is what each one does:

  • dot-dot. The request path is /files/../secrets/api_key.txt, so name is ../secrets/api_key.txt. The joined path keeps the .. in it, so when Windows opens the file it walks up out of public and into secrets.
  • dot-dot, percent-encoded. The same attack written as ..%2Fsecrets%2Fapi_key.txt. Servers decode %2F back into a slash before your code runs, so filtering the raw URL text for ../ would miss this form.
  • backslashes (Windows). On Windows a backslash separates folders just like a forward slash. Linux and macOS treat it as an ordinary character, so I would expect this row to return 404 there. It is a Windows-specific hole.
  • absolute path. The attacker sends the full path of the secret file. As you will see in Step 4, os.path.join throws away the base folder when it meets an absolute path.
  • sibling directory. ../public-backup/old.txt reaches a different folder that merely lives next to public.
  • through a symlink. shortcut/api_key.txt contains no dots and no drive letter, yet it lands in secrets, because the operating system follows the symlink.

Gotcha: your own test client may be quietly fixing the attack

Many people first try an attack like this with requests or curl, see a 404, and decide the endpoint is safe. That conclusion would be wrong, because both tools tidy up .. segments in the URL before sending anything. Save this as client_normalisation.py and run it while the server is still up:

"""Show that common clients collapse '..' before the request ever leaves your machine."""
import subprocess

import requests

URL = "http://127.0.0.1:8000/files/../secrets/api_key.txt"

prepared = requests.Request("GET", URL).prepare()
print("requests sends  :", prepared.url)

for flags in ([], ["--path-as-is"]):
    out = subprocess.run(
        ["curl", "-s", "-o", "-", "-w", " [HTTP %{http_code}]", *flags, URL],
        capture_output=True, text=True,
    ).stdout.strip().replace("\n", " ")
    print(f"curl {' '.join(flags) or '(default)':13}:", out)
python client_normalisation.py
requests sends  : http://127.0.0.1:8000/secrets/api_key.txt
curl (default)    : {"detail":"Not Found"} [HTTP 404]
curl --path-as-is : API_KEY=sk_live_demo_do_not_share  [HTTP 200]

The first line shows what requests would actually send: the .. was already collapsed, so the request asked for /secrets/api_key.txt, which does not exist. Plain curl got a 404 for the same reason. Only curl --path-as-is sent the path untouched and got the secret. The curl help text describes that flag as: “Do not squash .. sequences in URL path”. The http.client module in attack.py also sends the path exactly as written, which is why I used it. The lesson: when you test for traversal, make sure your client is not helping the server.

Step 4: Why the bug works, starting with what os.path.join really does

Why did the absolute-path attack work? Save this as join_demo.py. It uses a made-up base folder and never touches the disk:

import os

base = r"C:\srv\public"  # a made-up path: nothing here touches the disk
for part in ["report.txt", r"..\secret.txt", "/etc/passwd", r"D:\other\x.txt"]:
    print(f"{part:16} -> {os.path.join(base, part)}")
python join_demo.py
report.txt       -> C:\srv\public\report.txt
..\secret.txt    -> C:\srv\public\..\secret.txt
/etc/passwd      -> C:/etc/passwd
D:\other\x.txt   -> D:\other\x.txt

The Python documentation for os.path.join explains the surprising rows: “If a segment is an absolute path (which on Windows requires both a drive and a root), then all previous segments are ignored and joining continues from the absolute path segment.” It adds a Windows-specific rule: “On Windows, the drive is not reset when a rooted path segment (e.g., r'\foo') is encountered.” That is exactly what the third row shows. The segment /etc/passwd has a root but no drive, so the drive letter is kept and the result is C:/etc/passwd, a location at the top of the drive and nowhere near the base folder. The fourth row carries its own drive, so the base folder disappears entirely.

The deeper point is that join is string manipulation with a few special cases. It checks nothing. Your program builds a string, the operating system interprets it later, and everything between those two moments is unguarded.

Step 5: Removing :path is not a fix

A tempting shortcut is to remove :path so that slashes never reach the handler. Stop the server (Ctrl+C in the first terminal), copy app_v1.py to app_v1_plain.py, and change the route from /files/{name:path} to /files/{name}. Start it with python -m uvicorn app_v1_plain:app --port 8000 and run python attack.py again:

legit file                 200  'Q3 report: revenue is up'
legit subfolder            404  '{"detail":"Not Found"}'
dot-dot                    404  '{"detail":"Not Found"}'
dot-dot, percent-encoded   404  '{"detail":"Not Found"}'
backslashes (Windows)      200  'API_KEY=sk_live_demo_do_not_share'
absolute path              200  'API_KEY=sk_live_demo_do_not_share'
sibling directory          404  '{"detail":"Not Found"}'
through a symlink          404  '{"detail":"Not Found"}'

This looks like progress. The dot-dot rows, the sibling row and the symlink row now return 404, because the router refuses to match a path with a forward slash in it. But the legitimate subfolder request broke too (row 2), and two attacks still work: the backslash row and the absolute-path row. A Windows path made only of backslashes contains no forward slash for the router to object to, and a percent-encoded C:\... path is the same. The router is not a security control. It only stops one character. (I only tested Windows. On Linux and macOS an absolute path is full of forward slashes, so I would expect that row to return 404 there.) And the day someone needs subfolders and adds :path back, every attack from Step 3 returns. Stop this server before you continue.

Step 6: Five checks that look right, and only one is

Before writing the fix, it is worth seeing why fixing this is harder than it looks. Below are five checks, each only a few lines long, each reasonable at a glance. The script runs all five against nine payloads (three legitimate, six hostile) and reports whether each check allowed or blocked each path. It also asks the operating system the real question, does this path actually lead outside public once everything is resolved, so it can label a wrong answer. LEAK means the check allowed a path that escapes. wrong-block means it refused a harmless one.

The five functions at the top are the checks. The function escapes() is the referee: it joins the path the naive way, asks the operating system to resolve it completely with os.path.realpath, and reports whether the result is outside public. The function verdict() compares each check’s decision with the referee’s. Save this as validator_matrix.py:

"""Five plausible-looking path checks against nine payloads. Only one survives them all."""
import os
from pathlib import Path

import lab
from lab import PUBLIC, SECRET_FILE
from safe_paths import PathTraversalError, safe_join


class Blocked(Exception):
    pass


def v1_deny_dotdot(base: Path, user: str) -> Path:
    if ".." in user:
        raise Blocked
    return Path(os.path.join(base, user))


def v2_startswith(base: Path, user: str) -> Path:
    target = (base / user).resolve()
    if not str(target).startswith(str(base.resolve())):
        raise Blocked
    return target


def v3_lexical_relative(base: Path, user: str) -> Path:
    target = base / user
    if not target.is_relative_to(base):
        raise Blocked
    return target


def v4_abspath_commonpath(base: Path, user: str) -> Path:
    target = os.path.abspath(os.path.join(base, user))
    root = os.path.abspath(base)
    if os.path.commonpath([target, root]) != root:
        raise Blocked
    return Path(target)


def v5_resolve_relative(base: Path, user: str) -> Path:
    try:
        return safe_join(base, user)
    except PathTraversalError:
        raise Blocked


VALIDATORS = {
    "1 no '..'": v1_deny_dotdot,
    "2 startswith": v2_startswith,
    "3 lexical": v3_lexical_relative,
    "4 abspath": v4_abspath_commonpath,
    "5 resolve": v5_resolve_relative,
}

PAYLOADS = [
    ("legit file", "report.txt"),
    ("legit subfolder", "notes/2026/q3.txt"),
    ("legit, dot-dot inside", "notes/../report.txt"),
    ("dot-dot", "../secrets/api_key.txt"),
    ("nested dot-dot", "notes/../../secrets/api_key.txt"),
    ("backslashes", "..\\secrets\\api_key.txt"),
    ("absolute path", str(SECRET_FILE)),
    ("sibling 'public-backup'", "../public-backup/old.txt"),
    ("via symlink", "shortcut/api_key.txt"),
]


def escapes(user: str) -> bool:
    """Ground truth: does this path really lead outside PUBLIC once the OS resolves it?"""
    real = Path(os.path.realpath(os.path.join(PUBLIC, user)))
    return not real.is_relative_to(Path(os.path.realpath(PUBLIC)))


def verdict(check, user: str) -> str:
    try:
        check(PUBLIC, user)
        allowed = True
    except (Blocked, ValueError):
        allowed = False
    if allowed and escapes(user):
        return "LEAK"
    if not allowed and not escapes(user):
        return "wrong-block"
    return "allow" if allowed else "block"


if __name__ == "__main__":
    has_symlink = lab.build()
    payloads = [p for p in PAYLOADS if has_symlink or p[0] != "via symlink"]
    print(f"{'payload':26}" + "".join(f"{name:>14}" for name in VALIDATORS))
    leaks = dict.fromkeys(VALIDATORS, 0)
    for label, user in payloads:
        row = f"{label:26}"
        for name, check in VALIDATORS.items():
            result = verdict(check, user)
            leaks[name] += result == "LEAK"
            row += f"{result:>14}"
        print(row)
    print(f"{'LEAKS':26}" + "".join(f"{leaks[name]:>14}" for name in VALIDATORS))
python validator_matrix.py
payload                        1 no '..'  2 startswith     3 lexical     4 abspath     5 resolve
legit file                         allow         allow         allow         allow         allow
legit subfolder                    allow         allow         allow         allow         allow
legit, dot-dot inside        wrong-block         allow         allow         allow         allow
dot-dot                            block         block          LEAK         block         block
nested dot-dot                     block         block          LEAK         block         block
backslashes                        block         block          LEAK         block         block
absolute path                       LEAK         block         block         block         block
sibling 'public-backup'            block          LEAK          LEAK         block         block
via symlink                         LEAK         block          LEAK          LEAK         block
LEAKS                                  2             1             5             1             0

Only the last column has zero leaks. Here is why each of the others fails.

Check 1: refuse any path containing two dots

It blocks every dot-dot row, but it also refuses notes/../report.txt, a harmless path that stays inside public (the wrong-block cell). Worse, it lets the absolute path through, because that contains no dots at all, and it lets the symlink through for the same reason. A blacklist of bad patterns can only stop the patterns you thought of.

Check 2: resolve the path, then use startswith

Resolving first, so that .. and symlinks are handled, is the right instinct, and this check blocks five of the six attacks. It fails on ../public-backup/old.txt, because the string ...\lab\public-backup\old.txt begins with the string ...\lab\public. A string prefix is not the same thing as a parent folder. Compare path components, not characters.

Check 3: is_relative_to without resolving

The pathlib module has a method that sounds perfect, but its documentation says: “This method is string-based; it neither accesses the filesystem nor treats “..” segments specially.” So public/../secrets/api_key.txt counts as relative to public as far as the method is concerned, because it begins with public. Five of the six attacks sail through.

Check 4: abspath plus commonpath

This is the classic answer, and it is much better: os.path.abspath cleans up the dots and os.path.commonpath compares path components rather than characters. But abspath only rewrites the text of the path. It never asks the filesystem whether any folder along the way is a symlink. So it stops every dot-dot trick and leaks exactly one row, the symlink.

Check 5: resolve, then is_relative_to

The documentation for Path.resolve says: “Make the path absolute, resolving any symlinks.” It adds that “..” components are also eliminated, and that this is the only method that does so. After resolving, you hold the one true absolute location, and a component-wise comparison with is_relative_to is exactly the right question to ask. This is the check we will use.

One more Windows trap: resolve both sides

There is a subtler way to use the right check and still get the wrong answer. Save this as shortname_gotcha.py and run it:

import tempfile
from pathlib import Path

base = Path(tempfile.mkdtemp())  # on Windows this can contain an 8.3 short name
(base / "report.txt").write_text("hello\n")

candidate = (base / "report.txt").resolve()
print("base      :", base)
print("candidate :", candidate)
print("inside, base NOT resolved :", candidate.is_relative_to(base))
print("inside, base.resolve()    :", candidate.is_relative_to(base.resolve()))
python shortname_gotcha.py
base      : C:\Users\ADMINI~1\AppData\Local\Temp\tmpxojt31z8
candidate : C:\Users\Administrator\AppData\Local\Temp\tmpxojt31z8\report.txt
inside, base NOT resolved : False
inside, base.resolve()    : True

On my machine the temporary folder path contains ADMINI~1, an old-style 8.3 short name that Windows keeps for the user folder called Administrator. Calling resolve() expands short names into the real long names, so the resolved candidate no longer starts with the unresolved base, and a perfectly legitimate file is judged to be outside it. Resolving the base as well fixes it (the last line). Your output will differ if your paths contain no short names, but the rule is cheap to follow: always resolve both sides of the comparison.

Step 7: The fix, resolve first and then check containment

Now write the function. Save this as safe_paths.py:

from pathlib import Path


class PathTraversalError(ValueError):
    """The requested path would leave the allowed base directory."""


def safe_join(base, user_path: str) -> Path:
    """Join user_path onto base and guarantee the result stays inside base."""
    if "\x00" in user_path:  # resolve() does not reject NUL bytes, open() does
        raise PathTraversalError("NUL byte in path")
    base = Path(base).resolve()  # resolve the base too, not just the target
    try:
        candidate = (base / user_path).resolve()  # follows symlinks, removes ".."
    except (ValueError, OSError) as exc:  # embedded NUL, invalid Windows path, ...
        raise PathTraversalError(f"invalid path {user_path!r}") from exc
    if not candidate.is_relative_to(base):
        raise PathTraversalError(f"{user_path!r} escapes the base directory")
    return candidate

Walk through it once. First it rejects a NUL byte (the character with code 0) outright. Second it resolves the base folder, for the reason you just saw. Third it joins the user’s text onto the base and resolves the result, which removes every .. and follows every symlink. If the user text was an absolute path, pathlib discards the base, exactly like os.path.join does, but that is harmless here because the next line catches it. Fourth it asks is_relative_to whether the resolved location is inside the resolved base. Anything that goes wrong along the way, such as an invalid path, is reported as the same PathTraversalError, which is a subclass of ValueError. The function returns the resolved path, and you must open that returned path, not the string the user sent, because the resolved path is the one that was checked.

Why reject NUL bytes explicitly? Save nul_check.py and run it:

from pathlib import Path

from lab import PUBLIC

p = PUBLIC / "report.txt\x00.png"
print("resolve() ->", repr(p.resolve().name))
print("is_file() ->", p.is_file())
try:
    open(p)
except ValueError as exc:
    print("open()    -> ValueError:", exc)
python nul_check.py
resolve() -> 'report.txt\x00.png'
is_file() -> False
open()    -> ValueError: embedded null character

On Python 3.13.14, resolve() happily returned a path containing a NUL byte, is_file() quietly returned False, and only open() raised ValueError: embedded null character. Three functions, three different reactions. I would not want a security check to depend on which of them happens to run first, so the function rejects NUL bytes up front.

Now use it. Save this as app_v2.py:

from fastapi import FastAPI, HTTPException
from fastapi.responses import FileResponse

from lab import PUBLIC
from safe_paths import PathTraversalError, safe_join

app = FastAPI()


@app.get("/files/{name:path}")
def read_file(name: str):
    try:
        path = safe_join(PUBLIC, name)
    except PathTraversalError:
        # Same answer as a missing file, so attackers learn nothing about what exists.
        raise HTTPException(status_code=404, detail="Not found")
    if not path.is_file():
        raise HTTPException(status_code=404, detail="Not found")
    return FileResponse(path, media_type="text/plain")

Stop any server that is still running (Ctrl+C), start this one with python -m uvicorn app_v2:app --port 8000, and run the attack again:

legit file                 200  'Q3 report: revenue is up'
legit subfolder            200  'Q3 notes: ship the fix'
dot-dot                    404  '{"detail":"Not found"}'
dot-dot, percent-encoded   404  '{"detail":"Not found"}'
backslashes (Windows)      404  '{"detail":"Not found"}'
absolute path              404  '{"detail":"Not found"}'
sibling directory          404  '{"detail":"Not found"}'
through a symlink          404  '{"detail":"Not found"}'

Both legitimate requests return 200 and all six attacks return 404, including the backslash, absolute-path and symlink rows that beat most of the checks in Step 6. The handler answers a rejected path with exactly the same 404 it gives for a missing file, so an attacker cannot use the difference between “forbidden” and “not found” to learn which files exist.

Step 8: Or let your framework do it

If all you need is to serve a folder of static files, you do not have to write this logic at all. Starlette ships a StaticFiles class that contains its own containment check. Save this as app_v3.py:

from fastapi import FastAPI
from fastapi.staticfiles import StaticFiles

from lab import PUBLIC

app = FastAPI()
app.mount("/static", StaticFiles(directory=PUBLIC), name="static")

Stop the previous server (Ctrl+C), start this one with python -m uvicorn app_v3:app --port 8000, and attack the /static/ prefix instead:

python attack.py /static/
legit file                 200  'Q3 report: revenue is up'
legit subfolder            200  'Q3 notes: ship the fix'
dot-dot                    404  '{"detail":"Not Found"}'
dot-dot, percent-encoded   404  '{"detail":"Not Found"}'
backslashes (Windows)      404  '{"detail":"Not Found"}'
absolute path              404  '{"detail":"Not Found"}'
sibling directory          404  '{"detail":"Not Found"}'
through a symlink          404  '{"detail":"Not Found"}'

The same result, with no code of our own. Here is the method that makes the decision, copied from starlette/staticfiles.py in Starlette 1.7.0, the version FastAPI installed for this tutorial:

    def lookup_path(self, path: str) -> tuple[str, os.stat_result | None]:
        # Reject absolute paths so they cannot escape the served directory.
        if path.startswith(("/", "\\")):
            return "", None
        for directory in self.all_directories:
            joined_path = os.path.join(directory, path)
            if self.follow_symlink:
                full_path = os.path.abspath(joined_path)
                directory = os.path.abspath(directory)
            else:
                full_path = os.path.realpath(joined_path)
                directory = os.path.realpath(directory)
            if os.path.commonpath([full_path, directory]) != str(directory):
                # Don't allow misbehaving clients to break out of the static files directory.
                continue

It rejects names that start with a slash or a backslash, then resolves the joined path and compares components with os.path.commonpath. Notice the follow_symlink switch. By default the code takes the realpath branch, which follows symlinks, and that is why the symlink row returned 404 above. To see the other branch, I changed the mount line to StaticFiles(directory=PUBLIC, follow_symlink=True), which uses abspath instead, and ran the same attack:

legit file                 200  'Q3 report: revenue is up'
legit subfolder            200  'Q3 notes: ship the fix'
dot-dot                    404  '{"detail":"Not Found"}'
dot-dot, percent-encoded   404  '{"detail":"Not Found"}'
backslashes (Windows)      404  '{"detail":"Not Found"}'
absolute path              404  '{"detail":"Not Found"}'
sibling directory          404  '{"detail":"Not Found"}'
through a symlink          200  'API_KEY=sk_live_demo_do_not_share'

Every row is unchanged except the symlink row, which now returns the secret. That is the option doing what the source above says it does, and it is the same difference you saw between checks 4 and 5 in Step 6. Turn it on only if you really want symlinks inside the served folder to be able to point elsewhere. Use StaticFiles when you are serving a whole folder, and use safe_join when your own code decides which file to return, for example after an authorization check.

Step 9: The same bug in a different costume, archive extraction

Downloads are one way for a stranger’s text to become a path. Archives are another. A zip or tar file is a list of members, and each member carries a name chosen by whoever built the archive. If your code extracts the archive by gluing each name onto a destination folder, the archive’s author decides where files land. In 2018 Snyk gave this version of the bug a memorable name, and its research page describes Zip Slip as “a widespread arbitrary file overwrite critical vulnerability, which typically results in remote command execution.” A file written outside the intended folder can replace a script, a configuration file or a startup entry.

9.1 Build some hostile archives

To test defenses we need attackers, so save two small helpers. The first, hostile.py, builds a tar file with one harmless member plus one hostile member of a chosen kind: a .. name, a backslash version of it, an absolute path (which points at a file in a scratch folder next to the destination, so nothing outside your sandbox is touched) and a symlink member that points outside. It can also build a hostile zip.

"""Builds hostile archives, so we can prove our defences work."""
import io
import tarfile
import zipfile
from pathlib import Path


def _add_file(tar: tarfile.TarFile, name: str, data: bytes) -> None:
    info = tarfile.TarInfo(name)
    info.size = len(data)
    tar.addfile(info, io.BytesIO(data))


def build_hostile_tar(path: Path, kind: str, scratch: Path) -> None:
    """A tar with one harmless member plus one hostile member of the given kind."""
    with tarfile.open(path, "w") as tar:
        _add_file(tar, "docs/readme.txt", b"hello\n")
        if kind == "dot-dot":
            _add_file(tar, "../evil.txt", b"escaped\n")
        elif kind == "backslash":
            _add_file(tar, "..\\evil.txt", b"escaped\n")
        elif kind == "absolute":
            _add_file(tar, str(scratch / "evil.txt"), b"escaped\n")
        elif kind == "symlink":
            link = tarfile.TarInfo("shortcut")
            link.type = tarfile.SYMTYPE
            link.linkname = "../outside"
            tar.addfile(link)
        else:
            raise ValueError(kind)


def build_hostile_zip(path: Path) -> None:
    with zipfile.ZipFile(path, "w") as zf:
        zf.writestr("docs/readme.txt", "hello\n")
        zf.writestr("../evil.txt", "escaped\n")

The second, scratch.py, gives each demo a fresh scratch folder with an empty dest folder inside it. Anything that lands in scratch but outside dest is an escape.

import shutil
from pathlib import Path

SCRATCH = Path(__file__).resolve().parent / "scratch"


def fresh() -> Path:
    """Empty the scratch folder and return a new, empty 'dest' folder inside it."""
    shutil.rmtree(SCRATCH, ignore_errors=True)
    SCRATCH.mkdir()
    dest = SCRATCH / "dest"
    dest.mkdir()
    return dest

9.2 What tar extraction does

Save this as demo_tar_filters.py. For each hostile kind it extracts three times: with no filter argument at all, with filter="tar", and with filter="data", and it classifies what happened.

import sys
import tarfile
import warnings

from hostile import build_hostile_tar
from scratch import SCRATCH, fresh


def extract(archive, dest, flt) -> str:
    try:
        with tarfile.open(archive) as tar:
            if flt is None:
                tar.extractall(dest)  # no filter argument at all
            else:
                tar.extractall(dest, filter=flt)
    except Exception as exc:
        return f"refused: {type(exc).__name__}"
    if (SCRATCH / "evil.txt").exists():
        return "ESCAPED: wrote outside dest"
    if (dest / "shortcut").is_symlink():
        return "symlink pointing outside created"
    return "extracted inside dest"


print(sys.version.split()[0], "on", sys.platform)
first_warning = None
for kind in ("dot-dot", "backslash", "absolute", "symlink"):
    for flt in (None, "tar", "data"):
        dest = fresh()
        archive = SCRATCH / "attack.tar"
        build_hostile_tar(archive, kind, SCRATCH)
        with warnings.catch_warnings(record=True) as caught:
            warnings.simplefilter("always")
            result = extract(archive, dest, flt)
        if caught and first_warning is None:
            first_warning = f"{caught[0].category.__name__}: {caught[0].message}"
        print(f"{kind:10} filter={str(flt):5} -> {result}")
print("\nwarning printed when no filter is given:\n ", first_warning)
python demo_tar_filters.py
3.13.14 on win32
dot-dot    filter=None  -> ESCAPED: wrote outside dest
dot-dot    filter=tar   -> refused: OutsideDestinationError
dot-dot    filter=data  -> refused: OutsideDestinationError
backslash  filter=None  -> ESCAPED: wrote outside dest
backslash  filter=tar   -> refused: OutsideDestinationError
backslash  filter=data  -> refused: OutsideDestinationError
absolute   filter=None  -> ESCAPED: wrote outside dest
absolute   filter=tar   -> refused: AbsolutePathError
absolute   filter=data  -> refused: AbsolutePathError
symlink    filter=None  -> symlink pointing outside created
symlink    filter=tar   -> symlink pointing outside created
symlink    filter=data  -> refused: LinkOutsideDestinationError

warning printed when no filter is given:
  DeprecationWarning: Python 3.14 will, by default, filter extracted tar archives and reject files or modify their metadata. Use the filter argument to control this behavior.

What the default does on Python 3.13

The rows with filter=None are the lesson. With no filter, three of the four hostile archives wrote a file outside dest, and the fourth planted a symlink pointing outside it. Python also printed the warning shown at the bottom, which I captured from the run. (The backslash row is Windows-only, and if your account cannot create symlinks the symlink rows can look different.) The documentation explains the history: “Changed in version 3.12: Added the filter parameter. Changed in version 3.14: The filter parameter now defaults to 'data'.” Before 3.14 the default was “equivalent to fully_trusted”, so on Python 3.13 and older a plain extractall(dest) trusts the archive completely.

What the tar and data filters do

Both named filters refused the dot-dot, backslash and absolute-path archives, with OutsideDestinationError or AbsolutePathError, and nothing was written outside. The two filters differ on the symlink. The documentation for the tar filter says it will “Refuse to extract files whose absolute path (after following symlinks) would end up outside the destination.” That is about where files end up, so the link itself got through. The data filter adds: “Refuse to extract links (hard or soft) that link to absolute paths, or ones that link outside the destination.” It refused the link with LinkOutsideDestinationError. It also refuses device files and pipes. For data you do not fully trust, data is the filter to use, and it is the default from Python 3.14 on.

The documentation is direct about the limits: “Never extract archives from untrusted sources without prior inspection. Since Python 3.14, the default (data) will prevent the most dangerous security issues. However, it will not prevent all unintended or insecure behavior.” It recommends passing filter='data' explicitly so that your code stays safe on Python versions with a less secure default (3.13 and lower). PEP 706, the proposal that added filters, explains why they exist: “it’s quite tricky to do such an inspection correctly. As a result, many people don’t bother, or do the check incorrectly, resulting in security issues such as CVE-2007-4559.” That is exactly what the matrix in Step 6 showed for downloads.

9.3 A refused archive can still leave files behind

There is one more behavior to know about. Extraction happens member by member, so if the second member is refused, the first has already been written. The solution is to extract into a temporary staging folder and move it into place only if everything passed. Save this as safe_extract.py:

"""Extract untrusted archives all-or-nothing."""
import os
import shutil
import tarfile
import tempfile
import zipfile
from pathlib import Path

from safe_paths import PathTraversalError, safe_join


class ArchiveRejected(Exception):
    """The archive contained something we refuse to extract."""


def _staging_dir(dest: Path) -> Path:
    if dest.exists():
        raise FileExistsError(dest)
    return Path(tempfile.mkdtemp(dir=dest.parent, prefix=".extract-"))


def safe_extract_tar(archive: Path, dest: Path) -> None:
    """Extract into dest, or raise and leave nothing behind."""
    dest = Path(dest)
    staging = _staging_dir(dest)
    try:
        with tarfile.open(archive) as tar:
            tar.extractall(staging, filter="data")
        os.replace(staging, dest)  # only reached if every member passed the filter
    except (tarfile.TarError, ValueError) as exc:
        raise ArchiveRejected(f"{type(exc).__name__}: {exc}") from exc
    finally:
        shutil.rmtree(staging, ignore_errors=True)


def safe_extract_zip(archive: Path, dest: Path) -> None:
    """Same contract for zip files: reject bad names instead of silently rewriting them."""
    dest = Path(dest)
    staging = _staging_dir(dest)
    try:
        with zipfile.ZipFile(archive) as zf:
            for info in zf.infolist():
                safe_join(staging, info.filename)
            zf.extractall(staging)
        os.replace(staging, dest)
    except (zipfile.BadZipFile, PathTraversalError) as exc:
        raise ArchiveRejected(f"{type(exc).__name__}: {exc}") from exc
    finally:
        shutil.rmtree(staging, ignore_errors=True)

Read safe_extract_tar first. It creates a fresh staging folder next to the destination, extracts into it with filter="data", and only if every member passed does it rename the staging folder to the destination with os.replace. If anything raises, the finally block deletes the staging folder, so the caller sees either a complete destination or none at all. It catches tarfile.TarError, which is the base class of the filter errors as well as of corrupt-archive errors, plus ValueError. That last one is there because in an extra experiment on Windows a member with a drive-relative name (mine was D:drive_rel_evil.txt) made the filter raise ValueError: Paths don't have the same drive instead of a filter error. Failing closed on any exception is the safe behavior.

Now compare the two approaches. Save this as demo_leftovers.py:

import tarfile

from hostile import build_hostile_tar
from safe_extract import ArchiveRejected, safe_extract_tar
from scratch import SCRATCH, fresh

dest = fresh()
archive = SCRATCH / "attack.tar"
build_hostile_tar(archive, "dot-dot", SCRATCH)

print("== extractall with filter='data' straight into dest ==")
try:
    with tarfile.open(archive) as tar:
        tar.extractall(dest, filter="data")
except tarfile.FilterError as exc:
    print("refused:", type(exc).__name__)
print("left behind in dest:", sorted(p.relative_to(dest).as_posix() for p in dest.rglob("*") if p.is_file()))

print("\n== safe_extract_tar: all or nothing ==")
fresh()
build_hostile_tar(archive, "dot-dot", SCRATCH)
target = SCRATCH / "site"
try:
    safe_extract_tar(archive, target)
except ArchiveRejected as exc:
    print("ArchiveRejected:", str(exc)[:58], "...")
print("target exists:", target.exists())
print("stray staging dirs:", [p.name for p in SCRATCH.iterdir() if p.name.startswith(".extract-")])
print("escaped file:", (SCRATCH / "evil.txt").exists())
python demo_leftovers.py
== extractall with filter='data' straight into dest ==
refused: OutsideDestinationError
left behind in dest: ['docs/readme.txt']

== safe_extract_tar: all or nothing ==
ArchiveRejected: OutsideDestinationError: '../evil.txt' would be extracted  ...
target exists: False
stray staging dirs: []
escaped file: False

The first half extracts straight into dest. The archive is refused, but docs/readme.txt was already written and stays behind. The second half uses the helper: the archive is rejected, the target folder does not exist, no staging folders are left over, and nothing escaped.

9.4 The zipfile module plays by different rules

The zipfile module makes different choices. Its documentation warns: “Never extract archives from untrusted sources without prior inspection. It is possible that files are created outside of path, for example, members that have absolute filenames or filenames with “..” components. This module attempts to prevent that.” The extract() notes say that for an absolute member name the drive and leading slashes are stripped, and that all .. components are removed. Save this as demo_zip.py to see it happen:

import stat
import zipfile

from hostile import build_hostile_zip
from safe_extract import ArchiveRejected, safe_extract_zip
from scratch import SCRATCH, fresh

fresh()
names = ["docs/readme.txt", "../zip_evil.txt", "/zip_abs.txt", "..\\zip_bs.txt",
         "C:\\zip_drive.txt", "a/../../zip_nested.txt"]
archive = SCRATCH / "names.zip"
with zipfile.ZipFile(archive, "w") as zf:
    for name in names:
        zf.writestr(name, "x\n")
dest = SCRATCH / "zdest"
dest.mkdir()
with zipfile.ZipFile(archive) as zf:
    print("stored as:", zf.namelist())
    zf.extractall(dest)
print("inside dest :", sorted(p.relative_to(dest).as_posix() for p in dest.rglob("*") if p.is_file()))
print("outside dest:", sorted(p.name for p in SCRATCH.glob("*.txt")))

print("\n== a zip entry that claims to be a symlink ==")
fresh()
info = zipfile.ZipInfo("zlink")
info.create_system = 3  # Unix
info.external_attr = (stat.S_IFLNK | 0o777) << 16
archive = SCRATCH / "zlink.zip"
with zipfile.ZipFile(archive, "w") as zf:
    zf.writestr(info, "../outside_dir")
dest = SCRATCH / "zdest"
dest.mkdir()
with zipfile.ZipFile(archive) as zf:
    zf.extractall(dest)
link = dest / "zlink"
print("is_symlink:", link.is_symlink(), "| is_file:", link.is_file(), "| content:", link.read_text())

print("\n== safe_extract_zip: reject instead of rewrite ==")
fresh()
archive = SCRATCH / "attack.zip"
build_hostile_zip(archive)
target = SCRATCH / "site"
try:
    safe_extract_zip(archive, target)
except ArchiveRejected as exc:
    print("ArchiveRejected:", str(exc)[:60])
print("target exists:", target.exists())
python demo_zip.py
stored as: ['docs/readme.txt', '../zip_evil.txt', '/zip_abs.txt', '../zip_bs.txt', 'C:/zip_drive.txt', 'a/../../zip_nested.txt']
inside dest : ['a/zip_nested.txt', 'docs/readme.txt', 'zip_abs.txt', 'zip_bs.txt', 'zip_drive.txt', 'zip_evil.txt']
outside dest: []

== a zip entry that claims to be a symlink ==
is_symlink: False | is_file: True | content: ../outside_dir

== safe_extract_zip: reject instead of rewrite ==
ArchiveRejected: PathTraversalError: '../evil.txt' escapes the base directory
target exists: False

All six names ended up inside dest and nothing landed outside, so plain zipfile.extractall is safer than unfiltered tarfile.extractall here. Two more things are worth noticing. First, the names are silently rewritten: ../zip_evil.txt became zip_evil.txt. That is safe, but it hides the fact that someone tried something, and a rewritten name can collide with a legitimate one. For an upload endpoint I prefer to reject, which is what safe_extract_zip does by validating every member name with safe_join before extracting anything (the last section of the output). Second, a zip entry that claims to be a symlink was extracted as an ordinary file containing the link’s target text. In this run zipfile did not create a symlink.

Step 10: Lock it in with tests

Everything above was checked by eye. Turn it into an automated safety net so a future refactor cannot quietly reopen the hole. Save this as test_traversal.py:

import io
import os
import tarfile

import pytest

import lab
from hostile import build_hostile_tar, build_hostile_zip
from lab import PUBLIC, SECRET_FILE
from safe_extract import ArchiveRejected, safe_extract_tar, safe_extract_zip
from safe_paths import PathTraversalError, safe_join

HAS_SYMLINK = lab.build()
windows_only = pytest.mark.skipif(os.name != "nt", reason="backslash is a separator only on Windows")
needs_symlink = pytest.mark.skipif(not HAS_SYMLINK, reason="this account cannot create symlinks")


@pytest.mark.parametrize("name", ["report.txt", "notes/2026/q3.txt", "notes/../report.txt"])
def test_legitimate_paths_are_allowed(name):
    assert safe_join(PUBLIC, name).is_file()


@pytest.mark.parametrize("name", [
    "../secrets/api_key.txt",
    "notes/../../secrets/api_key.txt",
    "../public-backup/old.txt",
    str(SECRET_FILE),
    "report.txt\x00.png",
    pytest.param("..\\secrets\\api_key.txt", marks=windows_only),
    pytest.param("shortcut/api_key.txt", marks=needs_symlink),
])
def test_escapes_are_rejected(name):
    with pytest.raises(PathTraversalError):
        safe_join(PUBLIC, name)


@pytest.mark.parametrize("kind", ["dot-dot", "absolute", "symlink",
                                  pytest.param("backslash", marks=windows_only)])
def test_hostile_tar_is_rejected_and_leaves_nothing(tmp_path, kind):
    archive = tmp_path / "attack.tar"
    build_hostile_tar(archive, kind, tmp_path)
    with pytest.raises(ArchiveRejected):
        safe_extract_tar(archive, tmp_path / "site")
    assert not (tmp_path / "site").exists()
    assert not (tmp_path / "evil.txt").exists()
    assert not [p for p in tmp_path.iterdir() if p.name.startswith(".extract-")]


def test_hostile_zip_is_rejected_and_leaves_nothing(tmp_path):
    archive = tmp_path / "attack.zip"
    build_hostile_zip(archive)
    with pytest.raises(ArchiveRejected):
        safe_extract_zip(archive, tmp_path / "site")
    assert not (tmp_path / "site").exists()


def test_a_clean_archive_is_extracted(tmp_path):
    archive = tmp_path / "ok.tar"
    with tarfile.open(archive, "w") as tar:
        info = tarfile.TarInfo("docs/readme.txt")
        info.size = 6
        tar.addfile(info, io.BytesIO(b"hello\n"))
    safe_extract_tar(archive, tmp_path / "site")
    assert (tmp_path / "site" / "docs" / "readme.txt").read_text() == "hello\n"

The parametrize lists are the payload catalog: three legitimate paths that must work, and every escape from this tutorial that must be rejected. The windows_only and needs_symlink marks skip the cases that cannot apply on your machine instead of failing. The archive tests use pytest’s tmp_path fixture and assert three things after a rejected archive: the destination does not exist, no file escaped, and no staging folder was left behind. Run the suite:

python -m pytest -q
................                                                         [100%]
16 passed in 0.09s

All 16 tests passed on my Windows machine. On Linux or macOS the two Windows-only cases are skipped instead of passing.

Common mistakes and gotchas

Testing with a client that cleans the URL. As Step 3 showed, requests and plain curl collapse .. before sending. Use http.client or curl --path-as-is.

Checking the text instead of the location. Blacklisting .. misses absolute paths and symlinks and blocks harmless paths. Canonicalize with resolve(), then check.

Comparing strings instead of path components. startswith lets public-backup through a check meant for public. Use is_relative_to on resolved paths.

Resolving only one side. On Windows, 8.3 short names made a legitimate file look outside its own folder. Resolve the base and the candidate.

Assuming your operating system is everyone’s. The backslash and drive-letter attacks only work on Windows, and a check tested only on Linux never sees them.

Trusting the archive or the default. Before Python 3.14, tarfile.extractall trusted archives completely. Pass filter="data" explicitly. The same warning applies to shutil.unpack_archive, whose documentation says: “For zip files, filter is not accepted. For tar files, it is recommended to use 'data' (default since Python 3.14), unless using features specific to tar and UNIX-like filesystems.”

Extracting straight into the final folder. A refused archive can leave its earlier members behind. Stage first, then move into place.

Forgetting time. safe_join checks a path now and your code opens it a moment later. If untrusted users can create or replace files or symlinks inside the served folder, for example through an upload feature or by extracting an untrusted archive there, the check can be out of date by the time the file is opened. Keep served folders read-only for the web process, and never let user-controlled archives create links that point outside the folder, which is what the data filter refuses.

Solving it at the wrong level. OWASP’s first piece of advice is: “Prefer working without user input when using file system calls”, and it suggests, “Use indexes rather than actual portions of file names when templating or using language files”. If a download can be identified by an ID that your code maps to a file, the user never supplies a path at all, and there is nothing to traverse.

How to confirm everything works end to end

  1. Rebuild the lab with python lab.py.
  2. Start the fixed application in the first terminal with python -m uvicorn app_v2:app --port 8000, then run python attack.py in the second. You should see exactly two 200 responses (the legitimate files) and six 404 responses.
  3. Run python -m pytest -q. On Windows with symlink permission you should see 16 passed.
  4. Audit your own code. Search your project for extractall(, unpack_archive(, FileResponse(, send_file(, and any open( or os.path.join( whose arguments include a request value or an archive member name. Every hit should go through safe_join, a framework helper such as StaticFiles, or an explicit filter="data".

Next steps

Path traversal is one member of a family in which untrusted text crosses a boundary into something that interprets it. This site has tutorials on several others: preventing server-side request forgery, where user text becomes a request to the wrong place, preventing SQL injection with parameterized queries, stopping XSS with a Content Security Policy, and preventing argument injection in a Windows URI protocol handler, which shares this tutorial’s Windows-specific parsing surprises. If you would like an automated reviewer to help catch risky code before it is committed, see how to run Bandit in a pre-commit hook. Then take the payload catalog from Step 10 and run it against the real file-serving and upload code in your own projects. The hostile paths are already written, and the first time one of them gets through, you will be glad you looked.

Tags:

Application SecurityFastAPIPath TraversalpytestPython

Share

Two photos of the same wall outlet, one with open sockets and one with plastic safety covers fitted over the slots
Previous Post

Cloudflare’s EmDash 1.0 Turns the WordPress Plugin Problem Into a Permissions Decision

Yellow stenciled Irish text meaning 'Stay behind this line' painted beside a yellow boundary line at a railway platform edge, with tracks beyond
Next Post

Apple Patches a Meta-Reported CoreGraphics Zero-Day Possibly Exploited in Targeted iOS Attacks

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
30 Sep
How to Detect and Strip Invisible Unicode in Python to Stop ASCII Smuggling and Trojan Source
30 Sep
Two Cloudflare Agent Billing Betas Turn Web Monetization Into a Question of Who Holds the Meter
Trending
September 30, 2026
How to Detect and Strip Invisible Unicode in Python to Stop ASCII Smuggling and Trojan Source
September 30, 2026
Two Cloudflare Agent Billing Betas Turn Web Monetization Into a Question of Who Holds the Meter
September 30, 2026
Attackers Exploit a Hex-Encoding Bypass in Cisco SD-WAN Manager, and CISA Sets an October 3 Deadline
September 30, 2026
How to Enforce Guardrails on AI-Generated Terraform With Open Policy Agent and Rego
September 30, 2026
AMD Agrees to Buy Fei-Fei Li’s World Labs for $8.2 Billion to Steer Its Chip Roadmap
September 29, 2026
How to Build a Merkle Tree Certificate Issuer in Python to Keep Post-Quantum Certificates Small

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026