How to Prevent Path Traversal in Python File Downloads and Archive Extraction
This tutorial reproduces path traversal in a FastAPI download endpoint, shows four plausible-looking checks failing before a fifth holds, and builds a tested safe_join() function plus safe tar and...
Somewhere in almost every web application there is a route that hands a file back to the user: a report download, a profile picture, a log viewer, a template loader. Somewhere else there is code that unpacks a file the user uploaded. In both places the application takes a piece of text chosen by a stranger and turns it into a location on your disk. If that translation is careless, the stranger can steer it out of the folder you meant to expose and into one you did not, such as the folder that holds your API keys. The bug class is called path traversal (also known as directory traversal), and MITRE catalogs it as CWE-22, titled Improper Limitation of a Pathname to a Restricted Directory (‘Path Traversal’).
Table Of Content
- Prerequisites
- A note on operating systems
- What path traversal is, in plain language
- Step 1: Build a practice lab
- Step 2: Write a download endpoint with the classic bug
- Step 3: Attack it, and see why your first attempt may not work
- Gotcha: your own test client may be quietly fixing the attack
- Step 4: Why the bug works, starting with what os.path.join really does
- Step 5: Removing :path is not a fix
- Step 6: Five checks that look right, and only one is
- Check 1: refuse any path containing two dots
- Check 2: resolve the path, then use startswith
- Check 3: is_relative_to without resolving
- Check 4: abspath plus commonpath
- Check 5: resolve, then is_relative_to
- One more Windows trap: resolve both sides
- Step 7: The fix, resolve first and then check containment
- Step 8: Or let your framework do it
- Step 9: The same bug in a different costume, archive extraction
- 9.1 Build some hostile archives
- 9.2 What tar extraction does
- What the default does on Python 3.13
- What the tar and data filters do
- 9.3 A refused archive can still leave files behind
- 9.4 The zipfile module plays by different rules
- Step 10: Lock it in with tests
- Common mistakes and gotchas
- How to confirm everything works end to end
- Next steps
This is not a museum piece. This month sxz.io covered a maximum-severity GitLab flaw that GitLab itself described as a path traversal issue in the repository commits API, and a WordPress Core fix for an unauthenticated local file inclusion bug in page template resolution that had been present since 2016. Different products, closely related bug classes: in both, text an attacker could influence ended up naming a file the software should never have opened.
In this tutorial you will build a small download endpoint on purpose with the classic bug, attack it, and then fix it properly. Along the way you will watch four plausible-looking checks fail before a fifth one holds, learn why a check that works on Linux can quietly fail on Windows, and then apply the same thinking to archives, where the same bug is famous under the name Zip Slip. You will use the extraction filters that the Python tarfile module gained in version 3.12, and you will finish with a small, pytest-verified module you can adapt to your own project.
Prerequisites
- Python 3.12 or newer. I built and tested everything on Python 3.13.14 on Windows 11. Python 3.12 is the release that added the
filterargument to tar extraction, which the archive half of this tutorial relies on. - Four packages:
fastapi0.141.1,uvicorn0.54.0,requests2.34.2, andpytest9.1.1. Those are the versions I used. FastAPI installs Starlette (1.7.0 here), and we will look inside it later. - Two terminal windows open in the same project folder, one for the server and one for the attacker. One demo also uses the
curlcommand (I used curl 8.21.0), but you can skip it. - Basic Python and HTTP knowledge. You should know what a function is and what a URL path looks like. No security background is assumed, and every term is defined when it first appears.
A note on operating systems
Some behavior differs by platform, and I only ran this on Windows 11, so I mark every spot where Linux or macOS would differ. On those systems a backslash is an ordinary filename character rather than a folder separator, and there are no drive letters. Creating a symbolic link on Windows also needs extra permission. The Python documentation for os.symlink says: “On newer versions of Windows 10, unprivileged accounts can create symlinks if Developer Mode is enabled. When Developer Mode is not available/enabled, the SeCreateSymbolicLinkPrivilege privilege is required, or the process must be run as an administrator.” The lab script in Step 1 tries to create one symlink and reports whether it worked, and every later step copes if it could not. It worked on my machine because I ran as an administrator.
Create a project folder, then create and activate a virtual environment and install the packages:
python -m venv .venv
.venv\Scripts\activate # on macOS or Linux: source .venv/bin/activate
pip install fastapi uvicorn requests pytest
What path traversal is, in plain language
A filesystem path is just text that the operating system interprets when you open a file. Three details of that interpretation cause almost every traversal bug:
- The segment
..means “the parent folder”. Openingpublic/../secrets/api_key.txtclimbs out ofpublicand into its neighborsecrets. The operating system does this at the moment the file is opened, not when your program builds the string. - An absolute path starts from the top of a drive or filesystem instead of from a folder you chose, for example
C:\data\file.txtor/var/log/syslog. Many path-joining functions treat an absolute path as an instruction to forget everything that came before it. - A symbolic link (symlink) is a special file that acts as an alias for another location. A path can look perfectly innocent, with no dots at all, and still end up somewhere else because one of its folders is a symlink.
What your program wants is a promise, called containment: whatever the user asks for, I will only touch files inside this one base directory. The OWASP description of the attack puts it this way: “A path traversal attack (also known as directory traversal) aims to access files and directories that are stored outside the web root folder. By manipulating variables that reference files with “dot-dot-slash (../)” sequences and its variations or by using absolute file paths, it may be possible to access arbitrary files and directories stored on file system including application source code or configuration and critical system files.”
The plan for the rest of this tutorial is to establish containment correctly. The key idea is to canonicalize first: turn the messy, user-supplied text into the single absolute location the operating system would really open, with every .. removed and every symlink followed. Only then do you ask whether that location is inside your base directory.
Step 1: Build a practice lab
Save every file from this tutorial in the same project folder, starting with the following, which you should name lab.py. Every demo imports it. It builds a tiny tree of folders and files that we can attack safely, and it wipes and rebuilds that tree each time it runs, so you can always get a clean slate by running python lab.py.
"""Builds the small file tree that every demo in this tutorial attacks."""
import os
import shutil
from pathlib import Path
HERE = Path(__file__).resolve().parent
LAB = HERE / "lab"
PUBLIC = LAB / "public" # the only directory we mean to serve
SECRETS = LAB / "secrets" # a sibling we must never serve
BACKUP = LAB / "public-backup" # a sibling whose name STARTS with "public"
SECRET_FILE = SECRETS / "api_key.txt"
def build() -> bool:
"""(Re)create the lab tree. Returns True if the symlink could be created."""
if LAB.exists():
shutil.rmtree(LAB)
(PUBLIC / "notes" / "2026").mkdir(parents=True)
SECRETS.mkdir()
BACKUP.mkdir()
(PUBLIC / "report.txt").write_text("Q3 report: revenue is up\n")
(PUBLIC / "notes" / "2026" / "q3.txt").write_text("Q3 notes: ship the fix\n")
(BACKUP / "old.txt").write_text("OLD BACKUP: internal only\n")
SECRET_FILE.write_text("API_KEY=sk_live_demo_do_not_share\n")
try:
# A shortcut inside the public folder that points at the secrets folder.
os.symlink(SECRETS, PUBLIC / "shortcut", target_is_directory=True)
return True
except (OSError, NotImplementedError):
return False
if __name__ == "__main__":
linked = build()
for path in sorted(LAB.rglob("*")):
print(path.relative_to(HERE).as_posix())
print("symlink created:", linked)
Run it:
python lab.py
You should see this tree. The last line says False if your account cannot create symlinks, which is fine:
lab/public
lab/public/notes
lab/public/notes/2026
lab/public/notes/2026/q3.txt
lab/public/report.txt
lab/public/shortcut
lab/public-backup
lab/public-backup/old.txt
lab/secrets
lab/secrets/api_key.txt
symlink created: True
Here is what each piece is for. lab/public is the only folder our application is supposed to serve. lab/secrets holds api_key.txt, a fake key, and must never be served. lab/public-backup is a sibling of public whose name happens to start with the same letters, which will matter in Step 6. lab/public/shortcut is a symlink that points at lab/secrets. The recursive listing shows it as a single entry because it does not descend into symlinked folders, but anything that opens a path through it lands in the secrets folder.
Step 2: Write a download endpoint with the classic bug
Now write the vulnerable application. Save this as app_v1.py:
import os
from fastapi import FastAPI, HTTPException
from fastapi.responses import PlainTextResponse
from lab import PUBLIC
app = FastAPI()
@app.get("/files/{name:path}", response_class=PlainTextResponse)
def read_file(name: str) -> str:
path = os.path.join(PUBLIC, name) # VULNERABLE: trusts the user's path
try:
with open(path, encoding="utf-8") as f:
return f.read()
except (FileNotFoundError, IsADirectoryError):
raise HTTPException(status_code=404, detail="Not found")
Read it slowly, because the bug is one line. The route is /files/{name:path}. The :path part tells the router (Starlette, underneath FastAPI) that name may contain slashes, which is what lets a request like /files/notes/2026/q3.txt reach subfolders. The handler then calls os.path.join(PUBLIC, name), which glues the base folder and the user’s text together, and opens the result. Nothing checks where that path actually leads. The try block only turns a missing file into a clean 404.
Start the server in your first terminal and leave it running:
python -m uvicorn app_v1:app --port 8000
Step 3: Attack it, and see why your first attempt may not work
Now play the attacker. Save the following as attack.py. It sends a list of hostile paths to the server and prints the status code and the first few characters of each response. It uses http.client from the standard library on purpose, for a reason you will see in a moment.
"""Fire a list of hostile paths at a running demo server.
http.client sends the request path exactly as written. Many clients (curl,
requests) collapse ".." before sending, which would hide the bug from you.
"""
import http.client
import sys
from urllib.parse import quote
from lab import SECRET_FILE
def fetch(path: str, port: int) -> tuple[int, str]:
conn = http.client.HTTPConnection("127.0.0.1", port, timeout=5)
conn.request("GET", path)
resp = conn.getresponse()
body = resp.read().decode("utf-8", "replace").strip().replace("\n", " ")
conn.close()
return resp.status, body
PAYLOADS = {
"legit file": "report.txt",
"legit subfolder": "notes/2026/q3.txt",
"dot-dot": "../secrets/api_key.txt",
"dot-dot, percent-encoded": "..%2Fsecrets%2Fapi_key.txt",
"backslashes (Windows)": "..%5Csecrets%5Capi_key.txt",
"absolute path": quote(str(SECRET_FILE), safe=""),
"sibling directory": "../public-backup/old.txt",
"through a symlink": "shortcut/api_key.txt",
}
if __name__ == "__main__":
prefix = sys.argv[1] if len(sys.argv) > 1 else "/files/"
port = int(sys.argv[2]) if len(sys.argv) > 2 else 8000
for label, payload in PAYLOADS.items():
status, body = fetch(prefix + payload, port)
print(f"{label:26} {status} {body[:42]!r}")
In your second terminal, run it:
python attack.py
This is what I got:
legit file 200 'Q3 report: revenue is up'
legit subfolder 200 'Q3 notes: ship the fix'
dot-dot 200 'API_KEY=sk_live_demo_do_not_share'
dot-dot, percent-encoded 200 'API_KEY=sk_live_demo_do_not_share'
backslashes (Windows) 200 'API_KEY=sk_live_demo_do_not_share'
absolute path 200 'API_KEY=sk_live_demo_do_not_share'
sibling directory 200 'OLD BACKUP: internal only'
through a symlink 200 'API_KEY=sk_live_demo_do_not_share'
The first two rows are honest requests, and they work. Every other row is an attack, and every one of them returned real file contents. Here is what each one does:
- dot-dot. The request path is
/files/../secrets/api_key.txt, sonameis../secrets/api_key.txt. The joined path keeps the..in it, so when Windows opens the file it walks up out ofpublicand intosecrets. - dot-dot, percent-encoded. The same attack written as
..%2Fsecrets%2Fapi_key.txt. Servers decode%2Fback into a slash before your code runs, so filtering the raw URL text for../would miss this form. - backslashes (Windows). On Windows a backslash separates folders just like a forward slash. Linux and macOS treat it as an ordinary character, so I would expect this row to return 404 there. It is a Windows-specific hole.
- absolute path. The attacker sends the full path of the secret file. As you will see in Step 4,
os.path.jointhrows away the base folder when it meets an absolute path. - sibling directory.
../public-backup/old.txtreaches a different folder that merely lives next topublic. - through a symlink.
shortcut/api_key.txtcontains no dots and no drive letter, yet it lands insecrets, because the operating system follows the symlink.
Gotcha: your own test client may be quietly fixing the attack
Many people first try an attack like this with requests or curl, see a 404, and decide the endpoint is safe. That conclusion would be wrong, because both tools tidy up .. segments in the URL before sending anything. Save this as client_normalisation.py and run it while the server is still up:
"""Show that common clients collapse '..' before the request ever leaves your machine."""
import subprocess
import requests
URL = "http://127.0.0.1:8000/files/../secrets/api_key.txt"
prepared = requests.Request("GET", URL).prepare()
print("requests sends :", prepared.url)
for flags in ([], ["--path-as-is"]):
out = subprocess.run(
["curl", "-s", "-o", "-", "-w", " [HTTP %{http_code}]", *flags, URL],
capture_output=True, text=True,
).stdout.strip().replace("\n", " ")
print(f"curl {' '.join(flags) or '(default)':13}:", out)
python client_normalisation.py
requests sends : http://127.0.0.1:8000/secrets/api_key.txt
curl (default) : {"detail":"Not Found"} [HTTP 404]
curl --path-as-is : API_KEY=sk_live_demo_do_not_share [HTTP 200]
The first line shows what requests would actually send: the .. was already collapsed, so the request asked for /secrets/api_key.txt, which does not exist. Plain curl got a 404 for the same reason. Only curl --path-as-is sent the path untouched and got the secret. The curl help text describes that flag as: “Do not squash .. sequences in URL path”. The http.client module in attack.py also sends the path exactly as written, which is why I used it. The lesson: when you test for traversal, make sure your client is not helping the server.
Step 4: Why the bug works, starting with what os.path.join really does
Why did the absolute-path attack work? Save this as join_demo.py. It uses a made-up base folder and never touches the disk:
import os
base = r"C:\srv\public" # a made-up path: nothing here touches the disk
for part in ["report.txt", r"..\secret.txt", "/etc/passwd", r"D:\other\x.txt"]:
print(f"{part:16} -> {os.path.join(base, part)}")
python join_demo.py
report.txt -> C:\srv\public\report.txt
..\secret.txt -> C:\srv\public\..\secret.txt
/etc/passwd -> C:/etc/passwd
D:\other\x.txt -> D:\other\x.txt
The Python documentation for os.path.join explains the surprising rows: “If a segment is an absolute path (which on Windows requires both a drive and a root), then all previous segments are ignored and joining continues from the absolute path segment.” It adds a Windows-specific rule: “On Windows, the drive is not reset when a rooted path segment (e.g., r'\foo') is encountered.” That is exactly what the third row shows. The segment /etc/passwd has a root but no drive, so the drive letter is kept and the result is C:/etc/passwd, a location at the top of the drive and nowhere near the base folder. The fourth row carries its own drive, so the base folder disappears entirely.
The deeper point is that join is string manipulation with a few special cases. It checks nothing. Your program builds a string, the operating system interprets it later, and everything between those two moments is unguarded.
Step 5: Removing :path is not a fix
A tempting shortcut is to remove :path so that slashes never reach the handler. Stop the server (Ctrl+C in the first terminal), copy app_v1.py to app_v1_plain.py, and change the route from /files/{name:path} to /files/{name}. Start it with python -m uvicorn app_v1_plain:app --port 8000 and run python attack.py again:
legit file 200 'Q3 report: revenue is up'
legit subfolder 404 '{"detail":"Not Found"}'
dot-dot 404 '{"detail":"Not Found"}'
dot-dot, percent-encoded 404 '{"detail":"Not Found"}'
backslashes (Windows) 200 'API_KEY=sk_live_demo_do_not_share'
absolute path 200 'API_KEY=sk_live_demo_do_not_share'
sibling directory 404 '{"detail":"Not Found"}'
through a symlink 404 '{"detail":"Not Found"}'
This looks like progress. The dot-dot rows, the sibling row and the symlink row now return 404, because the router refuses to match a path with a forward slash in it. But the legitimate subfolder request broke too (row 2), and two attacks still work: the backslash row and the absolute-path row. A Windows path made only of backslashes contains no forward slash for the router to object to, and a percent-encoded C:\... path is the same. The router is not a security control. It only stops one character. (I only tested Windows. On Linux and macOS an absolute path is full of forward slashes, so I would expect that row to return 404 there.) And the day someone needs subfolders and adds :path back, every attack from Step 3 returns. Stop this server before you continue.
Step 6: Five checks that look right, and only one is
Before writing the fix, it is worth seeing why fixing this is harder than it looks. Below are five checks, each only a few lines long, each reasonable at a glance. The script runs all five against nine payloads (three legitimate, six hostile) and reports whether each check allowed or blocked each path. It also asks the operating system the real question, does this path actually lead outside public once everything is resolved, so it can label a wrong answer. LEAK means the check allowed a path that escapes. wrong-block means it refused a harmless one.
The five functions at the top are the checks. The function escapes() is the referee: it joins the path the naive way, asks the operating system to resolve it completely with os.path.realpath, and reports whether the result is outside public. The function verdict() compares each check’s decision with the referee’s. Save this as validator_matrix.py:
"""Five plausible-looking path checks against nine payloads. Only one survives them all."""
import os
from pathlib import Path
import lab
from lab import PUBLIC, SECRET_FILE
from safe_paths import PathTraversalError, safe_join
class Blocked(Exception):
pass
def v1_deny_dotdot(base: Path, user: str) -> Path:
if ".." in user:
raise Blocked
return Path(os.path.join(base, user))
def v2_startswith(base: Path, user: str) -> Path:
target = (base / user).resolve()
if not str(target).startswith(str(base.resolve())):
raise Blocked
return target
def v3_lexical_relative(base: Path, user: str) -> Path:
target = base / user
if not target.is_relative_to(base):
raise Blocked
return target
def v4_abspath_commonpath(base: Path, user: str) -> Path:
target = os.path.abspath(os.path.join(base, user))
root = os.path.abspath(base)
if os.path.commonpath([target, root]) != root:
raise Blocked
return Path(target)
def v5_resolve_relative(base: Path, user: str) -> Path:
try:
return safe_join(base, user)
except PathTraversalError:
raise Blocked
VALIDATORS = {
"1 no '..'": v1_deny_dotdot,
"2 startswith": v2_startswith,
"3 lexical": v3_lexical_relative,
"4 abspath": v4_abspath_commonpath,
"5 resolve": v5_resolve_relative,
}
PAYLOADS = [
("legit file", "report.txt"),
("legit subfolder", "notes/2026/q3.txt"),
("legit, dot-dot inside", "notes/../report.txt"),
("dot-dot", "../secrets/api_key.txt"),
("nested dot-dot", "notes/../../secrets/api_key.txt"),
("backslashes", "..\\secrets\\api_key.txt"),
("absolute path", str(SECRET_FILE)),
("sibling 'public-backup'", "../public-backup/old.txt"),
("via symlink", "shortcut/api_key.txt"),
]
def escapes(user: str) -> bool:
"""Ground truth: does this path really lead outside PUBLIC once the OS resolves it?"""
real = Path(os.path.realpath(os.path.join(PUBLIC, user)))
return not real.is_relative_to(Path(os.path.realpath(PUBLIC)))
def verdict(check, user: str) -> str:
try:
check(PUBLIC, user)
allowed = True
except (Blocked, ValueError):
allowed = False
if allowed and escapes(user):
return "LEAK"
if not allowed and not escapes(user):
return "wrong-block"
return "allow" if allowed else "block"
if __name__ == "__main__":
has_symlink = lab.build()
payloads = [p for p in PAYLOADS if has_symlink or p[0] != "via symlink"]
print(f"{'payload':26}" + "".join(f"{name:>14}" for name in VALIDATORS))
leaks = dict.fromkeys(VALIDATORS, 0)
for label, user in payloads:
row = f"{label:26}"
for name, check in VALIDATORS.items():
result = verdict(check, user)
leaks[name] += result == "LEAK"
row += f"{result:>14}"
print(row)
print(f"{'LEAKS':26}" + "".join(f"{leaks[name]:>14}" for name in VALIDATORS))
python validator_matrix.py
payload 1 no '..' 2 startswith 3 lexical 4 abspath 5 resolve
legit file allow allow allow allow allow
legit subfolder allow allow allow allow allow
legit, dot-dot inside wrong-block allow allow allow allow
dot-dot block block LEAK block block
nested dot-dot block block LEAK block block
backslashes block block LEAK block block
absolute path LEAK block block block block
sibling 'public-backup' block LEAK LEAK block block
via symlink LEAK block LEAK LEAK block
LEAKS 2 1 5 1 0
Only the last column has zero leaks. Here is why each of the others fails.
Check 1: refuse any path containing two dots
It blocks every dot-dot row, but it also refuses notes/../report.txt, a harmless path that stays inside public (the wrong-block cell). Worse, it lets the absolute path through, because that contains no dots at all, and it lets the symlink through for the same reason. A blacklist of bad patterns can only stop the patterns you thought of.
Check 2: resolve the path, then use startswith
Resolving first, so that .. and symlinks are handled, is the right instinct, and this check blocks five of the six attacks. It fails on ../public-backup/old.txt, because the string ...\lab\public-backup\old.txt begins with the string ...\lab\public. A string prefix is not the same thing as a parent folder. Compare path components, not characters.
Check 3: is_relative_to without resolving
The pathlib module has a method that sounds perfect, but its documentation says: “This method is string-based; it neither accesses the filesystem nor treats “..” segments specially.” So public/../secrets/api_key.txt counts as relative to public as far as the method is concerned, because it begins with public. Five of the six attacks sail through.
Check 4: abspath plus commonpath
This is the classic answer, and it is much better: os.path.abspath cleans up the dots and os.path.commonpath compares path components rather than characters. But abspath only rewrites the text of the path. It never asks the filesystem whether any folder along the way is a symlink. So it stops every dot-dot trick and leaks exactly one row, the symlink.
Check 5: resolve, then is_relative_to
The documentation for Path.resolve says: “Make the path absolute, resolving any symlinks.” It adds that “..” components are also eliminated, and that this is the only method that does so. After resolving, you hold the one true absolute location, and a component-wise comparison with is_relative_to is exactly the right question to ask. This is the check we will use.
One more Windows trap: resolve both sides
There is a subtler way to use the right check and still get the wrong answer. Save this as shortname_gotcha.py and run it:
import tempfile
from pathlib import Path
base = Path(tempfile.mkdtemp()) # on Windows this can contain an 8.3 short name
(base / "report.txt").write_text("hello\n")
candidate = (base / "report.txt").resolve()
print("base :", base)
print("candidate :", candidate)
print("inside, base NOT resolved :", candidate.is_relative_to(base))
print("inside, base.resolve() :", candidate.is_relative_to(base.resolve()))
python shortname_gotcha.py
base : C:\Users\ADMINI~1\AppData\Local\Temp\tmpxojt31z8
candidate : C:\Users\Administrator\AppData\Local\Temp\tmpxojt31z8\report.txt
inside, base NOT resolved : False
inside, base.resolve() : True
On my machine the temporary folder path contains ADMINI~1, an old-style 8.3 short name that Windows keeps for the user folder called Administrator. Calling resolve() expands short names into the real long names, so the resolved candidate no longer starts with the unresolved base, and a perfectly legitimate file is judged to be outside it. Resolving the base as well fixes it (the last line). Your output will differ if your paths contain no short names, but the rule is cheap to follow: always resolve both sides of the comparison.
Step 7: The fix, resolve first and then check containment
Now write the function. Save this as safe_paths.py:
from pathlib import Path
class PathTraversalError(ValueError):
"""The requested path would leave the allowed base directory."""
def safe_join(base, user_path: str) -> Path:
"""Join user_path onto base and guarantee the result stays inside base."""
if "\x00" in user_path: # resolve() does not reject NUL bytes, open() does
raise PathTraversalError("NUL byte in path")
base = Path(base).resolve() # resolve the base too, not just the target
try:
candidate = (base / user_path).resolve() # follows symlinks, removes ".."
except (ValueError, OSError) as exc: # embedded NUL, invalid Windows path, ...
raise PathTraversalError(f"invalid path {user_path!r}") from exc
if not candidate.is_relative_to(base):
raise PathTraversalError(f"{user_path!r} escapes the base directory")
return candidate
Walk through it once. First it rejects a NUL byte (the character with code 0) outright. Second it resolves the base folder, for the reason you just saw. Third it joins the user’s text onto the base and resolves the result, which removes every .. and follows every symlink. If the user text was an absolute path, pathlib discards the base, exactly like os.path.join does, but that is harmless here because the next line catches it. Fourth it asks is_relative_to whether the resolved location is inside the resolved base. Anything that goes wrong along the way, such as an invalid path, is reported as the same PathTraversalError, which is a subclass of ValueError. The function returns the resolved path, and you must open that returned path, not the string the user sent, because the resolved path is the one that was checked.
Why reject NUL bytes explicitly? Save nul_check.py and run it:
from pathlib import Path
from lab import PUBLIC
p = PUBLIC / "report.txt\x00.png"
print("resolve() ->", repr(p.resolve().name))
print("is_file() ->", p.is_file())
try:
open(p)
except ValueError as exc:
print("open() -> ValueError:", exc)
python nul_check.py
resolve() -> 'report.txt\x00.png'
is_file() -> False
open() -> ValueError: embedded null character
On Python 3.13.14, resolve() happily returned a path containing a NUL byte, is_file() quietly returned False, and only open() raised ValueError: embedded null character. Three functions, three different reactions. I would not want a security check to depend on which of them happens to run first, so the function rejects NUL bytes up front.
Now use it. Save this as app_v2.py:
from fastapi import FastAPI, HTTPException
from fastapi.responses import FileResponse
from lab import PUBLIC
from safe_paths import PathTraversalError, safe_join
app = FastAPI()
@app.get("/files/{name:path}")
def read_file(name: str):
try:
path = safe_join(PUBLIC, name)
except PathTraversalError:
# Same answer as a missing file, so attackers learn nothing about what exists.
raise HTTPException(status_code=404, detail="Not found")
if not path.is_file():
raise HTTPException(status_code=404, detail="Not found")
return FileResponse(path, media_type="text/plain")
Stop any server that is still running (Ctrl+C), start this one with python -m uvicorn app_v2:app --port 8000, and run the attack again:
legit file 200 'Q3 report: revenue is up'
legit subfolder 200 'Q3 notes: ship the fix'
dot-dot 404 '{"detail":"Not found"}'
dot-dot, percent-encoded 404 '{"detail":"Not found"}'
backslashes (Windows) 404 '{"detail":"Not found"}'
absolute path 404 '{"detail":"Not found"}'
sibling directory 404 '{"detail":"Not found"}'
through a symlink 404 '{"detail":"Not found"}'
Both legitimate requests return 200 and all six attacks return 404, including the backslash, absolute-path and symlink rows that beat most of the checks in Step 6. The handler answers a rejected path with exactly the same 404 it gives for a missing file, so an attacker cannot use the difference between “forbidden” and “not found” to learn which files exist.
Step 8: Or let your framework do it
If all you need is to serve a folder of static files, you do not have to write this logic at all. Starlette ships a StaticFiles class that contains its own containment check. Save this as app_v3.py:
from fastapi import FastAPI
from fastapi.staticfiles import StaticFiles
from lab import PUBLIC
app = FastAPI()
app.mount("/static", StaticFiles(directory=PUBLIC), name="static")
Stop the previous server (Ctrl+C), start this one with python -m uvicorn app_v3:app --port 8000, and attack the /static/ prefix instead:
python attack.py /static/
legit file 200 'Q3 report: revenue is up'
legit subfolder 200 'Q3 notes: ship the fix'
dot-dot 404 '{"detail":"Not Found"}'
dot-dot, percent-encoded 404 '{"detail":"Not Found"}'
backslashes (Windows) 404 '{"detail":"Not Found"}'
absolute path 404 '{"detail":"Not Found"}'
sibling directory 404 '{"detail":"Not Found"}'
through a symlink 404 '{"detail":"Not Found"}'
The same result, with no code of our own. Here is the method that makes the decision, copied from starlette/staticfiles.py in Starlette 1.7.0, the version FastAPI installed for this tutorial:
def lookup_path(self, path: str) -> tuple[str, os.stat_result | None]:
# Reject absolute paths so they cannot escape the served directory.
if path.startswith(("/", "\\")):
return "", None
for directory in self.all_directories:
joined_path = os.path.join(directory, path)
if self.follow_symlink:
full_path = os.path.abspath(joined_path)
directory = os.path.abspath(directory)
else:
full_path = os.path.realpath(joined_path)
directory = os.path.realpath(directory)
if os.path.commonpath([full_path, directory]) != str(directory):
# Don't allow misbehaving clients to break out of the static files directory.
continue
It rejects names that start with a slash or a backslash, then resolves the joined path and compares components with os.path.commonpath. Notice the follow_symlink switch. By default the code takes the realpath branch, which follows symlinks, and that is why the symlink row returned 404 above. To see the other branch, I changed the mount line to StaticFiles(directory=PUBLIC, follow_symlink=True), which uses abspath instead, and ran the same attack:
legit file 200 'Q3 report: revenue is up'
legit subfolder 200 'Q3 notes: ship the fix'
dot-dot 404 '{"detail":"Not Found"}'
dot-dot, percent-encoded 404 '{"detail":"Not Found"}'
backslashes (Windows) 404 '{"detail":"Not Found"}'
absolute path 404 '{"detail":"Not Found"}'
sibling directory 404 '{"detail":"Not Found"}'
through a symlink 200 'API_KEY=sk_live_demo_do_not_share'
Every row is unchanged except the symlink row, which now returns the secret. That is the option doing what the source above says it does, and it is the same difference you saw between checks 4 and 5 in Step 6. Turn it on only if you really want symlinks inside the served folder to be able to point elsewhere. Use StaticFiles when you are serving a whole folder, and use safe_join when your own code decides which file to return, for example after an authorization check.
Step 9: The same bug in a different costume, archive extraction
Downloads are one way for a stranger’s text to become a path. Archives are another. A zip or tar file is a list of members, and each member carries a name chosen by whoever built the archive. If your code extracts the archive by gluing each name onto a destination folder, the archive’s author decides where files land. In 2018 Snyk gave this version of the bug a memorable name, and its research page describes Zip Slip as “a widespread arbitrary file overwrite critical vulnerability, which typically results in remote command execution.” A file written outside the intended folder can replace a script, a configuration file or a startup entry.
9.1 Build some hostile archives
To test defenses we need attackers, so save two small helpers. The first, hostile.py, builds a tar file with one harmless member plus one hostile member of a chosen kind: a .. name, a backslash version of it, an absolute path (which points at a file in a scratch folder next to the destination, so nothing outside your sandbox is touched) and a symlink member that points outside. It can also build a hostile zip.
"""Builds hostile archives, so we can prove our defences work."""
import io
import tarfile
import zipfile
from pathlib import Path
def _add_file(tar: tarfile.TarFile, name: str, data: bytes) -> None:
info = tarfile.TarInfo(name)
info.size = len(data)
tar.addfile(info, io.BytesIO(data))
def build_hostile_tar(path: Path, kind: str, scratch: Path) -> None:
"""A tar with one harmless member plus one hostile member of the given kind."""
with tarfile.open(path, "w") as tar:
_add_file(tar, "docs/readme.txt", b"hello\n")
if kind == "dot-dot":
_add_file(tar, "../evil.txt", b"escaped\n")
elif kind == "backslash":
_add_file(tar, "..\\evil.txt", b"escaped\n")
elif kind == "absolute":
_add_file(tar, str(scratch / "evil.txt"), b"escaped\n")
elif kind == "symlink":
link = tarfile.TarInfo("shortcut")
link.type = tarfile.SYMTYPE
link.linkname = "../outside"
tar.addfile(link)
else:
raise ValueError(kind)
def build_hostile_zip(path: Path) -> None:
with zipfile.ZipFile(path, "w") as zf:
zf.writestr("docs/readme.txt", "hello\n")
zf.writestr("../evil.txt", "escaped\n")
The second, scratch.py, gives each demo a fresh scratch folder with an empty dest folder inside it. Anything that lands in scratch but outside dest is an escape.
import shutil
from pathlib import Path
SCRATCH = Path(__file__).resolve().parent / "scratch"
def fresh() -> Path:
"""Empty the scratch folder and return a new, empty 'dest' folder inside it."""
shutil.rmtree(SCRATCH, ignore_errors=True)
SCRATCH.mkdir()
dest = SCRATCH / "dest"
dest.mkdir()
return dest
9.2 What tar extraction does
Save this as demo_tar_filters.py. For each hostile kind it extracts three times: with no filter argument at all, with filter="tar", and with filter="data", and it classifies what happened.
import sys
import tarfile
import warnings
from hostile import build_hostile_tar
from scratch import SCRATCH, fresh
def extract(archive, dest, flt) -> str:
try:
with tarfile.open(archive) as tar:
if flt is None:
tar.extractall(dest) # no filter argument at all
else:
tar.extractall(dest, filter=flt)
except Exception as exc:
return f"refused: {type(exc).__name__}"
if (SCRATCH / "evil.txt").exists():
return "ESCAPED: wrote outside dest"
if (dest / "shortcut").is_symlink():
return "symlink pointing outside created"
return "extracted inside dest"
print(sys.version.split()[0], "on", sys.platform)
first_warning = None
for kind in ("dot-dot", "backslash", "absolute", "symlink"):
for flt in (None, "tar", "data"):
dest = fresh()
archive = SCRATCH / "attack.tar"
build_hostile_tar(archive, kind, SCRATCH)
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
result = extract(archive, dest, flt)
if caught and first_warning is None:
first_warning = f"{caught[0].category.__name__}: {caught[0].message}"
print(f"{kind:10} filter={str(flt):5} -> {result}")
print("\nwarning printed when no filter is given:\n ", first_warning)
python demo_tar_filters.py
3.13.14 on win32
dot-dot filter=None -> ESCAPED: wrote outside dest
dot-dot filter=tar -> refused: OutsideDestinationError
dot-dot filter=data -> refused: OutsideDestinationError
backslash filter=None -> ESCAPED: wrote outside dest
backslash filter=tar -> refused: OutsideDestinationError
backslash filter=data -> refused: OutsideDestinationError
absolute filter=None -> ESCAPED: wrote outside dest
absolute filter=tar -> refused: AbsolutePathError
absolute filter=data -> refused: AbsolutePathError
symlink filter=None -> symlink pointing outside created
symlink filter=tar -> symlink pointing outside created
symlink filter=data -> refused: LinkOutsideDestinationError
warning printed when no filter is given:
DeprecationWarning: Python 3.14 will, by default, filter extracted tar archives and reject files or modify their metadata. Use the filter argument to control this behavior.
What the default does on Python 3.13
The rows with filter=None are the lesson. With no filter, three of the four hostile archives wrote a file outside dest, and the fourth planted a symlink pointing outside it. Python also printed the warning shown at the bottom, which I captured from the run. (The backslash row is Windows-only, and if your account cannot create symlinks the symlink rows can look different.) The documentation explains the history: “Changed in version 3.12: Added the filter parameter. Changed in version 3.14: The filter parameter now defaults to 'data'.” Before 3.14 the default was “equivalent to fully_trusted”, so on Python 3.13 and older a plain extractall(dest) trusts the archive completely.
What the tar and data filters do
Both named filters refused the dot-dot, backslash and absolute-path archives, with OutsideDestinationError or AbsolutePathError, and nothing was written outside. The two filters differ on the symlink. The documentation for the tar filter says it will “Refuse to extract files whose absolute path (after following symlinks) would end up outside the destination.” That is about where files end up, so the link itself got through. The data filter adds: “Refuse to extract links (hard or soft) that link to absolute paths, or ones that link outside the destination.” It refused the link with LinkOutsideDestinationError. It also refuses device files and pipes. For data you do not fully trust, data is the filter to use, and it is the default from Python 3.14 on.
The documentation is direct about the limits: “Never extract archives from untrusted sources without prior inspection. Since Python 3.14, the default (data) will prevent the most dangerous security issues. However, it will not prevent all unintended or insecure behavior.” It recommends passing filter='data' explicitly so that your code stays safe on Python versions with a less secure default (3.13 and lower). PEP 706, the proposal that added filters, explains why they exist: “it’s quite tricky to do such an inspection correctly. As a result, many people don’t bother, or do the check incorrectly, resulting in security issues such as CVE-2007-4559.” That is exactly what the matrix in Step 6 showed for downloads.
9.3 A refused archive can still leave files behind
There is one more behavior to know about. Extraction happens member by member, so if the second member is refused, the first has already been written. The solution is to extract into a temporary staging folder and move it into place only if everything passed. Save this as safe_extract.py:
"""Extract untrusted archives all-or-nothing."""
import os
import shutil
import tarfile
import tempfile
import zipfile
from pathlib import Path
from safe_paths import PathTraversalError, safe_join
class ArchiveRejected(Exception):
"""The archive contained something we refuse to extract."""
def _staging_dir(dest: Path) -> Path:
if dest.exists():
raise FileExistsError(dest)
return Path(tempfile.mkdtemp(dir=dest.parent, prefix=".extract-"))
def safe_extract_tar(archive: Path, dest: Path) -> None:
"""Extract into dest, or raise and leave nothing behind."""
dest = Path(dest)
staging = _staging_dir(dest)
try:
with tarfile.open(archive) as tar:
tar.extractall(staging, filter="data")
os.replace(staging, dest) # only reached if every member passed the filter
except (tarfile.TarError, ValueError) as exc:
raise ArchiveRejected(f"{type(exc).__name__}: {exc}") from exc
finally:
shutil.rmtree(staging, ignore_errors=True)
def safe_extract_zip(archive: Path, dest: Path) -> None:
"""Same contract for zip files: reject bad names instead of silently rewriting them."""
dest = Path(dest)
staging = _staging_dir(dest)
try:
with zipfile.ZipFile(archive) as zf:
for info in zf.infolist():
safe_join(staging, info.filename)
zf.extractall(staging)
os.replace(staging, dest)
except (zipfile.BadZipFile, PathTraversalError) as exc:
raise ArchiveRejected(f"{type(exc).__name__}: {exc}") from exc
finally:
shutil.rmtree(staging, ignore_errors=True)
Read safe_extract_tar first. It creates a fresh staging folder next to the destination, extracts into it with filter="data", and only if every member passed does it rename the staging folder to the destination with os.replace. If anything raises, the finally block deletes the staging folder, so the caller sees either a complete destination or none at all. It catches tarfile.TarError, which is the base class of the filter errors as well as of corrupt-archive errors, plus ValueError. That last one is there because in an extra experiment on Windows a member with a drive-relative name (mine was D:drive_rel_evil.txt) made the filter raise ValueError: Paths don't have the same drive instead of a filter error. Failing closed on any exception is the safe behavior.
Now compare the two approaches. Save this as demo_leftovers.py:
import tarfile
from hostile import build_hostile_tar
from safe_extract import ArchiveRejected, safe_extract_tar
from scratch import SCRATCH, fresh
dest = fresh()
archive = SCRATCH / "attack.tar"
build_hostile_tar(archive, "dot-dot", SCRATCH)
print("== extractall with filter='data' straight into dest ==")
try:
with tarfile.open(archive) as tar:
tar.extractall(dest, filter="data")
except tarfile.FilterError as exc:
print("refused:", type(exc).__name__)
print("left behind in dest:", sorted(p.relative_to(dest).as_posix() for p in dest.rglob("*") if p.is_file()))
print("\n== safe_extract_tar: all or nothing ==")
fresh()
build_hostile_tar(archive, "dot-dot", SCRATCH)
target = SCRATCH / "site"
try:
safe_extract_tar(archive, target)
except ArchiveRejected as exc:
print("ArchiveRejected:", str(exc)[:58], "...")
print("target exists:", target.exists())
print("stray staging dirs:", [p.name for p in SCRATCH.iterdir() if p.name.startswith(".extract-")])
print("escaped file:", (SCRATCH / "evil.txt").exists())
python demo_leftovers.py
== extractall with filter='data' straight into dest ==
refused: OutsideDestinationError
left behind in dest: ['docs/readme.txt']
== safe_extract_tar: all or nothing ==
ArchiveRejected: OutsideDestinationError: '../evil.txt' would be extracted ...
target exists: False
stray staging dirs: []
escaped file: False
The first half extracts straight into dest. The archive is refused, but docs/readme.txt was already written and stays behind. The second half uses the helper: the archive is rejected, the target folder does not exist, no staging folders are left over, and nothing escaped.
9.4 The zipfile module plays by different rules
The zipfile module makes different choices. Its documentation warns: “Never extract archives from untrusted sources without prior inspection. It is possible that files are created outside of path, for example, members that have absolute filenames or filenames with “..” components. This module attempts to prevent that.” The extract() notes say that for an absolute member name the drive and leading slashes are stripped, and that all .. components are removed. Save this as demo_zip.py to see it happen:
import stat
import zipfile
from hostile import build_hostile_zip
from safe_extract import ArchiveRejected, safe_extract_zip
from scratch import SCRATCH, fresh
fresh()
names = ["docs/readme.txt", "../zip_evil.txt", "/zip_abs.txt", "..\\zip_bs.txt",
"C:\\zip_drive.txt", "a/../../zip_nested.txt"]
archive = SCRATCH / "names.zip"
with zipfile.ZipFile(archive, "w") as zf:
for name in names:
zf.writestr(name, "x\n")
dest = SCRATCH / "zdest"
dest.mkdir()
with zipfile.ZipFile(archive) as zf:
print("stored as:", zf.namelist())
zf.extractall(dest)
print("inside dest :", sorted(p.relative_to(dest).as_posix() for p in dest.rglob("*") if p.is_file()))
print("outside dest:", sorted(p.name for p in SCRATCH.glob("*.txt")))
print("\n== a zip entry that claims to be a symlink ==")
fresh()
info = zipfile.ZipInfo("zlink")
info.create_system = 3 # Unix
info.external_attr = (stat.S_IFLNK | 0o777) << 16
archive = SCRATCH / "zlink.zip"
with zipfile.ZipFile(archive, "w") as zf:
zf.writestr(info, "../outside_dir")
dest = SCRATCH / "zdest"
dest.mkdir()
with zipfile.ZipFile(archive) as zf:
zf.extractall(dest)
link = dest / "zlink"
print("is_symlink:", link.is_symlink(), "| is_file:", link.is_file(), "| content:", link.read_text())
print("\n== safe_extract_zip: reject instead of rewrite ==")
fresh()
archive = SCRATCH / "attack.zip"
build_hostile_zip(archive)
target = SCRATCH / "site"
try:
safe_extract_zip(archive, target)
except ArchiveRejected as exc:
print("ArchiveRejected:", str(exc)[:60])
print("target exists:", target.exists())
python demo_zip.py
stored as: ['docs/readme.txt', '../zip_evil.txt', '/zip_abs.txt', '../zip_bs.txt', 'C:/zip_drive.txt', 'a/../../zip_nested.txt']
inside dest : ['a/zip_nested.txt', 'docs/readme.txt', 'zip_abs.txt', 'zip_bs.txt', 'zip_drive.txt', 'zip_evil.txt']
outside dest: []
== a zip entry that claims to be a symlink ==
is_symlink: False | is_file: True | content: ../outside_dir
== safe_extract_zip: reject instead of rewrite ==
ArchiveRejected: PathTraversalError: '../evil.txt' escapes the base directory
target exists: False
All six names ended up inside dest and nothing landed outside, so plain zipfile.extractall is safer than unfiltered tarfile.extractall here. Two more things are worth noticing. First, the names are silently rewritten: ../zip_evil.txt became zip_evil.txt. That is safe, but it hides the fact that someone tried something, and a rewritten name can collide with a legitimate one. For an upload endpoint I prefer to reject, which is what safe_extract_zip does by validating every member name with safe_join before extracting anything (the last section of the output). Second, a zip entry that claims to be a symlink was extracted as an ordinary file containing the link’s target text. In this run zipfile did not create a symlink.
Step 10: Lock it in with tests
Everything above was checked by eye. Turn it into an automated safety net so a future refactor cannot quietly reopen the hole. Save this as test_traversal.py:
import io
import os
import tarfile
import pytest
import lab
from hostile import build_hostile_tar, build_hostile_zip
from lab import PUBLIC, SECRET_FILE
from safe_extract import ArchiveRejected, safe_extract_tar, safe_extract_zip
from safe_paths import PathTraversalError, safe_join
HAS_SYMLINK = lab.build()
windows_only = pytest.mark.skipif(os.name != "nt", reason="backslash is a separator only on Windows")
needs_symlink = pytest.mark.skipif(not HAS_SYMLINK, reason="this account cannot create symlinks")
@pytest.mark.parametrize("name", ["report.txt", "notes/2026/q3.txt", "notes/../report.txt"])
def test_legitimate_paths_are_allowed(name):
assert safe_join(PUBLIC, name).is_file()
@pytest.mark.parametrize("name", [
"../secrets/api_key.txt",
"notes/../../secrets/api_key.txt",
"../public-backup/old.txt",
str(SECRET_FILE),
"report.txt\x00.png",
pytest.param("..\\secrets\\api_key.txt", marks=windows_only),
pytest.param("shortcut/api_key.txt", marks=needs_symlink),
])
def test_escapes_are_rejected(name):
with pytest.raises(PathTraversalError):
safe_join(PUBLIC, name)
@pytest.mark.parametrize("kind", ["dot-dot", "absolute", "symlink",
pytest.param("backslash", marks=windows_only)])
def test_hostile_tar_is_rejected_and_leaves_nothing(tmp_path, kind):
archive = tmp_path / "attack.tar"
build_hostile_tar(archive, kind, tmp_path)
with pytest.raises(ArchiveRejected):
safe_extract_tar(archive, tmp_path / "site")
assert not (tmp_path / "site").exists()
assert not (tmp_path / "evil.txt").exists()
assert not [p for p in tmp_path.iterdir() if p.name.startswith(".extract-")]
def test_hostile_zip_is_rejected_and_leaves_nothing(tmp_path):
archive = tmp_path / "attack.zip"
build_hostile_zip(archive)
with pytest.raises(ArchiveRejected):
safe_extract_zip(archive, tmp_path / "site")
assert not (tmp_path / "site").exists()
def test_a_clean_archive_is_extracted(tmp_path):
archive = tmp_path / "ok.tar"
with tarfile.open(archive, "w") as tar:
info = tarfile.TarInfo("docs/readme.txt")
info.size = 6
tar.addfile(info, io.BytesIO(b"hello\n"))
safe_extract_tar(archive, tmp_path / "site")
assert (tmp_path / "site" / "docs" / "readme.txt").read_text() == "hello\n"
The parametrize lists are the payload catalog: three legitimate paths that must work, and every escape from this tutorial that must be rejected. The windows_only and needs_symlink marks skip the cases that cannot apply on your machine instead of failing. The archive tests use pytest’s tmp_path fixture and assert three things after a rejected archive: the destination does not exist, no file escaped, and no staging folder was left behind. Run the suite:
python -m pytest -q
................ [100%]
16 passed in 0.09s
All 16 tests passed on my Windows machine. On Linux or macOS the two Windows-only cases are skipped instead of passing.
Common mistakes and gotchas
Testing with a client that cleans the URL. As Step 3 showed, requests and plain curl collapse .. before sending. Use http.client or curl --path-as-is.
Checking the text instead of the location. Blacklisting .. misses absolute paths and symlinks and blocks harmless paths. Canonicalize with resolve(), then check.
Comparing strings instead of path components. startswith lets public-backup through a check meant for public. Use is_relative_to on resolved paths.
Resolving only one side. On Windows, 8.3 short names made a legitimate file look outside its own folder. Resolve the base and the candidate.
Assuming your operating system is everyone’s. The backslash and drive-letter attacks only work on Windows, and a check tested only on Linux never sees them.
Trusting the archive or the default. Before Python 3.14, tarfile.extractall trusted archives completely. Pass filter="data" explicitly. The same warning applies to shutil.unpack_archive, whose documentation says: “For zip files, filter is not accepted. For tar files, it is recommended to use 'data' (default since Python 3.14), unless using features specific to tar and UNIX-like filesystems.”
Extracting straight into the final folder. A refused archive can leave its earlier members behind. Stage first, then move into place.
Forgetting time. safe_join checks a path now and your code opens it a moment later. If untrusted users can create or replace files or symlinks inside the served folder, for example through an upload feature or by extracting an untrusted archive there, the check can be out of date by the time the file is opened. Keep served folders read-only for the web process, and never let user-controlled archives create links that point outside the folder, which is what the data filter refuses.
Solving it at the wrong level. OWASP’s first piece of advice is: “Prefer working without user input when using file system calls”, and it suggests, “Use indexes rather than actual portions of file names when templating or using language files”. If a download can be identified by an ID that your code maps to a file, the user never supplies a path at all, and there is nothing to traverse.
How to confirm everything works end to end
- Rebuild the lab with
python lab.py. - Start the fixed application in the first terminal with
python -m uvicorn app_v2:app --port 8000, then runpython attack.pyin the second. You should see exactly two 200 responses (the legitimate files) and six 404 responses. - Run
python -m pytest -q. On Windows with symlink permission you should see 16 passed. - Audit your own code. Search your project for
extractall(,unpack_archive(,FileResponse(,send_file(, and anyopen(oros.path.join(whose arguments include a request value or an archive member name. Every hit should go throughsafe_join, a framework helper such asStaticFiles, or an explicitfilter="data".
Next steps
Path traversal is one member of a family in which untrusted text crosses a boundary into something that interprets it. This site has tutorials on several others: preventing server-side request forgery, where user text becomes a request to the wrong place, preventing SQL injection with parameterized queries, stopping XSS with a Content Security Policy, and preventing argument injection in a Windows URI protocol handler, which shares this tutorial’s Windows-specific parsing surprises. If you would like an automated reviewer to help catch risky code before it is committed, see how to run Bandit in a pre-commit hook. Then take the payload catalog from Step 10 and run it against the real file-serving and upload code in your own projects. The hostile paths are already written, and the first time one of them gets through, you will be glad you looked.








No Comment! Be the first one.