How to Catch Hidden Bugs in Rust With Property-Based Testing and proptest
Learn property-based testing in Rust with proptest by installing Rust, writing a real function, and watching proptest catch and help you fix an actual bug.
A unit test only checks the specific inputs you thought to write down. If you test encode("aaabbbcc") and it comes back correct, that tells you nothing about what happens when the input is empty, a single character, or contains a digit. Property-based testing flips the approach: instead of picking examples by hand, you describe a property, a rule that should hold true for every valid input, and a tool generates hundreds of random inputs trying to break it. In Rust, the standard tool for this is proptest, the same technique Trail of Bits describes using to audit Rust codebases in their Testing Handbook.
Table Of Content
- What property-based testing catches that example-based tests miss
- Prerequisites
- Step 1: Install Rust
- Step 2: Create a new Rust project
- Confirm the project runs
- Step 3: Write a function worth testing
- Step 4: Start with example-based tests
- Why passing examples do not prove the code is correct
- Step 5: Add proptest to the project
- Step 6: Write your first property test
- Step 7: Run it and read a real failure
- What “shrinking” means, in practice
- The regression file proptest just created
- Step 8: Understand why the bug happens
- Step 9: Fix the bug
- Step 10: Confirm the fix
- Raise the number of cases for extra confidence
- Common mistakes and gotchas
- Verify everything works end-to-end
- Next steps
This tutorial installs Rust from scratch, writes a small text-compression function, tests it with ordinary hand-picked examples first, then adds a proptest property test that finds a genuine bug those examples missed entirely. You will watch proptest shrink a random failing input down to the simplest possible case, read what it saved to disk, fix the actual bug, and confirm the fix holds under thousands of random inputs. Every command below was run against a real installation of rustc 1.97.1, cargo 1.97.1, and proptest 1.11.0 on Ubuntu 26.04 LTS while writing this tutorial, and every behavior described was verified by actually triggering it, not assumed from documentation. Nothing here needs a server or special permissions: everything happens in one scratch directory you can delete when you are done.
What property-based testing catches that example-based tests miss
An ordinary Rust unit test looks like this: call a function with one specific input, assert the output matches one specific expected value. That is called example-based testing, and it is the right default for most code. Its weakness is that it only ever tells you about the exact inputs you wrote. If a bug only shows up for, say, strings containing a digit character, and you never happened to type a digit into a test string, an example-based test suite can pass at 100% while the bug ships anyway.
Property-based testing addresses this by inverting who picks the inputs. You write a property: a statement that should be true for any valid input, not just one. A strategy (proptest’s term for an input generator) then produces large numbers of random values, and proptest calls your property function with each one, watching for a failure. When it finds an input that breaks the property, it does not just report that random value, it performs shrinking: it repeatedly tries smaller or simpler variations of the failing input that still fail, converging on the smallest, most readable counterexample it can find. You will see this happen for real later in this tutorial.
A property test is not a replacement for example-based tests, it is a complement. Hand-picked examples are still the fastest way to document what a function should do for the cases you already know about; property tests are how you find out about the cases you did not think of.
Prerequisites
- A Linux or macOS machine with terminal access. This tutorial was verified on Ubuntu 26.04 LTS, but every command is identical on any modern Linux distribution or macOS, since Rust’s tooling is not distribution-specific.
- No Rust installation is required beforehand; Step 1 installs it. No administrator or root access is required either, rustup installs entirely inside your home directory.
- Basic comfort reading code and running shell commands. No prior testing framework experience is assumed; every Rust- and proptest-specific concept is explained the first time it appears.
- About 200 MB of free disk space for the Rust toolchain, and a few more megabytes for this tutorial’s project and its dependencies.
Step 1: Install Rust
Rust’s official installer is a small tool called rustup, which manages Rust toolchain versions and installs the two commands you will use constantly: rustc (the compiler) and cargo (the build tool, package manager, and test runner all in one). Install it with the command from the official Rust install page:
# Context: any Linux or macOS terminal, no existing Rust installation assumed.
# Purpose: download and run the official rustup installer.
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
The installer will ask a couple of questions about install options; the default (option 1, a standard installation) is correct for this tutorial. When it finishes, you will see output ending with something like this:
info: downloading installer
info: profile set to default
info: default host triple is x86_64-unknown-linux-gnu
info: syncing channel updates for stable-x86_64-unknown-linux-gnu
info: downloading 6 components
info: default toolchain set to stable-x86_64-unknown-linux-gnu
stable-x86_64-unknown-linux-gnu installed - rustc 1.97.1
Rust is installed now. Great!
Your exact version numbers may differ slightly since Rust ships a new stable release every six weeks, that is expected. The installer prints instructions to reload your shell’s PATH; the simplest way is to open a new terminal, or run the command it suggests: source "$HOME/.cargo/env". Confirm both tools are on your PATH:
rustc --version
cargo --version
Expected output (your version numbers will vary):
rustc 1.97.1 (8bab26f4f 2026-07-14)
cargo 1.97.1 (c980f4866 2026-06-30)
If either command is not found, close and reopen your terminal so it picks up the updated PATH, then try again.
Step 2: Create a new Rust project
Cargo organizes Rust code into packages (called crates), each with a manifest file describing its name, version, and dependencies. Create one now:
# Context: a scratch directory of your choosing.
# Purpose: scaffold a new Rust binary project named rle_proptest_demo.
cargo new rle_proptest_demo
cd rle_proptest_demo
Expected output:
Creating binary (application) package
This creates a small directory tree: Cargo.toml (the manifest, listing the package name, version, and dependencies) and src/main.rs (a starter program that prints Hello, world!).
Confirm the project runs
cargo run
Expected output (the first run also compiles, so you will see extra lines before this):
Running `target/debug/rle_proptest_demo`
Hello, world!
cargo run compiles your code and immediately executes the resulting binary. You will use this same command throughout the tutorial to check your work.
Step 3: Write a function worth testing
To have something meaningful to test, this tutorial implements run-length encoding (RLE), a simple, real compression technique: consecutive repeated characters are replaced with a count and the character, so "aaabbbcc" becomes "3a3b2c" (three a’s, three b’s, two c’s). RLE is a good subject for property-based testing because it has a natural, checkable property: decoding an encoded string should always reproduce the original string exactly, for any input, not just the ones you thought to try by hand.
Replace the contents of src/main.rs with the following:
fn encode(input: &str) -> String {
let mut result = String::new();
let mut chars = input.chars().peekable();
while let Some(current) = chars.next() {
let mut count = 1;
while chars.peek() == Some(¤t) {
chars.next();
count += 1;
}
result.push_str(&count.to_string());
result.push(current);
}
result
}
fn decode(input: &str) -> String {
let mut result = String::new();
let mut chars = input.chars().peekable();
while chars.peek().is_some() {
let mut digits = String::new();
while let Some(c) = chars.peek() {
if c.is_ascii_digit() {
digits.push(*c);
chars.next();
} else {
break;
}
}
let count: usize = digits.parse().unwrap_or(1);
if let Some(ch) = chars.next() {
for _ in 0..count {
result.push(ch);
}
}
}
result
}
fn main() {
let examples = ["aaabbbcc", "abc", "", "wwwwwwwwwwww"];
for s in examples {
let encoded = encode(s);
let decoded = decode(&encoded);
println!("{s:?} -> {encoded:?} -> {decoded:?}");
}
}
encode walks the input character by character using a Peekable iterator (one that lets you look at the next character without consuming it), counting how many times each character repeats in a row, then writes the count followed by the character. decode does the reverse: it reads leading digit characters as a count, then repeats the next character that many times. Run it:
cargo run
Expected output:
"aaabbbcc" -> "3a3b2c" -> "aaabbbcc"
"abc" -> "1a1b1c" -> "abc"
"" -> "" -> ""
"wwwwwwwwwwww" -> "12w" -> "wwwwwwwwwwww"
All four hand-picked examples round-trip correctly. Keep that output in mind; it is about to look a lot less reassuring.
Step 4: Start with example-based tests
Before reaching for property-based testing, write down the examples you already checked by hand as real, permanent tests. Add this to the bottom of src/main.rs:
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn encodes_simple_runs() {
assert_eq!(encode("aaabbbcc"), "3a3b2c");
}
#[test]
fn round_trips_hand_picked_examples() {
for s in ["aaabbbcc", "abc", "", "wwwwwwwwwwww", "a"] {
assert_eq!(decode(&encode(s)), s);
}
}
}
#[cfg(test)] tells the compiler to include this module only when running tests, not in a normal build. #[test] marks a function as a test case that cargo test should run. Run the suite:
cargo test
Expected output:
running 2 tests
test tests::encodes_simple_runs ... ok
test tests::round_trips_hand_picked_examples ... ok
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
Why passing examples do not prove the code is correct
Both tests pass. It would be reasonable, looking at only this output, to conclude the encode/decode pair is correct. It is not: there is a real bug in the code above, and none of these five hand-picked strings happen to trigger it. That is the exact failure mode example-based testing cannot protect you from: the tests only ever exercise the inputs someone thought to write. The next step hands input selection over to a tool that does not share your assumptions about what a “normal” string looks like.
Step 5: Add proptest to the project
Add proptest as a dev-dependency, a dependency that is only compiled in for tests and benchmarks, never for the release binary your users would run:
cargo add --dev proptest
Expected output (abbreviated):
Updating crates.io index
Adding proptest v1.11.0 to dev-dependencies
Locking 39 packages to latest Rust 1.97.1 compatible versions
...
This adds a line to Cargo.toml under a new [dev-dependencies] section: proptest = "1.11.0". proptest is dual-licensed MIT OR Apache-2.0 (confirmed directly from its published package manifest), the same permissive licensing scheme used by most of the Rust ecosystem, including Rust itself.
Step 6: Write your first property test
Add this inside the tests module, after the two existing #[test] functions:
use proptest::prelude::*;
proptest! {
#[test]
fn round_trips_any_string(s in any::<String>()) {
prop_assert_eq!(decode(&encode(&s)), s);
}
}
A few new pieces here:
proptest! { ... }is a macro that turns an ordinary-looking function into a property test. Instead of a fixed input,s in any::<String>()declaressas a variable that proptest should fill in using theany::<String>()strategy: proptest’s built-in generator for arbitrary RustStringvalues, which includes empty strings, single characters, digits, punctuation, and Unicode text, not just the tidy alphabetic examples a human tends to type first.prop_assert_eq!works like the standard library’sassert_eq!, but instead of panicking immediately, it returns a special failure value that proptest’s test runner catches. Functionally, panicking withassert_eq!and returning a failure withprop_assert_eq!have the same effect as far as proptest’s shrinking is concerned; the practical difference, according to proptest’s own documentation, is noise: a raw panic prints Rust’s full panic message (and a backtrace, if enabled) for every intermediate failing case during shrinking, whileprop_assert_eq!keeps that quiet and only prints the final, minimal counterexample.
Step 7: Run it and read a real failure
cargo test
This is the real, unedited output from running that command against the code above:
running 3 tests
test tests::encodes_simple_runs ... ok
test tests::round_trips_hand_picked_examples ... ok
test tests::round_trips_any_string ... FAILED
failures:
---- tests::round_trips_any_string stdout ----
proptest: Saving this and future failures in rle_proptest_demo/proptest-regressions/main.txt
proptest: If this test was run on a CI system, you may wish to add the following line to your copy of the file. (You may need to create it.)
cc 18f4f84737f93f5ec4cca00cfbb09b1d0411ada6a59bec5f465c6232f2408aeb
thread 'tests::round_trips_any_string' panicked at src/main.rs:69:5:
Test failed: assertion failed: `(left == right)`
left: `""`,
right: `"0"` at src/main.rs:72.
minimal failing input: s = "0"
successes: 4
local rejects: 0
global rejects: 0
test result: FAILED. 2 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
The two hand-written example tests still pass, exactly as before. The new property test fails immediately, and it fails on the simplest possible input: the one-character string "0". Neither of the two authors of this tutorial’s example list ever thought to test a bare digit; proptest found it by trying random strings until one broke the round-trip property.
What “shrinking” means, in practice
proptest did not necessarily fail on "0" on its very first attempt. Internally, it generates a random string, and if that fails, it does not just report it, it tries smaller and simpler variations (shorter strings, simpler characters) that still reproduce the failure, repeating until it cannot simplify any further. The line minimal failing input: s = "0" is the end result of that process: the smallest, most readable input proptest could find that still breaks the property. This is the practical payoff of property-based testing over pure random “fuzzing”: you do not just learn that something is broken, you get handed close to the simplest possible reproduction of the break.
The regression file proptest just created
Notice the line Saving this and future failures in rle_proptest_demo/proptest-regressions/main.txt. Open that file:
cat proptest-regressions/main.txt
Actual contents (your seed value will differ, the comment will not):
# Seeds for failure cases proptest has generated in the past. It is
# automatically read and these particular cases re-run before any
# novel cases are generated.
#
# It is recommended to check this file in to source control so that
# everyone who runs the test benefits from these saved cases.
cc 18f4f84737f93f5ec4cca00cfbb09b1d0411ada6a59bec5f465c6232f2408aeb # shrinks to s = "0"
According to proptest’s own source documentation, when this kind of failure persistence is enabled (the default), every case saved in this file is retested first, before any new random cases are generated, on every future cargo test run. That means once proptest finds a bug, that exact failing case becomes a permanent regression test automatically, with no extra code from you. This is exactly why the file’s own header comment recommends committing it to version control: without it, a fix that gets accidentally reverted later would only be caught again if proptest happened to re-roll the same failing string at random.
Step 8: Understand why the bug happens
Trace through what encode("0") actually produces. The input is a single character, '0', appearing once, so encode writes the count (1) followed by the character (0): the output is the two-character string "10".
Now trace decode("10"). The decoder’s first loop greedily reads every consecutive digit character it sees as part of the count, with no way to know where the “count” is supposed to end and the “actual character” is supposed to begin, since in this format both are drawn from the same character set. It reads '1', then keeps going and reads '0' too, because '0' is also an ASCII digit. The digit-reading loop consumes the entire string as the count (10), leaving nothing behind to treat as the repeated character, so the final if let Some(ch) = chars.next() finds nothing and the function returns an empty string.
The bug is a format ambiguity, not a typo: the original encoding scheme has no way to distinguish “a count digit” from “a payload character that happens to be a digit.” Any input containing a digit character was silently broken, and none of the five original hand-picked examples happened to contain one.
Step 9: Fix the bug
The fix is to make the boundary between the count and the character explicit, instead of relying on both being drawn from disjoint-looking character sets. Insert a fixed separator character between them. Update both functions:
fn encode(input: &str) -> String {
let mut result = String::new();
let mut chars = input.chars().peekable();
while let Some(current) = chars.next() {
let mut count = 1;
while chars.peek() == Some(¤t) {
chars.next();
count += 1;
}
result.push_str(&count.to_string());
result.push(':');
result.push(current);
}
result
}
fn decode(input: &str) -> String {
let mut result = String::new();
let mut chars = input.chars().peekable();
while chars.peek().is_some() {
let mut digits = String::new();
while let Some(c) = chars.peek() {
if c.is_ascii_digit() {
digits.push(*c);
chars.next();
} else {
break;
}
}
let count: usize = digits.parse().unwrap_or(1);
chars.next(); // consume the ':' separator
if let Some(ch) = chars.next() {
for _ in 0..count {
result.push(ch);
}
}
}
result
}
encode now writes the count, a literal :, then the character. decode reads the leading digits exactly as before (that boundary was never ambiguous, since : is not a digit), then unconditionally consumes and discards the next single character as the separator, then treats the character after that as the payload, whatever it is. This works even if the payload character is itself a digit, or even a literal :, because decode never searches for a matching delimiter character, it always consumes exactly one character in the separator’s position and exactly one character in the payload’s position, regardless of what those characters happen to be.
Update the one hand-picked assertion that checked the exact encoded format, since that format intentionally changed:
#[test]
fn encodes_simple_runs() {
assert_eq!(encode("aaabbbcc"), "3:a3:b2:c");
}
Step 10: Confirm the fix
cargo test
Expected output:
running 3 tests
test tests::encodes_simple_runs ... ok
test tests::round_trips_hand_picked_examples ... ok
test tests::round_trips_any_string ... ok
test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.04s
All three tests pass, including round_trips_any_string, which is now also silently replaying the exact "0" case saved in proptest-regressions/main.txt before generating any new random cases, confirming the specific bug you just fixed cannot regress unnoticed.
Raise the number of cases for extra confidence
By default, proptest runs 256 fresh random cases per property (confirmed directly from proptest’s own configuration source, plus every persisted regression case). You can raise that from the command line without touching any code, using the PROPTEST_CASES environment variable:
PROPTEST_CASES=2000 cargo test --release round_trips_any_string
Expected output (release mode compiles more slowly but runs faster, real timing from this tutorial’s machine):
running 1 test
test tests::round_trips_any_string ... ok
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 2 filtered out; finished in 0.02s
2,000 random strings, tested in roughly 20 milliseconds once compiled, with zero failures. That is a meaningfully stronger confidence signal than the five strings a human originally picked by hand, for less effort than writing even one more example-based test.
Common mistakes and gotchas
- Using
assert_eq!instead ofprop_assert_eq!inside aproptest!block. It still works, shrinking is unaffected, but every intermediate failing case during shrinking prints a full Rust panic message, which gets noisy fast. Prefer theprop_-prefixed macros inside property tests. - Forgetting that proptest is a dev-dependency. That is by design, not an oversight, since
cargo add --devkeeps testing-only code out of your actual release binary, but it does meanany::<String>()and friends are unavailable outside#[cfg(test)]code. That is expected and correct. - Not committing the
proptest-regressions/directory to version control. proptest’s own generated file explicitly recommends checking it in. Skipping this means a teammate who reverts your fix only gets caught again if proptest happens to randomly regenerate the same failing input. - Reaching for regex-based string strategies too early. proptest also supports generating strings that match a regular expression, useful when you need realistic-looking input instead of fully arbitrary text, documented at docs.rs/proptest. Since that regex is itself written as a Rust string literal, any backslash in the pattern needs to be doubled. It is easy to get this wrong the first time;
any::<String>(), used throughout this tutorial, sidesteps the issue entirely and is a reasonable default until you specifically need constrained input. - Treating every property failure as a bug in your code. Sometimes a property fails because the strategy generated an input your function was never meant to handle in the first place. When that happens, the fix is to narrow the strategy to your function’s actual documented input domain, not to contort the function into accepting garbage. In this tutorial’s case, the bug was real: any string is a legitimate input to a general-purpose text compressor, so the fix belonged in
encode/decode, not in the test.
Verify everything works end-to-end
To confirm the whole tutorial reproduces cleanly from nothing, the full sequence is: install Rust with rustup, cargo new a project, write encode/decode, add hand-picked tests, cargo add --dev proptest, add the property test, watch it fail and shrink to s = "0", fix the format with a separator character, then watch all three tests pass, including a replay of the saved regression case. If you followed along, running cargo test one final time should show 3 passed; 0 failed, exactly as in Step 10.
Nothing in this tutorial touched anything outside the project directory except the Rust toolchain itself. If you want to remove the example project, delete its directory; if you want to remove Rust entirely, rustup ships its own uninstaller: rustup self uninstall.
Next steps
The run-length encoder in this tutorial is deliberately small so the property-testing workflow stays front and center, but the same pattern, write a property, let proptest generate inputs, fix what it finds, scales to real parsers, serializers, and data structures. A few directions worth exploring next:
- Trail of Bits’ Testing Handbook chapter that inspired this tutorial covers considerably more Rust-specific testing ground: using Miri to detect undefined behavior, mutation testing, coverage measurement, and a checklist of real gotchas found while auditing production Rust code.
- The proptest book covers writing custom strategies for your own types, deriving the
Arbitrarytrait automatically, and tuning shrinking behavior. - For finding crashes and panics in code that parses untrusted input specifically (file formats, network protocols), cargo-fuzz is a complementary, coverage-guided fuzzing tool that pairs well with the property-testing habits from this tutorial.
- Try adding one property test to an existing project, ideally for a function that parses, serializes, or transforms data and round-trips through some inverse operation. Those round-trip properties tend to be the easiest to write and the most likely to catch a real, previously unknown bug on the first run.








No Comment! Be the first one.