Cloudflare Traces Turns Distributed Tracing Into a Trust Decision at the Edge
Cloudflare Traces entered open beta on October 2 with a setting that decides whether the edge joins a caller’s trace, a choice the W3C specification warns about for public services.
On October 2, Cloudflare put Cloudflare Traces into open beta, “extending automatic tracing beyond Workers to the rest of the request path.” Once tracing is enabled for a domain, each sampled request gets a trace with spans for the supported security rules, transformations, cache decisions, routing, Worker execution and origin handling that touched it. You can read the traces in the dashboard or export them over OTLP to another tool.
Table Of Content
- What the beta gives you
- Decision one: whether to believe a caller’s trace
- What the setting does
- What the W3C specification says about this setup
- What the documents do not say
- Decision two: sampling is decided before anything is known
- What a trace rule can see
- What a 1 percent baseline catches
- Reading the spans without blaming your origin
- The meter: what the beta costs from December 1
- One pool for logs and traces
- Where the pages disagree
- What to do before you turn it on
- Our reading
- What we checked, and what we did not
The waterfall view is the easy part to describe. The settings behind it are more interesting, because they force decisions with security and budget consequences: whether Cloudflare should believe a caller’s trace identity, how to choose which requests get traced before anyone knows which ones will fail, and how much of a shared logging allowance traces may use. This piece reads the launch post against Cloudflare’s own documentation, the W3C Trace Context specification and OpenTelemetry’s sampling guide. We did not run Traces itself, because that needs a Cloudflare account with the beta enabled on a hostname, so everything below comes from those documents. The one test they leave open is written out so you can run it.
What the beta gives you
The documentation breaks the beta into five areas, each with its own setting.
| Area | What the docs say | Where to read it |
|---|---|---|
| Spans | Spans for supported steps such as Rules, request routing, Cache, Workers and origin connections. “A span appears only when the corresponding operation runs.” | Overview, spans reference |
| Sampling | One default sample rate per domain. Trace Rules override it for matching requests, and the first matching rule wins. | Configuration |
| Storage and export | Persist traces in Cloudflare, send them to an account-level OTLP destination, or do both. | Configuration, OpenTelemetry export |
| Trace context | Incoming trace context (the default is Reject) and Forward to origin, configured separately. | Configuration |
| Price | Shared Observability pricing that starts on December 1, 2026. | Pricing |
Decision one: whether to believe a caller’s trace
What the setting does
Distributed tracing works because each system passes along a W3C traceparent header: a trace ID, the caller’s span ID and trace flags, one of which is the sampled flag. Cloudflare Traces can accept that header from an incoming request so its spans join a trace that began before the request reached Cloudflare. In the launch post’s words, “An incoming propagation policy controls whether Cloudflare accepts that context.”
The configuration page says what happens by default: “The default setting is Reject, which ignores the incoming context so Cloudflare does not join the caller’s trace.” It attaches a caution to the alternative: “Cloudflare does not verify that incoming trace context came from a trusted caller. If you accept it, treat the joined trace as untrusted.”
The launch post’s list of work after the beta names the missing middle: “Authenticated context propagation: Let trusted callers continue an existing trace without accepting context from every incoming request.” Until that ships, a public hostname has two settings: join no trace, or join any caller’s trace.
What the W3C specification says about this setup
The Level 1 specification, which is the version Cloudflare’s configuration page links to, has a security section that describes this situation without naming Cloudflare: “When distributed tracing is enabled on a service with a public API and naively continues any trace with the sampled flag set, a malicious attacker could overwhelm an application with tracing overhead, forge trace-id collisions that make monitoring data unusable, or run up your tracing bill with your SaaS tracing vendor.” It adds that tracing vendors “should account for these situations and make sure that checks and balances are in place,” and it offers one example: “One example of such protection may be different tracing behavior for authenticated and unauthenticated requests.” That is close to the split Cloudflare lists as future work. The text is in section 7.2 of the published Level 1 page and in the specification’s source repository.
What the documents do not say
The five Traces pages we read (overview, configuration, spans, pricing and OpenTelemetry export) do not mention the sampled flag, and neither does the launch post. So the documentation does not say whether a request that arrives with the flag set is traced even when the domain’s sample rate would have skipped it, which is the behavior the specification warns about. We cannot tell you that Cloudflare Traces does this. We can tell you the pages do not rule it out.
A staging hostname can settle it. Enable Traces there with Incoming trace context set to Accept and the lowest default sample rate the dashboard allows, then send three requests and note each response’s Ray ID. The docs say you can find one request’s trace by filtering on its Ray ID. We have not run this test.
HOST=https://staging.example.com/
TID=a1b2c3d4e5f60718293a4b5c6d7e8f90
curl -s -o /dev/null -D - -H "traceparent: 00-$TID-0102030405060708-01" "$HOST" | grep -i '^cf-ray'
curl -s -o /dev/null -D - -H "traceparent: 00-$TID-0102030405060708-00" "$HOST" | grep -i '^cf-ray'
curl -s -o /dev/null -D - "$HOST" | grep -i '^cf-ray'
The first request claims a sampled caller, the second an unsampled one, and the third sends no context. Read the dashboard like this: if only the first request has a trace, the flag overrides the domain’s rate; if that trace carries the ID a1b2c3d4e5f60718293a4b5c6d7e8f90, Cloudflare joined the caller’s trace; if none of the three has a trace, the rate won. Repeat with the setting on Reject to see the default behavior for comparison. Run it only against a hostname you own, and remember that the setting you leave on Accept for the test is the setting you are testing.
Decision two: sampling is decided before anything is known
What a trace rule can see
The configuration page describes the sampling model in a few sentences: “Default sample rate (%) sets how many requests Cloudflare traces. For example, a rate of 10% traces about 10 out of every 100 requests.” Trace Rules override that rate for matching requests, and “Cloudflare uses the first matching rule, so rule order matters.” OpenTelemetry’s sampling guide has a name for this design: “Head sampling is a sampling technique used to make a sampling decision as early as possible.”
The spans reference shows where that decision falls. The root span, cloudflare_request, “Starts after tracing configuration and sampling are resolved” and then includes the request processing, so the decision is made before the request is processed. That limits what a rule can know. The launch post’s examples all use properties of the incoming request: Trace Rules “use the same Cloudflare Rules language, so you can target paths, methods, headers, IP addresses, geographies, or combinations of those properties.” We found no rule field in the pages that depends on the response, such as a 5xx status. OpenTelemetry’s guide states the limit plainly: “For example, you cannot ensure that all traces with an error within them are sampled with head sampling alone.”
What a 1 percent baseline catches
The launch post suggests a low baseline: “You might trace 1% of requests during normal operation, giving you a continuous view of request behavior without collecting a trace for every request.” The chance that a baseline leaves at least one trace of a fault depends on how many requests the fault touches. If each request is sampled independently at rate p and a fault touches n requests, the chance of at least one trace is 1 - (1 - p) ** n.
| Requests the fault touches | 1 percent baseline | 5 percent baseline | 10 percent baseline |
|---|---|---|---|
| 10 | 9.6 percent | 40.1 percent | 65.1 percent |
| 20 | 18.2 percent | 64.2 percent | 87.8 percent |
| 50 | 39.5 percent | 92.3 percent | 99.5 percent |
| 100 | 63.4 percent | 99.4 percent | over 99.9 percent |
| 300 | 95.1 percent | over 99.9 percent | over 99.9 percent |
The table is our arithmetic, not Cloudflare’s, and it assumes independent sampling per request; real traffic and Cloudflare’s sampler may differ. Read it this way: a fault that touches 20 requests leaves a trace only 18.2 percent of the time at the 1 percent rate the post suggests.
That is the case for Trace Rules: raise the rate for a slice you already suspect. The launch post’s examples are tracing “100% of traffic for their hostname, source IP, or identifying request header” and “100% of requests carrying a temporary debug header.” A rule helps only if it exists before the failure, and the post lists “Ad hoc tracing: Capture a specific request on demand without changing the baseline sampling rate” as future work. One caution is our own reading: a rule keyed on a debug header is a switch that anyone who learns the header value can flip for their own requests. Combine it with a condition a stranger cannot satisfy, such as a source IP range, since the rules language lets you combine properties.
Reading the spans without blaming your origin
The spans reference spends as many words on what each duration does not measure as on what it does. These lines are worth keeping next to any waterfall.
| Span | What the reference says |
|---|---|
cloudflare_request |
“Starts after tracing configuration and sampling are resolved,” and it “excludes processing before tracing is activated.” |
cache |
“Includes the full response body, not just the cache lookup.” |
dynamic |
“It is not just the time your origin spends processing the request.” |
upstream |
“It is not just the network travel time between locations.” |
response |
“This includes the time spent to fully stream the response back to the client.” |
The launch post’s worked example says “there was a cache miss that went to origin and spent 527ms of the 539ms getting a response.” The post also tells readers to expand “nested cache, upstream, and origin spans” to see where the time went. The reference table lists no span for the origin fetch under that name; the nearest name, http_request_origin, is the Origin Rules phase. The span the reference describes as fetching an uncached response is dynamic, and it includes establishing a connection, sending the request and its body, waiting, and transferring the full response body. The post does not say which span carries the 527 ms. If it is dynamic, the figure is an upper bound on what your application did, not a measurement of it.
Workers spans have their own switch: “Workers tracing needs to be enabled to see Workers spans in a Cloudflare Trace.”
The meter: what the beta costs from December 1
One pool for logs and traces
Cloudflare’s pages put the start of billing at December 1, 2026: “persisted traces contribute to Cloudflare Observability ingestion and storage usage.” The pricing page says Cloudflare Traces share account-level allowances with Workers Logs, Workers Traces, Containers logs, R2 Data Access Logs, AI Gateway logs and Issues. On a Free account that matters, because “On Free, Cloudflare stops ingesting new data when the account reaches its daily limit. Ingestion resumes when the allowance resets at 00:00 UTC.” The Free allowance is 0.5 GB per day, so a burst of traces can use up the daily ingestion that the logs in the same pool also need. On Paid, “Paid usage continues automatically after an allowance is exhausted,” at $0.25 per GB of ingestion beyond 50 GB per billing cycle.
Billing is by volume, not by span count. The launch post says “Instead of charging by the number of spans/events, pricing is based on how much observability data you ingest and how long you retain it.” The pricing page defines ingestion as “the event body, the default enriched attributes, and custom attributes,” measured before compression and excluding “dropped events,” and it adds that “Sampling data before storage reduces both ingestion and storage usage.” None of the pages gives an average size for a trace, so you cannot yet turn a sample rate into gigabytes. Export a few days of traces to a receiver you control during the beta and count the bytes. The export page says Cloudflare does not support the binary OTLP format, so use a receiver that accepts the JSON encoding, and expect your count to differ from Cloudflare’s pre-compression one.
Two details are left open. The configuration page says persisted traces count toward usage and that you can turn off Persist to Cloudflare if you want traces only in an external tool, but no page says whether a trace that is exported and not persisted counts toward ingestion. And the Workers side has a different default: for Workers Traces, “The default sampling rate is 1, meaning 100% of requests will be traced if tracing is enabled.” Workers Traces draw on the same pool, so set head_sampling_rate explicitly on any Worker you trace before December 1.
Where the pages disagree
Two pages dated October 2 describe the plans differently.
| Detail | Launch post | Pricing page |
|---|---|---|
| Included storage on Paid | “10 GB-month of storage per billing cycle” | “12 GB-month per billing cycle” |
| Who moves on December 1 | A plan row labeled “Paid and Enterprise” | “Enterprise accounts move to this pricing when their contract renews.” |
The pricing page carries the formulas and a worked example, so it is the better page to budget from, but ask Cloudflare which storage figure is intended. Propagation is described differently too. The Traces configuration says Forward to origin lets Cloudflare include trace context in requests it sends to your origin, while the known-limitations page for Workers tracing says that “When exporting traces to external platforms, trace IDs are not propagated to services outside of Cloudflare,” and that Cloudflare is working on “automatic trace context propagation.” The two pages may describe different paths, a proxied request in one and a Worker’s own outbound calls in the other, but neither says so. Test the path your traffic actually takes before you promise anyone a joined trace.
What to do before you turn it on
- Leave Incoming trace context on Reject for any hostname that faces the public internet unless every caller is yours. If you accept it, follow the docs and treat the joined trace as untrusted: do not key access decisions or billing on trace IDs.
- Run the three-request test above on a staging hostname before enabling Accept anywhere that matters.
- Choose the baseline from the size of the incident you need to catch, using the table, and put the slices you already suspect into Trace Rules ahead of time.
- Keep the spans reference beside the waterfall, and do not read
dynamicas origin time. - Before December 1, estimate bytes per trace with an export to a receiver you control, ask Cloudflare whether export-only traces count toward ingestion, and set
head_sampling_rateon traced Workers.
Our reading
Cloudflare shipped the emitting half of distributed tracing before the trust half. The default, Reject, is the safe one, and nothing in the documentation says the beta is unsafe; the exposure begins when a team flips Accept to get a connected waterfall across its stack. The roadmap suggests Cloudflare sees the gap, and the specification describes the shape of the fix: different behavior for callers you can authenticate and callers you cannot. Until then, treat Accept on a public hostname as a decision to test and write down, not a checkbox.
The same release week produced Cloudflare’s agent billing betas, which we covered in Two Cloudflare Agent Billing Betas Turn Web Monetization Into a Question of Who Holds the Meter. Pay Per Use there leans on usage that the buyer reports about itself, so the question is similar: how far to trust what a caller says about itself. If you need the practical side of carrying a trace through your own Python services, our tutorial on fixing broken distributed traces between Python microservices shows how a traceparent header is extracted and injected.
What we checked, and what we did not
On October 4 we read the launch post, the Traces overview, configuration, spans and pricing pages, the OpenTelemetry export page, the Workers tracing pages, the Level 1 Trace Context security text and OpenTelemetry’s sampling guide. We did not test any behavior in a live Cloudflare account, so the sampled-flag question stays open until someone runs the test above. The probabilities are computed, not measured. The disagreements between pages are as they stood on October 4 and may be corrected.








No Comment! Be the first one.