How to Fix Broken Distributed Traces Between Python Microservices With OpenTelemetry
A hands-on Python tutorial on why distributed traces fragment across microservices and how to fix it with OpenTelemetry's W3C Trace Context propagation.
If your team runs more than one service, you have probably had this experience: you add tracing, you can see spans in your backend, and it still does not help during an incident, because every trace stops after one hop. Service A shows a complete little trace of its own work. Service B, which A just called, shows a completely separate trace of its own work. Nothing connects them. You have two traces where you needed one, and reconstructing what actually happened means manually matching timestamps across two disconnected views. A recent CNCF engineering write-up gave this problem a memorable name: zero plus zero equals two, meaning two services with individually perfect tracing can still add up to zero usable end-to-end visibility, and you are left counting two broken traces instead of one working one.
Table Of Content
- Understanding Traces, Spans, and Why They Split
- Prerequisites
- Step 1: Set Up the Project
- Step 2: Build a Shared Span Recorder
- Step 3: Build the Downstream Service
- Step 4: Build the Upstream Service and Watch the Trace Break
- Step 5: Understand the traceparent Header
- Step 6: Fix the Propagation
- Step 6a: Fix the Client
- The Half-Fix Gotcha
- Step 6b: Fix the Server
- Step 7: Visualize the Trace Tree
- Step 8: When Two Systems Speak Different Propagation Formats
- Common Mistakes and Gotchas
- How to Verify Everything Works End to End
- Next Steps
This tutorial teaches you the mechanism behind that failure, and the fix, using nothing but Python. You will build two tiny Flask services, deliberately break the trace between them, watch the exact failure happen with real captured output, then fix it using the same standard that real tracing backends and service meshes rely on: W3C Trace Context propagation. By the end you will understand not just how to fix this in your own Python services, but why the identical bug shows up in much bigger systems, including the Kubernetes and Istio service mesh setup described in the CNCF post that inspired this tutorial, where a proxy sidecar and an application disagreed on trace format and silently split every request into two disconnected traces.
Every command in this tutorial was actually run while writing it, on a real local machine with no Docker, no Kubernetes, and no cloud account required. Every block of output you see below is copied verbatim from that run, including one file where a leftover unused import was caught and removed before publishing.
Understanding Traces, Spans, and Why They Split
Before writing any code, it helps to know the four ideas that everything else in this tutorial is built from.
A span is a record of one unit of work: “the storefront service spent 40 milliseconds calling the inventory service.” A span has a name, a start time, an end time, and a set of key/value attributes describing what happened.
A trace is a collection of spans that together describe one end-to-end request, such as a single page load or API call, as it moves through every service that touched it. A healthy trace for “check whether an item is in stock” would contain one span for the storefront service and one span for the inventory service it called, linked together.
Every span carries a trace ID (which trace it belongs to) and a span ID (its own unique identifier). A span can also carry a parent span ID, pointing at the span that caused it to exist. A tracing backend reconstructs the tree you see in a UI like Jaeger or Grafana Tempo entirely from these three fields: it groups every span that shares a trace ID, then nests each span under the span whose ID matches its parent span ID.
That reconstruction step is exactly where things go wrong in a fragmented system. If Service A generates a brand new, random trace ID for its own span, and Service B also generates its own brand new, random trace ID when it handles A’s request instead of reusing A’s trace ID, then no backend on earth can stitch those two spans back into one trace. They do not share a trace ID. They are not fragments of one trace; they are two separate, complete, single-span traces that happen to have occurred around the same time. This is the literal mechanism behind “zero plus zero equals two,” and it is what you are about to reproduce and then fix.
The fix is called context propagation: before Service A calls Service B over HTTP, it stamps its current trace ID and span ID onto an outgoing request header. When Service B receives the request, instead of starting a fresh trace, it reads that header and starts its own span as a child of the span described in the header. The rest of this tutorial builds exactly that, twice: once without propagation so you can see it fail, and once with it so you can see it work.
Prerequisites
You will need:
- Python 3.10 or newer. This tutorial was written and tested against Python 3.13.14 on Windows, but every tool used here is cross-platform and works the same way on macOS and Linux.
pip, which ships with Python.- A terminal, and ideally two terminal windows or tabs so you can run two services side by side.
curl, or just a web browser, to make a test request.- Basic familiarity with what an HTTP request looks like (a URL, a response) and enough Python to read a short Flask app. You do not need any prior tracing or OpenTelemetry experience; that is what this tutorial teaches.
You do not need Docker, Kubernetes, or a real tracing backend like Jaeger. Everything here runs as plain Python processes on your own machine, and you will write a small script at the end that plays the role of a tracing backend, just enough to prove the concept.
Step 1: Set Up the Project
Create a new folder and a virtual environment, then install the OpenTelemetry SDK along with Flask and the requests library, which the two demo services will use:
mkdir otel-trace-demo
cd otel-trace-demo
python -m venv venv
Activate it. On Windows PowerShell:
.\venv\Scripts\Activate.ps1
On macOS or Linux:
source venv/bin/activate
Now install the packages:
pip install opentelemetry-api opentelemetry-sdk flask requests
Confirm the install worked and check versions:
python -c "import importlib.metadata as m; [print(pkg, m.version(pkg)) for pkg in ('opentelemetry-sdk', 'flask', 'requests')]"
Expected output (your exact version numbers may be newer):
opentelemetry-sdk 1.44.0
flask 3.1.3
requests 2.34.2
What each package is for: opentelemetry-api is the vendor-neutral interface for creating spans and propagating context; opentelemetry-sdk is the actual implementation (the API alone does nothing by itself, it needs the SDK registered to record and export real spans); Flask will play the role of two tiny backend services; requests is what one service uses to call the other over HTTP, exactly like fetch or any other HTTP client would in a real system.
Step 2: Build a Shared Span Recorder
OpenTelemetry’s SDK does not send spans anywhere on its own. You attach an exporter, a small class that decides what to do with every finished span. In production this is usually the OTLP exporter, sending spans over the network to a collector, which forwards them to a backend like Jaeger or Grafana Tempo. For this tutorial, you will write your own exporter that appends every span as one line of JSON to a shared file. This is intentionally close to what a real backend does internally: it receives a stream of individual, timestamped spans and only later groups them by trace ID.
Create tracing_setup.py:
import json
import os
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor, SpanExporter, SpanExportResult
SPANS_FILE = os.path.join(os.path.dirname(__file__), "spans.jsonl")
class FileSpanExporter(SpanExporter):
"""Appends every finished span to spans.jsonl as one JSON object per line."""
def export(self, spans):
with open(SPANS_FILE, "a", encoding="utf-8") as f:
for span in spans:
record = {
"service": span.resource.attributes.get("service.name"),
"name": span.name,
"trace_id": format(span.context.trace_id, "032x"),
"span_id": format(span.context.span_id, "016x"),
"parent_span_id": format(span.parent.span_id, "016x") if span.parent else None,
"attributes": dict(span.attributes or {}),
}
f.write(json.dumps(record) + "\n")
return SpanExportResult.SUCCESS
def shutdown(self):
pass
def configure_tracing(service_name: str):
provider = TracerProvider(resource=Resource.create({"service.name": service_name}))
provider.add_span_processor(SimpleSpanProcessor(FileSpanExporter()))
trace.set_tracer_provider(provider)
return trace.get_tracer(service_name)
What this does: configure_tracing() is a small helper both services will call once at startup. It creates a TracerProvider, OpenTelemetry’s central object for producing spans, tags it with a service.name resource attribute so you can tell which service a span came from later, and attaches your FileSpanExporter through a SimpleSpanProcessor, which calls export() synchronously every time a span finishes (a BatchSpanProcessor would batch spans up first; SimpleSpanProcessor is easier to reason about while learning, since output appears immediately). Inside export(), note span.parent.span_id if span.parent else None: a span’s parent is None when nothing set an incoming context for it, which is exactly the case that causes fragmentation, and you are about to see that happen for real.
Step 3: Build the Downstream Service
Create inventory_service.py, a minimal Flask service that answers “how many of this item do we have”:
from flask import Flask, jsonify
from tracing_setup import configure_tracing
tracer = configure_tracing("inventory-service")
app = Flask(__name__)
STOCK = {"widget": 42, "gadget": 7}
@app.route("/stock/<item>")
def get_stock(item):
with tracer.start_as_current_span("check_stock") as span:
span.set_attribute("item.name", item)
count = STOCK.get(item, 0)
span.set_attribute("item.count", count)
return jsonify({"item": item, "count": count})
if __name__ == "__main__":
app.run(port=5001)
What this does: every request to /stock/<item> opens a new span called check_stock, attaches a couple of attributes describing the request, and returns a JSON response. Nothing here reads anything from the incoming HTTP request except the URL, which is the root of the problem you are about to see.
Run it in its own terminal:
python inventory_service.py
Expected output:
* Serving Flask app 'inventory_service'
* Debug mode: off
WARNING: This is a development server. Do not use it in a production deployment. Use a production WSGI server instead.
* Running on http://127.0.0.1:5001
Press CTRL+C to quit
Leave it running. In a second terminal, confirm it works:
curl http://127.0.0.1:5001/stock/widget
{"count":42,"item":"widget"}
Flask’s development server logs the request in the first terminal, and a new line has appeared in a file called spans.jsonl in your project folder, the file your exporter just wrote to. That single line is one complete, valid span. The problem shows up once a second service is involved.
Step 4: Build the Upstream Service and Watch the Trace Break
Now create storefront_naive.py, representing a second service that needs to check inventory before showing a product page. This is the naive, unfixed version: it calls the inventory service over plain HTTP, with no thought given to tracing beyond wrapping its own work in a span:
import requests
from tracing_setup import configure_tracing
tracer = configure_tracing("storefront-service")
def check_widget_availability():
with tracer.start_as_current_span("check_widget_availability") as span:
response = requests.get("http://127.0.0.1:5001/stock/widget")
data = response.json()
span.set_attribute("item.count", data["count"])
print(f"storefront-service: widget count = {data['count']}")
if __name__ == "__main__":
check_widget_availability()
Stop and restart inventory_service.py so it starts with a clean log (or just note the current line count in spans.jsonl), then run the storefront script from your second terminal:
python storefront_naive.py
Expected output:
storefront-service: widget count = 42
That looks completely fine, the request succeeded and returned the right count. Now look at what actually landed in spans.jsonl for this one logical request:
{"service": "inventory-service", "name": "check_stock", "trace_id": "7f7ee7ca67bd7ad0b7e0bf363abffd97", "span_id": "487a72b97657fb42", "parent_span_id": null, "attributes": {"item.name": "widget", "item.count": 42}}
{"service": "storefront-service", "name": "check_widget_availability", "trace_id": "59039eb54a1c6dd8669dc1a40d7bb92c", "span_id": "486412fa93b554b4", "parent_span_id": null, "attributes": {"item.count": 42}}
Look closely at the two trace_id values: 7f7ee7ca... and 59039eb5.... They are completely different. Even though the storefront service’s call directly caused the inventory service’s span to exist, nothing on the wire told the inventory service that. From a tracing backend’s point of view, these are two unrelated, single-span traces that happened to occur around the same moment. This is “zero plus zero equals two” reproduced on your own machine: two individually correct services, two individually correct spans, and zero connection between them. Every service you add to a call chain like this multiplies the problem, not just adds to it, because none of them share a common trace ID with any of the others.
Step 5: Understand the traceparent Header
The fix is a single HTTP header, standardized by the W3C as Trace Context, called traceparent. It looks like this (a real example captured later in this tutorial):
traceparent: 00-7ae2cd25e37694729f91e606521b96e4-04de6f650e322894-03
It has four dash-separated parts:
- Version (
00): the format version of the header itself. In practice this is always00today. - Trace ID (
7ae2cd25e37694729f91e606521b96e4): a 32-character hex value, 16 bytes, identifying the whole trace. Every span that should belong to the same trace must carry this exact value. - Parent ID (
04de6f650e322894): a 16-character hex value, 8 bytes, identifying the specific span that made this request. The receiving service uses this as the parent of the new span it creates. - Trace flags (
03): an 8-bit field. Its lowest bit (0x01) is the sampled flag. The OpenTelemetry Python SDK installed in this tutorial also sets the second-lowest bit (0x02), the random-trace-id flag defined in the W3C Trace Context Level 2 draft, which simply asserts that the trace ID was generated with enough real randomness to be safely used as a sampling seed elsewhere in the pipeline.0x01 | 0x02 = 0x03, which is why you see03rather than01in output from a current OpenTelemetry SDK.
The OpenTelemetry API exposes exactly two functions for working with this header, and you do not need to build or parse the string yourself: propagate.inject(carrier) writes the current span’s trace ID and span ID into a dictionary (or any object that behaves like one) as a traceparent entry, and propagate.extract(carrier) reads a traceparent entry back out of a dictionary and turns it into a context object you can hand to a new span. Step 6 uses both.
Step 6: Fix the Propagation
Fixing this requires a change on both sides of the call: the caller has to send the header, and the callee has to read it. Doing only one side, which is a common half-fix, leaves you exactly where you started, and this section proves that with real output before showing the full fix.
Step 6a: Fix the Client
Create storefront_fixed.py:
import requests
from opentelemetry import propagate
from tracing_setup import configure_tracing
tracer = configure_tracing("storefront-service")
def check_widget_availability():
with tracer.start_as_current_span("check_widget_availability") as span:
# Stamp the *current* trace/span IDs onto an outgoing headers dict.
headers = {}
propagate.inject(headers)
print(f"storefront-service: outgoing headers = {headers}")
response = requests.get("http://127.0.0.1:5001/stock/widget", headers=headers)
data = response.json()
span.set_attribute("item.count", data["count"])
print(f"storefront-service: widget count = {data['count']}")
if __name__ == "__main__":
check_widget_availability()
What changed: propagate.inject(headers) reads OpenTelemetry’s notion of “the currently active span” (the check_widget_availability span created two lines earlier) and writes its trace ID and span ID into the headers dictionary as a traceparent string. That dictionary is then passed straight to requests.get(..., headers=headers), so it travels to the inventory service as a normal HTTP request header, the same way a browser sends a Cookie or Authorization header.
The Half-Fix Gotcha
Before fixing the server, try running this new, propagation-aware client against the original, unfixed inventory_service.py from Step 3, which is still ignoring incoming headers entirely:
python storefront_fixed.py
storefront-service: outgoing headers = {'traceparent': '00-7fdcf5fc0d9e8c704ed8c6ea235b2608-7dd5300e9f641438-03'}
storefront-service: widget count = 42
The client is now doing everything right, a real traceparent header goes out on the wire. But look at spans.jsonl:
{"service": "inventory-service", "name": "check_stock", "trace_id": "9888e31169c26afc54dc4083337a2543", "span_id": "24f5220c47e3c016", "parent_span_id": null, "attributes": {"item.name": "widget", "item.count": 42}}
{"service": "storefront-service", "name": "check_widget_availability", "trace_id": "7fdcf5fc0d9e8c704ed8c6ea235b2608", "span_id": "7dd5300e9f641438", "parent_span_id": null, "attributes": {"item.count": 42}}
Still two different trace IDs. The header arrived at the inventory service, but nothing there ever looked at it, so it was simply ignored, and the service generated a brand new trace ID exactly as before. Sending the header is necessary but not sufficient. This exact failure mode, a caller that propagates correctly meeting a callee that silently ignores the header, is precisely what the CNCF post’s authors found in their production Istio setup: their instrumented application was propagating W3C trace context correctly the entire time, but Envoy’s sidecar proxy had its own built-in tracer that, in their words, “starts a brand-new root span for every request. It never reads the incoming W3C traceparent header.” The fix on their side was a mesh configuration change; the fix on your side, for the plain Flask service you are building, is one more line of code.
Step 6b: Fix the Server
Create inventory_service_fixed.py:
from flask import Flask, jsonify, request
from opentelemetry import propagate
from tracing_setup import configure_tracing
tracer = configure_tracing("inventory-service")
app = Flask(__name__)
STOCK = {"widget": 42, "gadget": 7}
@app.route("/stock/<item>")
def get_stock(item):
# Read the traceparent header the caller sent, and turn it back into
# a Context object that says "continue this trace, don't start a new one".
parent_context = propagate.extract(request.headers)
with tracer.start_as_current_span("check_stock", context=parent_context) as span:
span.set_attribute("item.name", item)
count = STOCK.get(item, 0)
span.set_attribute("item.count", count)
return jsonify({"item": item, "count": count})
if __name__ == "__main__":
app.run(port=5001)
What changed: propagate.extract(request.headers) looks for a traceparent entry in the incoming Flask request’s headers and, if found, returns a context object representing “this request is continuing an existing trace.” That context is then passed as the context= argument to start_as_current_span(), which tells OpenTelemetry to make the new span a child of whatever was described in the header, instead of starting a fresh trace.
Stop the unfixed inventory service (Ctrl+C in its terminal) and start the fixed one instead:
python inventory_service_fixed.py
Now run the fixed client again:
python storefront_fixed.py
storefront-service: outgoing headers = {'traceparent': '00-7ae2cd25e37694729f91e606521b96e4-04de6f650e322894-03'}
storefront-service: widget count = 42
And the payoff, straight from spans.jsonl:
{"service": "inventory-service", "name": "check_stock", "trace_id": "7ae2cd25e37694729f91e606521b96e4", "span_id": "ec3f2517865610e4", "parent_span_id": "04de6f650e322894", "attributes": {"item.name": "widget", "item.count": 42}}
{"service": "storefront-service", "name": "check_widget_availability", "trace_id": "7ae2cd25e37694729f91e606521b96e4", "span_id": "04de6f650e322894", "parent_span_id": null, "attributes": {"item.count": 42}}
Both spans now share the same trace_id (7ae2cd25...), and the inventory service’s parent_span_id (04de6f65...) is an exact match for the storefront service’s own span_id. That is a correctly linked, two-span, single trace. A real backend receiving these two spans would draw one waterfall diagram with the inventory call nested visibly inside the storefront request, instead of showing you two disconnected, unhelpful fragments.
Step 7: Visualize the Trace Tree
A real tracing backend’s core job is exactly what you are about to build in miniature: read a flat stream of spans, group them by trace_id, and nest each one under the span whose ID matches its parent_span_id. Create reconstruct_trace.py:
import json
import sys
from collections import defaultdict
SPANS_FILE = sys.argv[1] if len(sys.argv) > 1 else "spans.jsonl"
def load_spans(path):
spans = []
with open(path, encoding="utf-8") as f:
for line in f:
line = line.strip()
if line:
spans.append(json.loads(line))
return spans
def print_tree(span, by_parent, spans_by_id, depth=0):
indent = " " * depth
print(f"{indent}- {span['service']}: {span['name']} (span_id={span['span_id'][:8]})")
for child_id in by_parent.get(span["span_id"], []):
print_tree(spans_by_id[child_id], by_parent, spans_by_id, depth + 1)
def main():
spans = load_spans(SPANS_FILE)
by_trace = defaultdict(list)
for span in spans:
by_trace[span["trace_id"]].append(span)
print(f"Found {len(spans)} span(s) across {len(by_trace)} trace(s) in {SPANS_FILE}\n")
for trace_id, trace_spans in by_trace.items():
spans_by_id = {s["span_id"]: s for s in trace_spans}
by_parent = defaultdict(list)
roots = []
for s in trace_spans:
if s["parent_span_id"] and s["parent_span_id"] in spans_by_id:
by_parent[s["parent_span_id"]].append(s["span_id"])
else:
roots.append(s)
print(f"trace_id={trace_id[:16]}... ({len(trace_spans)} span(s))")
for root in roots:
print_tree(root, by_parent, spans_by_id)
print()
if __name__ == "__main__":
main()
Copy the broken example’s output from Step 4 into a file called broken_spans.jsonl (or just save a copy of spans.jsonl right after running the naive version), and the fixed example’s output into fixed_spans.jsonl, then run the script against each:
python reconstruct_trace.py broken_spans.jsonl
Found 2 span(s) across 2 trace(s) in broken_spans.jsonl
trace_id=7f7ee7ca67bd7ad0... (1 span(s))
- inventory-service: check_stock (span_id=487a72b9)
trace_id=59039eb54a1c6dd8... (1 span(s))
- storefront-service: check_widget_availability (span_id=486412fa)
python reconstruct_trace.py fixed_spans.jsonl
Found 2 span(s) across 1 trace(s) in fixed_spans.jsonl
trace_id=7ae2cd25e3769472... (2 span(s))
- storefront-service: check_widget_availability (span_id=04de6f65)
- inventory-service: check_stock (span_id=ec3f2517)
Two traces became one, and one flat trace became a visible parent/child tree. This is the concrete, visual version of the difference the CNCF post’s authors describe in their own production fix: they went from seeing “one trace for the application and another for the sidecar” to seeing, in their words, “all spans together into a single trace that contains both,” turning clusters of roughly 14 or 2 spans into unified traces of 51 to 75 spans for their checkout service. The mechanism is identical at any scale, whether it is your two-service demo or a production Kubernetes deployment: one shared trace ID beats two disconnected ones.
Step 8: When Two Systems Speak Different Propagation Formats
There is a second, subtler way tracing breaks, and it is the specific issue the CNCF post spends most of its length on. It is not enough for both sides to propagate something; they have to propagate the same format. W3C Trace Context (the traceparent header you just used) is one format. B3, originally from Zipkin, is another, older format that many service mesh proxies, including Envoy’s built-in Zipkin tracer, speak by default. It uses entirely different header names: x-b3-traceid, x-b3-spanid, and x-b3-sampled, instead of a single traceparent string.
If your application only emits W3C headers and the infrastructure between your services only understands B3, you get exactly the same fragmentation you just fixed, except this time propagation code is running correctly on every side, and it is still broken, because the two ends are reading different vocabularies. You can reproduce this precisely. Install the B3 propagator package:
pip install opentelemetry-propagator-b3
Create mesh_format_mismatch.py:
from opentelemetry import propagate, trace
from opentelemetry.baggage.propagation import W3CBaggagePropagator
from opentelemetry.propagators.b3 import B3MultiFormat
from opentelemetry.propagators.composite import CompositePropagator
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.trace.propagation.tracecontext import TraceContextTextMapPropagator
provider = TracerProvider(resource=Resource.create({"service.name": "app"}))
tracer = provider.get_tracer("app")
print("--- Scenario 1: app only emits W3C tracecontext (the OTel Python default) ---")
propagate.set_global_textmap(TraceContextTextMapPropagator())
with tracer.start_as_current_span("app-span") as span:
headers = {}
propagate.inject(headers)
print("App sent headers:", headers)
# A B3-only mesh sidecar (like Envoy's default Zipkin/B3 tracer) tries to read it.
b3 = B3MultiFormat()
extracted_ctx = b3.extract(headers)
extracted_span = trace.get_current_span(extracted_ctx)
print("Sidecar (B3-only) sees parent span context:", extracted_span.get_span_context())
print(" -> is_valid:", extracted_span.get_span_context().is_valid)
print()
print("--- Scenario 2: app emits BOTH tracecontext and B3 (the CNCF fix) ---")
propagate.set_global_textmap(CompositePropagator([
TraceContextTextMapPropagator(),
W3CBaggagePropagator(),
B3MultiFormat(),
]))
with tracer.start_as_current_span("app-span-2") as span:
headers = {}
propagate.inject(headers)
print("App sent headers:", headers)
b3 = B3MultiFormat()
extracted_ctx = b3.extract(headers)
extracted_span = trace.get_current_span(extracted_ctx)
print("Sidecar (B3-only) sees parent span context:", extracted_span.get_span_context())
print(" -> is_valid:", extracted_span.get_span_context().is_valid)
same_trace = format(extracted_span.get_span_context().trace_id, "032x") == format(span.get_span_context().trace_id, "032x")
print(" -> trace_id matches app's:", same_trace)
Run it:
python mesh_format_mismatch.py
--- Scenario 1: app only emits W3C tracecontext (the OTel Python default) ---
App sent headers: {'traceparent': '00-0edf2d211bb0b0d7972b1d3e941b00ba-2c8d8b024caaf6e6-03'}
Sidecar (B3-only) sees parent span context: SpanContext(trace_id=0x00000000000000000000000000000000, span_id=0x0000000000000000, trace_flags=0x00, trace_state=[], is_remote=False)
-> is_valid: False
--- Scenario 2: app emits BOTH tracecontext and B3 (the CNCF fix) ---
App sent headers: {'traceparent': '00-91eae0d101979752772450de5d87db39-038eda391578a3de-03', 'x-b3-traceid': '91eae0d101979752772450de5d87db39', 'x-b3-spanid': '038eda391578a3de', 'x-b3-sampled': '1'}
Sidecar (B3-only) sees parent span context: SpanContext(trace_id=0x91eae0d101979752772450de5d87db39, span_id=0x038eda391578a3de, trace_flags=0x01, trace_state=[], is_remote=True)
-> is_valid: True
-> trace_id matches app's: True
In Scenario 1, the simulated B3-only sidecar cannot read a traceparent header at all; it comes back with an all-zeros, invalid span context, meaning it would start a brand new trace of its own, the exact behavior the CNCF post attributes to Envoy’s default OTel tracer. In Scenario 2, the application emits both formats side by side using a CompositePropagator, and the B3-only sidecar can now extract a valid, matching trace ID.
This is precisely the fix the CNCF post’s authors applied to their real Istio deployment: they configured their application services to emit B3 alongside the existing W3C headers (via the environment variable OTEL_PROPAGATORS=tracecontext,baggage,b3multi, which OpenTelemetry’s auto-instrumentation reads on startup), and they separately reconfigured Envoy to use its Zipkin tracer, which does understand B3 and can join an existing trace, instead of its default OTel tracer, which cannot. Two independent fixes, one on the application side and one on the mesh side, because the mismatch existed on both ends. Their full walkthrough, including the exact Istio meshConfig YAML for the Envoy side of that fix, is worth reading if you run Istio or a similar mesh in production; this tutorial’s Python demo reproduces the application side of the same underlying failure so you can see the mechanism directly, without needing a Kubernetes cluster to do it.
Common Mistakes and Gotchas
- Fixing only one side. Step 6 demonstrated this directly: a propagation-aware client talking to a naive server produces the exact same fragmentation as a naive client talking to a naive server. Both ends of every hop need to agree.
- Assuming propagation happens automatically. Creating a
TracerProviderin two separate services does not link their spans by itself. Nothing links two services’ traces until something explicitly callsinject()on the way out andextract()on the way in, whether you write that call yourself, as this tutorial did, or an auto-instrumentation library does it for you (see Next Steps). - Assuming “we added tracing” means propagation format is settled. As Step 8 showed, two systems can both be propagating context in good faith and still fail to connect if one speaks W3C tracecontext and the other only understands B3 (or vice versa). When a proxy, gateway, service mesh, or third-party SDK sits between two of your own services, check what propagation format it expects, do not assume it matches your application’s default.
- Forgetting that
parent_span_idonly helps within one trace. The reconstruction script in Step 7 only links a child span under a parent if both share the sametrace_idand the parent’sspan_idis present among the spans it has seen. A parent span ID pointing at a trace ID the backend never received (for example, because a span exporter crashed or a batch was dropped) will simply show up as an orphaned root in most real backends, not an error, which can be confusing to debug without understanding this trace ID plus parent ID relationship. - Running multiple copies of a test service on the same port. While preparing this tutorial, an earlier terminal’s
inventory_service.pywas left running (a failed process-kill command silently did nothing) while a second, different version of the same service was started on the same port. The result looked like a propagation bug, mismatched trace IDs, but was actually two different server processes answering requests inconsistently. If your own results ever look inexplicably wrong, confirm with your OS’s process or port tools (for exampleGet-Process pythonon Windows, orlsof -i :5001on macOS or Linux) that only the service you think is running is actually running.
How to Verify Everything Works End to End
Run through this checklist against your own setup:
- Start
inventory_service_fixed.py, then runstorefront_fixed.py. Confirm the printedwidget count = 42matches whatcurl http://127.0.0.1:5001/stock/widgetreturns directly. - Open
spans.jsonland confirm you see exactly two new lines from that run, oneservice: "storefront-service"and oneservice: "inventory-service". - Confirm their
trace_idfields are identical strings, and that the inventory span’sparent_span_idexactly matches the storefront span’sspan_id. - Run
python reconstruct_trace.py spans.jsonland confirm it prints one trace containing both spans, withinventory-servicenested one level understorefront-service, not two separate trace blocks. - As a negative check, stop
inventory_service_fixed.py, start the originalinventory_service.pyfrom Step 3 instead, runstorefront_fixed.pyagain, and confirm you are back to two different trace IDs. If you cannot reproduce the broken case on demand, something about your test setup (likely a leftover process on the port, per the gotcha above) is not what you think it is.
Next Steps
Everything in this tutorial used manual propagate.inject() and propagate.extract() calls on purpose, so you would see exactly where trace context lives and how little machinery is actually involved. In a real codebase, you do not normally write those calls by hand for every HTTP client and server; you install auto-instrumentation packages instead. Installing opentelemetry-instrumentation-flask and opentelemetry-instrumentation-requests, then calling FlaskInstrumentor().instrument_app(app) and RequestsInstrumentor().instrument() once at startup, reproduces the same fix automatically for every route and every outgoing request, and additionally generates extra spans carrying real HTTP attributes like status code and target URL, all without touching your route handlers. This was verified while preparing this tutorial: the same two-service demo, rebuilt with auto-instrumentation and zero manual propagation code, produced one correctly linked three-span trace on the first request.
From here, worthwhile next steps include:
- Point a real exporter, such as the OTLP exporter, at a local OpenTelemetry Collector and a backend like Jaeger or Grafana Tempo, so you get a real visual waterfall UI instead of this tutorial’s plain-text tree.
- If you run services behind a proxy, API gateway, or service mesh (Istio, Linkerd, Envoy standalone, or similar), check its documentation for which propagation formats it reads and writes by default, and confirm that matches what your application emits, using the same kind of test built in Step 8.
- If you run Istio specifically, read the full CNCF write-up that inspired this tutorial. It walks through the exact
IstioOperatorandmeshConfigchanges needed to switch Envoy’s sidecar tracer and align it with an OpenTelemetry-instrumented application, using the open source OpenTelemetry Demo application as a live example. - Explore baggage, a second, separate W3C-standardized propagation mechanism for carrying arbitrary key/value business data (like a user ID or feature flag) alongside trace context across the same service boundaries.








No Comment! Be the first one.