DigitalOcean’s Data Locality Tax Turns RAG Latency Into a Geography Problem
A DigitalOcean benchmark shows that placing a vector database one region away from its inference host adds far more latency than any index tuning can recover, and the cost multiplies with every...
Most RAG performance tuning fixates on the index: HNSW parameters, ef_search values, embedding dimensionality, which GPU to rent. A benchmark DigitalOcean published on August 14, 2026 argues that effort is aimed at the wrong variable. In most retrieval-augmented generation pipelines, the larger and more invisible cost is not the index at all. It is where the vector database physically sits relative to the machine doing the generating.
Table Of Content
The study, written by DigitalOcean Senior Technical Content Strategist and Team Lead Anish Singh Walia, measured retrieval latency across three placements of the same query, index, and corpus, then open-sourced the full test harness on GitHub under an MIT license so readers could reproduce the numbers on their own infrastructure rather than take the results on faith.
The Experiment
The article’s introduction describes the test broadly as running against “GPU Droplets and Vector Databases.” The actual methodology is more specific, and worth being precise about: the client was a Droplet in DigitalOcean’s NYC3 datacenter, querying Managed PostgreSQL 16 with the pgvector extension (the db-s-1vcpu-1gb plan) in both NYC3 and San Francisco (SFO3). The published GitHub repository lists the client instance as an s-2vcpu-4gb Droplet, a standard compute size, not a GPU instance. That distinction does not change the results, since the test measures network transit and database query time rather than model inference, but it is worth separating the recommended production setup (a GPU host running the embedding model and LLM) from what this specific benchmark run actually used to fire queries.
Three arms were tested:
- Arm A: the Droplet and the database in the same datacenter (NYC3), connected over the private VPC network.
- Arm B1: the Droplet in NYC3, the database in SFO3, connected over the public internet.
- Arm B2: the same NYC3-to-SFO3 path, but routed over a private, peered VPC connection instead of the public internet.
Every arm queried an identical corpus: 100,000 synthetic 768-dimensional vectors, indexed with HNSW (pgvector’s graph-based approximate nearest neighbor index) using cosine similarity. Each configuration ran 75 measured trials after 10 warmup queries, at both k=5 and k=100 nearest neighbors, plus a separate cold-first-call measurement that includes connection and TLS setup. A bare TCP probe to port 25060, DigitalOcean’s standard managed-database port, was measured with no query attached at all, to isolate network transit from database work.
The Physics Floor
Before measuring anything, the study derives a baseline that no amount of engineering can beat: the time it takes light to cross optical fiber. Fiber does not carry light at the vacuum speed of light. The glass core slows it to roughly 200,000 kilometers per second, close to two-thirds of light speed in a vacuum. Divide great-circle distance by that speed and the result is a hard floor for any round trip, regardless of how fast the database, the index, or the network hardware is.
For New York to San Francisco, that floor is 41.3 milliseconds round trip. New York to Frankfurt is 62.0 milliseconds. New York to Singapore is 153.3 milliseconds. Real routes run slower than this, since fiber follows cable routes and exchange points rather than great circles, but nothing recovers time below these numbers. They are lower bounds set by geometry, not benchmarking artifacts, which is why the study presents them as arithmetic a reader can check independently rather than a claim to take on trust.
What the Measurements Show
The measured results, from the August 10, 2026 run:
| Arm | Path | TCP probe (p50) | Retrieval p50, k=5 | Retrieval p95, k=5 | Cold first call |
|---|---|---|---|---|---|
| A | Same datacenter, private VPC (NYC3) | 2.43 ms | 1.90 ms | 3.51 ms | 111.34 ms |
| B1 | NYC3 to SFO3, public internet | 68.63 ms | 66.97 ms | 69.55 ms | 717.91 ms |
| B2 | NYC3 to SFO3, VPC peered | 67.73 ms | 69.90 ms | 70.57 ms | 668.87 ms |
Same-datacenter retrieval returned in 1.90 milliseconds at the median, for five nearest neighbors. Moving the identical query to a database one continent away raised that to 66.97 milliseconds, roughly a 35-fold increase. A bare TCP connection to the San Francisco database, with no query attached, already took 68.63 milliseconds on its own, meaning almost none of the cross-region time is spent inside the database doing search work. Once the two machines are far apart, the index stops being the bottleneck. The network is.
Peering Does Not Cancel Geography
The study directly tests an assumption teams often make: that routing traffic over a private, peered VPC connection instead of the public internet solves the problem. It barely helps. Arm B2, the peered path, measured 69.90 milliseconds at the median, within a few milliseconds of Arm B1’s public-internet 66.97 milliseconds. Peering keeps traffic off the public internet and changes how bandwidth is billed. It does not move San Francisco closer to New York. The distance is still the distance.
Cold Connections Make It Worse
The cold-first-call numbers show a second, compounding cost. On the same-datacenter arm, the first query, including connection setup, took 111.34 milliseconds. On the cross-region public arm, the same first call took 717.91 milliseconds. TCP setup costs one round trip, and TLS adds another round trip on TLS 1.3, or two more on TLS 1.2. On a 62-millisecond cross-region path, a client that opens a fresh connection per request instead of pooling them pays 124 to 186 milliseconds of pure handshake overhead, and pays it again on every new connection.
Why the Tax Compounds in Agentic RAG
For a single-shot chatbot that retrieves once per user question, a 62-millisecond round trip against a 2-second generation is roughly 3 percent overhead: real, but not something an engineering team would prioritize fixing. The math changes once retrieval is not a single call per task.
Agentic RAG patterns retrieve more than once: retrieve, reason about the result, retrieve again, rerank, fetch neighboring context, verify. The study treats five to ten sequential retrievals per task as an ordinary agentic pattern, not an edge case. Because those calls run sequentially rather than in parallel, the geography tax sums instead of averaging out. Using the measured NYC3-to-SFO3 figure, eight sequential hops add roughly 536 milliseconds of pure network transit before a single output token generates. Using nothing but the theoretical physics floor for New York to Frankfurt, eight hops still add about 496 milliseconds. Either way, a delay that was negligible in a single-shot chat becomes half a second of dead time in an agent loop, and users feel every millisecond of it, because retrieval blocks prefill: the model cannot start processing the prompt, let alone generate a response, until the retrieved context has arrived.
The Practical Takeaway
The study splits retrieval latency into two categories that respond to entirely different fixes. Server-side approximate nearest-neighbor search responds to index parameters, hardware, and corpus size, the knobs most teams already spend their time tuning. Network transit responds to exactly one thing: moving the two endpoints closer together. DigitalOcean’s own framing inverts the usual optimization order: co-locate data with compute first, then tune indexes second. Index tuning recovers milliseconds from the one part of the pipeline placement can never touch. Placement recovers tens to hundreds of milliseconds per call from the part index tuning can never touch.
That framing suggests a short checklist for teams building RAG or agentic retrieval pipelines: confirm the vector database and the inference host share a region before touching HNSW parameters; use connection pooling so TLS handshake costs are paid once rather than on every call; and if an agent workflow issues several sequential retrieval calls per task, treat the placement tax as multiplying by the hop count, not as a fixed constant that can be ignored. The full harness, including the SQL used to load the test corpus and the script that reproduces the benchmark, is published in the data-locality-tax repository under an MIT license, so any team can measure its own path instead of relying on someone else’s numbers.








No Comment! Be the first one.