Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
A Kubernetes blog says node swap packs up to three times more agent sandboxes on a node, but the raw benchmark data shows a latency bill, a baseline that differs by more than swap, and no test of...
On October 5, 2026 the Kubernetes blog published “Scaling Kubernetes Workloads with Node Swap” by Ocean Xie and Yuan Wang. The pitch is tidy. Agent sandboxes need a lot of memory to start and to run untrusted code, then sit idle waiting for the next prompt, and “That idle but resident memory is expensive”. Node swap, which reached general availability in Kubernetes v1.34 (KEP-2400 lists the stable milestone as v1.34), lets the kernel page that memory out to a fast local SSD. Across three workloads the authors report “density gains of up to 3×, often with little or no latency cost”.
Table Of Content
- What the post claims and what the repository measured
- Where the “3×” comes from
- What “healthy” means
- Requests set far below real use
- The latency bill
- A baseline that differs by more than swap
- What the short how-to leaves out
- The question nobody timed: waking a swapped-out sandbox
- Before you rely on it
- The bottom line
- Sources and method
I checked the pitch against the data the post links. The Kubernetes SIGs agent-sandbox repository keeps the benchmark READMEs and cluster scripts under examples/gke-swap and the test harnesses under test/e2e/extensions, and I read the READMEs, the deploy script, the swap configuration and both density harnesses as they stood on October 5. They support a narrower claim than the headline: swap lets a node ride out an overcommitted boot storm of agent sandboxes that crashes or degrades a node without it. They do not test the case the post uses to motivate swap, which is an idle sandbox that receives its next prompt. Four details change how the headline reads.
- The “3×” is the ratio of two sweep points, and the sweep ends at 240.
- On the isolating runtimes, gVisor and Kata, the extra density comes with a large latency bill.
- The “no swap” node is a different machine type from the swap node, so the comparison is not swap alone.
- The post’s how-to is a short kubelet configuration. The repository credits node tuning, kernel settings and specific request sizes that the post does not mention.
What the post claims and what the repository measured
The post’s table has four rows. The repository carries the numbers behind three of them, plus several the post leaves out. Latencies below are the repository’s own: start-to-ready times for the Chrome sweeps and execution time for the Python sweeps, each compared with the no-swap node at the highest density where it still passed. The multiples are my arithmetic.
| Workload | Post’s headline | No-swap node, highest density that passed | Swap node | Latency change (average, P99) |
|---|---|---|---|---|
Headless Chrome, Kata (kata-clh) |
40 to 50 pods, +25% | 40 pods: 34.54 s average, 60.99 s P99 | 50 pods: 71.02 s, 174.14 s | 2.1×, 2.9× |
| Headless Chrome, gVisor | 80 to 160 pods, +100% | 80 pods: 17.59 s, 41.07 s | 160 pods: 59.02 s, 214.95 s | 3.4×, 5.2× |
| Python sandbox, gVisor | 80 to 240 pods, +200% | 80 pods: 2.04 s, 4.97 s | 240 pods: 3.43 s, 12.83 s | 1.7×, 2.6× |
| Python sandbox, runc | Not in the post’s table | 80 pods: 1.62 s, 1.86 s | 240 pods: 1.39 s, 1.63 s | 0.86×, 0.88× |
| Headless Chrome, runc, 8 vCPU | Not in the post’s table; the README says +66% | 120 pods: P99 61 s | 200 pods: P99 268 s | P99 4.4× |
| Headless Chrome, runc, 32 vCPU | 512 to 768 pods, +50% (in the post’s text) | 512 pods: 97.53 s, 791.60 s | 768 pods: 6.00 s, 37.69 s | Lower, but see below |
The kernel-build row, a 600 MB memory limit cut to 300 MB with a 374 second build against 433 seconds, is the one I could not trace: the post links raw data for the Chrome and Python sweeps and I found no link for the build. The post also reports that squeezing the limit to 200 MB made the build take over 40% longer, and draws the lesson that swap is “an insurance policy for burst memory, not a replacement for active RAM”.
The last row is the exception, and it is the one the post’s text cites for runc. On a 32 vCPU, 120 GB node the no-swap pool had a P99 of 791.60 seconds at 512 pods, so its capacity of 512 includes a tail of about 13 minutes, while the swap pool at 768 pods sat at 37.69 seconds. The README says that result needed node tuning, and it credits part of it to disks rather than swap (both points are covered below).
Where the “3×” comes from
In the Python sweep the no-swap pool passed at 60 and 80 sandboxes and failed at 100, with the node going NotReady. The swap pool passed every point it was given, 60 through 240 in steps of 20, with a 100% pass rate. Two things follow. The baseline’s true limit lies somewhere from 80 to 99, and the swap pool’s limit was never found, because 240 is the last row of the sweep. The deploy script’s default of 256 maximum pods per node also sits just above that last row. The defensible statement is at least about 2.4×, with the upper end not measured, where 2.4 is 240 divided by 99.
What “healthy” means
The pass rule has no latency objective. In the Chrome harness a sandbox counts once its browser answers a DevTools version request, and the harness gives it three minutes after the pod is Ready. The Python harness waits up to eight minutes for the sandbox’s server and the execution call. At 200 Chrome sandboxes on the 8 vCPU node the swap pool passed 200 of 200 with a P99 start-to-ready time of 268 seconds, against 58 seconds at 120 sandboxes. “Density” in these tables means how many sandboxes started within the timeout, so the figure moves with whatever objective you would set for a real agent.
Requests set far below real use
The Python sandboxes declare a 100Mi memory request and a 15m CPU request with a 2Gi limit, while each one holds a DataFrame of about 375 MB (gVisor adds roughly 65 MB for its user-space kernel, per the README). That is overcommit by design, and the Python README says the small CPU request exists so the scheduler can “pack hundreds of sandboxes onto an 8-core node”. Without swap, the same README describes passing the memory ceiling this way: “Surpassing physical RAM capacity causes Kubelet to freeze and fail heartbeats”, a node crash rather than orderly OOM kills. In the Python sweep the swap node slowed down but kept answering all the way to 240, which is a real improvement over a crashed node. That is resilience, which is not the same thing as capacity at an acceptable latency, and the Kubernetes swap documentation warns that the scheduler “currently does not account for swap memory usage”.
The latency bill
The post says that at the maximum densities “the per-pod latency increase is driven mainly by pods competing for CPU, not by swap I/O”. The repository’s gVisor Chrome sweep partly supports that. CPU pressure is the highest of the three readings, at an average CPU PSI of 89.24% with 160 sandboxes on the swap node. But memory and I/O pressure are not small: 63.27% and 62.25%, against 0.01% and 31.65% for the no-swap node at 80. (PSI is the kernel’s pressure stall information, the share of time tasks were stalled waiting on a resource; the README prints a pair of figures for memory and I/O and I quote the first.) The README’s own failure analysis points at swap I/O too. In the runtimes README, “Scaling beyond 160 pods (to 180+) shifts the bottleneck from memory capacity to NVMe write queue depth saturation”, and at 180 sandboxes the node fell to an 85.56% pass rate.
Two runtime details are missing from the post. It says Kata microVMs went from 40 to 50 pods without naming the hypervisor. The runtimes README shows that figure is kata-clh (Cloud Hypervisor), while kata-qemu got no gain from swap at all: “Swap does not improve density for QEMU microVMs”, because when the host pages out VM memory, “guest page faults force uninterruptible disk reads (D-state)” and the main QEMU event loop stalls. If your Kata deployment uses QEMU, the Kata row is not yours.
The one place the “little or no latency cost” line holds cleanly is runc running the Python workload. Python sandboxes on the swap node ran at 1.39 seconds on average with 240 sandboxes, against 1.62 seconds on the no-swap node with 80.
A baseline that differs by more than swap
The post says swap lets you raise density “on the same infrastructure”, and the main README says “on the exact same hardware”. The deploy script creates the two pools from different machine types: c4-standard-8 for the baseline and c4-standard-8-lssd for the swap pool, which carries the local SSD that holds the swap. The 32 vCPU runs use the c4-standard-32 and c4-standard-32-lssd pair.
The data shows the gap is not only swap. At 60 sandboxes the swap node had used 0.00 GB of swap in both the runc Python sweep and the gVisor Chrome sweep, and it was still faster. Python execution averaged 1.36 seconds against 1.58 (14% lower), with about half the CPU stall time (6.12 seconds against 12.17). gVisor Chrome reached ready in 9.84 seconds on average against 11.66 (16% lower), with a P99 of 22.10 against 26.10. The main README gives a reason that has nothing to do with swapping: “Even when swap usage is 0B”, memory pressure can evict Chrome’s binaries from the page cache, and “On baseline PD pools, reading evicted binaries back from network disks saturates the 3,000 IOPS queue and causes watchdog timeouts”. It also reports how noisy single runs were: “In some runs, the baseline pool was able to survive up to 170 pods”. The headline +66% takes 120 pods as the baseline’s limit because, in the README’s words, “~120 pods is the limit for reliable deployments”, a qualification the post’s table does not carry.
What the data cannot say is which part of the gap is the swap, which is the local SSD under the page cache, and which is run-to-run noise. A third pool, the -lssd machine type with swap switched off, would separate them. None of the READMEs reports that arm.
What the short how-to leaves out
The post’s “How to use it” section is a kubelet configuration plus one instruction, to set “memory limits higher than its requests” so pods are Burstable. This is the snippet, which applies to Kubernetes v1.34 or later, where swap support is stable:
kind: KubeletConfiguration
apiVersion: kubelet.config.k8s.io/v1beta1
failSwapOn: false
memorySwap:
swapBehavior: LimitedSwap
The repository says the numbers came from a longer list:
| Setting | In the post | In the repository |
|---|---|---|
| Swap storage | “your cloud provider’s high-speed local disk” | A dedicated local SSD (dedicatedLocalSsdProfile with one disk) on the -lssd machine type |
| Kernel settings | Not mentioned | vm.swappiness: 100 and vm.watermark_scale_factor: 500, described as “just for reference”, with “tuned and more performant settings” expected in follow-up pull requests |
| Node tuning | Not mentioned | Kubelet and system daemons pinned to cores 0 and 1 (cores 0 to 7 for the 768 pod runs), static CPU manager policy, larger ARP cache thresholds, journald rate limits. The README says “CPU isolation is mandatory” |
| Requests | Limits above requests | 100Mi memory and 15m CPU requests with a 2Gi limit (Python); a 150 MiB request with a 2 GiB limit (Chrome, 8 vCPU) |
| Creation pace | Not mentioned | 1.8 seconds between sandboxes in the Python runs, credited with “eliminating Page Cache thrashing storms on boot”; 1000 ms in the 32 vCPU Chrome runs |
| Scheduling | Not mentioned | The Kubernetes docs say there is no swap-aware scheduling and advise tainting swap nodes |
| Security | “without compromising security boundaries” | Not tested; the Kubernetes docs make encrypted swap the administrator’s job |
On GKE 1.34.1-gke.1341000 or later, the minimum version the README states, the benchmark applies the repository’s swap-dedicated-lssd.yaml:
linuxConfig:
swapConfig:
enabled: true
dedicatedLocalSsdProfile:
diskCount: 1
sysctl:
vm.watermark_scale_factor: "500"
vm.swappiness: "100"
Nothing here is concealed: the post links the repository, and the repository documents these settings. But a reader who follows the how-to alone, on a node without a dedicated local SSD, with requests sized to real usage and no CPU reservation, is running a different experiment. The Kubernetes project’s own deep dive on tuning Linux swap explains why those kernel settings change when swapping starts.
The question nobody timed: waking a swapped-out sandbox
The scenario the post opens with is an agent that finishes a burst of work and then will “sit idle waiting for the next prompt”. The cost of that scenario is paid when the prompt arrives and the sandbox has to read its memory back from disk. Neither harness measures it.
The Python workload loads five million MovieLens rows, aggregates them, answers the single execution call, and then forks a child that sleeps for 600 seconds, which is how the Python README says it “retains the ~375 MB Pandas DataFrame resident in process memory for 10 minutes”. The workload never touches that memory again. The Go harness records the time from creating a sandbox to that one call returning, and the Chrome harness stops its clock when the browser’s DevTools endpoint first answers. No second request is sent to a sandbox after it goes quiet.
At 240 gVisor sandboxes the swap node held 48.78 GB in swap next to a 23.15 GB peak of node RAM. Treating those as simultaneous, about two thirds of the sandboxes’ memory was on disk, roughly 0.2 GB per sandbox (my arithmetic). Reading that back is the operation the Kubernetes documentation singles out: “swapping data back to memory is a heavy operation, sometimes slower by many orders of magnitude”. The Python README adds that the paging covers more than the Python heap, because “Linux kswapd pages out gVisor’s Sentry kernel memory out to disk alongside Python memory”.
The post calls swap an insurance policy, and the GKE page calls it a safety net. It says “Node memory swap is intended as a safety net for unpredictable memory spikes, not a replacement for sufficient physical memory”. The density result uses it as a capacity layer, with 48.78 GB of swap behind a node that has 27.0 GB of allocatable RAM. That is a legitimate design, but it needs a test the safety-net framing never required: how long does a sandbox take to answer after its memory has been parked?
That test is cheap to describe. Let the sandboxes sit idle until the kubelet’s container_swap_usage_bytes plateaus, send each one a second request that touches its whole working set, and record time to first byte. Repeat with 10%, 25% and 50% of the sandboxes waking together, and report percentiles against the same request on a warm sandbox. I did not run it: I have no cluster, so every figure in this piece comes from the published logs, the harness source and the documentation.
Before you rely on it
- Taint the swap nodes. The Kubernetes docs say the scheduler does not account for swap and that “Administrators can taint nodes with swap available to protect against this problem”. GKE’s page recommends a taint such as
gke-swap=enabled:NoSchedule. - Treat encryption as a decision. The Kubernetes docs say “It is the administrator’s responsibility to provision encrypted swap to mitigate this risk”. GKE “encrypts swap space by default using an ephemeral key”, and the benchmark’s swap device is named
/dev/mapper/encswap, which is consistent with that. The post’s line about “without compromising security boundaries” is not something the benchmarks test, and the mechanism pages out “anonymous memory”, which includes a sandbox’s heap. Secret volumes and memory-backedemptyDirmounts are kept out of swap by anoswaptmpfs option, which needs kernel 6.3 or a distribution backport. - Check which input sets the swap budget. Upstream Kubernetes computes a container’s swap limit as
(containerMemoryRequest / nodeTotalMemory) × totalPodsSwapAvailable, and the repository repeats that formula. On a 32 GiB node with 8 GiB of swap available to pods, a 100Mi request gets a swap limit of 25Mi however idle the container is (my arithmetic). The post says instead that the node “automatically rations fast swap space based on idle application memory usage”, which is not how the documented rule reads. GKE’s page says GKE “calculates container swap limits based on the container’s memory resource limits and the node’s total memory”. The three descriptions differ, so test which input moves the limit on your platform before you size requests. - Know who gets no swap. Only Burstable pods qualify. Guaranteed, BestEffort and high-priority pods do not, and a Burstable container opts out by setting its request equal to its limit. The Kubernetes project also recommends running control plane nodes without swap.
- Protect the node itself. The docs advise setting
memory.swap.max=0on the system slice and giving it I/O latency priority. On untuned nodes, the gVisor README says, syscall emulation threads starve the kubelet of CPU, it misses heartbeats, and GKE flags the node NotReady. - Watch pressure, not just swap bytes. The kubelet exposes
node_swap_usage_bytes,container_swap_usage_bytesandcontainer_swap_limit_bytes, andkubectl top nodes --show-swapprints swap use per node. The repository’s tables show why memory and I/O stall time belong next to those numbers. - Pick the runtime with the data in hand. gVisor and Kata Cloud Hypervisor gained density; Kata QEMU did not. For background on those isolation choices see our earlier pieces on kagent’s Agent-Substrate and Red Hat’s OpenShift Sandboxed Containers.
- Mind in-place resize. GKE lists, among its swap requirements, that to resize container memory you must set the container resize policy to
RestartContainer. That matters if you use the in-place resize feature we covered when v1.37 added scheduler preemption for it.
The bottom line
Node swap is a real feature, stable since v1.34, and the repository shows it can keep a node answering through a boot storm that crashes the no-swap pool. That is worth having on an agent platform. The headline figures are best read as how many sandboxes started within the harness timeout in a sweep, on a node with a faster disk, tuned for the test, at whatever latency resulted. Before treating 3× as a planning number, run the experiment the benchmark skipped, with real wake-ups and a latency objective, on your own runtime.
Sources and method
I read the Kubernetes blog post; the KEP-2400 metadata (stage stable, milestone v1.34); the agent-sandbox examples/gke-swap READMEs for the main example, the runtimes comparison, gVisor and Python density, plus the deploy script, the swap configuration and the node tuner; the source of the Python and Chrome density harnesses and the Python workload script; the Kubernetes swap memory management page (last modified September 1, 2026); GKE’s node memory swap page; and the Kubernetes project’s August 2025 post on tuning Linux swap. Repository content is as of October 5, 2026 and the branch can change. Multiples, percentages and the share of memory on disk are my arithmetic on the published figures, and the 2.4× lower bound divides 240 by 99. I did not run the benchmarks or inspect the raw per-run files they write; the tables are the repository’s own summaries.








No Comment! Be the first one.