Nvidia Details Vera CPU’s Custom Olympus Cores in a Direct Challenge to Intel and AMD
Nvidia has detailed the custom Olympus cores inside its standalone Vera CPU, a direct challenge to Intel and AMD in the datacenter, with Meta, Oracle, and Alibaba already signed on.
Nvidia has published new architectural detail on Vera, the custom Arm processor it is selling as a standalone datacenter CPU, independent of any GPU purchase. The disclosure, covered in depth by The Register on August 1, centers on Olympus, the custom core design Nvidia built specifically for Vera rather than licensing an off the shelf Arm core as it did for its earlier Grace CPU.
Table Of Content
Vera is the CPU half of the Vera Rubin platform generation Nvidia has been rolling out through 2026, but it can now be bought and deployed on its own, without an accompanying Rubin GPU order. That distinction is central to why the chip is being read as Nvidia’s most direct move yet into a datacenter CPU market long split between Intel and AMD. According to Nvidia’s own launch announcement, Alibaba Cloud, ByteDance, Meta, Oracle Cloud Infrastructure, CoreWeave, Lambda, Nebius, and Nscale have all signed on to deploy Vera, with Cloudflare, Crusoe, Together.AI, and Vultr named as planning to follow. “The CPU is no longer simply supporting the model; it’s driving it,” Nvidia founder and CEO Jensen Huang said in the announcement.
A Monolithic Core Surrounded by Chiplets
Unlike the multi-die chiplet designs used by x86 rivals, such as AMD’s Epyc processors, Nvidia packed all 88 Olympus cores onto a single monolithic compute die that The Register reports is fabricated on TSMC’s 3nm process. Surrounding that die are separate chiplets for memory and I/O: eight LPDDR5X memory controllers and what The Register describes as two dedicated I/O dies, one handling PCIe 6.4 and CXL 3.1 and the other dedicated to Nvidia’s NVLink Chip-to-Chip interface. Nvidia’s own engineering blog says the split is meant to deliver “strong single-thread performance, even when the socket is fully loaded.”
Each Vera CPU carries 176 threads (88 cores times two, via what Nvidia calls Spatial Multithreading), 164MB of shared L3 cache, and up to 1.5TB of SOCAMM2 LPDDR5X memory delivering 1.2TB/s of bandwidth per chip, up from 512GB/s of bandwidth and 480GB of capacity on Grace, according to specifications ServeTheHome gathered from Nvidia’s whitepaper. Configurable TDP ranges from 250W to 450W, per the same source.
In its dual-socket configuration, which Nvidia calls the Vera CPU Superchip and unveiled at GTC in March, two Vera chips connect over a 1.8TB/s NVLink-C2C link for a combined 176 cores, 352 threads, and 2.4TB/s of aggregate memory bandwidth, roughly twice the memory bandwidth of AMD’s 2024-era Turin Epyc processors, The Register reported. Nvidia’s agentic AI reference rack designs call for packing as many as 128 of these Superchips (256 CPUs total) into a single liquid-cooled rack, for a combined 22,528 cores and 384TB of memory.
Built for Agents, Not Just for Babysitting GPUs
Nvidia is positioning Vera for two distinct jobs: as the head node that manages GPUs inside its Vera Rubin rack systems, and as a host for AI agents themselves, which, unlike the large language models that power them, do not run directly on GPUs, The Register noted. Nvidia’s developer blog says agent workloads specifically need “sufficient memory bandwidth per core” and “predictable latency under concurrency,” the design goals it says shaped Olympus.
Inside each Olympus core, Nvidia built a 10-wide decoder capable of pulling up to 16 instructions into a 48-entry decode queue per cycle, backed by a 64KB instruction cache and a 96KB data cache, according to ServeTheHome’s review of the whitepaper. The Register’s own breakdown of the core’s block diagram lists eight integer ALUs, six vector and floating point pipelines using Arm’s SVE 128 extension, four load units, and two store units, a wider front and back end than the Zen 5 cores in AMD’s Turin Epyc chips or the Redwood Cove cores in Intel’s Granite Rapids Xeons. Nvidia also built a custom neural branch predictor into Olympus that can explore two branches simultaneously, predicting two code paths per clock cycle, per The Register’s analysis of the design.
Spatial Multithreading, Not Hyperthreading
For multithreading, Nvidia skipped the conventional simultaneous multithreading that Intel calls Hyperthreading. The Register describes Olympus’s approach, which Nvidia markets as Spatial Multithreading, as closer to splitting one core into two smaller ones that share a cache line: operators can run one core as a single fast thread or as two independent, lower-throughput threads, a tradeoff The Register compares to Nvidia’s own Multi-Instance GPU partitioning on its data center GPUs.
How Custom Is “Custom”
The Register calls Olympus Nvidia’s first fully custom CPU core, but its own analysis pushes back on how custom the design really is, noting that Olympus’s block diagram closely resembles a heavily modified version of Arm’s off the shelf Cortex X925 core, and that the design still leans on existing Arm intellectual property. Nvidia’s Grace CPU, by comparison, used unmodified Arm Neoverse V2, Neoverse V3, or Cortex X925 and A725 cores depending on the product variant, rather than a ground-up custom design.
Nvidia is backing Vera with a specific technical performance claim: up to 1.8 times the performance of x86 CPUs on agentic workloads, based on internal SPEC CPU 2026 testing the company ran in July, according to its developer blog. That figure comes from Nvidia’s own internal testing rather than an independent lab. As Vera reaches the hyperscalers that have signed on, independently run SPEC CPU 2026 scores, not Nvidia’s own whitepaper, will be the real test of how much of the standalone server CPU market the chip can take from Intel and AMD.








No Comment! Be the first one.