NVIDIA Vera Redefines Server CPU Performance For The Agentic AI Era
Nvidia has expanded well beyond GPUs over the years, building an increasingly more diverse and comprehensive AI platform spanning networking, DPUs, high-speed interconnects, rack and pod-scale systems and software. With its upcoming Vera CPU, the company is adding another strategic product, a custom Arm-based server processor featuring the company’s own “Olympus” microarchitecture, designed specifically for emerging agentic AI workloads .
Nvidia just shared additional architectural details regarding its Vera CPU, and it paints an interesting picture. Rather than chasing higher core counts or maximum throughput, the company says Vera was designed specifically for the demands of agentic AI, where fast single-threaded execution, predictable memory latency and efficient data movement can have a impact impact on overall system performance.
Optimizing The CPU For Agentic AI
Nvidia points out that agentic AI differs from inference and traditional data center server workloads. Rather than processing a single prompt and returning an answer, AI agents continuously shift between GPU inference and CPU-driven work such as code execution, database lookups, tool invocation and orchestration, to name a few. Nvidia refers to this as the "agent loop," where CPU responsiveness directly affects how quickly an AI workflow completes.
That isn't a new challenge for hyperscalers or AI developers, but Nvidia believes its Olympus microarchitecture used in Vera is best designed for this kind of work.
Modern cloud CPUs have largely been designed with maximizing throughput and core density for virtualized cloud workloads in mind. Nvidia argues that AI factories instead require high per-core performance, in combination with low memory latency and efficient movement of data between CPU cores and attached GPUs and other accelerators.
A Different Architectural Approach
Unlike Nvidia’s Grace CPU, which utilized Arm Neoverse cores, Vera leverages Nvidia's own custom CPU microarchitecture, code-named Olympus.
Rather than using chiplets, like many competitive high core-count server processors, Vera uses a monolithic die with Nvidia's second-generation Scalable Coherent Fabric linking the various IP. Using a monolithic design eliminates the latency penalties associated with traversing multiple chiplets, while simultaneously delivering roughly three times greater core-to-core bandwidth.
Memory is another area where Vera differs from conventional server CPUs. Instead of DDR5, Nvidia uses data center-class LPDDR5X capable of delivering up to 1.2TB/sec of memory bandwidth. The company claims the design provides roughly three times more memory bandwidth per core, five times greater bandwidth-per-watt efficiency and approximately 40% lower memory latency under load than competing server platforms.
Nvidia also disclosed a wide 10-way decode front end for the Olympus core, designed to sustain high single-threaded performance under heavy workloads. Combined with the coherent fabric and LPDDR5X memory subsystem, Vera is designed to reduce latency across the CPU complex while keeping attached AI accelerators fed with data, to maximize utilization.
Early Performance Claims Are Promising
Nvidia shared benchmark data intended to demonstrate Vera’s advantages across AI-centric workloads as well.
The company says the Olympus core was designed to deliver roughly 2X higher performance than today’s x86 processors. The data Nvidia and its partners shared showed approximately 2X faster agentic sandbox startup and job completion compared to current x86 platforms. Nvidia also reported as much as 6X higher performance in selected streaming data processing workloads developed with Redpanda and HPE, along with up to 7X faster performance in Los Alamos National Laboratory scientific computing workloads compared with an Intel Sapphire Rapids-based supercomputer.
Those are impressive numbers, but they are vendor-supplied benchmarks that highlight Vera's architectural strengths. Independent testing across a broader range of workloads will ultimately determine how the processor compares with current and next-generation offerings from AMD and Intel.
That said, the broader takeaway isn't whether every workload achieves a 2X or 6X improvement in my opinion. Nvidia is optimizing for end-to-end AI workflow performance by reducing the amount of time accelerators spend waiting on CPU-side execution, an increasingly important metric as agentic AI workloads become more widespread.
Another Step Toward A Fully Integrated AI Platform
While AMD and Intel continue building high-performance server CPUs intended to address a broad mix of enterprise, cloud and HPC workloads, Nvidia is taken a somewhat different approach. Nvidia is optimizing an increasingly integrated AI platform where CPUs, GPUs, networking, DPUs and system software are engineered and co-optimized to work together.
The company has also lined up notable ecosystem support, including OpenAI, Thinking Machines Lab, Perplexity, Los Alamos National Laboratory and multiple OEMs and cloud providers planning to deploy Vera-based systems. While those relationships don't guarantee widespread adoption, they suggest Nvidia has generated significant industry interest ahead of commercial availability.
Whether Vera establishes a new class of AI-focused server processor or is simply better tuned for certain latency-sensitive AI workloads remains to be seen. It is obvious, however, that Nvidia sees the CPU playing a much larger role in future AI systems moving forward. If agentic AI evolves as many expect, Vera could become another important building block in Nvidia's expanding AI infrastructure portfolio. More broadly, it underscores that the next phase of competition among Nvidia, AMD and Intel is likely to be defined less by individual chips and more by the capabilities of complete AI computing platforms.
Loading article...