Nvidia has unveiled more details about its Vera CPU, announced earlier this year and now in volume production. Vera appears differentiated in the crowded CPU market, with Nvidia claiming more efficiency and better performance especially for agentic AI workloads. The CPU is also gaining early traction in other areas besides AI, due to its memory bandwidth, power efficiency and single-core performance. Let’s take a deeper look. ( Disclosure: Nvidia is a client of Cambrian-AI Research, the author's firm. )

Nvidia Clams Vera CPU Delivers 2X Faster Core Performance

CPUs are back in fashion of late, thanks to the rise of agentic AI which needs fast and efficient CPU cores and memory. In agentic workloads, CPUs plan and execute a series of sequential steps to complete a task in a loop, calling on various AI models as needed. In executing a sequential workflow, multiple cores don’t help as much as they do for traditional cloud workloads. What matters are fast cores for agentic speed and latency, and efficient memory that can keep up with the cores.

The graph below shows the execution of a 10-step agentic loop, and you can see how the expensive GPUs are waiting for the agent (grey bars) to finish each step. Faster CPUs mean faster results, and less latency improves overall data center efficiency. And that is where the rubber meets the road as inference becomes a profit center for clouds and enterprises.

Anticipating the impact that slower CPUs have in the agentic work, Nvidia decided to re-engineer its arm-based CPU (Grace) to have fast single-threaded core performance which communicate efficiently without crossing chiplet boundaries which incur a performance hit. The SPEC CPU benchmark results are impressive but are only estimates at this time. More benchmarks are needed to fully evaluate this platform, and Nvidia promises they will be forthcoming. Much of the advantage that Vera brings to market is the memory fabric and the LPDDR5 memory, which was originally designed for mobile phones.

Nvidia engineers also brought out a new memory fabric on chip that lowered latencies by as much as 5-fold according to Nvidia. Once again, I personally have never seen such claims of improvement over decades of product management and analyst jobs.

The Vera CPU Business Case

Bank of America predicts that Nvidia could ship four to five million Vera CPUs in the 2nd half of 2026, compared with a total of only 2.5 million Grace CPUs shipped to date. BofA also estimates Vera could amount to $20B in revenue in 2026, and perhaps as much as half of that could be stand-alone CPUs. A fully configured rack with 256 Vera chips could cost “around $10 million,” implying an average effective price in the $30–40k range once you include memory and system costs. Nvidia estimates the Total Addressable Market (TAM) to be a $200B opportunity.

Vera CPU Application Examples Include HPC, not just AI Agents

Nvidia went on to share a few customer stories around Vera. Perplexity, in whom Nvidia has invested, is seeing 1.5-1.9X performance for Vera compared to x86 CPUs in agentic AI. After over 40 years working on CPUs, I do not recall seeing such dramatic improvements. A 50% lead in job completion time is probably enough to switch some data centers to consider purchasing a Vera CPU over an incumbent.

But there are some customer examples here beyond agentic AI workloads which could indicate Nvidia might carve more deeply into the Intel and AMD x86 CPU business. First off, the NYSE is seeing a six-fold improvement in P99 latency, demonstrating the consistency that can be achieved with the Vera CPU monolithic architecture (no chiplet boundaries between cores). “P99 latency” refers to the 99th percentile of response times in a system. In practical terms, it answers the question: How slow are the slowest 1% of requests? This consistent latency could also help large cloud service providers minimize the “jitter” they try so hard to avoid.

And in scientific simulation, the faster per-core performance and memory bandwidth are allowing Los Alamos National Labs to realize three to seven times better performance in their Veritas supercomputer, which uses Vera CPUs and Rubin GPUs.

Seizing on the opportunity Vera could present, Nvidia has built an ecosystem of customers, Supercomputing Centers, cloud providers, and system OEMs that could enable a significant upside to Nvidia’s existing GPU-centric data center business.

Vera CPU’s Strategic Impact on Nvidia

Vera finally allows Nvidia to claim a significant technology lead in every segment of the data center compute business, with top-shelf differentiated platforms including rack-scale systems, CPUs, GPUs, LPUs and networking of all scales (up, out, and across). Consequently, the $200B server CPU TAM is now open for Nvidia to address.

Disclosures: This article expresses the opinions of the author and is not to be taken as advice to purchase from or invest in the companies mentioned. My firm, Cambrian-AI Research, is fortunate to have many semiconductor firms as our clients, including Baya Systems BrainChip, Cadence, Cerebras Systems, D-Matrix, Flex, Groq, IBM, Infleqtion, Intel, Micron, NVIDIA, Qualcomm, SImA.ai, Synopsys, Taalas, Tenstorrent, Ventana Microsystems, and scores of investors. For more information, please visit our website at https://cambrian-AI.com .