
Nvidia’s data center CPU business has quietly scaled far beyond what most of the market expected. Ian Buck, the company’s vice president of hyperscale and high-performance computing and the inventor of CUDA, disclosed that Nvidia has shipped “hundreds of thousands” of Grace standalone servers, marking a decisive shift in how the company talks about its role inside AI data centers.
The figure puts actual deployment numbers behind a strategy that Nvidia has been building toward for years. In May, the company disclosed total Grace CPU shipments of over 2.5 million units, but those included Superchip configurations where Grace was paired with a GPU. Buck’s comments suggest that the standalone, CPU-only version is now a significant and growing share of that total.
The CPU pivot
The milestone arrives at a moment when the architecture of AI data centers is undergoing a structural change. Training large models still requires vast GPU clusters. But as AI applications move from training into production, the balance of compute shifts. Agentic AI workloads, where models call tools, run code, orchestrate pipelines and evaluate results, create CPU-intensive demands that GPU-only racks cannot economically serve.
“A few years ago you’d see eight GPUs per CPU for training. Now it’s moving toward one-to-one in some agentic deployments,” Buck told Tom’s Hardware.
Nvidia’s Grace CPU packs 72 Arm Neoverse V2 cores clocked at up to 3.35 GHz, with up to 480 gigabytes of LPDDR5X memory in single-chip configurations or 960 gigabytes in the dual-chip Superchip. The LPDDR5X memory delivers up to one terabyte per second of bandwidth while consuming roughly one-fifth the power of conventional DDR5, making it attractive for data-center operators facing energy constraints.
Vera on the horizon
The Grace numbers set the stage for Nvidia’s next-generation Vera CPU, which the company first unveiled at CES in January. Vera uses Nvidia’s first custom core design, called Olympus, with 88 cores and support for simultaneous multi-threading and confidential computing. The first Vera systems were hand-delivered to Anthropic, OpenAI, SpaceXAI and Oracle Cloud Infrastructure in May.
Nvidia sees CPUs in the data center as a $200 billion total addressable market opportunity, a notably rosier forecast than the industry consensus of roughly $120 billion (approximately £95 billion) by 2030. Morgan Stanley estimated in April that agentic AI alone could add as much as $60 billion (approximately £48 billion) to that figure.
Meta as key customer
Meta has been the most visible early adopter of standalone Grace servers. The social network deployed the CPUs in CPU-only systems for general-purpose and agentic AI workloads, and has already tested Vera on some workloads with “very promising” results, according to Buck.
Meta’s adoption runs counter to the broader hyperscaler trend toward custom Arm CPUs represented by Amazon Graviton and Google Axion. The company plans to deploy millions of Nvidia GB300 and Vera Rubin Superchips and has expanded its partnership with Nvidia to cover GPU, CPU and Spectrum-X networking components.
Sources: Nvidia has shipped ‘hundreds of thousands of Grace standalone servers’ (Tom’s Hardware, July 2026); Meta already deploying Nvidia’s standalone CPUs at scale (The Register, February 2026); Vera Arrives: NVIDIA’s First CPU Built for Agents (NVIDIA Blog, May 2026)

