NG Solution Team
Artificial Intelligence

Is NVIDIA’s Vera Rubin now in mass production with 300+ partners?

NVIDIA has announced the production ramp of its Vera Rubin NVL72 platform, saying the supply chain now spans more than 350 factories in 30 countries and that over 300 partners have already deployed racks tied to the platform. The announcement — made the day before a major rival AI event — emphasizes production scale, customer deployments and performance data aimed at agent-style AI workloads.

Why NVIDIA is betting on CPUs as AI evolves
NVIDIA argues that the rise of AI agents — systems that plan, invoke tools, execute code and orchestrate multi-step tasks — is changing the nature of server workloads. In this environment, CPUs are re-emerging as a critical bottleneck: they must handle many low-latency calls, coordinate task orchestration and quickly hand results back to GPUs. NVIDIA says this shift could reshape the server CPU market and estimates a potential addressable opportunity as large as $200 billion.

Vera architecture and claimed performance
NVIDIA positions the Vera processor as a CPU built specifically for agent-style workloads. It packs 88 in-house “Olympus” cores and, according to NVIDIA, delivers 1.2 TB/s of memory bandwidth. The company highlights architectural gains: significant single‑core performance uplift (NVIDIA cites up to 2× in some measurements), a threefold increase in inter‑core bandwidth and roughly 40% lower memory latency versus some competing chiplet designs. Customer tests cited by NVIDIA also show tangible improvements: DeepInfra reports up to 1.6× more concurrent agents and 2.2× faster orchestration, while other comparisons show up to 1.8× speedups on certain tasks versus x86 CPUs referenced in earlier announcements.

Vera Rubin: a complete infrastructure offering
NVIDIA presents Vera not as a standalone CPU but as a building block in a broader system. The Vera Rubin platform combines the Vera CPU, Rubin GPUs, networking chips and inference accelerators (including integrated Groq 3 LPX racks) to deliver racks optimized for training, real‑time inference and test‑time scaling. Ecosystem players — cloud providers and system integrators — are already deploying NVL72 racks; for example, Google Cloud offers A5X instances based on this configuration. CoreWeave reports a 10× increase in tokens generated per megawatt‑second on DeepSeek‑R1 compared with a prior generation.

Customers, volumes and business model
NVIDIA says it shipped Vera chips to several major customers as early as June, naming AI firms and companies in the space sector. The company is pursuing a vertically integrated model: selling complete systems and rack solutions through partners and cloud providers rather than only shipping standalone silicon.

Related posts

Will Reliance commission its first 120 MW AI infrastructure in Jamnagar by end-2026?

Emily Brown

APEC Digital Weeks 2026: How is Chengdu connecting culture and global AI?

James Smith

What initiatives did Satish Sharma launch at CUK?

Michael Johnson

This website uses cookies to improve your experience. We assume you agree, but you can opt out if you wish. Accept More Info

Privacy & Cookies Policy