Nvidia Vera Rubin AI platform on schedule, customer testing underway

Close-up of a hand holding a smartphone showing the NVIDIA logo on screen with a blurred background.

In brief

  • Vera Rubin integrates Vera CPU and Rubin GPU with seven specialized chips for unified architecture
  • Platform delivers up to 10x inference token cost reduction compared to Blackwell
  • Production shipments targeted for H2 2026; customer sampling already underway
  • CEO Huang confirmed manufacturing at scale in Taiwan and globally
  • Nvidia collaborating with OpenAI and Anthropic for platform integration

Performance and Architecture

Vera Rubin promises up to a 10x reduction in inference token costs compared to the current Blackwell architecture. The platform can also reduce the number of GPUs required for certain models by four times, addressing a critical pain point for hyperscalers managing massive inference workloads. It incorporates advanced liquid cooling systems to handle thermal demands at scale.

Manufacturing and Timeline

CEO Jensen Huang has personally swatted away rumors of production delays, confirming that Vera Rubin is being manufactured at scale both in Taiwan and globally. Full production ramp-up is set for May 2026, with production shipments targeted for the second half of 2026.

Customer Adoption and Partnerships

Customer sampling began earlier in the year, with partners like AWS, Google Cloud, and Microsoft expected to integrate Vera Rubin into their cloud infrastructure. Nvidia has also flagged collaborations with OpenAI and Anthropic, signaling broad industry backing for the platform's ecosystem.

Frequently asked questions

What is Vera Rubin?

Vera Rubin is Nvidia's next-generation AI platform that integrates the Vera CPU and Rubin GPU into a unified architecture with seven specialized chips. It's designed to reduce inference costs and GPU requirements for enterprise AI workloads.

How much can Vera Rubin reduce inference costs?

The platform promises up to a 10x reduction in inference token costs compared to Blackwell. It can also reduce the number of GPUs required for certain models by four times.

When will Vera Rubin be available?

Full production ramp-up is set for May 2026, with production shipments targeted for the second half of 2026. Customer sampling already began earlier in the year with AWS, Google Cloud, and Microsoft.