AIWiki
Malaysia
Back to all articles
ApplicationsNVIDIAGPURubin

NVIDIA Rubin

4 min readUpdated August 2026
NVIDIA Rubin
Type
AI GPU and data centre platform
Developer
NVIDIA
Announced
GTC, March 2025
Architecture
Vera CPU + Rubin GPU (NVLink 144)
Related
Blackwell, CUDA, HBM4

NVIDIA Rubin is a GPU architecture and data centre platform developed by NVIDIA, announced at GTC in March 2025 as the successor to the Blackwell generation and named after the American astronomer Vera Rubin.[1][2] The platform pairs the Vera CPU — NVIDIA's first custom-designed CPU, based on the Olympus core design — with the Rubin GPU, which features 288 GB of HBM4 memory, and interconnects them with NVLink 144.[1][3]

History

NVIDIA revealed the Rubin platform at GTC 2025, with CEO Jensen Huang describing the generation as built for the age of agentic AI and reasoning.[2][3] The Rubin GPU was specified with 288 GB of High Bandwidth Memory 4 (HBM4), approximately 4.2 times the memory capacity of the previous Grace CPU generation, and 2.4 times the memory bandwidth.[1] When paired with Vera, a single superchip was stated to deliver 50 petaflops of inference performance, more than double the 20 petaflops of the Blackwell generation.[1]

The Vera CPU itself contains 88 cores and was designed to deliver twice the overall performance of the Grace Blackwell generation.[1] NVIDIA also announced a successor, Vera Rubin Ultra, planned for 2027, which combines four GPU dies into a single unit and will scale to rack-level systems such as the NVL576 configuration.[1][3]

Key Concepts

The Rubin platform is organised around several design concepts:

  • Vera CPU: NVIDIA's first in-house CPU core design (codenamed Olympus), providing the host processing for the platform.[1][3]
  • Rubin GPU and HBM4: the GPU uses fourth-generation high-bandwidth memory, expanding capacity to 288 GB per GPU to support very large models and long-context inference.[1]
  • NVLink 144 scale-up: the platform interconnects 144 GPU compute dies within a rack-scale system, with the first-generation racks designated NVL144.[3]
  • Rubin CPX: announced in September 2025, a dedicated context processing accelerator for massive-context inference, integrated into the Vera Rubin NVL144 CPX platform, which NVIDIA states delivers 7.5 times the AI performance of a GB300 NVL72 system with 100 TB of fast memory per rack.[4]
  • Power scaling: rack power draw is expected to rise dramatically, with NVIDIA projecting up to 600 kW per rack by 2027, forcing data centre operators to redesign power delivery and cooling.[3][5]

Applications and Impact

The Rubin generation is aimed at AI training and, in particular, inference for reasoning and agentic AI, where long context windows and high memory bandwidth reduce cost per token.[2][4] The platform extends NVIDIA's CUDA ecosystem, which the company reports includes more than six million developers and nearly 6,000 CUDA applications.[4] Because of its higher rack power densities, Rubin-class hardware is a driver of data centre infrastructure changes, including liquid cooling and new power distribution systems.[3][5]

>See Also

References

🇲🇾Malaysian Context

Malaysia has become a significant deployment region for NVIDIA data centre hardware. The country's first NVIDIA-powered AI data centre, developed by YTL Power International in partnership with NVIDIA in Johor, uses liquid-cooled Grace Blackwell NVL72 GPUs and is powered by renewable energy from a 500 MW solar plant.[6] A 600 MW NVIDIA-powered facility at YTL's Green Data Center Park in Kulai, Johor, supports AI model training and deployment, and the partnership includes the development of Malaysia's first sovereign large language model.[7] As Malaysian operators plan upgrades beyond Blackwell, the Rubin platform's 600 kW rack projections reinforce the National Semiconductor Strategy's emphasis on advanced packaging and power infrastructure, and the 2026 Budget allocated RM5.9 billion to enhance AI infrastructure and adoption.[5][7]

References

  1. [NVIDIA GTC 2025: AI Reasoning, Blackwell Ultra, Vera Rubin — NADDOD](https://www.naddod.com/blog/nvidia-gtc-2025-ai-reasoning-blackwell-ultra-vera-rubin-cpo-dynamo-inference)
  2. [NVIDIA Vera Rubin Platform — official page](https://www.nvidia.com/en-us/data-center/technologies/rubin)
  3. [NVIDIA GTC 2025 — Built for Reasoning: Vera Rubin — SemiAnalysis](https://newsletter.semianalysis.com/p/nvidia-gtc-2025-built-for-reasoning-vera-rubin-kyber-cpo-dynamo-inference-jensen-math-feynman)
  4. [NVIDIA Unveils Rubin CPX — NVIDIA Newsroom](https://nvidianews.nvidia.com/news/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference)
  5. [NVIDIA Vera Rubin: 600kW racks by 2027 — Introl Blog](https://introl.com/blog/nvidia-vera-rubin-gpu-600kw-racks-2027)
  6. [Malaysia strengthens AI push with new Nvidia data centre in Johor — Mixmag Asia](https://mixmag.asia/read/malaysia-strengthens-ai-push-with-new-nvidia-data-centre-in-johor-local)
  7. [Malaysia Advances AI Sovereignty with Nvidia-Powered Data Center — Yahoo Finance](https://finance.yahoo.com/news/malaysia-advances-ai-sovereignty-nvidia-110000010.html)