- Type
- GPU architecture for AI
- Developed by
- NVIDIA
- Announced
- 2024 (GTC)
- Flagship GPU
- B200 (208 billion transistors, 180-192GB HBM3e)
- Rack system
- GB200 NVL72 (72 GPUs, liquid-cooled)
- Related
- GPU cluster, CUDA, AI data centres, sovereign AI
- Type
- GPU architecture for AI
- Developed by
- NVIDIA
- Announced
- 2024 (GTC)
- Flagship GPU
- B200 (208 billion transistors, 180-192GB HBM3e)
- Rack system
- GB200 NVL72 (72 GPUs, liquid-cooled)
- Related
- GPU cluster, CUDA, AI data centres, sovereign AI
NVIDIA Blackwell is a graphics processing unit (GPU) architecture developed by NVIDIA and announced in 2024, designed to train and serve the largest artificial intelligence models. Named after the mathematician David Blackwell, the architecture succeeds NVIDIA's Hopper generation and targets the demands of generative AI and large language models, where model sizes reaching trillions of parameters require enormous compute, memory, and interconnect bandwidth. Blackwell-based systems became the dominant high-end AI hardware deployed in data centres through 2025 and 2026.
Architecture Highlights
The flagship Blackwell GPU, the B200, uses a dual-die design housing approximately 208 billion transistors, far exceeding the single-die designs of previous generations. The two dies are connected by a high-bandwidth link so they operate effectively as one GPU. Each B200 includes high-bandwidth HBM3e memory — on the order of 180 to 192 gigabytes depending on configuration — with around 8 terabytes per second of memory bandwidth, addressing the memory capacity and speed bottlenecks that constrain large-model training and inference.
A defining feature of Blackwell is support for lower-precision numerical formats, notably FP4 (four-bit floating point) in its fifth-generation Tensor Cores. Reduced precision allows far higher throughput and lower energy per operation, which is especially valuable for inference. NVIDIA reported that Blackwell delivers roughly four times the training throughput and substantially greater inference performance compared with the prior Hopper H100 generation on transformer workloads, though exact gains depend on the task and precision used.
The Grace Blackwell Superchip
Blackwell is frequently deployed as part of the GB200 Grace Blackwell Superchip, which combines two B200 GPUs with an NVIDIA Grace CPU over a high-bandwidth NVLink chip-to-chip interconnect, reported at around 900 gigabytes per second. Pairing CPU and GPU tightly in one module reduces data-movement bottlenecks between processors and provides large, coherent memory accessible to the accelerators — an arrangement well suited to massive models and their associated data.
Rack-Scale Systems
To train and serve trillion-parameter models, individual GPUs must be linked into much larger systems. The GB200 NVL72 connects 36 Grace CPUs and 72 Blackwell GPUs in a single liquid-cooled rack, joined by fifth-generation NVLink so that the 72 GPUs function as one very large accelerator with a unified high-speed memory domain. NVIDIA describes this configuration as delivering on the order of an exaflop-scale of AI performance and tens of terabytes of fast memory, with large gains in real-time inference for the biggest language models. Liquid cooling is essential because the density of these racks produces heat that air cooling cannot efficiently remove. The NVL72 serves as a building block for larger DGX SuperPOD supercomputers.
NVIDIA subsequently introduced Blackwell Ultra products (such as the B300 generation) with increased memory and performance, continuing the architecture's roadmap.
Significance
Blackwell hardware sits at the centre of the global build-out of AI infrastructure. Its combination of large memory, low-precision throughput, tight CPU-GPU integration, and rack-scale interconnect made it the platform of choice for hyperscalers and AI labs training frontier models. Demand for Blackwell systems has shaped supply chains, data centre design (driving adoption of liquid cooling), and national AI strategies that depend on access to advanced accelerators.
NVIDIA Blackwell hardware is directly tied to Malaysia's emergence as a regional AI infrastructure hub. In late 2025, YTL Power completed what was reported as Malaysia's first NVIDIA-powered AI data centre in Kulai, Johor, built around liquid-cooled GB200 Grace Blackwell systems and operated as YTL AI Cloud — making advanced Blackwell-class compute available domestically for the first time.
Access to Blackwell-class accelerators underpins Malaysia's sovereign AI ambitions: training and serving local models such as Ilmu and MaLLaM, and hosting sensitive workloads on national soil, both depend on availability of frontier hardware within the country. The liquid-cooling and high power density of Blackwell systems also shape data centre design choices in Johor and the Klang Valley, connecting hardware decisions to the Green AI and energy considerations managed with Tenaga Nasional Berhad and the Energy Commission.
Hyperscalers expanding in Malaysia — including Microsoft, Google, Oracle, and AWS — deploy NVIDIA accelerators across their regional infrastructure, broadening access for Malaysian enterprises. Agencies such as MDEC and MIDA factor the availability of advanced compute into investment promotion, while export-control dynamics affecting advanced chips are a strategic consideration for national planners. For Malaysian developers and researchers, local Blackwell capacity reduces the need to rent overseas compute, lowering latency and supporting data residency requirements relevant under the PDPA and to regulators such as Bank Negara Malaysia.