- Type
- GPU/CPU interconnect
- Developer
- NVIDIA
- First released
- 2016 (1st generation)
- Latest generation
- NVLink 5 (Blackwell, 2025); NVLink 6 (2026)
- Peak bandwidth
- 1.8 TB/s per GPU (NVLink 5)
- Related
- NVSwitch, PCIe, NVLink Fusion, InfiniBand
- Type
- GPU/CPU interconnect
- Developer
- NVIDIA
- First released
- 2016 (1st generation)
- Latest generation
- NVLink 5 (Blackwell, 2025); NVLink 6 (2026)
- Peak bandwidth
- 1.8 TB/s per GPU (NVLink 5)
- Related
- NVSwitch, PCIe, NVLink Fusion, InfiniBand
History
NVLink debuted in 2016 alongside the Pascal-generation P100 GPU, offering several times the bandwidth of the PCIe Gen3 links then in use.[1] Early generations were exclusive to NVIDIA's own GPU-CPU pairings. In 2019 NVIDIA introduced NVSwitch, a switching device that lets every GPU in a server talk to every other GPU at full bandwidth, replacing mesh topologies with a non-blocking fabric.
The technology scaled up sharply with each architecture generation: the Hopper generation (H100) delivered up to 900 GB/s of bidirectional bandwidth per GPU, and the fifth-generation NVLink shipped with the Blackwell architecture in 2025 at 1.8 TB/s per GPU — about 14 times the bandwidth of PCIe Gen5.[1][4] In NVIDIA's GB200 NVL72 systems, 72 GPUs and their Grace CPUs are connected into a single NVLink domain spanning one liquid-cooled rack, so that the whole rack behaves like one massive GPU.[4]
In May 2025, at Computex, NVIDIA announced NVLink Fusion, which opens the interconnect to chips NVIDIA did not build — custom CPUs and accelerators from partners such as MediaTek, Marvell, Qualcomm and SiFive can be connected to NVIDIA GPUs, provided the system still contains at least one NVIDIA component.[2][3] NVIDIA expanded the program through 2025 and 2026, and the sixth generation of NVLink ships with the Vera Rubin platform, doubling per-GPU bandwidth again.[1]
Key Concepts and Technology
Modern AI training uses parallelism strategies — splitting a model's layers, attention heads or data across many processors — that constantly exchange intermediate activations between chips. When those links are slow, expensive GPUs sit idle waiting for tensors to arrive, so interconnect bandwidth directly determines real-world throughput.
NVLink addresses this with dedicated point-to-point lanes between chips plus NVSwitch fabric inside the rack. Its key characteristics are:
- Bandwidth far above PCIe: NVLink 5 provides 1.8 TB/s bidirectional per GPU versus roughly 128 GB/s for PCIe Gen5, so GPUs exchange activations at close to on-package memory speeds.[1]
- Coherent memory: GPUs can read each other's memory directly, enabling multi-GPU memory pooling so a model larger than one GPU's memory can be served across a rack.
- Low energy per bit: NVLink's signalling is markedly more energy-efficient per transferred bit than serial links, which matters at the scale of an AI data centre.[1]
- NVLink Fusion: third-party chips attach either through NVLink chip-to-chip IP for custom CPUs or through a UCIe bridge chiplet for custom AI accelerators, letting companies keep their own silicon while joining an NVLink rack.[3]
Applications or Impact
NVLink underpins the dominant configuration for training frontier models: clusters of thousands of GPUs organised into NVL72-style racks. It also matters for serving large mixture-of-experts and multimodal models, where weights must be sharded across many GPUs and the interconnect becomes part of the inference path. Economically, the interconnect is central to NVIDIA's platform strategy — by making NVLink the fabric of the data centre, NVIDIA keeps the ecosystem vertically integrated even as it opens the interface to rival silicon.[2]
>See Also
🇲🇾 Relevance to Malaysia: Malaysia's rapid data-centre build-out in Johor and the Klang Valley — including AI-focused facilities and local cloud offerings such as YTL AI Cloud — is acquiring rack-scale NVIDIA systems where NVLink architecture determines usable performance and energy efficiency. For local data-centre operators, the interconnect's bandwidth-per-watt profile feeds directly into the power-efficiency and green AI criteria that grid constraints and Malaysia's sustainability expectations place on new capacity.
System integrators and engineers supported through MDEC programmes and the National AI Office's talent initiatives increasingly design around NVL72-class racks rather than individual GPU servers, making NVLink topology knowledge part of the local skill set. Malaysian organisations handling personal data under the PDPA also benefit where NVLink racks allow large models to be served wholly onshore instead of routed to overseas inference APIs.
References
- ↑NVIDIA. NVLink — high-speed GPU interconnect. https://www.nvidia.com/en-us/data-center/nvlink/
- ↑CRN (2025). Nvidia Reveals Offering To Build Semi-Custom AI Systems With Hyperscalers. https://www.crn.com/news/ai/2025/nvidia-reveals-offering-to-build-semi-custom-ai-systems-with-hyperscalers2
- ↑NVIDIA. NVLink Fusion. https://www.nvidia.com/en-us/data-center/nvlink-fusion/
- ↑NVIDIA. NVIDIA GB200 NVL72 — rack-scale computing platform. https://www.nvidia.com/en-us/data-center/gb200-nvl72/