AIWiki
Malaysia
Back to all articles
Tools & Platformsrocmamdgpu computing

ROCm

4 min readUpdated October 2026
ROCm
Type
Open-source GPU compute software stack
Developer
Advanced Micro Devices (AMD)
Launched
2016 (Boltzmann Initiative announced 2015)
Licence
Open source (MIT, Apache 2.0 and related licences; GPU firmware excluded)
Key components
HIP, HIPIFY tools, rocBLAS, MIOpen, RCCL, Composable Kernel
Target hardware
AMD Instinct MI250X, MI300, MI350 accelerators; EPYC CPUs
Notable users
Frontier (first exascale supercomputer, 2022), El Capitan (TOP500 #1, 2024)
ROCm is Advanced Micro Devices' free and open-source software stack for graphics processing unit (GPU) computing, comprising drivers, compilers, runtime tools and libraries that enable artificial intelligence and high-performance computing (HPC) workloads on AMD hardware. The name originally stood for "Radeon Open Compute platform".[1]

History

AMD and its former ATI division had pursued GPU computing since the 2000s with the "Close to Metal" and AMD Stream toolkits. The Boltzmann Initiative, announced at SC15 in November 2015, committed AMD to a C++ programming model and CUDA-compatible compiler tools for its GPUs; ROCm 1.3 followed at SC16 in 2016, marking the platform's public launch.[2] Early releases focused on HPC research and the Heterogeneous System Architecture effort, and the stack accumulated components such as ROCr runtime, rocBLAS linear algebra libraries and the MIOpen deep-learning primitives.[1]

ROCm matured alongside AMD's Instinct accelerators. The platform powered Frontier at Oak Ridge National Laboratory, the first supercomputer to exceed the exascale threshold, which debuted at number one on the TOP500 list in November 2022 and also topped the Green500 energy-efficiency ranking.[3] AMD's acquisition of Xilinx (2022, US$49 billion) and ZT Systems (2024, US$4.9 billion) expanded its AI systems ambitions, and El Capitan at Lawrence Livermore National Laboratory, powered by Instinct MI300A accelerators and ROCm, took the TOP500 number-one position in November 2024.[3] In 2026 AMD announced a shift to a six-week ROCm release cadence, reflecting the platform's growing role in AI workloads.[1]

Key Concepts and Technology

ROCm's central abstraction is HIP (Heterogeneous-computing Interface for Portability), a C++ runtime API and kernel language with an interface closely modelled on NVIDIA's CUDA. The hipcc compiler wraps Clang with the open-source LLVM AMDGPU backend — and can also redirect to NVIDIA's compiler — while HIPIFY tools mechanically translate existing CUDA source code into HIP, easing migration of AI workloads.[4][5]

Beyond the compiler layer, ROCm provides a growing library ecosystem: rocBLAS and hipBLAS for linear algebra, rocFFT for Fourier transforms, MIOpen for deep-learning primitives, RCCL for multi-GPU collectives (the analogue of NVIDIA's NCCL), and the Composable Kernel library for building fused GPU kernels.[4] The stack integrates with major machine-learning frameworks, including official ROCm builds of PyTorch, TensorFlow and JAX, and serving systems such as vLLM.[1] Unlike NVIDIA's proprietary CUDA stack, ROCm components are developed in the open, with the exception of some GPU firmware binaries.[1]

Applications and Impact

ROCm's flagship applications are exascale scientific computing — Frontier and El Capitan support research in fusion energy, climate science and nuclear stockpile stewardship — and AI training and inference at hyperscale, where AMD's Instinct MI300X accelerators are deployed at operators including Microsoft Azure, Meta and Oracle.[3] ROCm's strategic significance lies in reducing dependence on a single vendor's proprietary software ecosystem: as CUDA compatibility layers and native ROCm ports of major AI frameworks have improved, organisations can now run large language models and training pipelines on AMD hardware at competitive cost, a shift that has grown AMD's AI accelerator revenue into the billions of US dollars annually.[1]

>See Also

🇲🇾Malaysian Context

🇲🇾 ROCm matters to Malaysia primarily as an alternative in the country's rapidly expanding AI data-centre landscape in Johor, Kulim and Cyberjaya, where NVIDIA's CUDA ecosystem currently dominates.[5] For Malaysian institutions — including public universities operating high-performance computing clusters and the developers working on sovereign AI initiatives — an open, royalty-free GPU software stack offers a hedge against single-vendor lock-in and a path to lower-cost compute, both concerns that sit at the centre of Malaysia's National AI Office (NAIO) agenda.[6]

AMD also maintains engineering presence in Malaysia's semiconductor ecosystem, and EPYC processors are common in local server deployments. The efficiency argument is relevant too: ROCm-powered Frontier ranked first on the Green500 list, and energy efficiency is a pressing issue as Malaysia's data-centre expansion meets scrutiny over power consumption and sustainability commitments. Training pathways through HRD Corporation grants and MDEC digital-talent programmes increasingly include open GPU-compute skills, and on-premise ROCm clusters allow Malaysian organisations to keep data within PDPA-compliant, in-country infrastructure.[3][6]

References

  1. ↑Wikipedia. ROCm — background, components and ecosystem. https://en.wikipedia.org/wiki/ROCm
  2. ↑AMD GPUOpen / AnandTech archive. AMD SC15: Boltzmann Initiative announced. https://web.archive.org/web/20151117085537/http://www.anandtech.com/show/9792/amd-sc15-boltzmann-initiative-announced-c-and-cuda-compilers-for-amd-gpus
  3. ↑AMD. ROCm: Powering the world's fastest supercomputers (Frontier and El Capitan). https://rocm.blogs.amd.com/ecosystems-and-partners/rocm-revisited-power/README.html
  4. ↑AMD. ROCm documentation. https://rocm.docs.amd.com/en/latest/
  5. ↑AMD. HIP — Heterogeneous-computing Interface for Portability. https://rocm.docs.amd.com/projects/HIP/en/latest/
  6. ↑TOP500. Green500 — most energy-efficient supercomputers. https://www.top500.org/lists/green500/