- Type
- Custom AI accelerator chips (ASICs)
- Developer
- Amazon Web Services (AWS)
- Key features
- NeuronCores, purpose-built for ML training/inference, AWS Neuron SDK
- Emerged
- 2019–2020 (first generations)
- Related
- NVIDIA, TPU, SambaNova, Groq
- Type
- Custom AI accelerator chips (ASICs)
- Developer
- Amazon Web Services (AWS)
- Key features
- NeuronCores, purpose-built for ML training/inference, AWS Neuron SDK
- Emerged
- 2019–2020 (first generations)
- Related
- NVIDIA, TPU, SambaNova, Groq
AWS Trainium and Inferentia are families of custom AI accelerator chips designed by Amazon Web Services (AWS) for machine learning workloads. Inferentia is optimized for running trained models (inference), while Trainium is purpose-built for training large deep-learning models, including generative AI and large language models; both are programmed through the AWS Neuron software development kit [1][2][3].
History
AWS announced Inferentia, its first custom ML chip, in 2018 as part of a strategy to reduce dependence on merchant GPUs, and it became generally available in 2019. Trainium, the training-focused sibling, was announced at AWS re:Invent in 2020 and reached general availability in 2023 on Amazon EC2 Trn1 instances, which can be connected in petabit-scale UltraClusters using the Elastic Fabric Adapter (EFA) network for distributed training [1][4].
Subsequent generations extended the family: Trainium2 powers the Trn2 instance generation, and Trainium3, a 3-nanometer design, arrived with the Trainium3 UltraServer announced in 2025 [3]. By March 2026 AWS reported more than 1.4 million Trainium chips deployed across all generations, with Anthropic and OpenAI among its anchor customers [3]. In June 2026, Bloomberg reported that Amazon was in early talks to sell Trainium chips directly to third-party data centres — a potential departure from its long-standing AWS-exclusive distribution model [3].
Key Concepts
Trainium and Inferentia are application-specific integrated circuits (ASICs) whose silicon is shaped around the three dominant pressures of AI workloads: matrix-heavy mathematics, memory bandwidth, and cluster communication [4]. Each chip contains a small number of large NeuronCores with specialized tensor engines, large on-chip SRAM placed close to the compute units, and dedicated hardware for inter-chip communication so that computation and data exchange happen in parallel [4].
Software support is provided by the AWS Neuron SDK, which compiles models from frameworks such as PyTorch into kernels that run on the NeuronCores [2][5]. The Neuron Kernel Interface (NKI) allows developers to write custom high-performance kernels, and in 2026 AWS introduced Neuron Agentic Development, an open-source suite of AI agents that author, debug, and profile NKI kernels from natural language inside agentic coding environments such as Claude Code [5]. AWS has also invested in the open-source ecosystem around Trainium through the $110 million Build on Trainium research-credit program [1].
Applications
AWS markets Trainium and Inferentia-based instances as delivering high performance at lower cost and higher energy efficiency than GPU instances for many workloads [2][3]. Trainium clusters are used to train large language models and diffusion models, while Inferentia instances serve production inference for chatbots, embeddings, and recommendation systems. Notable deployments include Anthropic's large-scale Trainium clusters for training its models [3]. The chips' energy-efficient design is also highlighted as supporting sustainability goals in AI infrastructure [1].
>See Also
🇲🇾 AWS launched its first infrastructure region in Malaysia in August 2024 — the AWS Asia Pacific (Malaysia) Region (ap-southeast-5) with three availability zones and a planned investment of about US$6.2 billion through 2037 — giving Malaysian organizations low-latency, in-country access to AWS AI services, including Trainium and Inferentia-based compute [6][7][8]. The region supports Malaysia's goals under MyDIGITAL and the New Industrial Master Plan (NIMP) 2030 and is expected to add roughly US$12.1 billion to Malaysia's GDP [7][8].
Malaysian customers already use AWS AI infrastructure: homegrown computer-vision startup Tapway runs its SamurAI platform on AWS machine learning services, the national broadcaster RTM uses Amazon Bedrock for its RTMKlik streaming platform, and telco Maxis collaborates with AWS on generative AI and 5G use cases [6][8]. For Malaysian startups and enterprises, Trainium and Inferentia offer a lower-cost path to building and serving AI models while keeping data within the country — a consideration that aligns with PDPA data-residency expectations and the National AI Office's sovereign-AI agenda [7][8].
References
- ↑[Amazon Science — Build on Trainium: Kernels for ML Acceleration call for proposals](https://www.amazon.science/research-awards/call-for-proposals/build-on-trainium-kernels-for-ml-acceleration-call-for-proposals-spring-2026)
- ↑[AWS Machine Learning Blog — Enhanced observability for AWS Trainium and AWS Inferentia with Datadog](https://aws.amazon.com/blogs/machine-learning/enhanced-observability-for-aws-trainium-and-aws-inferentia-with-datadog)
- ↑[Digital Applied — Amazon May Sell Its AI Chips: The Nvidia Challenge (2026)](https://www.digitalapplied.com/blog/amazon-custom-ai-chips-nvidia-challenge-2026-analysis)
- ↑[Medium — How AWS Trainium actually works (2026)](https://medium.com/@brookejamieson/how-aws-trainium-actually-works-2026-deb4142dae11)
- ↑[AWS Neuron Documentation — What's New in the AWS Neuron SDK](https://awsdocs-neuron.readthedocs-hosted.com/en/latest/about-neuron/whats-new.html)
- ↑[AWS Press Center — AWS Launches Infrastructure Region in Malaysia (August 2024)](https://press.aboutamazon.com/2024/8/aws-launches-infrastructure-region-in-malaysia)
- ↑[Digital News Asia — AWS launches Infrastructure Region in Malaysia with US$6.2bil investment](https://www.digitalnewsasia.com/digital-economy/aws-launches-infrastructure-region-malaysia-us62bil-investment-through-2037)
- ↑[Data Centre Dynamics — AWS launches Malaysian cloud region](https://www.datacenterdynamics.com/en/news/aws-launches-malaysian-cloud-region)