AWS Expands NVIDIA GPU Deployment Beyond 3 Million as AI Demand Surges
Amazon Web Services is dramatically expanding its NVIDIA infrastructure plans, with the cloud provider preparing to deploy 2 million additional NVIDIA GPUs across its global infrastructure in 2027 and 2028. The expansion builds on the more than 1 million NVIDIA GPUs AWS announced at GTC 2026 for deployment beginning this year, taking the planned total beyond 3 million accelerators as demand for AI computing continues to exceed earlier expectations. According to the official AWS and NVIDIA announcement, the additional deployment will include Blackwell Ultra, Rubin and Rubin Ultra GPUs supporting workloads ranging from agentic AI and scientific research to enterprise automation and physical AI.
The scale of the expansion is significant because AWS had only announced its initial plan for more than 1 million NVIDIA GPUs in March 2026. Just months later, the company says demand has already exceeded those expectations, resulting in another 2 million accelerators being added to the roadmap. The deployment also places AWS among the largest customers for NVIDIA's next generation data center platforms as Vera Rubin moves deeper into its 2026 rollout. NVIDIA Blackwell Ultra will continue expanding within AWS while Rubin and Rubin Ultra become increasingly important during the 2027 and 2028 deployment period.
The partnership extends far beyond GPU purchases. AWS and NVIDIA are also working to bring NVIDIA Vera CPU infrastructure into the AWS ecosystem, giving customers another CPU option for agentic AI workloads that require substantial general purpose compute alongside GPU acceleration. NVIDIA and Amazon's Annapurna Labs are also expanding their collaboration around NVLink Fusion and the recently introduced NVHBM memory architecture for future Trainium accelerators. This could allow Amazon's own AI silicon to access faster and more power efficient memory while integrating Trainium and NVIDIA GPUs within a common rack scale architecture.
Amazon therefore appears to be pursuing a heterogeneous AI infrastructure strategy rather than choosing between NVIDIA hardware and its internally developed Trainium accelerators. AWS can continue investing in its own silicon while simultaneously deploying NVIDIA hardware at enormous scale, giving customers access to different compute platforms depending on workload requirements. NVIDIA is also becoming more deeply integrated into AWS networking, storage, security and software infrastructure through technologies including Spectrum networking, the AWS Nitro System, Elastic Fabric Adapter, Nemotron models, CUDA libraries and NVIDIA's physical AI platform.
Another major component of the agreement targets government AI infrastructure. AWS and NVIDIA plan to build dedicated AI factories for the United States Government, including 100,000 NVIDIA GPUs operating on secure AWS infrastructure for federal and national security workloads. The companies say these systems are intended to support workloads classified at Impact Level 6 and above, showing how large scale GPU infrastructure is increasingly becoming part of government and national security computing strategies rather than remaining limited to commercial cloud services and frontier AI laboratories.
"NVIDIA and AWS have built one of the great growth engines of the AI era, and demand is running ahead of every forecast."
— Quote by: Jensen Huang, Founder and CEO of NVIDIA.
The timing also comes as the cost of deploying next generation AI infrastructure continues rising. Vera Rubin and Blackwell server pricing could increase by around 17% as memory and other component costs increase, meaning a deployment involving millions of accelerators could represent an enormous infrastructure commitment even before power, networking, cooling and data center construction are considered. Neither AWS nor NVIDIA disclosed the financial value of the expanded GPU agreement.
Going from more than 1 million NVIDIA GPUs to more than 3 million planned accelerators within months illustrates how quickly hyperscaler AI infrastructure requirements are changing. More interestingly, AWS is not abandoning Trainium to make room for NVIDIA. It is doing the opposite, integrating both ecosystems more closely through NVLink Fusion and NVHBM. The future of hyperscale AI may therefore be less about a single accelerator winning the entire data center and more about massive heterogeneous systems where GPUs, custom AI chips, CPUs, networking and memory architectures are optimized together.
With AWS planning more than 3 million NVIDIA GPUs while continuing to invest heavily in Trainium, could heterogeneous AI infrastructure become the dominant architecture for future cloud data centers?
