Amazon Web Services and Nvidia are dramatically expanding their infrastructure partnership, with AWS planning to deploy an additional two million Nvidia GPUs across its global infrastructure over 2027 and 2028. The expansion builds on AWS's previously announced plan to add more than one million Nvidia GPUs beginning in 2026, and reflects customer demand for AI infrastructure capacity that both companies say has exceeded their earlier expectations even after accounting for the scale of their existing commitments.
The deployment will include Nvidia's Blackwell Ultra, Rubin and Rubin Ultra chip generations, spanning both currently shipping and next-generation architectures still in development. That breadth signals a partnership structured for continuity across multiple successive Nvidia product cycles rather than a single, time-limited infrastructure commitment, giving AWS customers visibility into a multi-year capacity roadmap as they plan their own long-term AI infrastructure strategies.
The partnership extends meaningfully beyond raw GPU procurement. Nvidia's Vera CPU architecture will also come to AWS as part of the expanded collaboration, alongside deeper joint work on networking and data-processing infrastructure — areas that have increasingly emerged as critical bottlenecks in large-scale AI training and inference deployments, where the efficiency of moving data between compute nodes can matter as much as the raw processing power of individual chips.
The scale of this expansion reflects the broader industrial character AI infrastructure investment has assumed over the past several years, evolving from what was initially framed largely as a software competition into what increasingly resembles a capital-intensive industrial buildout comparable to historic infrastructure expansions in telecommunications or energy. Nvidia itself has indicated it expects sales growth of roughly 70 percent as demand for AI infrastructure continues climbing, a trajectory this AWS expansion helps substantiate.
For enterprise customers building AI applications on AWS, the expanded GPU deployment offers greater assurance of available compute capacity at a moment when GPU scarcity has periodically constrained AI development timelines across the industry, forcing some companies to delay training runs or product launches while awaiting sufficient hardware availability from cloud providers.
The competitive dynamics among major cloud providers — AWS, Microsoft Azure and Google Cloud — increasingly hinge on the scale and sophistication of AI infrastructure each can offer, making this kind of expanded chip commitment as much a competitive positioning statement as a straightforward capacity planning exercise. AWS's willingness to commit to such a substantial multi-year GPU deployment signals confidence that AI workload demand will continue its current trajectory rather than moderating, a bet that carries meaningful capital risk should AI infrastructure demand growth slow more than currently anticipated.

Nvidia's position at the centre of this expanding infrastructure buildout continues to reinforce its dominant position within the AI hardware supply chain, even as competitors including AMD and a growing field of custom AI chip developers, including cloud providers building their own silicon, attempt to erode that dominance over the coming product cycles. The depth and multi-generational scope of this AWS partnership suggests Nvidia's near-term competitive position within hyperscale cloud infrastructure remains robust despite these emerging competitive pressures.
As the deployment unfolds across 2027 and 2028, the practical execution challenges — data centre construction timelines, power availability, supply-chain coordination across two of technology's largest infrastructure operations — will likely prove as consequential to the partnership's ultimate success as the underlying chip technology itself, given the sheer physical and logistical scale involved in deploying millions of additional high-performance GPUs across a global data centre footprint.
Power availability, in particular, has emerged as an increasingly binding constraint on AI infrastructure expansion across the industry, with data centre operators and utilities in multiple regions reporting grid capacity limitations that could plausibly slow the pace at which AWS can bring newly deployed GPU capacity online, regardless of how quickly Nvidia can manufacture and ship the underlying chips themselves.
For the broader AI industry, this scale of committed future infrastructure capacity offers a degree of reassurance that compute scarcity — a persistent constraint on AI development and deployment timelines over the past several years — may ease somewhat as this expanded capacity comes online, though the compounding growth in AI model size and training compute requirements means demand could plausibly continue outpacing even this substantially expanded supply.
Investors evaluating both companies will likely scrutinise how this expanded commitment affects capital expenditure guidance and margin trajectories over the coming fiscal years, given the substantial upfront capital outlay required for data centre construction and chip procurement at this scale, weighed against the recurring revenue both companies expect to generate from AI infrastructure customers over the useful life of the deployed hardware. Rival cloud providers Microsoft Azure and Google Cloud will almost certainly face pressure to respond with comparable infrastructure commitments of their own, sustaining an already intense capital expenditure race across the hyperscale cloud industry that shows little sign of slowing as AI workload demand continues its rapid expansion. For enterprise customers watching from the sidelines, that competitive dynamic among the three major cloud providers could ultimately translate into more favourable pricing and greater compute availability than a less competitive market structure would likely produce. The partnership's ultimate success will also depend on how effectively both companies coordinate around emerging AI workload patterns, as inference-heavy applications increasingly supplement the training-focused compute demand that has dominated AI infrastructure planning over the technology's more experimental earlier years. That shift toward inference-heavy workloads carries its own distinct infrastructure requirements, favouring geographically distributed deployment closer to end users over the more centralised, training-optimised data centre architecture that has characterised much of the industry's build-out to date. As both companies navigate this evolving workload mix, their ability to flexibly reallocate capacity between training and inference use cases will likely prove as commercially significant as the raw scale of GPUs being deployed under this expanded partnership, particularly as enterprise customers increasingly demand cost-efficient inference at scale rather than raw training throughput alone. The years ahead will show whether this expanded capacity keeps pace with demand or whether AI infrastructure scarcity re-emerges as a recurring constraint on the industry's growth trajectory.



