A sprawling data center with glowing server racks and cooling pipes beneath a digital world map at sunset.
In 2026, NVIDIA’s named AI infrastructure partners—including AWS, Meta, Microsoft and OpenAI—are planning or deploying systems measured in millions of GPUs or gigawatts, while Microsoft’s already deployed Grace Blackwell fleet gives Azure customers a concrete sign of expansion without guaranteeing capacity in any particular region or service. The scale is real, but the numbers describe different things: installed hardware, future deployments and intended computing capacity. For enterprises, the useful question is which of those announcements translates into a cloud service they can actually buy.

Microsoft Azure has deployed Grace Blackwell; Vera Rubin is the next rollout​

Microsoft supplied one of the clearest statements about hardware already in place. On March 16, it said it had deployed hundreds of thousands of liquid-cooled NVIDIA Grace Blackwell GPUs across its global datacenter footprint in less than a year. Microsoft also said it had powered on a Vera Rubin NVL72 system in its labs and expected to roll that generation into modern, liquid-cooled Azure datacenters over the following months. The first figure describes a deployed fleet; the lab system and rollout statement describe a transition still underway.

That distinction is especially important for an Azure buyer. A GPU installed somewhere in Microsoft’s worldwide infrastructure does not tell an administrator which Azure region, service or customer allocation can use it. Microsoft’s figure establishes the scale of its Grace Blackwell investment, while its Vera Rubin statement establishes that it is preparing the next generation—not that every Azure customer can order a Rubin-backed instance today.

Microsoft’s Fairwater AI infrastructure illustrates why these projects are described as AI factories. The company describes datacenter sites linking hundreds of thousands of Blackwell GPUs across a large-scale network. At that size, the purchasing decision extends beyond accelerators to the connections, cooling and facilities needed to make them operate as a system. Fairwater is a description of Microsoft’s site architecture, not a disclosed total for every GPU Microsoft has bought.

NVIDIA’s fiscal Q2 2027 results show growth, not a customer league table​

NVIDIA reported $96.2 billion in revenue for its fiscal second quarter ended July 26, 2026, up 106% from a year earlier. Data Center revenue reached $89.0 billion, up 117%; that segment accounted for more than 92% of quarterly revenue. NVIDIA forecast approximately $108 billion in revenue for the following quarter, plus or minus 2%. Those results explain why large cloud deployments command attention, but segment revenue cannot be divided among named partners from the published figures.

The company’s SEC filing provides the firmer limit on any claim about its “biggest customers.” One unnamed direct customer accounted for 16% of total revenue in fiscal Q2 2027. Across the first half of that fiscal year, three unnamed direct customers individually accounted for 16%, 15% and 13%. NVIDIA’s definition of a direct customer includes intermediaries such as manufacturers and distributors as well as cloud providers and AI companies; an organization using NVIDIA systems can also acquire them through another company. The filing does not assign those shares to AWS, Meta, Microsoft or OpenAI.

InfotechLead brought these prominent partnerships together under a “biggest customers” framing. The defensible conclusion is narrower: they are among NVIDIA’s most visible named deployment partners, with substantial announced projects. Their public GPU counts and gigawatt plans cannot produce a reliable ranking by NVIDIA revenue, particularly when a model developer may use hardware supplied by a cloud provider.

AWS and Meta put millions of GPUs on different schedules​

AWS has the most specific forward GPU count in this group. On August 26, AWS and NVIDIA announced plans to deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure during 2027–2028, including Blackwell Ultra, Rubin and Rubin Ultra generations. The expansion follows an earlier AWS plan to add more than 1 million NVIDIA GPUs beginning in 2026. The August figure is an additional plan, not a statement that two million new accelerators had already been installed.

The AWS announcement also reaches beyond GPUs. The companies described work on NVIDIA Vera CPU-based infrastructure and networking, while AWS’s own account discusses how its cloud networking technologies fit the expanded fleet. That gives enterprise customers a reason to watch AWS’s eventual service offerings: the announcement concerns infrastructure AWS intends to turn into usable cloud capacity, but it supplies no blanket promise of a particular instance type in every region.

Meta’s February 17 agreement uses a broader measure. NVIDIA described a multiyear, multigenerational partnership spanning millions of Blackwell and Rubin GPUs, along with CPUs and Spectrum-X Ethernet networking, for infrastructure intended to serve both AI training and inference. The announcement gives neither an exact GPU total nor a delivery schedule. It therefore indicates the intended scale and breadth of Meta’s buildout, but it cannot be compared directly with AWS’s dated two-million-GPU expansion.

OpenAI’s gigawatts measure capacity, not a GPU purchase count​

OpenAI expressed its NVIDIA plans in a different unit. On February 27, 2026, it said an expanded collaboration included 3 gigawatts of dedicated inference capacity and 2 gigawatts of training capacity using Vera Rubin systems. Inference is the computing work that produces a model’s responses after training. OpenAI also said Hopper and Blackwell systems were already operating for it across Microsoft, Oracle Cloud Infrastructure and CoreWeave—evidence of its use of several infrastructure providers, not proof that it owns every GPU involved.

There is an earlier, larger headline figure, but it carries a different status. In September 2025, NVIDIA and OpenAI announced a letter of intent aimed at deploying at least 10 gigawatts of NVIDIA systems and said NVIDIA intended to invest up to $100 billion progressively as capacity was deployed. The February 2026 five-gigawatt description should not simply be added to that 10-gigawatt target: the announcements do not establish that the amounts are separate, non-overlapping installations. Neither figure is a report that all the stated capacity is operating.

A gigawatt commitment tells readers about the contemplated scale of computing infrastructure and its power requirements; it is not a universal conversion formula for GPU numbers. NVIDIA’s filing itself identifies access to land, power and datacenter facilities as material to its customers’ buildouts. For OpenAI, the practical story is therefore as much about obtaining usable capacity through partners as it is about selecting a chip generation.

Vera Rubin gives Oracle, CoreWeave and Google Cloud a place in the rollout​

NVIDIA’s Vera Rubin pitch is built around a platform rather than a GPU sold in isolation. Its published design combines processors with high-speed interconnects, networking and other system components. NVIDIA claims the platform can cut inference cost per token by up to 10 times compared with Blackwell under its comparison conditions. That is a vendor performance claim, not a price or throughput guarantee for an enterprise workload running in a cloud service.

The partner list also shows why counting only AWS, Meta, Microsoft and OpenAI misses part of the picture. NVIDIA named Google Cloud and Oracle Cloud Infrastructure among expected early providers of Vera Rubin-based instances, alongside Microsoft and AWS, and identified CoreWeave among its cloud partners. OpenAI separately identified Microsoft, OCI and CoreWeave as providers through which it was already using earlier NVIDIA systems. An early-deployment designation signals a provider’s place in the rollout; it does not specify when a particular customer can obtain a Rubin instance.

These companies occupy different positions. OCI and CoreWeave offer cloud capacity to customers such as OpenAI; Google Cloud and AWS also develop their own AI accelerators while working with NVIDIA. The coexistence is relevant to procurement: a cloud provider’s investment in alternative chips does not, by itself, remove NVIDIA-backed options from its plans. Buyers still need to compare the actual service and workload available to them, not infer a winner from an infrastructure announcement.

What this means for you​

Treat the announced buildouts as a reason to reassess future cloud-AI capacity, not as capacity already reserved for your project. For an Azure deployment in particular, Microsoft’s installed Grace Blackwell fleet is meaningful evidence of investment, while the purchasing decision still turns on the service, region and allocation a provider will offer.

  • Microsoft reported hundreds of thousands of Grace Blackwell GPUs deployed globally; its initial Vera Rubin NVL72 system was described as running in a lab ahead of a planned Azure rollout.
  • AWS’s two-million-GPU announcement describes additional deployment planned across 2027–2028, not two million GPUs already available for rental.
  • Meta’s “millions” figure belongs to a multiyear partnership without an exact public delivery schedule.
  • OpenAI’s five-gigawatt collaboration and earlier 10-gigawatt letter of intent are differently framed plans; do not add them together or translate them into a firm GPU count.
  • Before committing a workload, establish which accelerator-backed service, region, capacity allocation and commercial terms the provider is actually offering; none of these fleet-wide announcements supplies those project-level answers.

NVIDIA’s earnings and its partners’ disclosures point to an AI buildout increasingly planned as connected datacenters and power capacity, with Azure already operating a large Grace Blackwell fleet and several providers preparing Rubin deployments. The milestone for enterprise customers is more specific than another headline GPU count: it is when the announced hardware becomes available as a suitable, obtainable service for their workloads.