Microsoft says it will deploy AMD’s Helios rack-scale AI infrastructure at scale across its data centers, making the next-generation platform part of both Azure’s own AI services and capacity offered to cloud customers. The commitment, announced July 20 alongside AMD, is significant because it moves AMD’s MI455X accelerator from a future product roadmap into a named hyperscale rollout at a time when Azure says demand still exceeds the compute capacity it can bring online.
According to Tom’s Hardware, Microsoft plans to use Helios for frontier-model workloads internally as well as Azure AI infrastructure customers, including AI labs training models and serving inference. The systems are also intended to underpin managed enterprise deployments through Microsoft Foundry, extending the announcement beyond raw GPU rental into Microsoft’s increasingly important AI platform stack.
Neither company disclosed a purchase price, power commitment, deployment region, VM pricing, or first-availability date for Azure customers. That omission matters: “at scale” is a substantial signal of intent, but it does not yet tell customers when they can reserve capacity or how it will compare commercially with existing Nvidia-based Azure infrastructure.
The headline hardware is AMD’s Helios reference design, a double-wide rack-scale system that combines 72 Instinct MI455X accelerators with sixth-generation AMD EPYC processors code-named Venice, Pensando networking hardware, and the ROCm software stack. AMD describes Helios as a design blueprint for system vendors rather than a single boxed server product, with volume deployments expected in the second half of 2026.
That distinction is important for Azure. Hyperscale AI is increasingly defined by whether thousands of accelerators, networking, cooling, power delivery, firmware, drivers, and job schedulers operate as one predictable platform. A powerful accelerator in isolation is not enough for training frontier models or handling distributed inference at high volume.
AMD lists up to 31 TB of HBM4 memory across a Helios rack, along with claimed aggregate performance of 1.4 exaFLOPS at FP8 and 2.9 exaFLOPS at FP4. Each MI455X is specified with 432 GB of HBM4 memory and up to 19.6 TB/s of memory bandwidth. Those are vendor performance claims, not independent Azure benchmarks, and real-world outcomes will depend heavily on model architecture, precision, software maturity, interconnect behavior, and the proportion of time spent moving data rather than computing.
The more consequential figures may be the communication specifications. AMD says Helios targets 260 TB/s of scale-up bandwidth inside the rack using UALink over Ethernet, plus 43 TB/s of scale-out bandwidth between racks using Pensando networking. Large-model training and high-throughput inference are communication problems as much as they are GPU problems, so Azure’s ability to turn those figures into consistently usable cluster performance will determine whether Helios is a credible alternative for demanding workloads.
That is a notable expansion of the AMD relationship. AI services need CPUs for data preparation, orchestration, vector databases, storage pipelines, networking control planes, simulation, and inference tasks that do not belong on expensive accelerators. A cloud provider that can package CPU and GPU capacity around one platform has more freedom to tune performance, availability, and cost across different customer workloads.
Microsoft will additionally use its existing Pensando DPU deployment in Azure Boost, Microsoft’s infrastructure offload architecture for networking and storage. AMD’s Helios design includes Pensando Vulcano AI NICs for scale-out traffic and Salina DPUs for front-end networking, storage, and security services. The announced Azure Boost integration therefore connects a future AI rack to infrastructure Microsoft is already building into the cloud rather than treating Helios as an isolated accelerator island.
For Windows-focused IT teams, the immediate effect is indirect but real. Microsoft Foundry, Copilot services, Azure AI workloads, and the broader Microsoft cloud ecosystem all depend on available, economical compute. More supplier diversity at the hardware layer could eventually improve capacity access and reduce dependence on a single accelerator roadmap, even if end users never see “MI455X” in a portal.
Microsoft also reported that roughly two-thirds of its fiscal third-quarter capital spending went to short-lived assets, primarily GPUs and CPUs. This is not an experimental procurement cycle. The company needs enormous quantities of compute for first-party products, research and development, OpenAI-related requirements, and customers buying Azure AI services.
Microsoft’s April update on its OpenAI relationship adds further context. Microsoft remains OpenAI’s primary cloud partner, with OpenAI products shipping first on Azure unless Microsoft cannot or elects not to support the required capabilities. The amended arrangement grants both companies more flexibility, but it does not reduce Microsoft’s need to keep adding AI infrastructure rapidly.
AMD, meanwhile, needs wins that prove it can sell an integrated platform rather than only individual accelerator cards. Helios has already appeared in AMD’s plans with Meta, TCS, system builders, and manufacturing partners. Microsoft brings a different validation: a major public-cloud operator that must convert the hardware into durable, supportable services for external customers.
For customers, that promise has appeal. It suggests that a model and deployment workflow may be less tightly bound to a single vendor’s hardware and software ecosystem. It also gives Microsoft a potential way to diversify supply while retaining influence over the software, networking, and operations layers that turn accelerators into an Azure service.
But software compatibility is not the same as performance parity or operational parity. Enterprises migrating CUDA-tuned code, custom kernels, distributed training recipes, monitoring integrations, and inference stacks will need evidence that their workloads behave predictably on ROCm. Cloud customers will also expect mature images, drivers, SDKs, orchestration options, observability, support commitments, and clear service-level expectations—not merely theoretical framework support.
Microsoft’s involvement can help close that gap because a hyperscaler has strong incentives to harden the tools it exposes. Still, the measure of success will be customer deployments, published benchmarks, usable VM configurations, and the ability to obtain capacity without an extended wait.
The Venice-based HDv2 and HXv2 series likewise need fuller specifications. Agentic AI and data pipelines can be broad categories, while semiconductor design often imposes unusually demanding requirements around memory capacity, low-latency networking, EDA software certification, and licensing. Naming the series is a start; documenting their CPU counts, memory configurations, storage, networking, availability zones, and pricing will determine their practical value.
AMD’s Advancing AI event on July 22 and July 23 is the most immediate milestone for additional technical details. Until then, Microsoft’s Helios commitment is best understood as a major supply and platform announcement—not yet a new Azure instance type customers can deploy.
The report also says HXv2 will target EDA, simulation, and engineering workloads with 176 Venice cores running above 5GHz, nearly 4TB of memory, and 800Gbps InfiniBand networking. Microsoft has still not announced Azure regions, pricing, or customer availability for either series or the MI455X Helios capacity.
The report also adds that HXv2 will offer configurations with nearly 2TB or 4TB of memory, alongside up to 176 Venice EPYC cores above 5GHz and 800Gbps InfiniBand.
www.neowin.net
According to Tom’s Hardware, Microsoft plans to use Helios for frontier-model workloads internally as well as Azure AI infrastructure customers, including AI labs training models and serving inference. The systems are also intended to underpin managed enterprise deployments through Microsoft Foundry, extending the announcement beyond raw GPU rental into Microsoft’s increasingly important AI platform stack.
Neither company disclosed a purchase price, power commitment, deployment region, VM pricing, or first-availability date for Azure customers. That omission matters: “at scale” is a substantial signal of intent, but it does not yet tell customers when they can reserve capacity or how it will compare commercially with existing Nvidia-based Azure infrastructure.
Helios Is a Rack Design, Not Simply a New GPU Instance
The headline hardware is AMD’s Helios reference design, a double-wide rack-scale system that combines 72 Instinct MI455X accelerators with sixth-generation AMD EPYC processors code-named Venice, Pensando networking hardware, and the ROCm software stack. AMD describes Helios as a design blueprint for system vendors rather than a single boxed server product, with volume deployments expected in the second half of 2026.That distinction is important for Azure. Hyperscale AI is increasingly defined by whether thousands of accelerators, networking, cooling, power delivery, firmware, drivers, and job schedulers operate as one predictable platform. A powerful accelerator in isolation is not enough for training frontier models or handling distributed inference at high volume.
AMD lists up to 31 TB of HBM4 memory across a Helios rack, along with claimed aggregate performance of 1.4 exaFLOPS at FP8 and 2.9 exaFLOPS at FP4. Each MI455X is specified with 432 GB of HBM4 memory and up to 19.6 TB/s of memory bandwidth. Those are vendor performance claims, not independent Azure benchmarks, and real-world outcomes will depend heavily on model architecture, precision, software maturity, interconnect behavior, and the proportion of time spent moving data rather than computing.
The more consequential figures may be the communication specifications. AMD says Helios targets 260 TB/s of scale-up bandwidth inside the rack using UALink over Ethernet, plus 43 TB/s of scale-out bandwidth between racks using Pensando networking. Large-model training and high-throughput inference are communication problems as much as they are GPU problems, so Azure’s ability to turn those figures into consistently usable cluster performance will determine whether Helios is a credible alternative for demanding workloads.
Microsoft Is Buying a Full AMD Stack
Microsoft’s deployment is not limited to accelerators. The companies also said Azure will introduce two VM series based on the forthcoming Venice EPYC CPUs: HDv2, intended for agentic AI and data-pipeline work, and HXv2, targeted at semiconductor design workflows.That is a notable expansion of the AMD relationship. AI services need CPUs for data preparation, orchestration, vector databases, storage pipelines, networking control planes, simulation, and inference tasks that do not belong on expensive accelerators. A cloud provider that can package CPU and GPU capacity around one platform has more freedom to tune performance, availability, and cost across different customer workloads.
Microsoft will additionally use its existing Pensando DPU deployment in Azure Boost, Microsoft’s infrastructure offload architecture for networking and storage. AMD’s Helios design includes Pensando Vulcano AI NICs for scale-out traffic and Salina DPUs for front-end networking, storage, and security services. The announced Azure Boost integration therefore connects a future AI rack to infrastructure Microsoft is already building into the cloud rather than treating Helios as an isolated accelerator island.
For Windows-focused IT teams, the immediate effect is indirect but real. Microsoft Foundry, Copilot services, Azure AI workloads, and the broader Microsoft cloud ecosystem all depend on available, economical compute. More supplier diversity at the hardware layer could eventually improve capacity access and reduce dependence on a single accelerator roadmap, even if end users never see “MI455X” in a portal.
Azure’s Capacity Problem Explains the Timing
Microsoft’s most recent earnings call laid out why a platform deal such as this matters. The company said Azure demand continued to exceed available capacity, even as it accelerated infrastructure delivery. It expected to spend roughly $190 billion in calendar 2026 capital expenditures, including higher component costs, and said it would remain capacity-constrained at least through the end of the year.Microsoft also reported that roughly two-thirds of its fiscal third-quarter capital spending went to short-lived assets, primarily GPUs and CPUs. This is not an experimental procurement cycle. The company needs enormous quantities of compute for first-party products, research and development, OpenAI-related requirements, and customers buying Azure AI services.
Microsoft’s April update on its OpenAI relationship adds further context. Microsoft remains OpenAI’s primary cloud partner, with OpenAI products shipping first on Azure unless Microsoft cannot or elects not to support the required capabilities. The amended arrangement grants both companies more flexibility, but it does not reduce Microsoft’s need to keep adding AI infrastructure rapidly.
AMD, meanwhile, needs wins that prove it can sell an integrated platform rather than only individual accelerator cards. Helios has already appeared in AMD’s plans with Meta, TCS, system builders, and manufacturing partners. Microsoft brings a different validation: a major public-cloud operator that must convert the hardware into durable, supportable services for external customers.
Openness Is the Pitch; Software Is the Test
AMD’s competitive case rests heavily on open standards. Helios is based on the Open Compute Project’s Open Rack Wide form factor and uses UALink and Ultra Ethernet Consortium-oriented networking rather than a wholly proprietary rack fabric. AMD also positions ROCm as an open software environment that supports frameworks and tools including PyTorch, TensorFlow, JAX, vLLM, Triton, and ONNX Runtime.For customers, that promise has appeal. It suggests that a model and deployment workflow may be less tightly bound to a single vendor’s hardware and software ecosystem. It also gives Microsoft a potential way to diversify supply while retaining influence over the software, networking, and operations layers that turn accelerators into an Azure service.
But software compatibility is not the same as performance parity or operational parity. Enterprises migrating CUDA-tuned code, custom kernels, distributed training recipes, monitoring integrations, and inference stacks will need evidence that their workloads behave predictably on ROCm. Cloud customers will also expect mature images, drivers, SDKs, orchestration options, observability, support commitments, and clear service-level expectations—not merely theoretical framework support.
Microsoft’s involvement can help close that gap because a hyperscaler has strong incentives to harden the tools it exposes. Still, the measure of success will be customer deployments, published benchmarks, usable VM configurations, and the ability to obtain capacity without an extended wait.
The First Public Details Will Need to Answer Operational Questions
The companies have established the strategic direction, but Azure users now need the operational details. Microsoft has not identified the regions that will host Helios, the Azure VM families carrying MI455X capacity, whether access will start as private preview, or whether the systems will initially be reserved for major model providers and Microsoft’s own services.The Venice-based HDv2 and HXv2 series likewise need fuller specifications. Agentic AI and data pipelines can be broad categories, while semiconductor design often imposes unusually demanding requirements around memory capacity, low-latency networking, EDA software certification, and licensing. Naming the series is a start; documenting their CPU counts, memory configurations, storage, networking, availability zones, and pricing will determine their practical value.
AMD’s Advancing AI event on July 22 and July 23 is the most immediate milestone for additional technical details. Until then, Microsoft’s Helios commitment is best understood as a major supply and platform announcement—not yet a new Azure instance type customers can deploy.
Update: Additional details (July 20, 2026)
Neowin reports that the planned Venice-based HDv2 VM will offer nearly 500 physical EPYC cores, 4TB of RAM, 32TB of local NVMe storage, and 400Gbps Azure Boost networking. Microsoft is positioning it for data preparation, search, reinforcement learning, and agent coordination—not just GPU-adjacent work.The report also says HXv2 will target EDA, simulation, and engineering workloads with 176 Venice cores running above 5GHz, nearly 4TB of memory, and 800Gbps InfiniBand networking. Microsoft has still not announced Azure regions, pricing, or customer availability for either series or the MI455X Helios capacity.
AMD's Helios platform ships H2 2026 with Microsoft Azure as first major deployment
AMD will ship its Helios rack-scale AI platform in H2 2026, delivering 1.4 exaFLOPS per rack with Microsoft Azure, Meta, and Oracle as early deployers.
cryptobriefing.com
Update: Additional details (July 20, 2026)
Techgenyz reports that Microsoft has named the Helios-backed Azure GPU family: Azure ND MI455X v7. The planned instances are positioned for production-scale AI inference, including reasoning, search, and agentic workloads. This fills in the customer-facing VM-family detail that was not included in the initial announcement, although Microsoft still has not provided regions, pricing, configuration sizes, or an availability date.The report also adds that HXv2 will offer configurations with nearly 2TB or 4TB of memory, alongside up to 176 Venice EPYC cores above 5GHz and 800Gbps InfiniBand.
Microsoft and AMD expand partnership to deploy ... | Pluang
Microsoft and AMD have extended their strategic partnership to deploy AMD's next-generation Helios AI infrastructure on Microsoft Azure. This includes rolling out AMD Helios rack-scale systems designed for large-scale AI inference workloads, combining AMD's latest GPUs, processors...
pluang.com
Update: Additional details (July 20, 2026)
Microsoft’s description of HXv2 adds that its 176-core Venice EPYC configuration will include 3D V-Cache and support 800Gbps InfiniBand for distributed MPI workloads. The series remains aimed at EDA, scientific simulation, and other technical-computing deployments.Update: HXv2 adds larger per-core cache claim (July 20, 2026)
Microsoft says Azure HXv2 will provide 50% more addressable cache per core than the prior HX generation, alongside its 176-core Venice EPYC configuration, 3D V-Cache, up to 4TB of memory, and 800Gbps InfiniBand.Microsoft expands Azure AI infrastructure with AMD's next-generation GPUs and CPUs - Neowin
Alongside its existing NVIDIA systems and proprietary Maia silicon, Microsoft is positioning Azure as a highly flexible AI cloud environment with the new expanded partnership with AMD.
References
- Primary source: Tom's Hardware
Published: 2026-07-20T13:05:00+00:00
Microsoft will deploy AMD’s Helios rack-scale AI accelerator ‘at scale’ on Azure – Radeon Instinct MI455X and Epyc Venice power will be available through Redmond’s cloud infrastructure | Tom's Hardware
But it’s not clear just how much AMD AI compute Microsoft is buyingwww.tomshardware.com - Related coverage: ir.amd.com
AMD and TCS to bring state-of-the-art ‘Helios’ rack-scale AI architecture to India :: Advanced Micro Devices, Inc. (AMD)
News Highlights: Enterprises across India will gain access to a new 200MW deployment of the AMD “Helios” rack-scale AI architecture, supporting…...ir.amd.com - Related coverage: supermicro.com
- Related coverage: d1io3yog0oux5.cloudfront.net
AMD and TCS to bring state-of-the-art Helios rack-scale AI architecture to India
PDF documentd1io3yog0oux5.cloudfront.net
- Related coverage: techradar.com
AMD details Instinct MI500 architecture and memory plans ahead of the AI accelerator's 2027 debut | TechRadar
AMD MI500 is planned around AMD’s CDNA 6 architecture with 2nm manufacturing and HBM4E memorywww.techradar.com - Official source: news.microsoft.com
Last edited:






