Nvidia has denied that Rubin Ultra is being cut back, telling DigiTimes that the next-generation AI accelerator will be offered through flexible SKU options rather than a single fixed configuration. The response addresses a specific and increasingly consequential report: that Nvidia has tested Rubin Ultra variants with dramatically less high-bandwidth memory, potentially replacing HBM4E with HBM4 and reducing capacity from the 1TB configuration it publicly showed in March.

The denial is important, but it does not settle the underlying issue. Nvidia has not released an updated Rubin Ultra product table, confirmed a final memory configuration, or said whether the four-compute-die package demonstrated at GTC 2026 remains the shipping design. “Flexible SKUs” can describe normal product segmentation; it can also describe a way to ship available hardware when a flagship configuration depends on memory and packaging capacity that are difficult to secure at volume.

For enterprise buyers, cloud customers and Microsoft’s AI infrastructure planners, the practical question is no longer whether Nvidia can attach a different amount of HBM to Rubin Ultra. It is whether the company can deliver the rack-scale performance and memory capacity customers used when planning deployments around the Kyber NVL144 platform.

A futuristic AI accelerator board dominates a glowing data center, comparing HBM4E and HBM4 memory systems.The dispute is about Rubin Ultra, not Vera Rubin shipping this year​

Nvidia’s public roadmap separates the Vera Rubin platform expected to reach partners in the second half of 2026 from the later Rubin Ultra generation. The distinction has been lost in some of the recent coverage, but it is central to reading Nvidia’s response correctly.

Nvidia has said Vera Rubin is in production and named AWS, Google Cloud, Microsoft, Oracle Cloud Infrastructure, CoreWeave, Lambda, Nebius and Nscale among the early deployment partners. Microsoft has separately been identified by Nvidia as a customer for Vera Rubin NVL72 rack-scale systems, including future Fairwater AI superfactory sites. Those systems use the standard Rubin generation and an NVL72 rack design.

Rubin Ultra was presented as the higher-density follow-on: a four-compute-chiplet accelerator with 1TB of HBM4E, deployed in a new Kyber rack holding 144 GPU packages. That combination — more compute per package, much more memory per package, and twice the package count in a rack — was the basis for Nvidia’s very large performance claims for the platform.

The current dispute concerns that later product. It should not be read as evidence that Nvidia’s 2026 Vera Rubin rollout has been cancelled or delayed. Nvidia’s own announcements still put partner availability of Rubin products in the second half of 2026. But it does mean that buyers should stop treating “Rubin” as a single immutable specification. The initial Vera Rubin systems and Rubin Ultra/Kyber systems face different engineering, supply and deployment risks.


A flexible SKU is not necessarily the original 1TB design​

DigiTimes reports that Nvidia rejects the characterization of a Rubin Ultra downgrade and points instead to SKU flexibility. In the ordinary sense, that is plausible. Nvidia has long sold multiple configurations across a platform generation, varying memory capacity, interconnect, form factor and power envelope for different customers.

What makes this case different is the size of the reported gap between the configuration Nvidia displayed and the alternatives now said to be under test. The Information, as summarized by Tom’s Hardware, reported that Nvidia has evaluated Rubin Ultra designs with 192GB or 256GB of memory, fewer than the 16 previously discussed HBM stacks, and HBM4 rather than HBM4E. Tom’s Hardware also reported that the tests were tied to concerns over obtaining enough of the more advanced HBM4E supply.

Those details remain reporting, not a confirmed Nvidia specification. Nvidia has not published a part number, a definitive stack count, bandwidth figure, board layout, or customer-availability date for any lower-memory Rubin Ultra variant. But the company’s wording matters: it denies a downgrade, not necessarily the existence of multiple designs.

That is a carefully drawn line. A product family can retain its launch schedule while shipping a lower-capacity initial SKU, reserving the full-memory part for later availability, or offering both products at once to different customers. Each outcome can fit under “flexible SKUs.” None automatically delivers the system Nvidia demonstrated at GTC 2026.

For customers, the difference is not cosmetic. Large language model training, long-context inference and multi-user inference services are constrained by memory capacity and bandwidth as much as raw accelerator compute. A lower-HBM Rubin Ultra could remain extremely fast in compute-bound workloads while handling smaller models, shorter contexts, fewer concurrent sessions or more aggressive model partitioning than customers expected from a 1TB accelerator.

HBM4E is the harder part of the story​

The rumored move from HBM4E to HBM4 would be more consequential than merely offering a lower-capacity model. Rubin’s base generation already depends on HBM4. Rubin Ultra was positioned around HBM4E, a more advanced version expected to provide faster signaling and a customizable base logic die.

Nvidia’s original Rubin Ultra presentation was ambitious because it paired four compute chiplets with 16 HBM4E stacks in one package. That combination demanded more than memory supply. It required sufficient yields across leading-edge GPU dies, stacked HBM, advanced packaging and the massive interposer or related packaging approach needed to bind the components together. Every weak link can constrain output.

Memory suppliers have been expanding HBM4 production for Vera Rubin. Nvidia and the major memory makers have publicly emphasized HBM4 supply progress, and Micron has announced volume HBM4 production intended for Vera Rubin. That is not the same thing as demonstrating adequate HBM4E supply for the later Ultra generation at the capacities Nvidia showed in March.

The significance of the DigiTimes account is therefore less about whether Nvidia can create a 192GB, 256GB or 1TB part. Nvidia obviously has the engineering resources to segment products. It is about whether SKU flexibility is being used to absorb a supply constraint that otherwise would force a delay, a lower-volume launch, or a reduced flagship configuration.

Nvidia’s statement does not answer that. The company has not disclosed which HBM4E suppliers are qualified for Rubin Ultra, how many stacks its production package will use, whether HBM4 models are fallback products or mainstream offerings, or what performance and capacity differences customers should expect between configurations.


The earlier four-die and Kyber reports are still unresolved​

This week’s memory dispute follows two earlier reports that Nvidia has not fully addressed. In late June, SemiAnalysis reported that Nvidia had abandoned the original four-die Rubin Ultra design for a smaller two-die package because of manufacturing execution concerns. Tom’s Hardware reported the claim at the time, while noting that Nvidia had not confirmed it.

In early July, SemiAnalysis also reported that Kyber NVL144 — the rack intended to house Rubin Ultra — could slip from 2027 into 2028 because of manufacturing problems involving its orthogonal backplane or PCB midplane. The design replaces vast cable harnesses with a large rigid board carrying the rack’s copper NVLink fabric. Nvidia replied to Tom’s Hardware with a short statement that its roadmap remained intact, but did not clarify whether that meant the original Kyber schedule and design were unchanged.

Those reports should not be promoted into facts simply because they are repeated. Both depend substantially on SemiAnalysis’s reporting, and Nvidia has not supplied enough public technical detail to independently verify the alleged two-die change, the claimed Kyber delay, or the reported cancellation of a stopgap NVL72x2 arrangement.

But the reports are no longer isolated speculation either. The newer memory reporting describes a separate potential pressure point: HBM4E availability. Nvidia’s “flexible SKU” response is the first public indication that it sees configurable product options as relevant to the Rubin Ultra discussion.

The company’s public position has become clear at the broad level — the roadmap is still intact. Its public record is still thin at the level that matters to procurement teams: what exact accelerator, memory capacity and rack architecture will be available on which date.

What Microsoft, cloud buyers and Windows AI teams should watch​

Microsoft’s interest in Vera Rubin makes this more than a semiconductor-industry argument. The company is among the cloud providers Nvidia has named for Vera Rubin deployments, and Rubin-derived capacity is expected to feed the next generation of Azure-hosted AI infrastructure. That does not mean a Rubin Ultra configuration change would translate into an immediate change to Windows PCs, Copilot availability or Azure pricing.

It could, however, affect the economics and scheduling of the infrastructure behind enterprise AI services. Cloud providers budget data center power, cooling, networking and rack space years ahead. They also size AI clusters around expected accelerator memory. If an early Rubin Ultra offering has less HBM per package than planned, a provider may need more accelerators for the same aggregate model capacity, use a different model-sharding strategy, or prioritize the highest-capacity hardware for only its largest customers.

Windows administrators and developers should also keep the product categories straight. Rubin Ultra is not a successor to an RTX workstation card and is not destined for a desktop Windows upgrade cycle. It is a data center accelerator and rack platform. The direct effect will be on cloud capacity, availability of GPU-backed virtual machines, AI service throughput and the price of training or hosting large models — especially where Microsoft, Nvidia Cloud Partners and enterprise AI vendors compete for the same hardware.

The next useful milestone is not another broad roadmap reassurance. It is a production-level Rubin Ultra disclosure: final compute-die count, HBM type, memory capacity, bandwidth, rack compatibility, named system builders and a customer shipment window. Until Nvidia publishes those details, its “flexible SKU” explanation should be read as a commitment to keep Rubin Ultra moving — **not confirmation that every customer will receive the 1TB HBM4E, four-die accelerator shown on stage in March.