Nvidia’s Vera Rubin production ramp is reshaping the AI infrastructure conversation from a narrow race for faster GPUs into a broader contest over tokens, watts, racks, networks, cooling loops, and grid access. The key performance claim around Vera Rubin NVL72 is not merely that the new hardware executes more AI operations; it is that an integrated system can produce dramatically more useful inference output within a fixed power envelope—a shift that could redefine how cloud providers, data-center operators, and investors judge the economics of artificial intelligence infrastructure. Nvidia’s July 21 announcement says production is ramping with systems operating at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius.
That distinction matters because the AI industry is increasingly running into constraints that cannot be solved by adding accelerators alone. Electricity availability, transformer capacity, data-center construction schedules, liquid-cooling systems, networking latency, and the ability to manufacture complete racks at volume have become central to deployment. Vera Rubin’s arrival therefore strengthens Nvidia’s position—but it also expands the set of infrastructure layers that determine whether AI capacity can be installed, powered, and monetized.
For much of the generative AI boom, the investment narrative was easy to state: buy the companies supplying the fastest AI accelerators. Nvidia became the defining beneficiary because its GPUs, software ecosystem, and high-speed interconnect technology formed the foundation of many frontier-model training and inference clusters.
That narrative remains partly true. Nvidia still occupies the most strategically valuable point in the AI compute stack, because it provides the core accelerator architecture, system-level networking, and software needed to turn hardware into productive AI capacity. But the practical unit of deployment is changing from an individual GPU or server into a rack-scale AI factory building block.
Vera Rubin NVL72 illustrates that transition. CoreWeave describes each system as containing 72 Rubin GPUs and 36 Vera CPUs, connected through a sixth-generation NVLink fabric delivering 260 TB/s of bandwidth across the rack. CoreWeave’s production-validation announcement frames the platform as a complete compute system rather than an assortment of independently optimized chips.
This is a meaningful evolution. An AI rack cannot deliver its advertised performance if memory bandwidth becomes a bottleneck, if the network leaves accelerators waiting for data, if power delivery cannot support the load, or if cooling infrastructure forces lower operating limits. In a tightly coupled system, performance depends on the weakest link.
The broader electricity backdrop makes that system-level focus especially important. The International Energy Agency estimates that global data-center electricity use reached roughly 415 TWh in 2024 and projects consumption to rise to around 945 TWh by 2030, with AI acting as the principal driver of growth. The agency also warns that AI-focused data centers can have electricity requirements comparable to energy-intensive industrial facilities, while being concentrated in a relatively small number of locations. The IEA’s analysis of energy and AI makes clear why access to power is becoming a strategic asset rather than a background operating expense.
For cloud providers selling inference, this may be more economically useful than quoting peak floating-point performance. A customer does not buy FLOPS in isolation; they buy an AI service that responds quickly, handles sufficient context, meets quality requirements, and does so at an acceptable cost. The operator, in turn, has to fit that service inside constrained power, cooling, networking, and real-estate budgets.
CoreWeave’s initial measured-silicon comparison ran the DeepSeek-R1 reasoning model on Vera Rubin NVL72 and Nvidia GB200 NVL72 systems. At matched interactivity targets—measured as tokens per second per user—CoreWeave reported that Vera Rubin produced 10 times more tokens per second per megawatt on the same workload. CoreWeave’s benchmark write-up says the test enabled large-scale expert parallelism, NVFP4 precision, multi-token prediction, and disaggregated prefill and decode through Nvidia TensorRT-LLM and Nvidia Dynamo.
That is an important qualification. The 10x result is not a universal statement that every AI model, precision format, context length, software stack, or deployment will perform exactly ten times better. It is a benchmark outcome from a specific reasoning workload and serving configuration. The comparison remains highly relevant, but it should be read as evidence of the industry’s new optimization direction rather than a blanket multiplier for all AI use cases.
Nvidia’s own product material similarly labels relevant large-language-model inference performance as subject to change and ties its token-cost comparison to a particular model configuration. The company says Vera Rubin NVL72 can achieve one-tenth the cost per million tokens relative to GB200 NVL72 in highly interactive deep-reasoning workloads, while also claiming up to 10x more tokens per megawatt. Nvidia’s Vera Rubin NVL72 overview provides useful context, but the model-specific caveats are as important as the headline numbers.
CoreWeave argues that a reasoning model such as DeepSeek-R1 places particular importance on interactivity because every token must move through the memory system and, at scale, across the NVLink domain. Its Vera Rubin analysis describes this as a workload where responsiveness and system throughput can matter more than raw compute in isolation.
That changes the equation for AI operators. A system with more theoretical compute but poor data movement, insufficient memory bandwidth, or congested networking may deliver weaker real-world service economics than a more balanced architecture. The relevant objective becomes useful token output at a target latency and quality level, not maximum component-level speed.
For investors, the implication is straightforward: the value chain expands. The GPU remains central, but it is no longer the only component capable of limiting the monetization of an AI deployment.
The distinction between a component and a platform is commercially meaningful. In a traditional server supply chain, an enterprise or systems integrator may select processors, memory, networking, storage, and cooling from multiple vendors, then assemble a solution. Nvidia’s rack-scale strategy increasingly moves the design center toward a pre-validated, vertically integrated architecture.
At the same time, the Vera CPU is not simply a host processor in the conventional sense. Nvidia positions it for data movement and agentic reasoning workloads, where low-latency coordination, deterministic behavior, and energy-efficient compute can have an outsized effect on overall system utilization. Nvidia’s Vera Rubin platform page presents the CPU and GPU as parts of a shared infrastructure design rather than separate products.
This is why Nvidia may capture more value from each deployed AI system than in earlier accelerator cycles. The company is selling a growing share of the critical compute, networking, software, and architecture around the GPU. That does not eliminate the role of external suppliers, but it may reduce the extent to which customers can freely substitute components in the highest-performance configurations.
Nvidia says its NVLink 6 switches provide 3.6 TB/s of all-to-all scale-up bandwidth per GPU, while ConnectX-9 SuperNICs deliver 1.6 Tb/s of per-GPU bandwidth for GPU-direct networking. The Vera Rubin NVL72 specification page identifies these links as fundamental to keeping large pools of accelerators working efficiently together.
The scale-out side matters just as much. Nvidia argues that conventional Ethernet was created primarily for enterprise-oriented north-south traffic, not the synchronized collective communications patterns common in giant AI clusters. Spectrum-6 and ConnectX-9 are designed as part of its Spectrum-X architecture to address that challenge, including support for both pluggable and co-packaged optics as well as liquid cooling. Nvidia’s Spectrum-6 overview outlines the company’s attempt to turn Ethernet from a commodity layer into a differentiated AI infrastructure product.
This does not mean every AI deployment will require Nvidia’s entire networking stack. Many enterprises will continue using mixed environments, conventional Ethernet designs, InfiniBand configurations, and specialized networking approaches. But at the gigascale end of the market, an integrated network can materially affect the return on expensive accelerator capital.
Nvidia says Vera Rubin spans more than 350 factory sites in 30 countries, calling it its largest and most mature rack-scale supply chain. The company’s July production-ramp announcement is notable not just for the deployment list but for what it reveals about manufacturing complexity. A platform at this scale requires coordination across silicon, substrates, memory, packaging, power, liquid cooling, networking, optics, mechanical assembly, validation, and field deployment.
For the supply chain, that raises the importance of memory manufacturing yields, advanced packaging capacity, substrate availability, and co-design between logic and memory. The opportunity is real, but it is not frictionless. Advanced packages are difficult to manufacture, involve complex thermal and electrical requirements, and can create bottlenecks even when demand for the underlying GPU is strong.
The key investor distinction is between content growth and profit capture. A component may become more essential to each system while its supplier still faces pricing pressure, capacity constraints, customer concentration, or higher capital spending. Rising technical importance does not automatically translate into superior margins.
This creates potential demand for optical transceivers, photonic integration, switch components, cables, connectors, and specialized networking equipment. Yet this category also deserves disciplined analysis. The supply base is broad, technology transitions can arrive quickly, and Nvidia’s increasing system-level influence may limit the pricing freedom of some downstream vendors.
In other words, the most attractive connectivity exposure may not necessarily be the company that ships the highest number of optical parts. It may be the supplier with a difficult-to-replace technology, validated qualification status, strong manufacturing capacity, or a position inside a higher-value portion of the design.
The IEA’s data-center outlook explains why this layer is attracting so much attention. It projects that data-center electricity demand will more than double by 2030 and notes that AI-focused facilities are geographically concentrated enough to create major local impacts even when they represent a smaller share of total national consumption. The IEA’s Energy and AI report identifies power availability as a decisive constraint on AI expansion.
This is where the “tokens per megawatt” metric becomes more than a technical benchmark. If a cloud provider has secured a fixed power allocation, a more efficient AI platform can increase revenue-generating output without waiting for a new substation, transmission upgrade, generator, or data-center campus. That can make performance per watt economically valuable even if the platform itself carries a significant upfront price.
The counterpoint is important: higher rack density can shift, rather than eliminate, infrastructure challenges. A facility may need less energy for a given amount of inference output, yet its individual racks can still require much more concentrated power and cooling capability. Grid constraints, construction timelines, and electrical-equipment lead times remain material risks.
The U.S. Department of Energy’s data-center design guidance describes direct-liquid cooling as including cold-plate and immersion approaches, and notes that high-performance computing environments were early adopters because of rising rack power densities. The DOE’s energy-efficient data-center guide provides useful context for why cooling is not a peripheral concern in AI infrastructure.
Nvidia’s own network roadmap reinforces that trend: Spectrum-6 is designed to support liquid cooling, allowing cooling considerations to extend across the AI factory rather than stopping at the compute tray. Nvidia’s Spectrum-6 material signals an increasingly holistic approach in which compute, network, and thermal design are jointly optimized.
For suppliers, cooling is attractive because it represents a growing share of infrastructure complexity. But it is also a field where execution matters intensely. Reliability failures, leaks, maintenance demands, water constraints, and integration problems can quickly offset theoretical efficiency gains. The winners will likely be firms that can prove performance in production installations rather than merely market a high-density cooling concept.
That brings several advantages:
Real deployments vary by model architecture, context window, batching strategy, precision, latency target, utilization rate, and software maturity. Investors should therefore treat tokens per megawatt as a useful operational metric, not a shortcut that can replace workload-level diligence.
That makes “picks and shovels” investing more nuanced than it appears. A supplier may enjoy strong AI demand yet still deliver disappointing financial results if content is commoditized, margins are competed away, or capital expenditure rises faster than returns.
A more efficient rack can improve the value of a fixed megawatt. It cannot guarantee that the next megawatt will be available where a cloud provider needs it.
Nvidia remains the obvious central beneficiary because it owns the platform architecture across compute, networking, and software. Yet the Vera Rubin production ramp also elevates the importance of HBM4 memory, advanced packaging, high-speed interconnects, optics, liquid cooling, electrical distribution, and rack-level manufacturing.
The most important investment question is therefore not simply who sells into AI. It is which companies provide the components and infrastructure that become more essential, more difficult to replace, and more valuable per deployed megawatt as AI factories move from GPU clusters toward fully integrated rack-scale systems.
That is why Vera Rubin changes the AI trade. It does not make the GPU less important. It makes everything required to keep that GPU productive far more important than before.
That distinction matters because the AI industry is increasingly running into constraints that cannot be solved by adding accelerators alone. Electricity availability, transformer capacity, data-center construction schedules, liquid-cooling systems, networking latency, and the ability to manufacture complete racks at volume have become central to deployment. Vera Rubin’s arrival therefore strengthens Nvidia’s position—but it also expands the set of infrastructure layers that determine whether AI capacity can be installed, powered, and monetized.
Background: Why AI Infrastructure Is No Longer a GPU-Only Story
For much of the generative AI boom, the investment narrative was easy to state: buy the companies supplying the fastest AI accelerators. Nvidia became the defining beneficiary because its GPUs, software ecosystem, and high-speed interconnect technology formed the foundation of many frontier-model training and inference clusters.That narrative remains partly true. Nvidia still occupies the most strategically valuable point in the AI compute stack, because it provides the core accelerator architecture, system-level networking, and software needed to turn hardware into productive AI capacity. But the practical unit of deployment is changing from an individual GPU or server into a rack-scale AI factory building block.
Vera Rubin NVL72 illustrates that transition. CoreWeave describes each system as containing 72 Rubin GPUs and 36 Vera CPUs, connected through a sixth-generation NVLink fabric delivering 260 TB/s of bandwidth across the rack. CoreWeave’s production-validation announcement frames the platform as a complete compute system rather than an assortment of independently optimized chips.
This is a meaningful evolution. An AI rack cannot deliver its advertised performance if memory bandwidth becomes a bottleneck, if the network leaves accelerators waiting for data, if power delivery cannot support the load, or if cooling infrastructure forces lower operating limits. In a tightly coupled system, performance depends on the weakest link.
The broader electricity backdrop makes that system-level focus especially important. The International Energy Agency estimates that global data-center electricity use reached roughly 415 TWh in 2024 and projects consumption to rise to around 945 TWh by 2030, with AI acting as the principal driver of growth. The agency also warns that AI-focused data centers can have electricity requirements comparable to energy-intensive industrial facilities, while being concentrated in a relatively small number of locations. The IEA’s analysis of energy and AI makes clear why access to power is becoming a strategic asset rather than a background operating expense.
The New Metric: Tokens Per Megawatt
The most consequential Vera Rubin claim is its focus on tokens per megawatt. That measure asks a practical question: how much AI output can an operator deliver from a fixed amount of electrical infrastructure?For cloud providers selling inference, this may be more economically useful than quoting peak floating-point performance. A customer does not buy FLOPS in isolation; they buy an AI service that responds quickly, handles sufficient context, meets quality requirements, and does so at an acceptable cost. The operator, in turn, has to fit that service inside constrained power, cooling, networking, and real-estate budgets.
CoreWeave’s initial measured-silicon comparison ran the DeepSeek-R1 reasoning model on Vera Rubin NVL72 and Nvidia GB200 NVL72 systems. At matched interactivity targets—measured as tokens per second per user—CoreWeave reported that Vera Rubin produced 10 times more tokens per second per megawatt on the same workload. CoreWeave’s benchmark write-up says the test enabled large-scale expert parallelism, NVFP4 precision, multi-token prediction, and disaggregated prefill and decode through Nvidia TensorRT-LLM and Nvidia Dynamo.
That is an important qualification. The 10x result is not a universal statement that every AI model, precision format, context length, software stack, or deployment will perform exactly ten times better. It is a benchmark outcome from a specific reasoning workload and serving configuration. The comparison remains highly relevant, but it should be read as evidence of the industry’s new optimization direction rather than a blanket multiplier for all AI use cases.
Nvidia’s own product material similarly labels relevant large-language-model inference performance as subject to change and ties its token-cost comparison to a particular model configuration. The company says Vera Rubin NVL72 can achieve one-tenth the cost per million tokens relative to GB200 NVL72 in highly interactive deep-reasoning workloads, while also claiming up to 10x more tokens per megawatt. Nvidia’s Vera Rubin NVL72 overview provides useful context, but the model-specific caveats are as important as the headline numbers.
Why Tokens Matter More for Agentic AI
The transition toward agentic AI makes token economics especially significant. Traditional chatbot interactions may involve a relatively limited prompt and response. Reasoning systems, coding agents, research agents, and tool-using enterprise workflows can generate many more intermediate tokens as they plan, call tools, inspect results, revise outputs, and continue across multi-step tasks.CoreWeave argues that a reasoning model such as DeepSeek-R1 places particular importance on interactivity because every token must move through the memory system and, at scale, across the NVLink domain. Its Vera Rubin analysis describes this as a workload where responsiveness and system throughput can matter more than raw compute in isolation.
That changes the equation for AI operators. A system with more theoretical compute but poor data movement, insufficient memory bandwidth, or congested networking may deliver weaker real-world service economics than a more balanced architecture. The relevant objective becomes useful token output at a target latency and quality level, not maximum component-level speed.
For investors, the implication is straightforward: the value chain expands. The GPU remains central, but it is no longer the only component capable of limiting the monetization of an AI deployment.
Vera Rubin Is a Rack-Scale Platform, Not Just a New Accelerator
Nvidia describes Vera Rubin as a platform built around multiple chips and rack components engineered to operate together. Its architecture incorporates the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet switch, alongside the associated software and physical infrastructure. Nvidia’s Spectrum-6 announcement explicitly positions those elements as part of one coordinated AI-factory design.The distinction between a component and a platform is commercially meaningful. In a traditional server supply chain, an enterprise or systems integrator may select processors, memory, networking, storage, and cooling from multiple vendors, then assemble a solution. Nvidia’s rack-scale strategy increasingly moves the design center toward a pre-validated, vertically integrated architecture.
The Compute Layer Remains Nvidia’s Core Advantage
The Rubin GPU remains the most visible part of the system. Nvidia says it combines HBM4 memory with a 50 PF NVFP4 Transformer Engine, aiming at the precision formats and memory behavior required by next-generation AI models. Nvidia’s platform specifications underscore the degree to which inference performance is tied to both arithmetic capability and memory architecture.At the same time, the Vera CPU is not simply a host processor in the conventional sense. Nvidia positions it for data movement and agentic reasoning workloads, where low-latency coordination, deterministic behavior, and energy-efficient compute can have an outsized effect on overall system utilization. Nvidia’s Vera Rubin platform page presents the CPU and GPU as parts of a shared infrastructure design rather than separate products.
This is why Nvidia may capture more value from each deployed AI system than in earlier accelerator cycles. The company is selling a growing share of the critical compute, networking, software, and architecture around the GPU. That does not eliminate the role of external suppliers, but it may reduce the extent to which customers can freely substitute components in the highest-performance configurations.
NVLink and Ethernet Are Becoming Performance Products
Networking used to be treated as supporting infrastructure. In large AI clusters, it has become a first-order performance determinant.Nvidia says its NVLink 6 switches provide 3.6 TB/s of all-to-all scale-up bandwidth per GPU, while ConnectX-9 SuperNICs deliver 1.6 Tb/s of per-GPU bandwidth for GPU-direct networking. The Vera Rubin NVL72 specification page identifies these links as fundamental to keeping large pools of accelerators working efficiently together.
The scale-out side matters just as much. Nvidia argues that conventional Ethernet was created primarily for enterprise-oriented north-south traffic, not the synchronized collective communications patterns common in giant AI clusters. Spectrum-6 and ConnectX-9 are designed as part of its Spectrum-X architecture to address that challenge, including support for both pluggable and co-packaged optics as well as liquid cooling. Nvidia’s Spectrum-6 overview outlines the company’s attempt to turn Ethernet from a commodity layer into a differentiated AI infrastructure product.
This does not mean every AI deployment will require Nvidia’s entire networking stack. Many enterprises will continue using mixed environments, conventional Ethernet designs, InfiniBand configurations, and specialized networking approaches. But at the gigascale end of the market, an integrated network can materially affect the return on expensive accelerator capital.
The Supply-Chain Implications Go Well Beyond GPUs
The Vera Rubin ramp creates a broader set of beneficiaries than a simple GPU shipment cycle. The challenge is separating the suppliers with genuine rising content per rack from those merely adjacent to the narrative.Nvidia says Vera Rubin spans more than 350 factory sites in 30 countries, calling it its largest and most mature rack-scale supply chain. The company’s July production-ramp announcement is notable not just for the deployment list but for what it reveals about manufacturing complexity. A platform at this scale requires coordination across silicon, substrates, memory, packaging, power, liquid cooling, networking, optics, mechanical assembly, validation, and field deployment.
HBM4 and Advanced Packaging
High-bandwidth memory is essential to the design. Nvidia specifically identifies the Rubin GPU as an HBM4-based product, and that matters because AI models often need to move enormous quantities of weights, activations, and key-value cache data. Nvidia’s Vera Rubin platform materials put memory bandwidth alongside compute capability, reinforcing that the two are inseparable for advanced inference.For the supply chain, that raises the importance of memory manufacturing yields, advanced packaging capacity, substrate availability, and co-design between logic and memory. The opportunity is real, but it is not frictionless. Advanced packages are difficult to manufacture, involve complex thermal and electrical requirements, and can create bottlenecks even when demand for the underlying GPU is strong.
The key investor distinction is between content growth and profit capture. A component may become more essential to each system while its supplier still faces pricing pressure, capacity constraints, customer concentration, or higher capital spending. Rising technical importance does not automatically translate into superior margins.
Optical Components and High-Speed Connectivity
As AI clusters expand beyond a single rack, optical connectivity becomes increasingly relevant. Spectrum-6’s support for pluggable and co-packaged optics illustrates the direction of travel: network power efficiency and bandwidth density are becoming part of the compute economics. Nvidia’s Spectrum-6 announcement connects optics directly to end-to-end AI-factory cooling and efficiency goals.This creates potential demand for optical transceivers, photonic integration, switch components, cables, connectors, and specialized networking equipment. Yet this category also deserves disciplined analysis. The supply base is broad, technology transitions can arrive quickly, and Nvidia’s increasing system-level influence may limit the pricing freedom of some downstream vendors.
In other words, the most attractive connectivity exposure may not necessarily be the company that ships the highest number of optical parts. It may be the supplier with a difficult-to-replace technology, validated qualification status, strong manufacturing capacity, or a position inside a higher-value portion of the design.
Power Delivery Is Becoming a Competitive Constraint
A rack-scale AI system is also a power-delivery system. The larger the compute density, the more demanding the requirements become for electrical distribution, voltage conversion, protection, backup systems, switchgear, transformers, and site-level grid interconnection.The IEA’s data-center outlook explains why this layer is attracting so much attention. It projects that data-center electricity demand will more than double by 2030 and notes that AI-focused facilities are geographically concentrated enough to create major local impacts even when they represent a smaller share of total national consumption. The IEA’s Energy and AI report identifies power availability as a decisive constraint on AI expansion.
This is where the “tokens per megawatt” metric becomes more than a technical benchmark. If a cloud provider has secured a fixed power allocation, a more efficient AI platform can increase revenue-generating output without waiting for a new substation, transmission upgrade, generator, or data-center campus. That can make performance per watt economically valuable even if the platform itself carries a significant upfront price.
The counterpoint is important: higher rack density can shift, rather than eliminate, infrastructure challenges. A facility may need less energy for a given amount of inference output, yet its individual racks can still require much more concentrated power and cooling capability. Grid constraints, construction timelines, and electrical-equipment lead times remain material risks.
Liquid Cooling Moves Toward the Center of the Design
At very high rack densities, air cooling alone becomes less practical. Direct liquid cooling, cold plates, coolant distribution units, heat exchangers, pumps, pipes, manifolds, and monitoring systems become part of the deployment-critical bill of materials.The U.S. Department of Energy’s data-center design guidance describes direct-liquid cooling as including cold-plate and immersion approaches, and notes that high-performance computing environments were early adopters because of rising rack power densities. The DOE’s energy-efficient data-center guide provides useful context for why cooling is not a peripheral concern in AI infrastructure.
Nvidia’s own network roadmap reinforces that trend: Spectrum-6 is designed to support liquid cooling, allowing cooling considerations to extend across the AI factory rather than stopping at the compute tray. Nvidia’s Spectrum-6 material signals an increasingly holistic approach in which compute, network, and thermal design are jointly optimized.
For suppliers, cooling is attractive because it represents a growing share of infrastructure complexity. But it is also a field where execution matters intensely. Reliability failures, leaks, maintenance demands, water constraints, and integration problems can quickly offset theoretical efficiency gains. The winners will likely be firms that can prove performance in production installations rather than merely market a high-density cooling concept.
Nvidia’s Strength: Capturing the Architecture, Not Just the Chip
Vera Rubin strengthens Nvidia’s strategic position because the company is not merely supplying a GPU that sits inside someone else’s generic system. It is increasingly defining the architecture of the entire AI deployment.That brings several advantages:
- Higher system-level content: Nvidia participates in accelerators, CPUs, scale-up interconnects, SuperNICs, DPUs, Ethernet switches, and software.
- Tighter software-hardware integration: TensorRT-LLM and Dynamo optimization are part of the practical performance story behind the CoreWeave benchmark. CoreWeave’s test methodology demonstrates that software configuration is inseparable from hardware results.
- Faster deployment through validation: Rack-level systems can reduce integration uncertainty for customers that need to bring large AI clusters online quickly.
- Ecosystem control: A larger share of the architecture can reinforce Nvidia’s software ecosystem and make platform substitution more difficult.
Risks That the Vera Rubin Narrative Does Not Remove
The production ramp is a major milestone, but it does not make the AI infrastructure trade risk-free. In fact, a more integrated platform can create new forms of concentration and complexity.Benchmark Results Require Context
The 10x tokens-per-megawatt figure is powerful, but it comes from a DeepSeek-R1 comparison at matched interactivity with a highly optimized serving stack. CoreWeave’s benchmark details explain the conditions behind the result.Real deployments vary by model architecture, context window, batching strategy, precision, latency target, utilization rate, and software maturity. Investors should therefore treat tokens per megawatt as a useful operational metric, not a shortcut that can replace workload-level diligence.
More Integration Can Squeeze Some Suppliers
Nvidia’s end-to-end platform strategy may expand total infrastructure spending while also concentrating more value inside Nvidia-designed subsystems. Server makers, integrators, and component vendors could see higher shipment volumes but lower control over system architecture or pricing.That makes “picks and shovels” investing more nuanced than it appears. A supplier may enjoy strong AI demand yet still deliver disappointing financial results if content is commoditized, margins are competed away, or capital expenditure rises faster than returns.
Power Is Not a Problem Hardware Can Solve Alone
Efficiency improvements help stretch limited electricity capacity, but they do not create new transmission lines, transformers, generation facilities, or local permits. The IEA expects data centers to account for a substantial share of U.S. electricity-demand growth through 2030, highlighting how the bottleneck can move outside the data-center fence line. The IEA’s electricity-demand assessment makes clear that grid readiness will remain central to AI deployment plans.A more efficient rack can improve the value of a fixed megawatt. It cannot guarantee that the next megawatt will be available where a cloud provider needs it.
The Bottom Line: The AI Trade Is Becoming an Infrastructure Trade
Vera Rubin NVL72 marks a shift in the way AI performance should be understood. The unit that increasingly matters is not a GPU benchmark in isolation, but a production system’s ability to generate responsive, profitable AI output under real constraints on power, cooling, networking, and deployment time.Nvidia remains the obvious central beneficiary because it owns the platform architecture across compute, networking, and software. Yet the Vera Rubin production ramp also elevates the importance of HBM4 memory, advanced packaging, high-speed interconnects, optics, liquid cooling, electrical distribution, and rack-level manufacturing.
The most important investment question is therefore not simply who sells into AI. It is which companies provide the components and infrastructure that become more essential, more difficult to replace, and more valuable per deployed megawatt as AI factories move from GPU clusters toward fully integrated rack-scale systems.
That is why Vera Rubin changes the AI trade. It does not make the GPU less important. It makes everything required to keep that GPU productive far more important than before.
References
- Primary source: dqindia.com
Published: 2026-07-27T00:30:09.370999
Nvidia Vera Rubin shifts the AI trade beyond GPUs!
Integrates CPUs, GPUs, memory, networking, optical communications, power delivery, liquid cooling and rack assembly into a coordinated AI infrastructure platform
www.dqindia.com
- Related coverage: tomshardware.com
Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack | Tom's Hardware
FLOPS are only the startwww.tomshardware.com - Related coverage: coreweave.com
First-Ever Measured Vera Rubin NVL72 Silicon Performance Stats | CoreWeave Blog
NVIDIA Vera Rubin NVL72 on CoreWeave delivers 10x tokens-per-second per megawatt compared to NVIDIA Blackwell NVL72 on the same DeepSeek R1 workload.www.coreweave.com - Related coverage: blogs.nvidia.com
NVIDIA Vera Rubin Driving Performance Per Watt, Lower Token Costs for Partners Worldwide | NVIDIA Blog
Backed by 300 global partners, Vera Rubin is ramping up worldwide. NVIDIA partners CoreWeave, Google Cloud, Microsoft Azure and Mistral are among many deploying Vera Rubin, which delivers benchmark leadership on performance per watt and lowest token costs.blogs.nvidia.com - Related coverage: tweaktown.com
First NVIDIA Vera Rubin NVL72 benchmarks show 10X improvement over Grace Blackwell
CoreWeave ran the same DeepSeek-R1 benchmark on both Vera Rubin NVL72 and Grace Blackwell NVL72 systems and found a 10X improvement in one key area.www.tweaktown.com
- Related coverage: nvidia.com
NVIDIA Vera Rubin NVL72
NVIDIA Vera Rubin NVL72 is a rack-scale AI supercomputer unifying 72 Rubin GPUs and 36 Vera CPUs to power agentic reasoning AI and the AI industrial revolution.www.nvidia.com