The desk
Die Brief

DIE.BRIEF

Hardware desk
Systems CONFIRMED

Vera Rubin NVL144: the number the vendor page does not print

The Die Brief Desk, after nvidianews.nvidia.com

The NVIDIA Newsroom page names NVL72 with 72 Rubin GPUs and 36 Vera CPUs; the NVL144 with 144 R200 GPUs is a wire-sourced configuration the vendor page does not print.

NVIDIA Vera Rubin Opens Agentic AI Frontier — Picture from nvidianews.nvidia.com
NVIDIA Vera Rubin Opens Agentic AI Frontier — Picture from nvidianews.nvidia.com
Vera Rubin Platform
GPUs (NVL72)72 Rubin
CPUs (NVL72)36 Vera
InterconnectNVLink 6
LPX processors256
LPX on-chip SRAM128 GB
LPX scale-up bandwidth640 TB/s
AvailabilityH2 2026
DSX Max-Q30% more in fixed-power DC

The NVIDIA Newsroom page dated March 16, 2026, names the NVL72 as the rack configuration: 72 Rubin GPUs, 36 Vera CPUs, NVLink 6, ConnectX-9 SuperNICs, BlueField-4 DPUs. The page does not name an NVL144. The wire at awesomeagents.ai describes a 144-GPU configuration built from 72 Superchips, each pairing one 88-core Vera CPU with two R200 GPUs and 288 GB of HBM4. That is the number the assignment asks me to open on, but the vendor page in hand stops at 72.

The wire states 3.6 ExaFLOPS in NVFP4 across the full rack, roughly 25 PFLOPS per R200 GPU, and 1.2 ExaFLOPS in FP8. The vendor page does not print a peak FLOPS figure for any configuration. It does say the NVL72 trains large mixture-of-experts models with one-fourth the GPUs compared with the Blackwell platform and achieves up to 10x higher inference throughput per watt at one-tenth the cost per token. Those are relative claims against a prior generation, not absolute numbers. The 3.6 ExaFLOPS figure is a vendor projection carried by the wire, not a measurement. No independent lab has run a rack of 144 R200s, and the vendor page does not state whether the peak assumes sparsity.

The stretch between the two configurations is the gap between what the vendor page names and what the wire describes. The NVL72 is the configuration the vendor page details: 72 GPUs, 36 CPUs, NVLink 6. The NVL144 doubles the GPU count to 144 and the CPU count to 72, packaged as 72 Superchips. The wire says the NVL144 carries 20,736 GB of HBM4, about 20 TB, at up to 20.5 TB/s per GPU. The vendor page does not state a total HBM figure for any rack. It does not state a per-GPU HBM capacity. The 144 GB per R200 and 288 GB per Superchip are wire numbers, not vendor-page figures.

The vendor page lists seven chips in full production: Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 Ethernet switch, and the Groq 3 LPU. Five rack types follow from those chips. The Vera CPU rack holds 256 Vera CPUs. The Groq 3 LPX rack holds 256 LPU processors with 128 GB of on-chip SRAM and 640 TB/s of scale-up bandwidth. The BlueField-4 STX rack handles KV cache storage. The Spectrum-6 SPX rack handles east-west traffic. The vendor page says the LPX rack delivers up to 35x higher inference throughput per megawatt when deployed with NVL72. That is a system-level claim, not a per-chip measurement.

The DSX platform is where the vendor page gets specific about power. DSX Max-Q enables dynamic power provisioning across the AI factory, and NVIDIA says it results in 30% more AI infrastructure within a fixed-power data center. DSX Flex is described as making AI factories "grid-flexible assets," unlocking 100 gigawatts of stranded grid power. The 100 GW figure is a capacity claim, not a deployment figure. No data center has been built to that specification yet. The 30% figure is a design target from the reference architecture, not a measured outcome.

The vendor page names 80 MGX ecosystem partners and 200 data center infrastructure partners. Cloud providers listed include AWS, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. System manufacturers include Cisco, Dell, HPE, Lenovo, Supermicro, and others. AI labs named include Anthropic, Meta, Mistral AI, and OpenAI. The wire notes the GB200 NVL72 was reportedly in the $2-3M range and says the NVL144 "will almost certainly cost more." That is a wire inference, not a vendor figure.

Picture from awesomeagents.ai
Picture from awesomeagents.ai

The vendor page does not mention AMD. AMD's Helios MI450X is a separate rack-scale system and gets its own article. The Vera Rubin platform is positioned against NVIDIA's own prior generation, Blackwell, not against a competitor's rack.

The buyer who walks into a data center planning meeting in H2 2026 will be constrained by what the vendor page does not say. There is no per-rack power draw. There is no HBM4 bandwidth figure in the vendor page. The 3.6 ExaFLOPS figure is a theoretical peak from the wire, and the vendor page does not confirm it, does not state whether it assumes sparsity, and does not state the clock frequency at which it is achieved. The 10x inference throughput per watt is relative to Blackwell, which means the absolute number depends on what Blackwell was actually delivering in production, and that number is not in this document. The buyer cannot size a power budget from this page. The buyer cannot compare against a competitor's rack because the vendor page does not name one. The 30% more infrastructure in a fixed-power data center is a design claim from the DSX reference architecture, and the buyer will need to see it in a live deployment before trusting it.

What the vendor page does prove is the production status. The partner list is long and specific. The NVLink 6 interconnect is named. The ConnectX-9 and BlueField-4 are named. They are shipping components. The Groq 3 LPU integration is new, and the vendor page says it is available in the second half of this year. That is the same H2 2026 window. The LPX rack is the least proven element in the stack because it introduces a new processor type into the rack, and the 35x throughput claim is the largest single number in the document.

No independent benchmark exists for any Vera Rubin rack. No power measurement has been published. No per-GPU or per-rack price is stated. The 3.6 ExaFLOPS figure, the 20 TB HBM4 total, and the 20.5 TB/s per-GPU bandwidth are all wire-sourced and unverified. The 30% DSX Max-Q figure and the 100 GW DSX Flex figure are design targets, not measured outcomes. The buyer in H2 2026 will be buying on a vendor projection and a partner list, not on a benchmark.

The buyer cannot size a power budget or compare against a competitor from this page; the 3.6 ExaFLOPS peak is a wire-sourced projection the vendor does not confirm, and the 10x throughput claim is relative to Blackwell, whose production numbers are not in the document. The NVL144 configuration, the 20 TB HBM4 total, and the 20.5 TB/s per-GPU bandwidth are all unverified against the vendor page, which names only the NVL72.

No independent benchmark, per-GPU TDP, HBM4 bandwidth figure, or rack price appears in the vendor page or any document in the packet.

The vendor page confirms seven chips in production and five rack types, but the NVL144 configuration, the 3.6 ExaFLOPS FP4 peak, the 20 TB HBM4 total, and the 20.5 TB/s per-GPU bandwidth are all wire-sourced and unverified against the primary document. The 30% DSX Max-Q figure is a design target, not a measured outcome. The buyer is constrained by the absence of per-GPU TDP, per-rack power draw, HBM4 bandwidth, and pricing. The 10x inference claim is relative to Blackwell, whose production throughput is not stated in the document.

After nvidianews.nvidia.com. We did not report this. The pictures, if any, are theirs. Also filed by awesomeagents.ai.

October 11, 2026 · 9 min
Copy link Share to X