The desk
Die Brief

DIE.BRIEF

Hardware desk
GPU Architecture SINGLE-SOURCE

Rubin's FP4 headline outruns its memory table

The Die Brief Desk, after blog.barrack.ai

The Barrack AI breakdown lists 288 GB of HBM4 per Rubin GPU, 22 TB/s bandwidth, and 50 PFLOPS of NVFP4 inference.

Picture from blog.barrack.ai
Picture from blog.barrack.ai

The Barrack AI breakdown puts 288 GB of HBM4 on each Rubin GPU and says the stack delivers up to 22 TB/s. The wire page's opening and table do not say the same thing: the opening line pushes 50 petaFLOPS of FP4 inference per chip, while its table labels the same figure as NVFP4 inference and pairs it with 35 PFLOPS for training. The phrase "50 petaFLOPS of FP4 inference per chip" is the weaker claim, because it drops the NVFP4 qualifier and hides the training number. The trusted document is the Barrack AI table, and the number that decides it is 22 TB/s.

Die Brief's read is that the memory path, not the FP4 headline, is the constraint. A 22 TB/s figure matters only if the 288 GB HBM4 stacks can be populated, binned, and cooled at 1,800 to 2,300 W, and the page does not prove that the 50 PFLOPS number is sustained because it gives no kernel, batch, or power state. The 35 PFLOPS training row in the table is the more honest anchor, because it separates inference from training. For a buyer, the useful comparison is bandwidth per watt, not peak FP4.

What is still unmeasured is the per-stack HBM4 bandwidth, the thermal headroom at Max-P, and the software path that would make NVFP4 compression real. It leaves out yield, interposer area, and the power state behind the 50 PFLOPS figure.

The memory path, not the FP4 headline, is the constraint. The 35 PFLOPS training row is the more honest anchor, because it separates inference from training.

The document leaves out per-stack HBM4 bandwidth, yield, interposer area, and the power state behind the 50 PFLOPS figure.

The 22 TB/s HBM4 path and the 1,800 to 2,300 W envelope set the real limit. The 35 PFLOPS training row is the safer anchor than the 50 PFLOPS inference headline.

288 GB HBM4 per GPU · 22 TB/s memory bandwidth · 336 billion transistors · 50 PFLOPS NVFP4 inference · 35 PFLOPS NVFP4 training · 1,800 to 2,300 W TDP

After blog.barrack.ai. We did not report this. The pictures, if any, are theirs.

October 08, 2026 · 4 min
Copy link Share to X