Engineering

NVIDIA Vera Rubin NVL72: Specs, Price & Why There's No H300

NVIDIA Vera Rubin NVL72NVIDIA H300Vera Rubin SpecsRubin GPU HBM4NVL72NVLink 6GPU CloudAI Infrastructure
NVIDIA Vera Rubin NVL72: Specs, Price & Why There's No H300

Short version: there is no NVIDIA H300. If a search for "H300" brought you here, you are looking for the NVIDIA Rubin GPU inside Vera Rubin NVL72. This guide covers what NVIDIA actually publishes about that system, what it will cost, when it ships, and what to rent in the meantime. R100 pre-orders are open on the R100 pre-order page.

Every spec in this guide is taken from NVIDIA's own Vera Rubin NVL72 spec table and its press releases, checked on 16 August 2026. NVIDIA labels that table "preliminary information, all values are up to and subject to change", and several widely-copied numbers on other sites are now out of date, so this page states its sources inline.

There Is No NVIDIA H300. Here's What You Actually Want

This needs saying plainly, because a lot of pages get it wrong, including an earlier version of this one.

NVIDIA has never announced a GPU called the H300. The Hopper H-series ended at H200. The lineage runs:

GenerationArchitectureData center GPUs
2022-2024HopperH100, H200
2025-2026Blackwell / Blackwell UltraB200, B300
2026 onwardRubinRubin GPU (in Vera Rubin NVL72)

There is no H300 anywhere in that sequence. Searching NVIDIA's newsroom and product pages returns zero results for the term. The four NVIDIA press releases covering Rubin, from CES in January 2026 through GTC Taipei in May 2026, all use one name: "Rubin GPU".

The name almost certainly came from extrapolation. H100 was followed by H200, so a next step called H300 looked plausible, and once a handful of sites published it, other sites copied it. NVIDIA instead changed architecture family after H200, which is why the sequence jumps to B200.

Two other names you will encounter:

  • R100 and R200. Supply-chain and press shorthand for the Rubin part. Not NVIDIA marketing names. Spheron uses "R100" on its pre-order page because that is the term buyers search, but NVIDIA's own material says "NVIDIA Rubin GPU". Our Rubin R100 chip guide covers the per-die architecture, and the Rubin vs Blackwell vs Hopper comparison puts three generations side by side.
  • VR200. Partner and supply-chain designation for the system. Again, not what NVIDIA's product pages call it.

Vera Rubin NVL144 is a different thing. NVIDIA originally counted GPU dies (144) and later renamed the flagship to count GPU packages (72), so today "Vera Rubin NVL72" is the correct name for the flagship rack. "Vera Rubin NVL144 CPX" still exists, but it is the separate Rubin CPX long-context variant, not the flagship.

Vera Rubin NVL72 Specs, Straight From NVIDIA's Table

NVIDIA publishes three columns: the full rack, the two-GPU Vera Rubin Superchip, and a single Rubin GPU. Reproduced here in full, because most secondary coverage quotes only a subset and several quote superseded numbers.

SpecVera Rubin NVL72 (rack)Vera Rubin SuperchipRubin GPU
Configuration72 Rubin GPU, 36 Vera CPU2 Rubin GPU, 1 Vera CPU1 Rubin GPU
NVFP4 inference3,600 PFLOPS100 PFLOPS50 PFLOPS
NVFP4 training (dense)2,520 PFLOPS70 PFLOPS35 PFLOPS
FP8/FP6 training (dense)1,260 PFLOPS35 PFLOPS17.5 PFLOPS
INT818 POPS500 TOPS250 TOPS
FP16/BF16288 PFLOPS8 PFLOPS4 PFLOPS
TF32144 PFLOPS4 PFLOPS2 PFLOPS
FP642,400 TFLOPS67 TFLOPS33 TFLOPS
GPU memory / bandwidth20.7 TB HBM4 / 1,580 TB/s576 GB / 44 TB/s288 GB HBM4 / 22 TB/s
NVLink generationSixthSixthSixth
NVLink bandwidth260 TB/s7.2 TB/s3.6 TB/s
NVLink-C2C65 TB/s1.8 TB/sn/a
CPU cores3,168 Olympus (Arm compatible)88 Olympusn/a
CPU memory54 TB LPDDR5X1.5 TB LPDDR5Xn/a
Scale-out networking28.8 TB/s0.8 TB/s0.4 TB/s

Source: NVIDIA Vera Rubin NVL72 product page, specifications section, retrieved 16 Aug 2026.

Numbers to stop repeating

Several figures circulating widely are wrong or stale. If you are cross-referencing other articles:

  • "13 TB/s HBM4 per GPU" is superseded. NVIDIA's current table says 22 TB/s.
  • "250 TB/s NVLink 6" is superseded. The current figure is 260 TB/s.
  • "336 billion transistors", "TSMC N3", "2,300 W TDP", "166 kW rack" are not published by NVIDIA for this part. An earlier version of this page carried them as estimates; they have been removed rather than restated, because there is no primary source for any of them.
  • "75 TB of fast memory" is fine, and checks out: 20.7 TB HBM4 plus 54 TB LPDDR5X is 74.7 TB.

NVIDIA also does not publish a rack power figure for Vera Rubin. Supply-chain analysis circulating in 2026, originating with Ming-Chi Kuo, puts it around 190 kW in a Max-Q profile and 230 kW in Max-P. Note that both figures trace back to a single originator, so they are one estimate reported in several places rather than several independent estimates. NVIDIA's own engineering blog does describe Dynamic Max-Q and Static Max-P power provisioning for the platform, which corroborates the framing even though the wattages are not NVIDIA-confirmed. Treat those as reported, not as spec.

What Is Actually New: Seven Chips, Not One

The part most coverage misses is that Vera Rubin NVL72 is a seven-chip platform, not a GPU refresh. NVIDIA ships the rack with:

  • Rubin GPU, HBM4 and a 50 PFLOPS NVFP4 Transformer Engine.
  • Vera CPU, 88 custom NVIDIA Olympus cores per CPU, Arm-compatible. This is NVIDIA's own core design, not licensed Neoverse as in Grace.
  • NVLink 6 Switch, 3.6 TB/s of all-to-all scale-up bandwidth per GPU.
  • ConnectX-9 SuperNIC, 1.6 Tb/s per GPU, double ConnectX-8's 800 Gb/s in GB300.
  • BlueField-4 DPU for storage, networking and security offload.
  • Spectrum-X Ethernet with co-packaged optics, which NVIDIA credits with 5x better power efficiency than pluggable transceivers.
  • NVIDIA Groq 3 LPU, the inference accelerator. An LPX rack holds 256 LPUs with 128 GB SRAM and 40 PB/s of memory bandwidth. Our Groq 3 LPU explainer covers why a non-GPU inference chip ended up inside NVIDIA's flagship platform.

NVIDIA claims the pairing of Vera Rubin NVL72 with LPX racks delivers up to 35x higher throughput per megawatt for trillion-parameter models versus Blackwell.

There is also a smaller sibling, Vera Rubin NVL4: four Rubin GPUs on a second-generation NVLink bridge with two Vera CPUs, aimed at scientific computing rather than frontier LLM training. NVIDIA quotes up to 4x the simulation performance of Grace Hopper. If you are sizing between the two, the NVL4 vs NVL72 decision guide covers where the split makes sense.

Rubin vs Blackwell Ultra vs Hopper

Comparing per-GPU figures using NVIDIA's published rack tables divided by 72, so the three columns stay on the same basis.

SpecH100 SXM5H200 SXM5B200B300Rubin GPU
ArchitectureHopperHopperBlackwellBlackwell UltraRubin
Memory80 GB HBM3141 GB HBM3e192 GB HBM3e288 GB HBM3e288 GB HBM4
Memory bandwidth3.35 TB/s4.8 TB/s8 TB/s8 TB/s22 TB/s
NVLink per GPU900 GB/s900 GB/s1.8 TB/s1.8 TB/s3.6 TB/s
NVLink generation44556
Dense FP4n/an/a10 PFLOPS15 PFLOPS35 PFLOPS
Scale-out per GPU400 Gb/s400 Gb/s400 Gb/s800 Gb/s1.6 Tb/s
StatusAvailableAvailableAvailableAvailableNot yet shipping

Dense FP4 for B200 and B300 is derived from NVIDIA's rack tables: GB200 NVL72 is 720 PFLOPS dense across 72 GPUs, GB300 NVL72 is 1,080 PFLOPS dense across 72. That gives 10 and 15 PFLOPS per GPU respectively.

The headline is bandwidth, not compute. B200 and B300 share the same 8 TB/s HBM3e; Rubin's HBM4 is 22 TB/s, about 2.75x. Dense FP4 rises 2.33x from B300. That ratio matters: NVIDIA gave this generation proportionally more memory bandwidth than compute, which is the opposite of the last two generations and tells you what they think the bottleneck is.

For decode-heavy serving this is the whole story. Token generation at high batch sizes is memory-bandwidth-bound, not compute-bound, so a 2.75x bandwidth increase translates far more directly into tokens per second than a 2.33x FP4 increase does. The same logic is why H200 outperforms H100 on large-model serving despite identical compute. Our GPU memory requirements guide covers how to work out which side of that line your workload sits on.

Availability: What NVIDIA Has Actually Committed To

The dates have shifted through 2026, so here is the sequence with sources:

DateStatementSource
5 Jan 2026 (CES)Rubin in full production; partner products available H2 2026NVIDIA newsroom
16 Mar 2026 (GTC)Vera Rubin products available from partners starting H2 2026NVIDIA newsroom
31 May 2026 (GTC Taipei)Production shipments begin "starting this fall"NVIDIA newsroom
4 Aug 2026"In full production, on track to ship in the second half of 2026"NVIDIA developer blog

Launch cloud partners named by NVIDIA: AWS, Google Cloud, Microsoft Azure and OCI, plus CoreWeave, Crusoe, Lambda, Nebius, Nscale and Together AI.

Nothing is rentable today. No provider on that list has published an hourly rate, because customer shipments have not started. Rubin CPX is a further step out, expected at the end of 2026.

For planning purposes: first allocations go to hyperscalers and frontier labs, and specialist GPU clouds have historically received new NVIDIA silicon several months after the initial cohort. If your roadmap depends on Rubin capacity, plan around 2027 and treat anything earlier as upside.

Vera Rubin Price: What Is Known and What Is Guesswork

NVIDIA has never confirmed a list price for any NVL72 product, and that has not changed.

What is reported:

  • Vera Rubin NVL72: roughly $5M to $7M per rack, per Tom's Hardware reporting in March 2026, syndicated here. That figure reportedly includes around $1M of 3D NAND storage.
  • A widely-repeated "$8.8M per rack" figure is misattributed. Reading the source, the $7M to $8.8M range applies to NVL144 VR300 (Rubin Ultra), a later product that had not taped out at the time and is associated with 2028. Several outlets ran headlines conflating the two. Do not use it for Vera Rubin NVL72.
  • A Morgan Stanley bill-of-materials estimate of roughly $7.8M circulated in mid-2026. That is a component cost estimate, not a sale price, and should not be compared against rack prices directly.

There is no cloud hourly price for Vera Rubin. Any per-hour number you find today is invented. For context on what current rack-scale actually costs when it is bookable, CoreWeave publishes $42.00/hr for a 4-GPU GB200 NVL72 instance, which works out to $10.50 per GPU-hour; GB300 NVL72 is quote-only at every provider we checked. The GB300 vs GB200 pricing breakdown has the detail, the GB200 NVL72 architecture guide covers the rack it replaces, and the Vera Rubin cloud availability and cost-per-token guide tracks who is expected to ship it first.

What to Rent Instead, Today

Live Spheron on-demand and spot rates, pulled from the public pricing API on 16 August 2026:

GPUVRAMOn-demand $/hrSpot $/hr
A100 80G PCIe80 GB HBM2e$1.43$1.19
A100 80G SXM480 GB HBM2e$1.82$1.14
H100 PCIe80 GB HBM3$2.98$2.20
GH200 PCIe96 GB HBM3$3.02n/a
H100 SXM580 GB HBM3$3.98$2.91
H200 SXM5141 GB HBM3e$4.79$3.31
B300 SXM6288 GB HBM3e$9.08$5.81
B200 SXM6192 GB HBM3e$9.36$5.34

Pricing fluctuates based on GPU availability. The prices above are based on 16 Aug 2026 and may have changed. Check current GPU pricing → for live rates.

Those rates are aggregated across data center partners rather than run from a single fleet, which is why the Spheron GPU marketplace can list Hopper, Ada and Blackwell generations side by side instead of pushing whichever generation it happens to own. The full catalogue, including H100 GPU rental and A100 on Spheron, is in the GPU rental catalog.

Two things worth noticing in that table.

B300 currently costs less on-demand than B200 ($9.08 versus $9.36) while carrying 50% more memory and 50% more dense FP4. If you were defaulting to B200 out of habit, check the live rate before you provision. The B300 vs B200 cost-per-token analysis covers where that translates into real savings and where it does not.

B200 spot is cheaper than B300 spot ($5.34 versus $5.81), so the ordering flips on interruptible capacity. Spot is reclaimable at any time without notice, so it suits checkpointed training and batch inference rather than latency-sensitive serving.

If you are sizing for a model that does not fit in 141 GB, B300's 288 GB per GPU is the same capacity a Rubin GPU will offer, at HBM3e bandwidth rather than HBM4. That is the closest thing to Rubin available to rent today, and for capacity-bound rather than bandwidth-bound workloads the difference is small.

Should You Wait for Rubin?

Wait only if all of these are true:

  • Your workload is genuinely memory-bandwidth-bound, not capacity-bound or compute-bound. If B300's 288 GB at 8 TB/s already serves your model, Rubin's extra bandwidth buys latency, not feasibility.
  • Your timeline extends into 2027. Specialist clouds will not have meaningful Rubin capacity before then.
  • You can absorb launch pricing. Every generation has launched at a premium and compressed. H100 launched around $8-10/hr and now sits at $3.98 on-demand here.

Rent Blackwell now if any of these are true:

  • You need to ship in 2026.
  • Your model fits in 288 GB per GPU, which covers essentially everything below frontier scale.
  • Your utilisation is below about 60%. Buying more of the GPU you have beats waiting for a faster one you cannot get.
  • Cost per token already meets target on spot.

For most teams the honest answer is that Rubin changes nothing about 2026 planning. It matters for capacity decisions in 2027, and for anyone whose serving economics are dominated by decode bandwidth on very large models.

Summary

  • There is no NVIDIA H300. The name is a widely-copied error. The real product is the NVIDIA Rubin GPU inside Vera Rubin NVL72.
  • Per NVIDIA's published table: 72 Rubin GPUs and 36 Vera CPUs per rack, 288 GB HBM4 at 22 TB/s per GPU, 20.7 TB and 1,580 TB/s per rack, 260 TB/s NVLink 6, 3,600 PFLOPS NVFP4 inference. All marked preliminary.
  • It is a seven-chip platform: Rubin GPU, Vera CPU, NVLink 6 switch, ConnectX-9, BlueField-4, Spectrum-X co-packaged optics, and the Groq 3 LPU.
  • Production shipments begin in autumn 2026 per NVIDIA's 31 May statement. Nothing is rentable by the hour today, and no provider has published a rate.
  • Reported rack cost is roughly $5M to $7M. The "$8.8M" figure in circulation belongs to the later Rubin Ultra NVL144, not this system.
  • Bandwidth, not compute, is the real generational jump: 2.75x on memory versus 2.33x on dense FP4.

Rubin is not bookable yet, and will not be for a while. B300 gives you the same 288 GB per GPU today, currently at a lower on-demand rate than B200, with per-minute billing and no commitment.

Pre-order R100 → | B300 GPU pricing → | Check H200 availability → | Get started on Spheron →

FAQ / 06

Frequently Asked Questions

No. NVIDIA has never announced, released, or documented a GPU called the H300. The Hopper H-series ended at H200. NVIDIA's line went H100 and H200 (Hopper), then B200 and B300 (Blackwell and Blackwell Ultra), and now the Rubin GPU inside Vera Rubin NVL72. The name H300 appears to be pattern-extrapolation from H100 to H200, repeated by low-quality sites until it looked real. If you searched for H300, the product you almost certainly want is the NVIDIA Rubin GPU.

Per NVIDIA's published spec table, Vera Rubin NVL72 pairs 72 Rubin GPUs with 36 Vera CPUs. Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s, giving 20.7 TB and 1,580 TB/s across the rack. The rack delivers 3,600 PFLOPS of NVFP4 inference, 2,520 PFLOPS of dense NVFP4 training, and 1,260 PFLOPS of dense FP8/FP6 training. NVLink 6 provides 260 TB/s of rack fabric at 3.6 TB/s per GPU. NVIDIA marks all of these as preliminary and subject to change.

You cannot rent it today. NVIDIA said at CES in January 2026 that Rubin was in full production with partner availability in H2 2026, and in its 31 May 2026 GTC Taipei release said production shipments begin in the fall. No cloud provider has published an hourly rate for Vera Rubin, because the hardware is not yet in general customer hands. Treat any Vera Rubin price you see quoted today as speculation.

NVIDIA has never published a list price for any NVL72 product. Press reporting in March 2026 put Vera Rubin NVL72 in the region of 5 to 7 million dollars per rack. Be careful with the 7 to 8.8 million dollar figure circulating in headlines: that range refers to the later NVL144 Rubin Ultra generation, whose silicon had not taped out at the time of reporting, not to Vera Rubin NVL72.

The biggest jump is memory bandwidth. Both carry 288 GB, but B300 uses HBM3e at 8 TB/s while Rubin uses HBM4 at 22 TB/s, roughly 2.75x. Dense NVFP4 goes from 15 PFLOPS per B300 GPU to 35 PFLOPS per Rubin GPU, and NVLink doubles from 1.8 TB/s to 3.6 TB/s per GPU. For memory-bandwidth-bound decode and long-context serving, the bandwidth difference matters more than the compute difference.

Rent Blackwell now for anything you need to ship in 2026. Vera Rubin is not orderable by the hour from any provider, first allocations go to hyperscalers and large labs, and specialist clouds historically receive next-generation silicon months later. B200 and B300 are available today on Spheron with per-minute billing, and B300 currently prices below B200 on-demand.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute with no contracts and no minimum. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min