NVIDIA Rubin R100 GPU: 288GB HBM4 Specs, Pricing & Rental.
Pre-Order R100 GPU on Spheron
288 GB HBM4 · 22 TB/s · 50 PFLOPS FP4 · NVLink 6. Pre-order R100 GPU rentals ahead of general availability.
The NVIDIA Rubin R100, which NVIDIA itself calls simply the Rubin GPU, is the generational successor to the Blackwell B300. Rubin pairs 288GB of HBM4 with up to 22 TB/s of memory bandwidth, 50 PFLOPS of FP4 dense compute, and NVLink 6 at 3.6 TB/s. It is purpose-built for trillion-parameter inference at FP4 precision and multi-node training runs where Blackwell memory bandwidth becomes the bottleneck.
No pricing published yet. NVIDIA announced the Rubin architecture at GTC 2024 with R100 sampling beginning Q4 2026 and broad cloud GA expected in 2027. Register your interest to be notified first when Spheron capacity opens.
Register Interest
Join the pipeline for R100 access. We'll reach out with pricing and availability as soon as capacity goes live.
The NVIDIA Rubin R100, NVIDIA's own name for which is simply the Rubin GPU, is the generational successor to B300 Blackwell Ultra. It ships with 288GB HBM4 at up to 22 TB/s bandwidth (2.75x faster than B300), 50 PFLOPS FP4 compute (3.33x B300), NVLink 6 at 3.6 TB/s per GPU, and ConnectX-9 networking. First cloud availability is H2 2026 for AWS, Google Cloud, Azure, and specialist providers. Spheron is onboarding R100 capacity and will contact registered teams with pricing as soon as it is confirmed. For workloads running today, B300 and B200 are available now.
NVIDIA GPU generation roadmap
Where R100 sits in the NVIDIA GPU generation stack. All generations prior to Rubin are available on Spheron today.
R100 GPU specifications
R100 specs are taken from NVIDIA's published Vera Rubin NVL72 specification table: 288GB HBM4 at 22 TB/s, 50 PFLOPS NVFP4 inference, 17.5 PFLOPS dense FP8/FP6, NVLink 6 at 3.6 TB/s, and ConnectX-9. NVIDIA labels that table preliminary and subject to change, and publishes no transistor count or TDP for the part, so neither is quoted here.
R100 vs B300 vs B200 vs H100
| Spec | R1002026 | B300 | B200 | H100 |
|---|---|---|---|---|
| Architecture | Rubin | Blackwell Ultra | Blackwell | Hopper |
| VRAM | 288 GB HBM4 | 288 GB HBM3e | 192 GB HBM3e | 80 GB HBM3 |
| Memory Bandwidth | Up to 22 TB/s | 8 TB/s | 8 TB/s | 3.35 TB/s |
| FP4 Compute | 50 PFLOPS | 15 PFLOPS | 9 PFLOPS | N/A |
| FP8 Throughput | 17,500 TFLOPS | 7,000 TFLOPS | 4,500 TFLOPS | ~2,000 TFLOPS |
| Interconnect | NVLink 6 (3.6 TB/s) | NVLink 5 (1.8 TB/s) | NVLink 5 (1.8 TB/s) | NVLink 4 (900 GB/s) |
| Transistors | Not published | 208 billion | 208 billion | 80 billion |
| Cloud Availability | H2 2026 (first cohort) | Available now | Available now | Available now |
R100 per-GPU figures derived from NVIDIA's published Vera Rubin NVL72 rack specification divided by 72 GPUs. NVIDIA marks those values preliminary and subject to change. Availability reflects the first cohort; broader availability from additional providers follows.
Workloads built for R100
Trillion-Parameter FP4 Inference
At 50 PFLOPS FP4 and 288GB HBM4, a single R100 can hold and serve a 200B parameter model in FP4 with 88GB of headroom for KV cache. Multi-GPU setups handle 400B+ models that currently require 4x B300 at FP8.
Frontier Model Pre-Training
3.33x the FP4 compute of B300 and 2.75x the memory bandwidth. Training runs that require 8x B300 nodes may fit on fewer R100 nodes, reducing inter-node communication overhead and wall time.
Rack-Scale NVL72 Workloads
In the Vera Rubin NVL72 configuration, 72 R100 GPUs share 260 TB/s NVLink fabric and 20.7TB of aggregate HBM4. Models that require multi-node sharding on Blackwell may fit in a single NVL72 rack.
High-Bandwidth Inference Serving
22 TB/s memory bandwidth is 2.75x B300 on a single GPU. Decode-phase throughput scales nearly linearly with bandwidth for memory-bound LLM serving. At equivalent batch sizes, R100 serves 2.5–3x more tokens per second than B300.
When to pick the R100
Pick R100 if
You're training or serving 400B+ parameter models and B300's 8 TB/s memory bandwidth is the bottleneck. R100's 22 TB/s HBM4 is 2.75x faster, and 50 PFLOPS FP4 compute is 3.33x B300. If your workload is memory-bandwidth-bound, R100 is the first GPU where bandwidth stops being the ceiling.
Pick B300 instead if
Your timeline is 2025 or early 2026. B300 ships now with 288GB HBM3e and 15 PFLOPS FP4. For most 200B–400B workloads, B300 is the practical choice today. R100 is worth waiting for if you have flexible timelines and need the bandwidth or compute ceiling.
Pick B200 instead if
Your model fits in 192GB and you want the most widely available Blackwell option at lower cost. B200 handles most 70B–200B workloads, has better spot pricing, and is available on Spheron today.
Pick R100 for NVL72 if
You're running trillion-parameter workloads bottlenecked on memory bandwidth at rack scale. 72 R100 GPUs in NVL72 share 20.7 TB of unified HBM4, delivering 1,584 TB/s aggregate HBM bandwidth (~2.75x GB300 NVL72) and 260 TB/s NVLink 6 fabric (2x GB300 NVL72). The same capacity ceiling, with bandwidth that lifts the wall for memory-bound 10T+ parameter serving and training.
Available now on Spheron
R100 ships H2 2026. For workloads that need to run now, single-GPU Blackwell and Hopper options bill per-minute, while GB300 and GB200 NVL72 racks are available by reservation.
Blackwell Ultra. Same VRAM as R100, available now.
Blackwell. Most workloads under 200B parameters.
Hopper. Best value for inference and long-context serving.
R100 Guides & Benchmarks
More GPU Deep Dives guides →NVIDIA R100 Specs
288GB HBM4, 50 PFLOPS FP4, 22 TB/s bandwidth. What changed from Blackwell to Rubin.
NVIDIA Rubin CPX Explained
What NVIDIA's million-token GPU concept was, and what replaced it at GTC 2026.
NVIDIA B200 Specs & Benchmarks
The current-generation Blackwell flagship R100 is set to succeed.
R100 Pre-Order FAQ
NVIDIA confirmed H2 2026 availability for the first cloud cohort (AWS, Google Cloud, Azure, CoreWeave, Lambda, Nebius, Nscale). Spheron is targeting broader availability in 2027 as supply scales. Register your interest now to be notified when R100 capacity goes live.
Per NVIDIA's published Vera Rubin NVL72 spec table, each Rubin GPU has 288GB of HBM4 at 22 TB/s, 50 PFLOPS of NVFP4 inference, 35 PFLOPS of dense NVFP4 training, 17.5 PFLOPS of dense FP8/FP6, NVLink 6 at 3.6 TB/s, and ConnectX-9 at 1.6 Tb/s per-GPU networking. NVIDIA marks these as preliminary and subject to change, and does not publish a transistor count or TDP for the part. It is the generational successor to the B300 (Blackwell Ultra).
No official pricing has been published. At launch, hyperscalers typically price next-gen GPUs at a 30–50% premium over the prior generation. Specialist GPU clouds like Spheron historically run lower than that. Register your interest and we'll send you R100 pricing the moment it's confirmed.
R100 outpaces B300 on every axis except capacity. Memory: 288GB HBM4 vs 288GB HBM3e (same size, but HBM4 hits up to 22 TB/s vs 8 TB/s). Compute: 50 PFLOPS vs 15 PFLOPS FP4. Interconnect: NVLink 6 at 3.6 TB/s vs NVLink 5 at 1.8 TB/s. For compute-dense FP4 inference on 400B+ parameter models, R100 delivers roughly 3.3x the throughput of B300.
NVIDIA officially calls it the 'Rubin GPU' and publishes no alphanumeric SKU for it. R100 and R200 are supply-chain and press shorthand, R200 reflecting the dual-die package design. H300 is not an NVIDIA name and no cloud provider brands it that way: it appears to be extrapolation from H100 and H200, since NVIDIA changed architecture family after H200 rather than shipping an H300. We keep the term on this page because people search for it, but the product you want is the Rubin GPU.