GPU Deep Dives

19 guides in this topic

Every GPU Spheron rents gets its own datasheet here: memory bandwidth, tensor core throughput, NVLink generation, and the number that actually matters, how many tokens per second it pushes on a real model. These are reference posts, the kind you bookmark and come back to when you need the exact HBM3e bandwidth on an H200 or the FP4 TFLOPS on a B200 without digging through a vendor PDF.

The cluster spans single-GPU cards (H100, H200, RTX 5090, L40S, L40, RTX 6000 Ada) up to full rack systems (GB200 NVL72), plus the Grace Hopper superchip and early R100 Rubin specs as they become public. Where a benchmark exists, we ran it or cited MLPerf; where NVIDIA hasn't published a number yet, we say so instead of guessing.

The pillar guide, NVIDIA B200 Specs & Benchmarks, is the most complete of the set: full spec table, MLPerf v6.0 throughput against H100, and live rental pricing so the datasheet and the dollar figure sit in one place. Read this cluster when you already know which GPU you want and need the hard numbers to size a cluster, write a proposal, or check a vendor's claims.

Start Here

All GPU Deep Dives Guides

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute with no contracts and no minimum. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min