The H200 comes in two form factors that share identical memory and bandwidth but almost nothing else. NVL is a 600W, air-cooled PCIe card built for standard enterprise racks. SXM5 is a 700W module built for full NVSwitch fabric in HGX and DGX nodes. Both carry 141GB of HBM3e at 4.8TB/s, so if you're picking H200 NVL vs SXM5 based on memory capacity, you're solving the wrong problem. The decision is about your rack's power and cooling budget, and how many GPUs your workload actually needs to talk to each other at once.
That's a different question than the one H100 buyers had to answer. Our H100 NVL vs SXM5 vs PCIe guide was largely about memory: NVL's 94GB against SXM5's 80GB was the headline. On H200, NVIDIA closed that gap. Both variants get the same 141GB, so the tradeoff moved entirely to power envelope and interconnect topology. If you want the full TFLOPS and MIG breakdown before you get into form factors, see the complete H200 specs datasheet.
H200 NVL vs SXM5: Full Spec Comparison (Memory, Bandwidth, TDP, MIG)
Here's the side-by-side, straight from NVIDIA's data sheet:
| Specification | H200 SXM5 | H200 NVL |
|---|---|---|
| Architecture | Hopper (GH100) | Hopper (GH100) |
| VRAM | 141GB HBM3e | 141GB HBM3e |
| Memory Bandwidth | 4.8 TB/s | 4.8 TB/s |
| FP8 Tensor Core (sparse) | 3,958 TFLOPS | 3,341 TFLOPS |
| Max TDP | Up to 700W | Up to 600W |
| Form Factor | SXM module | PCIe dual-slot, air-cooled |
| NVLink | Full NVSwitch, 900GB/s per GPU, up to 8 GPUs | 2-way or 4-way bridge, 900GB/s per GPU |
| Secondary Interconnect | N/A (all GPUs on NVSwitch) | PCIe Gen5, 128GB/s (non-bridged GPUs) |
| Server Configuration | HGX H200, 4 or 8 GPUs | MGX H200 NVL, up to 8 GPUs |
| MIG Instances | Up to 7 x 18GB | Up to 7 x 16.5GB |
(Source: NVIDIA H200 datasheet.)
Two things stand out. First, memory and bandwidth are a wash, unlike the H100 generation where NVL carried more VRAM than SXM5. Second, SXM5 keeps a real compute edge: 3,958 TFLOPS FP8 sparse versus NVL's 3,341 TFLOPS, roughly 18% more raw throughput per GPU. That gap comes straight from the higher 700W power budget on SXM5, which lets the die run at a higher sustained clock. If your workload is compute-bound rather than memory-bound, SXM5 wins on paper before you even get to interconnect.
MIG capacity follows the same pattern as memory: nearly identical, 18GB versus 16.5GB per slice across 7 instances either way. If you're partitioning H200s for multi-tenant inference, the practical difference between the two form factors here is marginal.
NVLink Bandwidth Is the Real Difference: 2/4-Way Bridge vs Full NVSwitch
The number that actually separates these two cards is 900GB/s per GPU on both, but delivered through very different topologies. SXM5 puts every GPU in an 8-GPU HGX node on the same NVSwitch fabric: all-to-all, 900GB/s, no exceptions. H200 NVL bridges GPUs in pairs or quads, 900GB/s per GPU within that bridged group, and drops to PCIe Gen5 at 128GB/s, about 7x slower, for any traffic between non-bridged GPUs in the same MGX server.
That distinction determines what scales and what doesn't. A tensor-parallel job running TP=2 or TP=4 on a bridged NVL group gets NVSwitch-class bandwidth for its all-reduce operations. The moment you need TP=8, or any communication pattern that crosses outside the bridge, you're on PCIe Gen5, and NVL was never built for that. SXM5's NVSwitch has no such boundary; every GPU pair in the node talks at 900GB/s regardless of which two you pick.
Why H200 NVL's Bridge Beats H100 NVL's
H200 NVL's bridge is a meaningful upgrade over its predecessor, not just a memory refresh. H100 NVL connected exactly two GPUs at a fixed 600GB/s total across the pair. H200 NVL's bridge runs at 900GB/s per GPU in a 2-way or 4-way configuration, a 50% bandwidth increase over H100 NVL's 600GB/s total bridge across its fixed 2-GPU pair, and it adds a 4-way option that H100 NVL never had. NVIDIA states H200 NVL delivers up to 1.7x faster large language model inference and 1.3x more performance on HPC applications versus H100 NVL, with a 1.5x memory increase and 1.2x bandwidth increase driving most of that gain.
For a deeper look at how NVLink bridging compares to full switch fabrics and other interconnect approaches, see what NVLink actually is and how its bandwidth is measured.
Pricing by Form Factor: On-Demand and Spot Rates
Spheron's live GPU pricing API returned the following on 19 Jul 2026, split by instanceType before computing any per-GPU rate:
| Form Factor | On-Demand ($/GPU/hr) | Spot ($/GPU/hr) |
|---|---|---|
| H200 SXM5 | $4.54 | $3.31 |
| H200 NVL | See /pricing/ | See /pricing/ |
H200 SXM5's on-demand floor comes from single-GPU dedicated offers; spot bottoms out on an 8-GPU spot bundle at $26.514/hr, or $3.31 per GPU. Both figures move with data center availability, so treat them as a snapshot, not a fixed rate.
Across the wider market, H200 SXM on-demand pricing spans a wide range depending on the provider. GMI Cloud lists around $2.60/hr, Lambda sits near $4.49/hr, RunPod ranges $2.69 to $3.59/hr, and CoreWeave charges $6.31/hr sold only in 8-GPU bundles, with no single-GPU option (source: GMI Cloud's H200 provider pricing comparison). Nebius lists H200 at $4.50/hr on-demand and $2.45/hr preemptible (source: Nebius pricing), and Azure's ND H200 v5 tops the range at $13.78/hr per GPU. These are directional; hyperscaler and neocloud rates change often enough that you should re-check before budgeting a production deployment.
Pricing fluctuates based on GPU availability. The prices above are based on 19 Jul 2026 and may have changed. Check current GPU pricing → for live rates.
Why Standalone H200 NVL Rental Listings Are Rare
You won't find many public cloud rental SKUs labeled "H200 NVL," and that's not an oversight. NVL is sold primarily as an enterprise on-prem card through OEM systems partners rather than as a common cloud instance type. Spheron's API currently lists H200 SXM5 offers only, with no standalone NVL SKU in the response, the same gap documented for H100 NVL availability a generation earlier. When a cloud aggregator does show an "H200 NVL" listing, check the specs closely; some listings on third-party sites appear to relabel standard H200 SXM5 on-demand pricing as "NVL" rather than reflecting a genuine NVL SKU with the lower 600W envelope and PCIe form factor. If a listing doesn't specify the interconnect topology and TDP, assume it's SXM5 under a different name until proven otherwise.
Which Workloads Need SXM5 Scaling vs a Single H200 NVL Card
The short version: if your job fits on 1 to 4 GPUs and doesn't need to cross a bridge boundary, NVL is a reasonable, cheaper-to-operate fit. If it needs 8-way all-to-all communication, SXM5 is the only option that actually delivers it.
Single-GPU and 2-4 GPU Inference on NVL
A single H200 NVL card holds the same 141GB as SXM5, which means the same 70B-at-FP16-on-one-GPU math applies: Llama 3.1 70B at FP16 needs roughly 140GB for weights alone, so it fits on one H200 card of either form factor with a thin KV cache margin. For FP8 serving at around 70GB, you get real headroom for KV cache and concurrent long-context requests on a single NVL GPU, without needing tensor parallelism at all. See the H200 deployment guide for the KV cache sizing math and multi-model colocation patterns that apply identically to NVL and SXM5, since compute and memory per GPU are close enough that the deployment playbook doesn't change.
Where NVL specifically helps is 2-way or 4-way tensor-parallel inference for models that need more than one GPU but don't need eight. A 4-way bridged NVL group gets the full 900GB/s per GPU for its all-reduce traffic, the same bandwidth SXM5 offers, at a lower power draw and in a server that doesn't need a custom NVLink board or liquid cooling. For teams running inference-only workloads at 1 to 4 GPU scale, NVL's air-cooled, standard-rack profile is the more practical fit.
8-GPU Training and Large-Scale Serving on SXM5
Anything that needs FSDP or tensor-parallel training across more than 4 GPUs has to run on SXM5. NVSwitch gives every GPU in an 8-GPU HGX H200 node full 900GB/s all-to-all bandwidth, which is what gradient synchronization and pipeline parallelism at scale actually require. NVL's bridge tops out at a 4-way group; anything larger falls back to PCIe Gen5's 128GB/s between groups, a real bottleneck for the frequent, large all-reduce operations that distributed training generates.
The same logic applies to large-scale serving of frontier-class MoE models. DeepSeek V3-class models running TP=8 need every GPU on the same fabric, not split across bridge pairs connected by PCIe. If your serving stack scales past what a 4-way NVL group can hold, SXM5 with NVSwitch is the correct form factor, not a workaround.
Air-Cooled Enterprise Racks: Why NVIDIA Built H200 NVL
NVIDIA didn't build H200 NVL to compete with SXM5 on performance. It built it to fit into racks that SXM5 physically can't. NVIDIA says roughly 70% of enterprise racks operate at 20kW and below with air cooling (source: NVIDIA's H200 NVL announcement), and SXM5's 700W-per-GPU, NVSwitch-fabric design generally needs higher power density and, for sustained peak workloads, liquid cooling that most of those racks don't have. NVL solves that by dropping to a 600W envelope, using a standard PCIe dual-slot card, and running in ordinary air-cooled servers without a custom NVLink board.
That design choice also buys configuration flexibility SXM5 doesn't have. NVL supports 1, 2, 4, or 8 GPUs per MGX server, while HGX H200 SXM5 ships as a fixed 4- or 8-GPU NVSwitch baseboard. An enterprise team that wants to start with a single H200 in an existing rack, then add more later, can do that on NVL. Doing the same on SXM5 means committing to the fixed baseboard configuration up front.
H200 NVL became available beginning December 2024 from NVIDIA's global systems partners, including Dell, HPE, Lenovo, Supermicro, ASRock Rack, ASUS, GIGABYTE, and MSI, putting it directly into the standard enterprise server channel rather than requiring a specialized HGX/DGX purchase. Dropbox's VP of Infrastructure, Ali Zafar, said at launch: "Dropbox handles large amounts of content, requiring advanced AI and machine learning capabilities. We're exploring H200 NVL to continually improve our services and bring more value to our customers." (Source: NVIDIA's H200 NVL announcement.)
H200 NVL vs H100 NVL: What Changed Generation-Over-Generation
| Spec | H100 NVL | H200 NVL |
|---|---|---|
| VRAM | 94GB HBM3 | 141GB HBM3e |
| Memory Bandwidth | ~3,938 GB/s | 4.8 TB/s |
| FP8 (sparse) | 3,341 TFLOPS | 3,341 TFLOPS |
| TDP | 350-400W per GPU | Up to 600W per GPU |
| NVLink Bridge | 600GB/s total, fixed 2-GPU pair | 900GB/s per GPU, 2-way or 4-way |
| Max GPUs per server | 2 | 8 |
| Form Factor | PCIe dual-slot | PCIe dual-slot, air-cooled |
Two changes matter more than the rest. First, H200 NVL closed the memory-per-GPU gap that made H100 NVL distinctive: it went from 94GB to 141GB, a 1.5x increase, matching SXM5 instead of exceeding it. Second, the bridge itself scaled up: from a fixed 2-GPU pairing at 600GB/s total to a flexible 2-way or 4-way bridge at 900GB/s per GPU, plus support for up to 8 GPUs per MGX server versus H100 NVL's hard cap at 2. FP8 compute held flat at 3,341 TFLOPS across both generations of NVL, so the generational win here is entirely about memory capacity and interconnect flexibility, not raw throughput.
Practically, this means H200 NVL is a more capable building block for enterprise inference clusters than H100 NVL ever was. Where H100 NVL was really a 2-GPU appliance, H200 NVL can scale to an 8-GPU server with the same air-cooled, standard-rack profile. If you're comparing the two Hopper generations more broadly, including SXM5 and PCIe variants across both, see the full H100 vs H200 breakdown. And if your workload has already outgrown 141GB in either form factor, the H200 vs B200 vs GB200 decision framework covers what Blackwell adds beyond memory.
Deploy H200 on Spheron
Spheron runs H200 SXM5 on-demand and spot with per-second billing and no minimum commitment, sourced from multiple data center partners. NVSwitch fabric is included on every SXM5 node, so 8-way tensor-parallel training and large MoE serving jobs get full 900GB/s all-to-all bandwidth without any special provisioning. H200 NVL availability varies by data center partner and is not currently listed as a standalone SKU; check the H200 rental page for what's live right now. If you're working through the actual deployment, not just the hardware choice, the H200 deployment guide covers KV cache sizing, multi-model colocation, and tensor-parallel configuration, and the best NVIDIA GPUs for LLMs guide puts H200 in context against the rest of the current lineup for cost-per-token.
Whether your rack can run 700W-per-GPU liquid cooling or needs to stay air-cooled at 600W, Spheron gives you H200 SXM5 access without a long-term contract or an OEM procurement cycle.
Frequently Asked Questions
Both carry the same 141GB of HBM3e memory and the same 4.8TB/s bandwidth. The difference is power and interconnect. H200 SXM5 runs up to 700W per GPU on a full NVSwitch fabric at 900GB/s per GPU across up to 8 GPUs. H200 NVL runs up to 600W per GPU, is air-cooled and PCIe-slotted, and connects GPUs in 2-way or 4-way NVLink bridge pairs at 900GB/s per GPU, with PCIe Gen5 (128GB/s) between non-bridged GPUs in an 8-GPU server.
Standalone H200 NVL is not a common public cloud rental SKU. It ships mainly as an enterprise on-prem card through OEM partners like Dell, HPE, and Supermicro. Spheron's live GPU API lists H200 SXM5 on-demand from $4.54/GPU/hr and spot from about $3.31/GPU/hr as of 19 Jul 2026; check current GPU pricing for live rates and NVL availability by data center partner.
No. SXM5's full NVSwitch fabric gives every GPU in an 8-GPU node 900GB/s all-to-all bandwidth, which is what FSDP and tensor-parallel training need. H200 NVL's bridge only connects 2 or 4 paired GPUs directly; anything beyond that pairing falls back to PCIe Gen5 at 128GB/s, which bottlenecks multi-GPU training.
NVIDIA says roughly 70% of enterprise racks are 20kW and below and use air cooling, and SXM5's 700W-per-GPU, NVSwitch-fabric design typically needs liquid cooling and higher-density power delivery than those racks support. H200 NVL drops to 600W per GPU, ships as a standard PCIe card, and slots into existing air-cooled enterprise servers from Dell, HPE, Lenovo, Supermicro, and other OEM partners without a custom NVLink board. (Source: NVIDIA's H200 NVL announcement, blogs.nvidia.com.)
Yes. Spheron lists H200 SXM5 on-demand and spot instances from multiple data center partners with per-second billing and no minimum commitment. H200 NVL availability varies by partner and is not currently surfaced as a standalone SKU.
