The result worth reading closely is buried further down the results table: at matched small scale, GB300 NVL72's advantage over the prior-generation HGX B300 shrinks to roughly 11-12%, a very different number from the 1.3-1.6x gap NVIDIA reports at 8,192 GPUs. If you're deciding what to rent for an 8- or 64-GPU training run, that second number is the one that should drive the decision.
This is the training-side counterpart to MLPerf Inference v6.0, which covered the same round's inference results back in April. Training and inference are scored on entirely different benchmarks and entirely different hardware configurations, so don't expect the same rankings to carry over.
TL;DR: What Do MLPerf Training v6.0 Results Mean for GPU Buyers?
- What's new: MLCommons published MLPerf Training v6.0 on 16 June 2026, adding a first 671B-parameter MoE benchmark, DeepSeek-V3, plus single-node GPT-OSS-20B.
- Headline number: GB300 NVL72, via CoreWeave, trained DeepSeek-V3 671B in 2.02 minutes on 8,192 GPUs, MLCommons's fastest published time.
- The number that matters more: at matched 8 GPUs, Nebius found GB300 NVL72 only ~12% faster than HGX B300 on Llama-3.1-8B, ~11% on GPT-OSS-20B.
- Field diversity: 95 systems, 24 organizations, 13 accelerators; 60% of entries multi-node.
- AMD's milestone: AMD's first multi-node entry ran FLUX.1-schnell on 64 MI325X GPUs across 8 nodes; Oracle submitted 512 MI300X separately.
- For renters: Spheron covers H100 through B300 by the hour, including B200 training cluster pricing.
What Changed in v6.0: The First 671B MoE Training Benchmark
MLPerf Training has run since 2018 on a steady rotation of dense models: ResNet, BERT, GPT-3-scale pretraining, Llama 2 70B LoRA. v6.0 is the first round to add a genuinely frontier-scale mixture-of-experts model to the training suite, and it added a second, much smaller one alongside it.
DeepSeek-V3 671B and GPT-OSS-20B: Why MLCommons Added Two MoE Benchmarks at Once
DeepSeek-V3 has 671 billion total parameters but activates only 37 billion per token, MoE routing's whole point: pay compute for a fraction of the model on any given forward pass, and pay memory and communication cost for all of it. Training this pattern well means testing something dense-model benchmarks never touch, expert routing and all-to-all communication across a cluster, not just matmul throughput on a fixed set of weights. GPT-OSS-20B (21 billion total parameters, 3.6 billion activated) is the same architecture pattern at a scale that fits inside a single 8-GPU node, giving MLCommons a MoE datapoint that doesn't require rack-scale hardware to even attempt. Pairing a frontier-scale MoE benchmark with a single-node one lets the suite score both ends of the market that actually trains these architectures: hyperscale labs pretraining from scratch, and everyone else fine-tuning or training smaller MoE variants on rented capacity.
The Submission Field: 95 Systems, 24 Organizations, 60% Multi-Node
The v6.0 round drew 95 unique systems from 24 organizations, spanning 13 different hardware accelerators and 19 host processors, which MLCommons describes as a record for hardware diversity in the Training suite. Sixty percent of all submissions were multi-node systems, and cloud-provider submissions more than doubled compared to the prior v5.1 round six months earlier. First-time submitters included Inventec, Netweb Technologies India, TTA, and Vultr, evidence the benchmark is pulling in infrastructure vendors beyond the usual hyperscaler and chipmaker roster.
GB300 NVL72's Sweep, and What It Doesn't Tell You at Sub-8,192-GPU Scale
The Headline: 2.02 Minutes on 8,192 GPUs, and How CoreWeave Got There
CoreWeave's DeepSeek-V3 671B submissions scaled close to linearly as GPU count grew: 5.54 minutes at 2,048 GPUs, 3.09 minutes at 4,096 GPUs, and 2.02 minutes at 8,192 GPUs. Near-linear scaling at that GPU count, on a MoE model whose expert routing generates heavy all-to-all traffic, is the harder engineering problem than the raw time-to-train number itself. Making a cluster that size behave close to linearly depends on the network fabric moving expert activations between GPUs without becoming the bottleneck, which is exactly the kind of workload our InfiniBand vs RoCE vs Spectrum-X decision guide walks through for teams sizing their own multi-node interconnect.
Where the GB300 Advantage Shrinks: Nebius's Matched-Scale 8-GPU Numbers
Here's the number that changes the buying decision. NVIDIA's official round shows GB300 NVL72 running 1.3x to 1.6x faster per GPU than the prior generation at CoreWeave's 8,192-GPU scale. But Nebius submitted its own comparison at matched small scale, GB300 NVL72 against its own HGX B300, both at 8 GPUs, and found GB300 NVL72 only about 12% faster on Llama-3.1-8B and about 11% faster on GPT-OSS-20B. That's not a rounding difference from the hyperscale figure. It's a different regime entirely: at 8,192 GPUs, GB300's advantages in NVLink domain size, rack-level power delivery, and cluster-wide interconnect compound into a large multiplier. At 8 GPUs, none of that infrastructure is in play yet, so what's left is closer to the raw silicon difference between B300 and the earlier Blackwell generation, and that gap is roughly 11-12%, not 30-60%.
For a mid-size team sizing an 8- to 64-GPU training run, the honest read is: GB300 NVL72's headline speedup is a real result for CoreWeave-scale clusters, and it is not the number to expect from renting a small B300 or GB300 node. Our Vera Rubin NVL4 vs NVL72 decision guide covers the same form-factor tradeoff for NVIDIA's next generation: a full 72-GPU rack is frequently more capacity than a workload actually needs, and the smaller form factor clears the bar for most training jobs outside frontier pretraining.
Software Closed Part of the Gap Without New Silicon
Some of GB300's hyperscale-scale gain didn't come from hardware at all. NVIDIA's own writeup on the round credits a roughly 1.3x rise in per-GPU DeepSeek-V3 throughput, from 1,298 to 1,648 TFLOPS/GPU, to software changes over the roughly three months between submissions, with no hardware change: CuTe DSL kernel fusions and full-iteration CUDA graphs that eliminate CPU-GPU synchronization overhead in token-dropless MoE routing. That's a useful data point independent of which GPU you're renting: a meaningful share of MoE training throughput on any given hardware generation is left on the table by software that hasn't caught up yet, and it keeps arriving well after a chip ships.
AMD's First Multi-Node Submission: MI325X on FLUX.1, Read Correctly
AMD's v6.0 submissions are a smaller story than NVIDIA's sweep, but they mark AMD's first appearance in MLPerf Training at multi-node scale, which is the scale that actually stresses an interconnect and a software stack rather than a single card.
64 MI325X GPUs Across 8 Nodes, and Oracle's 512-GPU Submission Alongside It
AMD submitted FLUX.1-schnell text-to-image pretraining on 64 MI325X GPUs spread across 8 nodes, its first-ever multi-node MLPerf Training entry. Oracle Cloud Infrastructure submitted separately, working with AMD, scaling the same FLUX.1 workload to 512 MI300X GPUs across 64 nodes. Across the round, AMD's submissions covered three chips, MI325X, MI350X, and MI355X, spread across three benchmarks: Llama 2 70B LoRA fine-tuning, Llama 3.1 8B pretraining, and FLUX.1-schnell pretraining.
What "Within 5-6% of B200" Actually Covers (and Doesn't)
AMD's vendor-stated MI355X numbers land within 5% of B200 on Llama 2-70B LoRA fine-tuning and within 6% on Llama 3.1-8B pretraining against B200's smaller HBM3e footprint. Read the scope carefully before generalizing it: those figures are MI355X, not the MI325X that ran the multi-node FLUX.1 submission, they cover two specific benchmarks rather than the full v6.0 suite, and they're AMD's own reported comparison rather than a head-to-head MLCommons scored directly. It's a genuine narrowing of the gap on the workloads it covers, not a claim that MI355X matches B200 across the board.
Best GPU for LLM Training in Late 2026, By Scale
The scale of your training run, not a single "best GPU" spec sheet, is what should drive the hardware decision. This isn't a two-GPU-generation comparison; it's a scale ladder built directly from what v6.0 actually measured, from a single fine-tuning node up to the 8,192-GPU cluster that set the headline record. Here's how the results and current rental availability map onto four scale tiers:
| Scale tier | Best-fit hardware | What it's for | Rentable by the hour today? |
|---|---|---|---|
| Single node (1-8 GPUs) | H100 SXM5, H200, RTX PRO 6000 | LoRA/QLoRA fine-tuning, small pretraining, GPT-OSS-20B-class MoE | Yes, on-demand and spot on Spheron |
| Multi-node (8-64 GPUs) | H100, H200, B200, or B300 clusters with InfiniBand | Full fine-tunes, mid-size pretraining, the scale Nebius's matched-scale comparison actually measured | Yes, Spheron custom clusters with InfiniBand on request |
| Rack / large cluster (72-512+ GPUs) | B200 or B300 clusters, GB200 NVL72 | Multi-billion-parameter fine-tunes, larger MoE pretraining | Spheron custom clusters scale to 512+ GPUs; GB200 NVL72 is quote-only at most providers, GB300 NVL72 quote-only everywhere we checked |
| Hyperscale (1,000-8,192+ GPUs) | GB300 NVL72, large MI325X clusters | Frontier pretraining runs like MLPerf's DeepSeek-V3 671B submission | Hyperscaler-only (CoreWeave, Azure, Oracle); not hourly rentable anywhere |
The middle two rows are where most teams renting GPUs actually live, and they're also the rows where MLPerf's headline GB300 numbers stop applying directly. If your workload sits at the single-node tier and the decision is a straight two-generation cost trade-off rather than a scale question, our H100 vs A100 cost guide is the more direct answer than this ladder: it goes deeper on the per-GPU dollar math for exactly that comparison, covering when the newer generation's throughput actually pays for its higher hourly rate.
What This Means If You're Renting 8-64 GPUs, Not 8,192
Two things are true at once after v6.0: NVIDIA's newest rack-scale system is the fastest training platform MLCommons has measured, and that speed advantage is concentrated at a cluster size a small or mid-size AI team doesn't operate at. Sizing hardware off the 8,192-GPU headline number, when your actual job runs on 8 or 32 GPUs, means paying for infrastructure whose main advantage doesn't show up at your scale.
Translating MLPerf Minutes Into GPU-Hours and Dollars
CoreWeave's 2.02-minute DeepSeek-V3 run used 8,192 GPUs, which works out to roughly 276 GPU-hours of raw compute time (8,192 GPUs multiplied by 2.02 minutes, converted to hours), a figure derived from the submission's own numbers rather than one MLCommons publishes directly. That math is the same math you'd run on a training job at any scale: GPU count times wall-clock hours times your hourly rate. At the 8- to 64-GPU tier this post is actually about, that means a formula like 8 GPUs times your job's runtime in hours times $5.73/hr per GPU-hour for an 8x H200 node, or the $8.38/hr rate if you're running on B200. Our cost to train a 70B parameter LLM from scratch post walks through this same GPU-hours-to-dollars conversion in full for a real pretraining budget, with the assumptions frozen and dated so the math stays checkable.
Pricing fluctuates based on GPU availability. Spheron rates above are live as of 29 Sep 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.
What's Actually Rentable by the Hour Today vs Reserve-Only
GB300 NVL72, the system behind every headline number in this round, is not an hourly SKU anywhere we've found, Spheron included, where it's listed as future availability rather than on-demand inventory. If your project genuinely needs that literal rack, budget for a reservation process and a sales conversation, not a same-day rental. Our GB300 NVL72 vs GB200 NVL72 pricing guide covers what that quote-only landscape looks like across providers in detail. If your workload fits in the 8- to 64-GPU tier, though, H100, H200, B200, and B300 are all rentable by the hour or on spot pricing right now, without a sales cycle, and non-rack clusters at that scale still clear the bar for the overwhelming majority of production fine-tuning and mid-size pretraining. With data center GPU lead times still running long for teams trying to buy hardware outright, our GPU shortage guide covers the broader strategies for securing training capacity without waiting on a rack-scale reservation.
MLPerf's hyperscale numbers are a real result for the clusters that ran them. They are not a sizing guide for a rental at 8 or 64 GPUs, where B200 and B300 close most of the gap.
Frequently Asked Questions
MLPerf Training v6.0 is the latest round of industry-standard AI training benchmarks [published by MLCommons on 16 June 2026](https://mlcommons.org/2026/06/mlperf-training-v6-0-results). It introduces the suite's first mixture-of-experts training benchmarks, DeepSeek-V3 671B (37B activated parameters per token) and GPT-OSS-20B (21B total, 3.6B activated, sized to fit one 8-GPU node), and drew 95 systems from 24 organizations across 13 different accelerators.
It depends on scale, not just raw speed. For single-node fine-tuning (1-8 GPUs), H100 or H200 on-demand covers most LoRA and QLoRA jobs. For 8-64 GPU multi-node runs, the scale MLPerf Training v6.0's Nebius comparison actually measured, B200 and B300 clusters deliver most of GB300 NVL72's benchmark speed at a fraction of the availability friction. GB300 NVL72 itself only earns its premium at rack scale (72+ GPUs) or above, and it is quote-only at every provider we've checked, Spheron included. Check $3.96/hr for H100 and $8.38/hr for B200 on-demand, live as of 29 Sep 2026.
Not anywhere we've found. GB300 NVL72 is quote-only or reserved-capacity industry-wide, including on Spheron's marketplace, where it's listed as future availability rather than an hourly SKU. If your workload fits in a non-rack B200 or B300 cluster, that's the faster path to actually renting capacity this week instead of entering a sales cycle.
It depends entirely on scale. NVIDIA's own MLPerf Training v6.0 numbers show GB300 NVL72 delivering roughly 1.3x to 1.6x the per-GPU throughput of the prior generation at CoreWeave's 8,192-GPU scale. But [Nebius's matched-scale submission, GB300 NVL72 against its own HGX B300 at an identical 8 GPUs each, found only about a 12% speed advantage on Llama-3.1-8B and about 11% on GPT-OSS-20B](https://nebius.com/blog/posts/mlperf-training-v6-0-results). The bigger gap is a hyperscale-cluster story, not a per-GPU hardware story.
Yes, and it was [AMD's first-ever multi-node MLPerf Training submission: FLUX.1-schnell text-to-image pretraining on 64 MI325X GPUs across 8 nodes](https://rocm.blogs.amd.com/artificial-intelligence/mlperf-training-v6.0/README.html). Oracle Cloud Infrastructure separately submitted a 512-GPU MI300X FLUX.1 run working with AMD. Across the round, AMD's submissions spanned MI325X, MI350X, and MI355X on three benchmarks: Llama 2 70B LoRA fine-tuning, Llama 3.1 8B pretraining, and FLUX.1-schnell.






