Comparison

Azure GB300 Pricing 2026: What ND GB300 v6 Actually Costs vs Spheron

Azure GB300 PricingAzure ND GB300 v6Azure GB300 NVL72 CostGB300 NVL72 Azure AvailabilityND GB300 v6 PricingAzure GB300 NVL72 Production ClusterBlackwell Ultra Azure
Azure GB300 Pricing 2026: What ND GB300 v6 Actually Costs vs Spheron

Azure has not published a per-hour price for ND GB300 v6 anywhere, not on its Linux VM pricing page, not in its GB300 NVL72 cluster announcement, not in the ND GB300 v6 series documentation. If you've been searching for an Azure GB300 pricing number to plug into a budget spreadsheet, it doesn't exist yet in list form. What does exist is a confirmed spec sheet, a confirmed production deployment, and a sales-gated path to get a quote. This post covers all three, plus what comparable rack-scale Blackwell Ultra capacity costs on the providers that do publish a number.

All information in this article is current as of 9 Aug 2026 and can change as Azure rolls out broader GB300 access. Check current GPU pricing for live self-serve rates.

Azure GB300 Pricing 2026: Quick Answer

There is no self-serve hourly rate for ND GB300 v6 (Standard_ND128isr_GB300_v6). Azure's own Linux VM pricing page lists no per-hour figure for the GB300 v6 series at all, unlike H100 or H200 SKUs which show list prices directly. The page routes GB300-class inquiries to "Request a pricing quote" or the pricing calculator instead. What's confirmed is the hardware: 4 B300 GPUs and 2 Grace CPUs per VM, 18 VMs per rack for a full GB300 NVL72, and a production cluster of more than 4,600 GPUs already running for OpenAI. Third-party cost estimator CloudPricer models the VM at roughly $28.12/hr in West US 3 and East US 2, but that's a modeled estimate, not an Azure list price, and it swings as high as $168.72/hr in other regions.

What's confirmedWhat's not
4x B300 GPU, 2x Grace CPU per VMPer-hour VM rate
18 VMs (72 GPUs) per NVL72 rackPer-GPU or per-rack price
4,600+ GPU production cluster for OpenAI (Oct 9, 2025)Self-serve region/SKU availability
Sales-gated quote processReserved vs pay-as-you-go tiers

GB300 NVL72 Azure Availability: What's Confirmed and What Isn't

GB300 NVL72 Azure availability breaks into two very different facts. First: the hardware is real, deployed, and running in production today. Second: there is no published path for a team outside that one deployment to self-serve a region, a quota, or a checkout page for it. Both are true at once, and conflating them is where most confusion about "is GB300 available on Azure" comes from.

On October 9, 2025, Microsoft announced its first at-scale production cluster built on NVIDIA GB300 NVL72, more than 4,600 GPUs, purpose-built to serve OpenAI's largest models. This is the deployment that proved GB300 could run at supercomputer scale in production rather than in a lab benchmark, and it's the reason Azure GB300 NVL72 availability shows up in search at all right now: the cluster exists, it's running, and it's the clearest public signal that GB300 capacity is real Azure inventory, not a roadmap slide.

The cluster runs on NVIDIA Quantum-X800 InfiniBand at 800 Gbps of cross-rack bandwidth per GPU, arranged in a full fat-tree, non-blocking topology designed to scale to tens of thousands of GPUs. Ian Buck, NVIDIA's VP of Hyperscale and HPC, put it plainly: "This co-engineered system delivers the world's first at-scale GB300 production cluster, providing the supercomputing engine needed for OpenAI to serve multitrillion-parameter models."

That framing matters for anyone evaluating GB300 NVL72 Azure availability for their own workload. This cluster was built for one customer at a scale almost no other team will approach. It confirms the hardware works at production scale; it says nothing about whether a mid-sized team can get 8 or 72 GPUs of it without an OpenAI-sized commitment behind the request.

Azure's own ND GB300 v6 documentation, updated as recently as July 2026, lists full host specs for Standard_ND128isr_GB300_v6 but no region list and no quota-request flow distinct from the general Azure capacity request process. That's a meaningful gap: for an established SKU like H100, Azure publishes region availability directly. For ND GB300 v6, the only publicly documented access route right now is the same enterprise sales and pricing-quote path covered in the next section, not a self-serve region picker. CoreWeave deployed the first GB300 NVL72 industry-wide in July 2025 and AWS reached general availability with P6e-GB300 UltraServers in December 2025, and both gate access the same way Azure does, so this isn't an Azure-specific bottleneck. It's the current state of rack-scale Blackwell Ultra across every major cloud.

ND GB300 v6 Specs: What One VM and One Rack Actually Contain

One Standard_ND128isr_GB300_v6 VM has two NVIDIA Grace CPUs and four B300 (Blackwell Ultra) GPUs, each carrying 288 GB of HBM3e. Add 128 vCPUs and 864 GiB of LPDDR5X system memory, connected by 4x1.8 TB/s of NVLink bandwidth per VM. Each B300 GPU also gets its own 800 Gbps of InfiniBand, twice the 400 Gbps per GPU that ND GB200 v6 ships with, which is the single biggest networking jump between the two generations.

Eighteen of these VMs make a full rack: 72 B300 GPUs, 36 Grace CPUs, roughly 37 TB of fast memory (about 20 TB HBM3e plus about 17 TB CPU memory), and 130 TB/s of intra-rack NVLink bandwidth. At rack scale, Microsoft states ND GB300 v6 delivers up to 1.44 exaFLOPS of FP4 Tensor Core performance and has demonstrated roughly 1.1 million tokens/second of LLM inference throughput per rack, about 27% higher than ND GB200 v6 on the same benchmark. Against ND GB200 v6 chip for chip, GB300 brings 1.5x the FP4 compute, 50% more HBM capacity (288 GB vs 192 GB per GPU), and 2x the back-end network bandwidth per GPU.

For the deeper architecture walkthrough of the NVL72 rack design that both GB200 and GB300 share, see the GB200 NVL72 guide. If you want the single-GPU spec breakdown of the B300 chip itself before comparing rack-level numbers, the B300 Blackwell Ultra guide has it.

Why GB300 Pricing Isn't Self-Serve Yet

This isn't unique to Azure. Every major GB300 NVL72 deployment announced in the last year has shipped the same way: confirmed specs, a confirmed customer story, and a sales conversation instead of a checkout button. AWS made EC2 P6e-GB300 UltraServers generally available on December 2, 2025, and its own announcement tells customers to "contact your AWS sales representative" rather than listing a rate. Together AI's GB300 NVL72 product page shows the same pattern at the pricing-tier level: on-demand is a blank dash, and every reserved tier from 7 days upward reads "Contact us." Rack-scale Blackwell Ultra capacity is still new enough, and demand-constrained enough, that no major provider has moved it to a self-serve price list.

The practical read for a buyer: if you're comparing Azure GB300 pricing against a competitor's published number, you're likely comparing a real Azure quote against another provider's list price for a different, smaller SKU. That's not really a comparison until you have both numbers in hand.

What Azure's Own Pricing Page Shows (and Doesn't)

Azure's Linux VM pricing page lists per-hour rates for its established GPU SKUs, H100 and H200 both show a list price directly on the page. GB300 v6 does not. Instead, the page directs GB300-class inquiries to "Request a pricing quote" or the standalone pricing calculator, the same gated path Azure uses for other capacity that isn't broadly self-serve yet. That's a meaningful signal on its own: Azure differentiates between SKUs it's ready to sell at scale (published rate) and SKUs it's still allocating carefully (quote only), and GB300 v6 is currently in the second bucket.

Third-party estimator CloudPricer models Standard_ND128isr_GB300_v6 at $28.12/hr in West US 3 and East US 2, with estimates running as high as $168.72/hr in other regions. Treat that range as a modeling exercise built from public rate cards and comparable SKUs, not a confirmed Azure number, especially given how wide the regional spread is. A 6x gap between the cheapest and most expensive modeled region is a sign the underlying data is thin, not that Azure's actual regional pricing varies that dramatically.

What Comparable Rack-Scale Capacity Costs Elsewhere

No cloud provider currently publishes a self-serve, checkout-ready GB300 NVL72 rate. What does exist are adjacent data points that triangulate the range: analyst rack-cost estimates, one hyperscaler's published Capacity Blocks rate for standalone B300 accelerator time, and a third-party VM estimate for Azure specifically.

SourceFigureWhat it actually measures
Loop Capital (analyst estimate)$3.7M-$4.0M per rackEstimated GB300 NVL72 rack hardware cost
SemiAnalysis (analyst estimate)~$3.1M bare / ~$3.9M all-inEstimated GB200 NVL72 rack cost, prior generation, for comparison
AWS EC2 Capacity Blocks for ML$14.04/accelerator/hr (effective Jul 1, 2026)Standalone P6-B300 accelerator time, not NVL72 rack capacity
CloudPricer (third-party model)$28.12/hr (West US 3, East US 2)Modeled ND128isr_GB300_v6 VM rate, unconfirmed by Azure

The Loop Capital and SemiAnalysis figures both price the physical rack, not an hourly rental rate, and neither is an official NVIDIA or cloud-provider number; they're the two most-cited analyst estimates and put the GB300 premium over GB200 at roughly 10-30%, not double. For the full breakdown of that rack-price gap and where each estimate comes from, see the GB300 NVL72 vs GB200 NVL72 pricing and availability guide.

AWS P6e-GB300 UltraServers and CoreWeave's GB300 Rate

AWS made P6e-GB300 UltraServers generally available on December 2, 2025, delivering the same 1.5x GPU memory and 1.5x FP4 compute jump over P6e-GB200 that Azure's ND GB300 v6 shows over ND GB200 v6. Like Azure, AWS doesn't list a self-serve UltraServer rate; the GA announcement points customers to an AWS sales representative. The closest AWS gets to a public GB300-class number is its EC2 Capacity Blocks for ML pricing, which set the standalone P6-B300 accelerator rate at $14.04/hr effective July 1, 2026, after a broader roughly 20% price increase across its GPU instance lineup. That figure prices a single B300 accelerator reservation, not a full NVL72 rack, so it's a useful anchor point rather than a direct GB300 NVL72 comparison.

CoreWeave powered on the first GB300 NVL72 deployment industry-wide in July 2025, ahead of both Azure's and AWS's production announcements, and it remains the provider with the longest GB300 track record. CoreWeave hasn't published a self-serve GB300 NVL72 hourly rate either; access there runs through the same enterprise sales motion as Azure and AWS. What CoreWeave does have that Azure doesn't yet is a longer operating history on the hardware, which matters if track record is part of your evaluation, but it doesn't get you closer to a checkout page.

Self-Serve Blackwell Alternatives While You Wait

Waiting on an Azure GB300 NVL72 sales quote doesn't mean waiting on Blackwell hardware. Spheron has B200 on-demand from about $7.35/GPU/hr, with spot from about $3.89/GPU/hr, and B300 on spot from about $5.81/GPU/hr, all self-serve, no sales call, provisioned in minutes rather than a quote-and-approval cycle. B300 on-demand capacity isn't live at the moment; if your workload needs guaranteed on-demand B300, check current availability before committing to a timeline.

Pricing fluctuates based on GPU availability. The prices above are based on 9 Aug 2026 and may have changed. Check current GPU pricing → for live rates.

For rack-scale needs specifically, Spheron also keeps GB300 NVL72 and GB200 NVL72 open for reservation outside the AWS, Azure, and CoreWeave sales process, sized from a single 8-GPU node up to a full rack. Pricing is quote-based and depends on quantity, commitment length, region, and networking, the same variables that drive every rack-scale quote industry-wide, but the process doesn't require an enterprise account relationship to start. If a full rack is more than your workload needs right now, the GPU cloud pricing comparison covers where B200 and B300 non-rack capacity sits against every other major provider, and the top cloud GPU providers guide is a starting point if Azure specifically isn't a hard requirement for your stack.

For multi-node setup once you've picked a configuration, the distributed training guide covers PyTorch DDP and DeepSpeed across multi-GPU clusters.

Azure's ND GB300 v6 spec sheet is public; the price isn't, not yet. If your team needs Blackwell Ultra capacity on a timeline shorter than an enterprise sales cycle, self-serve B200 and B300 are live today.

Check B300 availability on Spheron →

FAQ / 04

Frequently Asked Questions

Azure has not published a per-hour rate for ND GB300 v6 (Standard_ND128isr_GB300_v6) anywhere, including its own Linux VM pricing page, which routes GB300 inquiries to a sales quote instead of a list price. Third-party estimator CloudPricer models it at roughly $28.12/hr in West US 3 and East US 2, up to $168.72/hr in other regions, but that is a third-party model, not an Azure-confirmed figure.

Yes, but only through enterprise sales. Microsoft announced its first at-scale GB300 NVL72 production cluster, more than 4,600 GPUs built for OpenAI, on October 9, 2025. There is no self-serve console flow for ND GB300 v6; access runs through an Azure account team the same way AWS gates its P6e-GB300 UltraServers.

Standard_ND128isr_GB300_v6 has two NVIDIA Grace CPUs, four B300 (Blackwell Ultra) GPUs at 288 GB HBM3e each, 128 vCPUs, 864 GiB of LPDDR5X system memory, and 4x1.8 TB/s NVLink bandwidth per VM. Each B300 GPU also gets 800 Gbps of InfiniBand, double the 400 Gbps per GPU on ND GB200 v6.

Spheron has B200 on-demand from about $7.35/GPU/hr and B300 on spot from about $5.81/GPU/hr, both self-serve with no sales call, as of 9 Aug 2026. Spheron also has GB300 NVL72 and GB200 NVL72 rack capacity open for reservation outside the hyperscaler sales process, sized from a single node up to a full rack.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute with no contracts and no minimum. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min