Comparison

H100 Price Per Hour in 2026: What You'll Actually Pay

h100 price per hourh100 gpu price per hour 2026h100 cloud rental priceh100 on-demand pricingh100 spot price per hourh100 hourly rateh100 reserved pricing 2026
H100 Price Per Hour in 2026: What You'll Actually Pay

The H100 price per hour depends entirely on which provider and purchase model you're looking at. On-demand runs $2.64/hr per GPU on Spheron right now, versus $6.88/hr on AWS, $10.98/hr on GCP, and $12.29/hr on Azure, a nearly 5x spread for the same silicon. Spot and preemptible capacity drops further still, down to roughly $2.04/hr in places. This post gives you the current rate table across the providers that matter, explains why on-demand, reserved, and spot behave so differently right now, and shows you how to turn a headline $/hr number into a real workload budget.

H100 Price Per Hour: Current Rates Across Major GPU Clouds (2026)

ProviderTierPer-GPU $/hrNotes
SpheronH100 PCIe on-demand$2.64Per-minute billing, no node minimum
SpheronH100 SXM5 spot$2.04Interruptible, reclaimable without notice
SpheronH100 PCIe spot$2.20Interruptible
SpheronH100 SXM5 on-demand$3.388-way HGX, NVLink
CoreWeaveHGX H100 spot~$2.468-GPU bundle, ~60% off on-demand
CoreWeaveHGX H100 on-demand~$6.16$49.24/hr for 8 GPUs, no single-GPU option
AWSp5.48xlarge on-demand~$6.88$55.04/hr for 8x H100 SXM5
GCPA3 High preemptible~$3.30-$3.69a3-highgpu-8g, ~30s reclaim notice
GCPA3 High on-demand~$10.98$87.83/hr for 8 GPUs
AzureND H100 v5 on-demand~$12.29$98.32/hr for 8x H100 SXM5, quota required

That $2.64-$12.29 range for on-demand alone is close to a 5x spread. Widen the lens to individual listings across the market and it stretches further: an industry roundup covering 15+ providers found a $1.49/hr promotional low on one marketplace against a $6.98/hr regional high on Azure, with most verified on-demand listings clustering closer to $2.89-$5.19/hr. You can check live H100 pricing on Spheron at any time rather than working from a table that's already a week stale.

Pricing fluctuates based on GPU availability. The prices above are based on 24 Aug 2026 and may have changed. Check current GPU pricing → for live rates.

For rates across the other six GPUs Spheron tracks alongside H100, see the full GPU cloud pricing comparison; this post stays focused on H100 alone. Two providers deserve a deeper look before you pick a rate off this table. AWS bundles its H100 exclusively inside the p5 instance family, which comes with a service quota process and capacity block minimums you won't hit on a marketplace. Azure's ND H100 v5 pricing carries the highest on-demand rate here, largely because the quota and reservation overhead is priced into every hour. GCP's A3 H100 instance pricing sits in between, with a genuinely useful preemptible tier if your job can checkpoint.

On-Demand vs Reserved vs Spot H100 Pricing

Three purchase models produce three different prices for identical hardware, and they're moving in different directions right now. Know which one you're actually comparing before you decide a rate is "expensive."

On-demand: no commitment, pay per hour or per minute

On-demand is the default: you spin up a GPU, pay for what you use, and walk away with zero remaining obligation. It's the rate every table above is built on. On-demand H100 pricing has fallen hard industrywide, from $8-10/hr at the 2024 peak to $1.80-3.50/hr by Q2 2026, a 64-75% decline according to Value Add VC's chip-shortage tracker. If your workload is bursty, exploratory, or you don't yet know your steady-state GPU count, on-demand is the right default even at the higher end of that range.

Reserved and committed contracts: lower rate, locked-in term

Reserved capacity trades flexibility for a discount, and in 2026 that discount has been shrinking rather than growing. SemiAnalysis reports that 1-year H100 reserved contract pricing rose almost 40%, from a low of $1.70/hr/GPU in October 2025 to $2.35/hr/GPU by March 2026, moving in the opposite direction from on-demand and spot over the same stretch. That's a genuinely odd market: the walk-up rate got cheaper while the committed rate got more expensive. SemiAnalysis attributes part of this to hoarding: "On-Demand GPU rental capacity is sold out across all GPU types, those that have locked up on-demand instances are not willing to relinquish this capacity back into the pool despite recent price hikes." Reserved still beats on-demand at Azure's list prices, up to 60% off the $12.29/hr rate, but the math only works if you can actually keep the GPU busy for the term. Idle reserved capacity is a sunk cost either way.

Spot and preemptible: cheapest, but reclaimable without notice

Spot is the cheapest tier for a reason: the provider can take the GPU back whenever it needs the capacity for a higher-paying customer, typically with seconds to no notice at all. Spheron's H100 SXM5 spot sits at $2.04/hr against $3.38/hr on-demand, roughly a 40% discount. GCP's A3 preemptible tier runs $3.30-3.69/hr against $10.98/hr on-demand, a steeper 66-70% cut. CoreWeave's spot pricing lands near $2.46/hr against $6.16/hr on-demand, about 60% off. Spot is the right call for checkpointed training runs, offline batch inference, and anything you can afford to lose and resume. It is the wrong call for a production API where a reclaimed GPU means a dropped request mid-stream.

What Drives H100 Price Differences Between Providers

The rate table above isn't noise. Three structural factors explain almost all of the spread.

Hyperscaler overhead vs neocloud margin

Hyperscalers price H100 capacity inside a managed instance stack: quota approval processes, capacity blocks, VPC networking, and enterprise support all get folded into the hourly rate whether you use them or not. Marketplace and neocloud providers sell closer to bare metal, with fewer layers between you and the card. That's most of the gap between Azure's $12.29/hr and Spheron's $2.64/hr for the same GPU generation, not raw hardware cost.

SXM vs PCIe form factor and InfiniBand networking

The cheapest "H100 price" you'll see quoted is almost always a PCIe card. PCIe runs at lower power and skips the 900 GB/s NVLink fabric that SXM5 boards carry, which is why SXM5 rents for meaningfully more, $3.38/hr vs $2.64/hr on-demand on Spheron. If your job is single-GPU inference, PCIe is usually the right call. Multi-node training that needs NVLink or InfiniBand between GPUs needs SXM5, and that premium is buying real interconnect bandwidth, not just a bigger number on a spec sheet.

Supply cycles: Blackwell ramp pushing H100 into mid-tier pricing

H100 used to be the top-of-market card; it isn't anymore. As Blackwell (B200, GB200) supply ramps for the largest training runs, H100 gets repositioned as the mid-tier, inference-heavy workhorse, and providers reprice it to stay competitive in that slot. Layer on top of that the more than 300 new neocloud GPU providers that entered the market in 2025, fragmenting demand across a much larger supply pool, and you get sustained downward pressure on on-demand and spot rates even while reserved pricing climbs.

How to Turn H100 Price Per Hour Into Total Workload Cost

A $/hr rate isn't your budget. It's one input into it.

GPU-hours x rate, plus storage and egress

The core formula is simple: total GPU-hours multiplied by your per-GPU rate, plus anything the provider bills separately. Say you're fine-tuning a 70B-parameter model on a single H100 SXM5 for 40 hours. At Spheron's $3.38/hr on-demand rate, that's $135.20 in raw compute. Add persistent storage for checkpoints (typically $0.08-$0.15/GB/month on neoclouds, more on hyperscalers) and, if you're on a hyperscaler, egress fees of $0.08-$0.12/GB every time you pull a checkpoint or dataset out. On a large model, egress alone can rival the compute line. Neoclouds, including Spheron, more commonly bundle bandwidth into the hourly rate with no separate egress charge, which is worth checking before you commit to a provider based on the headline $/hr number alone.

Per-minute vs per-hour billing rounding

Billing granularity is an easy line item to miss. A provider that bills per hour rounds every job up to the next full hour, so a 20-minute test run costs a full hour regardless. Over hundreds of short jobs, an experimentation-heavy team, notebook sessions, quick fine-tune sweeps, CI runs against a GPU, that rounding adds up fast. Spheron bills per minute, so a 20-minute job costs 20 minutes. If your workload involves a lot of short or variable-length runs rather than long steady jobs, per-minute billing is worth more than a marginally lower headline rate on an hourly plan.

Cheaper Alternatives to H100 for Similar Performance

H100 isn't always the right GPU for the job. Before you commit to its rate, check whether a cheaper card actually covers your workload.

A100 80GB for training and inference where H100's extra throughput isn't needed

The A100 vs H100 comparison is the one to read before defaulting to H100 out of habit. A100 is more than adequate for 70B-class fine-tuning, INT8 inference under 70B parameters, and any stack still running on an older CUDA toolchain, and it rents for meaningfully less than H100 across every provider in the table above. If your workload doesn't lean on H100's Transformer Engine or FP8 throughput, A100 often wins on cost-per-token.

L40S and RTX 5090 for inference-only workloads

For pure inference on models that fit comfortably in 24-48GB VRAM, L40S and RTX 5090 both undercut H100 substantially while holding up fine on latency for single-request or lightly batched serving. Neither has H100's HBM bandwidth or NVLink, so they're the wrong pick for large-batch training, but for a serving endpoint that's exactly the capacity you're paying for and not using on H100.

H200 and B200 spot as a price-competitive step up

Counterintuitively, stepping up a generation on spot can land close to H100 on-demand while giving you more VRAM headroom. Spheron's H200 SXM5 spot runs about $2.53/hr, close to H100 SXM5 on-demand at $3.38/hr, with 141GB of HBM3e against H100's 80GB. If your model is memory-bound rather than compute-bound, that trade is often worth taking, provided the workload can tolerate spot's interruption risk.


H100 PCIe and SXM5 are both live on Spheron with per-minute billing, from $2.64/hr on-demand as of 24 Aug 2026, no quota queue and no 8-GPU minimum.

H100 capacity →

FAQ / 04

Frequently Asked Questions

As of 24 Aug 2026, H100 on-demand runs from about $2.64/hr per GPU on Spheron (PCIe) up to $12.29/hr on Azure's ND H100 v5. Spot and preemptible capacity is cheaper across the board: Spheron H100 SXM5 spot is $2.04/hr, GCP preemptible A3 runs $3.30-$3.69/hr, and CoreWeave spot is about $2.46/hr per GPU. Individual listings across the wider market run from a $1.49/hr promotional low to a $6.98/hr regional high across 15+ providers. Rates move with availability, so treat any single number as a snapshot.

Not for anything that can't tolerate an interruption. Spot and preemptible H100 capacity gets reclaimed with little to no notice when the provider needs the hardware back, which is why it prices 30-70% below on-demand. It's a good fit for batch training with checkpointing, offline inference, and queue-based jobs. For a live serving endpoint where a dropped GPU means dropped requests, pay the on-demand rate or lock in reserved capacity instead.

1-year reserved H100 contracts rose almost 40% industry-wide, from a low of $1.70/hr/GPU in October 2025 to $2.35/hr/GPU by March 2026, according to SemiAnalysis, even as on-demand and spot rates kept falling over the same window. Azure's ND H100 v5 lists up to 60% off its $12.29/hr on-demand rate for a 1-year term. Reserved pricing only pays off if you can keep the GPU busy for most of the contract; idle reserved capacity is money burned either way.

Two things moved at once. More than 300 new neocloud GPU providers entered the market in 2025, according to Value Add VC, which fragmented demand across a much bigger supply pool and pushed on-demand and spot rates down. At the same time, Blackwell (B200/GB200) supply started absorbing the highest-end training workloads, which pushes H100 further into mid-tier, inference-focused pricing. On-demand H100 rates fell from roughly $8-10/hr at the 2024 peak to $1.80-3.50/hr by Q2 2026, a 64-75% decline.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute with no contracts and no minimum. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min