Comparison

AMD MI300X & MI355X Pricing 2026: The Real On-Demand Rates

mi300x pricingmi355x pricingmi455x pricerent amd mi300x cloudAMD MI300XAMD MI355XGPU Cloud PricingAMD GPU Rental
AMD MI300X & MI355X Pricing 2026: The Real On-Demand Rates

AMD's Instinct lineup rents for anywhere between $0.95 and $8.60 an hour per GPU, and the provider you pick matters more than the chip you pick. An MI300X on a bare-metal hyperscaler shape can cost eight times what the same GPU costs on a spot marketplace. This guide breaks down real 2026 rental prices across providers, explains why the spread is so wide, and works through the cost-per-token math against NVIDIA's H100 and H200.

For the hardware comparison behind these numbers, see our AMD MI300X vs NVIDIA H100, AMD MI300X vs NVIDIA H200, and AMD MI350X vs NVIDIA B200 breakdowns. If you're deciding whether to train on AMD hardware at all, our ROCm pretraining guide covers the software side in depth. If you need more memory than MI300X's 192GB but aren't ready for MI355X, our MI325X pricing and availability breakdown covers the 256GB chip that sits between them.

MI300X and MI355X On-Demand Pricing Across Providers (2026)

MI300X: $2.59/hr On-Demand, or $1.11/hr if You Can Take a Reclaim

MI300X ships with 192 GB of HBM3 memory and 5.3 TB/s of bandwidth on AMD's CDNA 3 architecture, and it's the more mature of the two chips in terms of provider coverage. Fluence's 2026 provider survey put spot and marketplace listings as low as $0.95/GPU-hr and hyperscaler rates at $6.00-$7.86/GPU-hr. We re-checked the named providers directly rather than relying on aggregate surveys, because several widely-cited MI300X rates do not survive contact with the vendor's own page.

Two figures to drop from your comparisons entirely: the "$2.39 on RunPod" number that circulates is unsupported, because RunPod's pricing tables list no AMD SKU at all, and the "$1.99 AMD Developer Cloud" rate traces to June 2025 reporting and is contradicted by DigitalOcean's current $2.59. Here is what the vendors actually publish today, with billing mode called out:

ProviderBilling mode$/GPU-hrNotes
TensorWaveNot stated by vendor$1.71"Starting at" price, quote-driven
VultrPreemptible$1.85$1.75 on 24-month prepay. Reclaimable, not on-demand
Azure ND96isr_MI300X_v5Spot$1.11Verified via the Azure Retail Prices API
DigitalOceanOn-demand$2.59$20.72/hr for 8 GPUs; $1.91 on a 12-month reservation
Hot AisleOn-demand$2.99Per-minute, self-serve, 1/2/4-GPU VMs
Crusoe CloudOn-demand$3.458-GPU node
CirrascaleMonthly commitment$3.85$22,499/month minimum
Azure ND96isr_MI300X_v5On-demand$6.00$48/hr per node, eastus2 and westus3
Oracle Cloud BM.GPU.MI300X.8On-demand bare metal$6.00$48/hr per node

Worth flagging: Fluence's April 2026 survey had Azure's MI300X rate at $7.86/GPU-hr ($62.85/hr for the 8-GPU node), noticeably higher than the $6.00/GPU-hr the Azure Retail Prices API returns now. Hyperscaler GPU pricing moves, and a rate you read in one article can already be stale by the time you check out. Always confirm current pricing directly with the provider before committing. For the full rate card on that Azure instance, including regional deltas and why it lacks a reserved discount tier, see our Azure MI300X pricing breakdown.

Read that table by billing mode, not by price. The cheapest genuinely on-demand MI300X is DigitalOcean at $2.59/GPU-hr, then Hot Aisle at $2.99. Everything below that is either preemptible (Vultr), spot (Azure at $1.11), or a quote anchor rather than a rate card (TensorWave). Azure is the clearest illustration: the same instance is $6.00 on-demand and $1.11 on spot, a 5.4x spread on identical hardware, which is why quoting "Azure MI300X pricing" without the billing mode is meaningless.

The same trap shows up across chip vendors, not just within one. Vultr's preemptible 8x MI300X cluster at $14.80/hr works out to $1.85/GPU-hr, which reads as cheaper than the $2.19/hr Thunder Compute charges for an on-demand H100, but that's a preemptible rate against an on-demand one. Match the billing mode before you match the number.

MI355X: One Verified On-Demand Rate, and a Lot of Preemptible Capacity

MI355X is AMD's CDNA 4 part: 288 GB of HBM3E per GPU at 8 TB/s, 10.1 PFLOPS of dense FP4, 256 compute units, 185 billion transistors on TSMC 3nm, and a 1,400 W total board power. Those figures are from AMD's own MI355X product page.

Here is where the published pricing actually sits, with billing mode stated explicitly, because that is where nearly every comparison of this chip goes wrong:

ProviderBilling mode$/GPU-hrNotes
Oracle OCIOn-demand$8.60BM.GPU.MI355X.8 at $68.80/hr for the node. The only verified true on-demand rate.
TensorWaveNot stated by vendor$2.95"Starting at"; CTA is "Book a Call", so treat as a quote anchor, not a rate card
VultrPreemptible$2.59Vultr's own page says "on-demand preemptible instances". Reclaimable.
DigitalOceanSpot only$4.50No on-demand MI355X SKU exists. Pricing page notes new pricing effective 1 Aug 2026
CrusoeContact salesn/aNo published price

The $2.59 figure is not an on-demand rate. Vultr's own pricing page describes these as preemptible instances, which means they can be reclaimed. Comparing that number against an on-demand MI300X or H100 rate is not a like-for-like comparison, and it is the single most common error in MI355X pricing articles. An earlier version of this post made it too.

Two more traps worth naming, because we hit both while checking:

  • Vultr's own two pages disagree on the 48-month prepaid rate ($2.650 on the pricing page, $2.290 on the product page). When a vendor contradicts itself, don't pick the number that flatters your argument. Our Vultr cloud GPU pricing breakdown covers how its preemptible tier is structured and where the contract rates actually land.
  • DigitalOcean's docs list $11.19 where its pricing page lists $4.50. The $11.19 figure is byte-identical to DigitalOcean's B300 spot price, which suggests a copy error rather than a real rate.

Availability as of August 2026: generally available on Oracle OCI and Vultr, spot-only public preview on DigitalOcean, contact-sales at TensorWave and Crusoe. Azure is not offering MI355X at all, having moved to the MI455X generation. Nscale has dropped AMD entirely. RunPod, Vast.ai and Together AI do not list it.

Pricing fluctuates based on GPU availability. The prices above are based on 16 Aug 2026 and may have changed. Check current GPU pricing → for live rates.

Why the Price Spread Is So Wide

A single GPU model shouldn't cost 8x more from one provider to the next for the same silicon. Four factors explain most of the gap.

The sparsity trap in MI355X benchmark numbers

Before comparing any MI355X throughput figure, check whether it is dense or sparse. Oracle's MI355X launch blog quotes FP16 at 5 PFLOPS, FP8 at 10.1 PFLOPS and FP4 at 20.1 PFLOPS. Those are with-sparsity numbers presented without the label. AMD's own dense figures are exactly half: 2.5 PFLOPS FP16, 5.0 PFLOPS FP8, 10.1 PFLOPS FP4. Same silicon, different convention, and mixing the two makes AMD look twice as fast as it is.

A second one: some spec aggregators list MI355X at 80.5 PFLOPS FP4 with 2,048 compute units and 2.3 TB of memory. Those are 8-GPU platform totals mislabelled as per-GPU. Divide by eight.

SpecMI300XMI355X
Memory192 GB HBM3288 GB HBM3E
Memory bandwidth5.3 TB/s8 TB/s
FP4 densenot supported10.1 PFLOPS
FP8 dense / sparse2.61 / 5.22 PFLOPS5.0 / 10.1 PFLOPS
FP16 dense / sparse1.3 / 2.61 PFLOPS2.5 / 5.0 PFLOPS
FP64 vector81.7 TFLOPS78.6 TFLOPS
Total board power750 W1,400 W
ArchitectureCDNA 3, 304 CU, 153B transistorsCDNA 4, 256 CU, 185B transistors

Note the FP64 line: MI355X is slightly slower than MI300X at double precision despite being the newer part, and it draws nearly twice the power. Like NVIDIA's Blackwell Ultra, CDNA 4 is an inference-optimised design.

VRAM Configuration: 192GB MI300X vs 288GB MI355X

MI355X's extra 96 GB of HBM3e over MI300X isn't free. The newer memory stacks cost more to source, yield rates on a new chip are lower in the first year of production, and providers pass that through directly. It's the same dynamic that made H200 cost more than H100 at launch: more memory per GPU, higher bill of materials, higher rental rate, at least until supply catches up with demand.

Bare Metal vs Virtualized Instances

Oracle's MI300X and MI355X offerings are both bare metal: you get the entire physical node, no hypervisor overhead, no noisy-neighbor risk from other tenants. Azure's ND MI300X v5 and Vultr's MI355X pods are virtualized, sharing the underlying hardware layer with other customers even when you're billed for the whole node. Bare metal costs more to provision and operate, and hyperscalers price that into the rate. It's not purely a markup; you're paying for isolation.

Reserved and Annual Commitments vs On-Demand

CoreWeave's MI355X drops from $7.20/hr on-demand to $4.25/hr with a 1-year reservation, a 41% cut. Cirrascale skips on-demand billing entirely and requires a $22,499/month minimum commitment to get its $3.85/GPU-hr MI300X rate. Both are the same trade: you give up the ability to walk away in exchange for a lower rate. That math only works if you know your workload will run continuously for the length of the commitment. If your usage is bursty or you're still validating a model architecture, on-demand or spot pricing is worth the premium for the flexibility.

Neocloud vs Hyperscaler Markup

The clearest pattern in both tables above: specialist neoclouds (TensorWave, Vultr, Crusoe) consistently undercut hyperscalers (Oracle, Azure) by 2-4x on the same hardware. Hyperscalers carry enterprise support contracts, compliance certifications, and integration with their broader cloud stack (IAM, VPC peering, managed storage) that neoclouds don't offer. If you need that, the markup buys something real. If you just need GPU-hours for a training or inference job, you're paying for infrastructure you won't touch.

When MI300X/MI355X Beats an H100 or H200 on Cost Per Token

The short answer: AMD tends to win when your model is large enough that memory capacity, not raw throughput, is the bottleneck. Market rates put MI300X on-demand pricing roughly 15-40% below H100 SXM5 pricing for comparable inference throughput. That gap alone often decides it, but the memory story is what makes it decisive for large models.

The Memory Advantage: Fewer GPUs Per Model

A single MI300X carries 192 GB of HBM3, enough to host models in FP16 that would need a 2x H100 (80 GB each) setup on NVIDIA hardware. Cutting a two-GPU node down to one doesn't just save GPU-hours, it removes the NVLink/interconnect complexity and the multi-GPU serving overhead that comes with tensor parallelism. For the sizing math behind where that threshold falls for specific model sizes, see our GPU memory requirements guide.

Here's a worked example that keeps both sides on the same footing: same source, same billing mode, on-demand only, no hyperscaler markup or spot floor on either side. That 16% gap from Thunder Compute's own MI300X-versus-H100 comparison above ($1.85/GPU-hr against $2.19/GPU-hr) survives once throughput enters the picture. ROCm delivers 90-95% of CUDA throughput on standard PyTorch/vLLM inference; at 92% of H100's estimated 19,500 tokens/sec (from our ROCm vs CUDA throughput benchmarks), an MI300X running near 17,900 tokens/sec at $1.85/GPU-hr works out to about $0.029 per million tokens, versus about $0.031 per million tokens on the $2.19/GPU-hr H100. That's a modest, single-digit-percent edge, not the multiple you'd get by pairing a spot-market AMD listing against a premium NVIDIA one. Treat the throughput figures as estimates, not measured benchmarks, and rerun the math against your own provider quotes before committing budget.

MI355X posted AMD's strongest MLPerf Inference 6.0 result to date, published April 1, 2026: 92% of NVIDIA B300 throughput in offline mode, 93% in server mode, and it actually beat B300 at 104% in interactive mode. B200 wasn't part of that submission, so B300 is the real comparison point, and a single-digit gap against AMD's toughest NVIDIA competitor is a meaningfully closer result than prior generations posted.

Where H100/H200 Still Win

The memory and price advantages don't hold at every batch size. At batch size 1-4 (low-latency, single-request serving), H100 with TensorRT-LLM holds a 20-30% throughput edge over ROCm. That gap narrows to 5-10% at batch size 64-128 (high-throughput serving with many concurrent users), where PyTorch and vLLM on ROCm close most of the distance. If your workload is latency-critical chat or single-user inference rather than high-concurrency batch serving, H100 or H200 is still the safer default. Our best NVIDIA GPUs for LLMs guide ranks H100, H200, and B200 by use case if you're weighing that trade-off directly.

ROCm Software Readiness Before You Commit

AMD's HIP translation layer converts most CUDA code automatically, and PyTorch, vLLM, and SGLang all carry official ROCm support in 2026. For standard inference workloads, ROCm reaches 90-95% of CUDA throughput on both MI300X and MI355X. The gap widens for anything that depends on TensorRT-LLM or FlashAttention 3, which don't have full ROCm equivalents yet, and for teams running custom CUDA kernels that need manual porting rather than an automatic HIP conversion.

The practical rule: if your stack is PyTorch, vLLM, or SGLang and you're running standard transformer inference, ROCm compatibility is not the blocker it was two years ago. If you're leaning on TensorRT-LLM-specific optimizations or hand-written CUDA kernels, budget engineering time before you move workloads over. Our full ROCm vs CUDA comparison covers the framework compatibility matrix in more detail.

MI455X Pricing Outlook: What Comes After MI355X

The next tier up has now launched. AMD announced the MI400 series on 23 July 2026 at Advancing AI 2026, with the MI455X as the shipping SKU: 432 GB of HBM4 at 23.3 TB/s, 40.3 PFLOPS of dense FP4, 320 billion transistors on TSMC 2nm and 3nm, CDNA 5. Specs are on AMD's MI455X page. AMD does not publish a TDP for it, so the 900 W figure circulating is third-party only and should not be relied on.

The rack-scale system is Helios: 72 MI455X plus 18 EPYC Venice CPUs per rack, 2.9 exaflops of FP4, 31 TB of HBM4, with first shipments in Q3 2026 ramping into 2027. OEM partners are Bull, HPE, Lenovo and Supermicro; Dell is frequently listed and is not on AMD's Helios OEM list. Cloud availability still lags silicon by 3-6 months while providers qualify hardware and drivers, so expect MI455X cloud instances in early-to-mid 2027. No provider has published MI455X rental rates, so treat any dollar figure you see today as speculation. If MI300X or MI355X pricing above doesn't fit your memory budget and you can wait, the AMD Helios and MI455X deployment guide covers the architecture, UALink fabric, and what fractional MI400 capacity will look like when it lands. For a head-to-head against NVIDIA's current flagship, see AMD MI400 vs NVIDIA B300.

How to Pick a Provider

Run through these four checks before you commit to a rate:

  1. Match billing mode to your usage pattern. Spot and marketplace listings ($0.95-$4.80/hr) make sense for interruptible batch jobs and experimentation. On-demand ($1.71-$8.60/hr) is right for anything that needs to run reliably to completion. Annual reservations only pay off once you're confident the workload runs continuously for the full term.
  2. Decide if you need bare metal. If you're running multi-tenant-sensitive workloads or need consistent low-level performance, bare metal (Oracle, Cirrascale) is worth the premium. If you just need GPU-hours, a virtualized instance from a neocloud is usually cheaper for the same chip.
  3. Confirm the memory tier fits your model. 192 GB (MI300X) covers most models up to roughly 90B parameters at FP16 without multi-GPU sharding. 288 GB (MI355X) pushes that further, but at current pricing it only wins on cost if you shop the low end (Vultr) rather than the bare-metal ceiling (Oracle).
  4. Test your actual framework before committing budget. ROCm compatibility varies by framework and kernel. Run your real inference or training script on a short-term rental before signing a monthly or annual commitment.

As of 16 Aug 2026, Spheron's own catalog doesn't list AMD MI-series GPUs; it's NVIDIA-only, with H100 SXM5 on-demand at $3.98/hr, H200 SXM5 at $4.79/hr, B300 SXM6 at $9.08/hr and B200 SXM6 at $9.36/hr, all live-tracked on the pricing page and aggregated across multiple data center partners. If your cost-per-token math above points to H100 or H200 rather than AMD, that's where to compare next, and the GPU cloud pricing comparison puts those NVIDIA rates side by side with every other major provider.


If the math above points you toward NVIDIA rather than AMD for your workload, Spheron runs H100 and H200 on-demand and spot with per-minute billing and no long-term contracts.

Compare live GPU pricing on Spheron →

FAQ / 04

Frequently Asked Questions

Spot and marketplace listings go as low as $0.95/GPU-hr, and on-demand rates from smaller neoclouds like TensorWave start around $1.71/GPU-hr. Hyperscaler on-demand rates (Oracle, Azure) run $6.00-$7.86/GPU-hr, roughly four to eight times more for the same hardware.

Yes, on a like-for-like basis. The cheap MI355X rates in circulation are not on-demand. Vultr's $2.59/GPU-hr is explicitly preemptible and DigitalOcean's MI355X is spot-only, so neither is comparable to an on-demand MI300X rate. The only verified on-demand MI355X price is Oracle at $8.60/GPU-hr, against roughly $2.99 to $6.00 for on-demand MI300X. Compare billing modes before comparing numbers.

It depends on the model and provider pair you compare. AMD's per-GPU rate is typically 15-40% below H100 SXM5 for comparable inference throughput, and a single MI300X's 192GB can replace a 2x H100 setup for large models, which cuts node count in half. For latency-critical, low-batch serving, H100 with TensorRT-LLM still wins on raw throughput per dollar.

For standard PyTorch and vLLM inference workloads, yes: ROCm reaches roughly 90-95% of CUDA throughput on MI300X and MI355X. The gap widens for workloads that depend on TensorRT-LLM or FlashAttention 3, which don't have full ROCm equivalents yet.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute with no contracts and no minimum. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min