Comparison

AMD MI300X & MI355X Pricing 2026: The Real On-Demand Rates

Back to BlogWritten by Published Jul 8, 2026Updated
mi300x pricingmi355x pricingmi455x pricerent amd mi300x cloudAMD MI300XAMD MI355XGPU Cloud PricingAMD GPU Rental
AMD MI300X & MI355X Pricing 2026: The Real On-Demand Rates

AMD's Instinct lineup rents for anywhere between $0.95 and $8.60 an hour per GPU, and the provider you pick matters more than the chip you pick. An MI300X on a bare-metal hyperscaler shape can cost eight times what the same GPU costs on a spot marketplace. This guide breaks down real 2026 rental prices across providers, explains why the spread is so wide, and works through the cost-per-token math against NVIDIA's H100 and H200.

For the hardware comparison behind these numbers, see our AMD MI300X vs NVIDIA H100, AMD MI300X vs NVIDIA H200, and AMD MI350X vs NVIDIA B200 breakdowns. If you're deciding whether to train on AMD hardware at all, our ROCm pretraining guide covers the software side in depth. If you need more memory than MI300X's 192GB but aren't ready for MI355X, our MI325X pricing and availability breakdown covers the 256GB chip that sits between them.

TL;DR: AMD MI300X & MI355X Pricing 2026

AMD MI300X rents on-demand at $3.59/hr per GPU on Spheron as of 06 Oct 2026, against $2.59/GPU-hr at DigitalOcean and $6.00 at Azure and Oracle on 16 Aug 2026. MI355X has one verified on-demand rate: Oracle at $8.60/GPU-hr. Compare live GPU rates.

ProviderMI300X rateDeployable on 17 Sep 2026
Spheron$3.59/hr on-demand, as of 06 Oct 2026Yes, self-serve 1x or 2x
Runpod$2.39/hr listed in its consoleNo, console showed Unavailable
DigitalOcean$2.59/hr on-demand (16 Aug 2026)Not checked
Azure ND MI300X v5$6.00/hr on-demand, $1.11/hr spot (16 Aug 2026)Not checked; quota request, 8-GPU VM

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 06 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

MI300X and MI355X On-Demand Pricing Across Providers (2026)

MI300X: $2.59/hr On-Demand, or $1.11/hr if You Can Take a Reclaim

MI300X ships with 192 GB of HBM3 memory and 5.3 TB/s of bandwidth on AMD's CDNA 3 architecture, and it's the more mature of the two chips in terms of provider coverage. Fluence's 2026 provider survey put spot and marketplace listings as low as $0.95/GPU-hr and hyperscaler rates at $6.00-$7.86/GPU-hr. We re-checked the named providers directly rather than relying on aggregate surveys, because several widely-cited MI300X rates do not survive contact with the vendor's own page.

Two figures need care before they go into a comparison. The "$2.39 on Runpod" number is real: Runpod's console lists MI300X at $2.39/hr, but it showed the GPU as Unavailable on 17 Sep 2026, so the rate existed and could not be deployed that day. The "$1.99 AMD Developer Cloud" rate is one to drop entirely, since it traces to June 2025 reporting and is contradicted by DigitalOcean's current $2.59. Here is what the vendors published when we checked on 16 Aug 2026, plus Spheron's live rate, with billing mode called out:

ProviderBilling mode$/GPU-hrNotes
TensorWaveNot stated by vendor$1.71"Starting at" price, quote-driven
VultrPreemptible$1.85$1.75 on 24-month prepay. Reclaimable, not on-demand
Azure ND96isr_MI300X_v5Spot$1.11Verified via the Azure Retail Prices API
DigitalOceanOn-demand$2.59$20.72/hr for 8 GPUs; $1.91 on a 12-month reservation
SpheronOn-demand$3.59/hrPer-minute after 20-min minimum (60 min on 2x), 1x and 2x, self-serve
Hot AisleOn-demand$2.99Per-minute, 1x/2x/4x VMs; 8x bare metal $3.39/hr, 1-month minimum
Crusoe CloudOn-demand$3.458-GPU node
CirrascaleMonthly commitment$3.85$22,499/month minimum
Azure ND96isr_MI300X_v5On-demand$6.00$48/hr per node, eastus2 and westus3
Oracle Cloud BM.GPU.MI300X.8On-demand bare metal$6.00$48/hr per node

Worth flagging: Fluence's April 2026 survey had Azure's MI300X rate at $7.86/GPU-hr ($62.85/hr for the 8-GPU node), noticeably higher than the $6.00/GPU-hr the Azure Retail Prices API returns now. Hyperscaler GPU pricing moves, and a rate you read in one article can already be stale by the time you check out. Always confirm current pricing directly with the provider before committing. For the full rate card on that Azure instance, including regional deltas and why it lacks a reserved discount tier, see our Azure MI300X pricing breakdown.

Read that table by billing mode, not by price. Of the rate cards we checked on 16 Aug 2026, the genuinely on-demand ones are DigitalOcean at $2.59/GPU-hr, Crusoe at $3.45, and Azure and Oracle at $6.00 each. Spheron's self-serve MI300X GPU rental joined that on-demand group on 10 Sep 2026 and runs $3.59/hr per GPU as of 06 Oct 2026, rented as 1x or 2x VMs. Vultr's $1.85 is preemptible, Azure's $1.11 is spot, and TensorWave's $1.71 is a quote anchor rather than a rate card. Azure is the clearest illustration: the same instance is $6.00 on-demand and $1.11 on spot, a 5.4x spread on identical hardware, which is why quoting "Azure MI300X pricing" without the billing mode is meaningless.

The same trap shows up across chip vendors, not just within one. Vultr's preemptible 8x MI300X cluster at $14.80/hr works out to $1.85/GPU-hr, which reads as cheaper than the $2.19/hr Thunder Compute charges for an on-demand H100, but that's a preemptible rate against an on-demand one. Match the billing mode before you match the number.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 06 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

MI355X: One Verified On-Demand Rate, and a Lot of Preemptible Capacity

MI355X is AMD's CDNA 4 part: 288 GB of HBM3E per GPU at 8 TB/s, 10.1 PFLOPS of dense FP4, 256 compute units, 185 billion transistors on TSMC 3nm, and a 1,400 W total board power. Those figures are from AMD's own MI355X product page.

Here is where the published pricing actually sits, with billing mode stated explicitly, because that is where nearly every comparison of this chip goes wrong:

ProviderBilling mode$/GPU-hrNotes
Oracle OCIOn-demand$8.60BM.GPU.MI355X.8 at $68.80/hr for the node. The only verified true on-demand rate.
TensorWaveNot stated by vendor$2.95"Starting at"; CTA is "Book a Call", so treat as a quote anchor, not a rate card
VultrPreemptible$2.59Vultr's own page says "on-demand preemptible instances". Reclaimable.
DigitalOceanSpot only$4.50No on-demand MI355X SKU exists. Pricing page notes new pricing effective 1 Aug 2026
CrusoeContact salesn/aNo published price

The $2.59 figure is not an on-demand rate. Vultr's own pricing page describes these as preemptible instances, which means they can be reclaimed. Comparing that number against an on-demand MI300X or H100 rate is not a like-for-like comparison, and it is the single most common error in MI355X pricing articles. An earlier version of this post made it too.

Two more traps worth naming, because we hit both while checking:

  • Vultr's own two pages disagree on the 48-month prepaid rate ($2.650 on the pricing page, $2.290 on the product page). When a vendor contradicts itself, don't pick the number that flatters your argument. Our Vultr cloud GPU pricing breakdown covers how its preemptible tier is structured and where the contract rates actually land.
  • DigitalOcean's docs list $11.19 where its pricing page lists $4.50. The $11.19 figure is byte-identical to DigitalOcean's B300 spot price, which suggests a copy error rather than a real rate.

Availability as of August 2026: generally available on Oracle OCI and Vultr, spot-only public preview on DigitalOcean, contact-sales at TensorWave and Crusoe. Azure is not offering MI355X at all, having moved to the MI455X generation. Nscale has dropped AMD entirely. Runpod, Vast.ai and Together AI do not list it.

Pricing fluctuates based on GPU availability. The prices above are based on 21 Aug 2026 and may have changed. Check current GPU pricing → for live rates.

Why the Price Spread Is So Wide

A single GPU model shouldn't cost 8x more from one provider to the next for the same silicon. Four factors explain most of the gap.

The sparsity trap in MI355X benchmark numbers

Before comparing any MI355X throughput figure, check whether it is dense or sparse. Oracle's MI355X launch blog quotes FP16 at 5 PFLOPS, FP8 at 10.1 PFLOPS and FP4 at 20.1 PFLOPS. Those are with-sparsity numbers presented without the label. AMD's own dense figures are exactly half: 2.5 PFLOPS FP16, 5.0 PFLOPS FP8, 10.1 PFLOPS FP4. Same silicon, different convention, and mixing the two makes AMD look twice as fast as it is.

A second one: some spec aggregators list MI355X at 80.5 PFLOPS FP4 with 2,048 compute units and 2.3 TB of memory. Those are 8-GPU platform totals mislabelled as per-GPU. Divide by eight.

SpecMI300XMI355X
Memory192 GB HBM3288 GB HBM3E
Memory bandwidth5.3 TB/s8 TB/s
FP4 densenot supported10.1 PFLOPS
FP8 dense / sparse2.61 / 5.22 PFLOPS5.0 / 10.1 PFLOPS
FP16 dense / sparse1.3 / 2.61 PFLOPS2.5 / 5.0 PFLOPS
FP64 vector81.7 TFLOPS78.6 TFLOPS
Total board power750 W1,400 W
ArchitectureCDNA 3, 304 CU, 153B transistorsCDNA 4, 256 CU, 185B transistors

Note the FP64 line: MI355X is slightly slower than MI300X at double precision despite being the newer part, and it draws nearly twice the power. Like NVIDIA's Blackwell Ultra, CDNA 4 is an inference-optimised design.

VRAM Configuration: 192GB MI300X vs 288GB MI355X

MI355X's extra 96 GB of HBM3e over MI300X isn't free. The newer memory stacks cost more to source, yield rates on a new chip are lower in the first year of production, and providers pass that through directly. It's the same dynamic that made H200 cost more than H100 at launch: more memory per GPU, higher bill of materials, higher rental rate, at least until supply catches up with demand.

Bare Metal vs Virtualized Instances

Oracle's MI300X and MI355X offerings are both bare metal: you get the entire physical node, no hypervisor overhead, no noisy-neighbor risk from other tenants. Azure's ND MI300X v5 and Vultr's MI355X pods are virtualized, sharing the underlying hardware layer with other customers even when you're billed for the whole node. Bare metal costs more to provision and operate, and hyperscalers price that into the rate. It's not purely a markup; you're paying for isolation.

Reserved and Annual Commitments vs On-Demand

CoreWeave's MI355X drops from $7.20/hr on-demand to $4.25/hr with a 1-year reservation, a 41% cut. Cirrascale skips on-demand billing entirely and requires a $22,499/month minimum commitment to get its $3.85/GPU-hr MI300X rate. Both are the same trade: you give up the ability to walk away in exchange for a lower rate. That math only works if you know your workload will run continuously for the length of the commitment. If your usage is bursty or you're still validating a model architecture, on-demand or spot pricing is worth the premium for the flexibility.

Neocloud vs Hyperscaler Markup

The clearest pattern in both tables above: specialist neoclouds (TensorWave, Vultr, Crusoe) consistently undercut hyperscalers (Oracle, Azure) by 2-4x on the same hardware. Hyperscalers carry enterprise support contracts, compliance certifications, and integration with their broader cloud stack (IAM, VPC peering, managed storage) that neoclouds don't offer. If you need that, the markup buys something real. If you just need GPU-hours for a training or inference job, you're paying for infrastructure you won't touch.

When MI300X/MI355X Beats an H100 or H200 on Cost Per Token

The short answer: AMD tends to win when your model is large enough that memory capacity, not raw throughput, is the bottleneck. The third-party rates we surveyed on 16 Aug 2026 put MI300X on-demand pricing roughly 15-40% below H100 SXM5 pricing for comparable inference throughput. Spheron's own MI300X and H100 rates are live and not part of that survey, so check both on the day before assuming the gap holds. Where the gap holds, price alone often decides the choice, but the memory story is what makes it decisive for large models.

The Memory Advantage: Fewer GPUs Per Model

A single MI300X carries 192 GB of HBM3, enough to host models in FP16 that would need a 2x H100 (80 GB each) setup on NVIDIA hardware. Cutting a two-GPU node down to one doesn't just save GPU-hours, it removes the NVLink/interconnect complexity and the multi-GPU serving overhead that comes with tensor parallelism. For the sizing math behind where that threshold falls for specific model sizes, see our GPU memory requirements guide.

Here's a worked example that keeps both sides on the same footing: same source, same billing mode, on-demand only, no hyperscaler markup or spot floor on either side. That 16% gap from Thunder Compute's own MI300X-versus-H100 comparison above ($1.85/GPU-hr against $2.19/GPU-hr) survives once throughput enters the picture. ROCm delivers 90-95% of CUDA throughput on standard PyTorch/vLLM inference; at 92% of H100's estimated 19,500 tokens/sec (from our ROCm vs CUDA throughput benchmarks), an MI300X running near 17,900 tokens/sec at $1.85/GPU-hr works out to about $0.029 per million tokens, versus about $0.031 per million tokens on the $2.19/GPU-hr H100. That's a modest, single-digit-percent edge, not the multiple you'd get by pairing a spot-market AMD listing against a premium NVIDIA one. Treat the throughput figures as estimates, not measured benchmarks, and rerun the math against your own provider quotes before committing budget.

It is worth pairing that arithmetic with a job somebody actually timed. Our H100 vs MI300X fine-tuning cost benchmark found a roughly 20% cheaper Runpod MI300X hourly rate turning into a job that cost about 51% more, because the same 375-step LoRA run took about 34 minutes against the H100's 18 on 13 Sep 2026. Cost per token and cost per job are not the same number.

MI355X posted AMD's strongest MLPerf Inference 6.0 result to date, published April 1, 2026: 92% of NVIDIA B300 throughput in offline mode, 93% in server mode, and it actually beat B300 at 104% in interactive mode. B200 wasn't part of that submission, so B300 is the real comparison point, and a single-digit gap against AMD's toughest NVIDIA competitor is a meaningfully closer result than prior generations posted.

Where H100/H200 Still Win

The memory and price advantages don't hold at every batch size. At batch size 1-4 (low-latency, single-request serving), H100 with TensorRT-LLM holds a 20-30% throughput edge over ROCm. That gap narrows to 5-10% at batch size 64-128 (high-throughput serving with many concurrent users), where PyTorch and vLLM on ROCm close most of the distance. If your workload is latency-critical chat or single-user inference rather than high-concurrency batch serving, H100 or H200 is still the safer default. Our best NVIDIA GPUs for LLMs guide ranks H100, H200, and B200 by use case if you're weighing that trade-off directly.

ROCm Software Readiness Before You Commit

AMD's HIP translation layer converts most CUDA code automatically, and PyTorch, vLLM, and SGLang all carry official ROCm support in 2026. For standard inference workloads, ROCm reaches 90-95% of CUDA throughput on both MI300X and MI355X. The gap widens for anything that depends on TensorRT-LLM or FlashAttention 3, which don't have full ROCm equivalents yet, and for teams running custom CUDA kernels that need manual porting rather than an automatic HIP conversion.

The practical rule: if your stack is PyTorch, vLLM, or SGLang and you're running standard transformer inference, ROCm compatibility is not the blocker it was two years ago. If you're leaning on TensorRT-LLM-specific optimizations or hand-written CUDA kernels, budget engineering time before you move workloads over. Our full ROCm vs CUDA comparison covers the framework compatibility matrix in more detail, and HipKittens landing as an official AITER backend is the newest reason that "hand-written CUDA kernel" gap is smaller on MI355X than it was even six months ago.

MI455X Pricing Outlook: What Comes After MI355X

The next tier up has now launched. AMD announced the MI400 series on 23 July 2026 at Advancing AI 2026, with the MI455X as the shipping SKU: 432 GB of HBM4 at 23.3 TB/s, 40.3 PFLOPS of dense FP4, 320 billion transistors on TSMC 2nm and 3nm, CDNA 5. Specs are on AMD's MI455X page. AMD does not publish a TDP for it, so the 900 W figure circulating is third-party only and should not be relied on.

The rack-scale system is Helios: 72 MI455X plus 18 EPYC Venice CPUs per rack, 2.9 exaflops of FP4, 31 TB of HBM4, with first shipments in Q3 2026 ramping into 2027. OEM partners are Bull, HPE, Lenovo and Supermicro; Dell is frequently listed and is not on AMD's Helios OEM list. Cloud availability still lags silicon by 3-6 months while providers qualify hardware and drivers, so expect MI455X cloud instances in early-to-mid 2027. No provider has published MI455X rental rates, so treat any dollar figure you see today as speculation. If MI300X or MI355X pricing above doesn't fit your memory budget and you can wait, the AMD Helios and MI455X deployment guide covers the architecture, UALink fabric, and what fractional MI400 capacity will look like when it lands. For a head-to-head against NVIDIA's current flagship, see AMD MI400 vs NVIDIA B300.

How to Pick a Provider

Run through these four checks before you commit to a rate:

  1. Match billing mode to your usage pattern. Spot and marketplace listings ($0.95-$4.80/hr) make sense for interruptible batch jobs and experimentation. On-demand ($1.71-$8.60/hr) is right for anything that needs to run reliably to completion. Annual reservations only pay off once you're confident the workload runs continuously for the full term. Spheron's MI300X sits in the on-demand bucket: it listed no spot tier and no reservation option as of 17 Sep 2026.
  2. Decide if you need bare metal. If you're running multi-tenant-sensitive workloads or need consistent low-level performance, bare metal (Oracle, Cirrascale) is worth the premium. If you just need GPU-hours, a virtualized instance from a neocloud is usually cheaper for the same chip. Spheron's MI300X is a VM, not bare metal, and runs in a single US region (Michigan).
  3. Confirm the memory tier fits your model. 192 GB (MI300X) covers most models up to roughly 90B parameters at FP16 without multi-GPU sharding. 288 GB (MI355X) pushes that further, but at current pricing it only wins on cost if you shop the low end (Vultr) rather than the bare-metal ceiling (Oracle).
  4. Test your actual framework before committing budget. ROCm compatibility varies by framework and kernel. Run your real inference or training script on a short-term rental before signing a monthly or annual commitment. A 1x MI300X on Spheron fits that test: it bills per minute after a 20-minute minimum, and you destroy it when the run ends, since these instances can't be stopped.

Spheron added the MI300X on 10 Sep 2026, on-demand at $3.59/hr per GPU as of 06 Oct 2026, in 1x and 2x VMs on a fixed Ubuntu ROCm image. MI355X is still not on Spheron. The NVIDIA side of the catalog runs H100 SXM5 on-demand at $2.65/hr, H200 SXM5 at $5.51/hr, B300 SXM6 at $12.42/hr and B200 SXM6 at $11.06/hr, all live-tracked on the pricing page and aggregated across multiple data center partners. If your cost-per-token math above points to H100 or H200 rather than AMD, that's where to compare next, and the GPU cloud pricing comparison puts those NVIDIA rates side by side with every other major provider.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 06 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.


MI300X is now bookable on Spheron: rent one or two GPUs on-demand with per-minute billing, and keep H100 and H200 on the same account if the cost-per-token math points you back to NVIDIA.

Rent MI300X GPU →

FAQ / 04

Frequently Asked Questions

Spot and marketplace listings go as low as $0.95/GPU-hr, and smaller neoclouds like TensorWave quote from around $1.71/GPU-hr. Hyperscaler on-demand rates (Oracle, Azure) run $6.00-$7.86/GPU-hr, roughly four to eight times more for the same hardware. Runpod's console lists MI300X at $2.39/hr, but it showed the GPU as Unavailable on 17 Sep 2026, so that rate could not be deployed that day. For a self-serve on-demand option, Spheron rents MI300X at $3.59/hr per GPU as of 07 Oct 2026, one or two GPUs at a time, billed per minute after a 20-minute minimum (60 minutes on 2x).

Yes, on a like-for-like basis. The cheap MI355X rates in circulation are not on-demand. Vultr's $2.59/GPU-hr is explicitly preemptible and DigitalOcean's MI355X is spot-only, so neither is comparable to an on-demand MI300X rate. The only verified on-demand MI355X price is Oracle at $8.60/GPU-hr, against $2.59 to $6.00 for on-demand MI300X at DigitalOcean, Azure and Oracle in August 2026. Compare billing modes before comparing numbers.

It depends on the model and provider pair you compare. In third-party rates surveyed on 16 Aug 2026, AMD's per-GPU rate ran roughly 15-40% below H100 SXM5 for comparable inference throughput. Spheron's live MI300X and H100 rates were not part of that survey, so compare them on the day. A single MI300X's 192GB can replace a 2x H100 setup for large models, which cuts node count in half. For latency-critical, low-batch serving, H100 with TensorRT-LLM still wins on raw throughput per dollar.

For standard PyTorch and vLLM inference workloads, yes: ROCm reaches roughly 90-95% of CUDA throughput on MI300X and MI355X. The gap widens for workloads that depend on TensorRT-LLM or FlashAttention 3, which don't have full ROCm equivalents yet.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min