Research

GPU Refresh Cycle Rental: How Fast H100s and H200s Lose Value

Back to BlogWritten by Published Oct 9, 2026
GPU Refresh Cycle RentalGPU Depreciation 2026H100 Depreciation ScheduleGPU Resale ValueH100 vs H200 DepreciationB200 DepreciationGPU Cost OptimizationGPU Cloud Pricing
GPU Refresh Cycle Rental: How Fast H100s and H200s Lose Value

A GPU refresh cycle rental decision comes down to a gap most budgets never model: the depreciation schedule your finance team books and the price a buyer will actually pay for that GPU on the secondary market are two different numbers, and the gap gets wider every quarter a generation ages. Treat depreciation as one curve and you'll plan a refresh around the wrong year. Treat it as four separate curves, one per generation, with a book number and a market number for each, and the refresh timing gets a lot clearer.

TL;DR: How Fast H100s and H200s Lose Value

  • A100 resale: same tracker shows 85-95% at year 1.
  • Book vs market: Microsoft, Alphabet, and Meta stretched server useful life to 5.5-6 years for accounting between 2023 and 2025, well past what resale value supports.
  • Cadence: NVIDIA confirmed an annual GPU release rhythm at GTC 2025, so today's GPU has a named successor within about 12 months.
  • Spheron: on-demand H100 runs $2.64/hr as of 11 Oct 2026 on Spheron's H100 GPU rental, no resale risk when the next SKU ships.

Why the GPU Refresh Cycle Compressed in 2026

Until 2024, NVIDIA's data center GPUs moved on a roughly two-year architecture cycle: Ampere, then Hopper, then Blackwell, each with a couple of years to earn out before the next one showed up. That cadence is gone. At its GTC 2025 keynote, NVIDIA said outright: "NVIDIA will follow an annual rhythm for the buildout of AI infrastructure. Each year will bring new GPUs, CPUs and accelerated computing advancements, including the upcoming NVIDIA Vera Rubin architecture, designed to drive performance gains and efficiency improvements in AI data centers."

A buyer who signed a multi-year hardware commitment in 2023 was planning against a two-year architecture window. A buyer doing the same math today is planning against a one-year window, which is the single biggest reason a GPU refresh cycle rental plan needs its own depreciation curve rather than a rule of thumb borrowed from the server-refresh playbook that preceded it.

That compression doesn't mean every GPU is obsolete the day a successor ships. It means the clock on "when does this card stop being the best available option" now resets every 12 months instead of every 24, and the resale market reacts to that clock faster than most procurement teams expect.

Depreciation Curves by Generation: A100, H100, H200, B200

GPU depreciation is not one line, it's four, and they diverge the moment each generation passes its first birthday. A100 and H100 both have multi-year tracked resale data; H200 and B200 are too recent for an independent multi-year curve to exist yet, which is its own useful signal for anyone timing a refresh around them.

GPU GenerationLaunch EraYear 1 Retained ValueYear 2-3 Retained ValueBeyond Year 3
A100 (Ampere)2020~95% at year 1~85% at year 2, ~70% at year 3~55% at year 4, ~40% at year 5, 30-35% at year 6
H100 (Hopper)2022-202395-100% at 0-12 months85-95% at 12-18 months, 75-85% at 18-24 months, 60-75% at 24-36 months45-60% at 36-48 months, 25-35% floor past 60 months
H200 (Hopper)2024No independent multi-year tracker published yetShares H100's architecture and TSMC process; expect a similar curve lagged by launch dateNot yet observed
B200 (Blackwell)2024-2025Too recent for a tracked multi-year curve; supply constraints make a premium-over-list scenario plausibleNot yet observedNot yet observed

A100 and H100 figures from Mercatus's H100 depreciation analysis. H200 and B200 rows are analytical, not tracked.

The A100 and H100 lines are the ones worth memorizing, because they're the only two generations with enough transaction history to trust. A100 loses value a little faster than H100 in the early years and then flattens out around the same mid-30s floor by year six, which tracks with it being the generation hyperscalers have had the longest window to cycle out of. H100 holds up better in year one (it was still the fastest card you could buy through most of 2023 and 2024) before the curve steepens hard once H200 and Blackwell both reached volume.

H200 shares the Hopper die and the same manufacturing process as H100, so its curve should track H100's shape with a lag roughly equal to the gap between their launch dates, likely discounted a bit faster given its memory premium shrinks as more HBM3e capacity reaches the market. B200 is under two years old as of this post, which means there simply isn't a tracked multi-year resale curve for it yet, accounting or otherwise, and any specific B200 depreciation percentage you see quoted deserves real skepticism until a tracker like Mercatus's has two or three years of actual transactions behind it.

B200 is also the one generation on this table where the usual assumption, that a GPU discounts the moment a newer SKU exists, might not hold at all for a while. A100 and H100 both depreciated into markets where supply eventually caught up with demand, which is what let a discount curve form in the first place. Blackwell's rollout has instead been defined by the opposite condition: capacity has stayed tight well past launch, and a chip that's hard to get doesn't trade like a chip that's easy to get. If that scarcity persists, book depreciation and market reality could invert from every other row on this table, a card holding near, or even above, its original price well into its second year, instead of sliding the way A100 and H100 did. Nobody has the multi-year transaction data to confirm that yet, which is exactly why it belongs in this post as an open question rather than a settled curve: don't plan a B200 refresh on the assumption that it will discount on H100's schedule, because the forces setting its price right now are supply-driven, not age-driven.

What Drives the Curve: Annual SKU Cadence and Falling Cloud Rates

Two forces push a GPU's value down the curve, and they compound. The first is NVIDIA's annual cadence itself: every new architecture resets the performance-per-dollar bar, so last generation's card has to discount to stay competitive against both the new SKU and whatever cloud rate the new SKU commands. The second is what happens to the cloud rental rate for the older chip once that newer option exists, since rental price and resale value move together more often than they diverge.

That second force isn't a one-way slide, which is a detail most depreciation write-ups skip. Cloud list prices for the same chip can rise, not just fall, when a provider's own supply tightens. Nebius raised GPU rental rates across H100 through B300 by as much as 21% in 2026, which is the opposite of what a simple "older chip, cheaper rate" model predicts. A refresh plan built on the assumption that rental rates only go down will misread a supply-driven price spike as a demand signal to hold onto older hardware longer than the resale curve actually supports.

As of 11 Oct 2026, current on-demand rates on Spheron span the generations this post covers: A100 at $1.48/hr, H100 at $2.64/hr, H200 at $5.52/hr, and B200 at $11.06/hr. Spot rates run lower on each: $1.24/hr for A100, $2.20/hr for H100, $3.35/hr for H200, and $5.22/hr for B200.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 11 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

That spread between generations is the rental-side mirror of the resale curve: the newest chip commands the steepest premium per hour, same as it commands the highest resale floor. A team deciding between H100 and A100 for a training workload is really deciding how much of that premium is worth paying for a generation with more runway left on its own depreciation curve.

Book Depreciation vs Market Resale Reality

This is the gap most TCO models never show, because they only have one of the two numbers. Book depreciation is an accounting assumption about useful life, chosen to smooth reported expense. Market resale value is what a specific card actually sells for on a given day. They are built for different jobs, and recent filings show exactly how far apart they can get.

Microsoft extended server and network equipment useful life from 4 to 6 years effective fiscal year 2023, a change that cut its depreciation expense by about $3.7 billion and raised net income by about $3.0 billion. Alphabet made the same move the same year, shifting servers from a 4-year to a 6-year schedule (and certain network equipment from 5 to 6 years), which reduced its depreciation by roughly $3.9 billion. Meta followed in 2025, moving the majority of its servers and network assets to a 5.5-year useful life effective January 29, 2025, cutting its depreciation expense by about $2.92 billion.

Read those three changes against the H100 resale curve above and the disconnect is obvious. A hyperscaler now books an H100 as losing value in a straight line over 5.5 to 6 years. The tracked market says that same card is already down to a 60-75% retained value range by its third year and a 25-35% floor past its fifth, well before the books call it fully depreciated. Neither number is wrong, exactly: one describes an accounting policy, the other describes what a buyer will pay. But only one of them tells you what your GPU is actually worth if you needed to sell it next quarter, and accounting useful-life schedules and observed secondary-market resale value are two different numbers that frequently diverge, with hyperscalers' 5-6 year book schedules not tracking actual resale trajectories.

The dollar scale of that gap is real money, not a rounding error. Mercatus's modeling shows a 100-H100 cluster's 3-year total cost of ownership swinging by about $1.7 million between a conservative and an optimistic depreciation and resale assumption. That's the cost of guessing wrong about which curve, book or market, actually governs your exit plan. It also shows up in secondary-market pricing directly: used H100 cards that sold near $40,000-$50,000 at mid-2024 peak scarcity now trade for roughly $15,000-$28,000 in 2026, about 60-70% of launch-era value depending on SKU and form factor. Form factor matters here too; a bare GPU and a full HGX or DGX node don't resell the same way, which is worth understanding before you assume a server's chassis format carries the same liquidity as the GPU inside it.

Here's what that swing looks like scaled down to a size most teams actually buy at, instead of a 100-GPU hyperscaler fleet. An 8-GPU H100 SXM5 node is an 8% slice of Mercatus's modeled cluster, so the same $1.7 million three-year TCO swing works out to roughly $136,000 of planning uncertainty on that single node, before a single hour of workload revenue is counted. Run the same node through the retained-value bands above and the two exit scenarios look nothing alike: at 30 months, the 60-75% band says the node is still worth 60-75% of what it cost; IntuitionLabs' own 2026 resale figures put real listed H100 cards at $15,000-$28,000 each, which can land below even the pessimistic end of that percentage math once scarcity-era buy-in prices are factored in. That gap, between what the percentage curve implies and what an actual dealer will cut a check for, is the number a refresh budget either plans around or gets blindsided by.

How Utilization Changes the Effective Depreciation Rate

Every curve above describes a GPU sitting in inventory, not one earning revenue. Utilization changes the number that actually matters: how much of the book value loss a given card has offset before it hits the market. A card running at 80% utilization on paid workloads for 18 months has recovered a large share of its original cost before it ever gets listed for resale. A card sitting idle in a rack, waiting for the next workload to show up, is pure book loss with nothing to show for it, and it's still sliding down the same curve either way.

That's the practical argument for not letting GPU capacity sit idle through a refresh window, whether you're the one who bought the hardware or you're deciding whether to buy at all. For a GPU owner or data center holding aging H100, H200, or B200 capacity between workloads, listing that idle inventory to demand is a different exit than a forced resale discount: Spheron's partner program lets a data center or neocloud list H100 through B300 capacity, with onboarding run in about one to two weeks and monthly settlement, and the supplier setting the price floor and approving each deal rather than taking whatever a broker liquidation offers. That's a lever for recovering value on hardware that's already aging on the books, not a substitute for the resale market when you actually need to sell the physical card. For infrastructure operators weighing that math directly, the partner page covers what onboarding capacity looks like in practice.

The same logic runs the other direction for a renter. Paying for GPU time only while a workload is actually running, instead of buying hardware and hoping utilization stays high enough to outrun the depreciation curve, sidesteps the idle-capacity problem entirely: there's no card sitting in your own rack losing book value between jobs, because you never owned one.

GPU Refresh Cycle Rental vs Owning: A Decision Framework

The right answer depends on how long your workload actually needs a given generation and how confident you are in your own utilization, not on which option sounds cheaper per hour in isolation.

Rent on-demand or spot if:

  • Your workload is bursty, project-based, or still finding its steady-state size, and an idle GPU between jobs would just be book depreciation with nothing to show for it.
  • You want to capture a generation's useful life (A100 through B200 on Spheron's current catalog) without ever carrying resale risk when NVIDIA's annual cadence ships the next SKU.
  • You're comfortable with spot's tradeoff: instances can be reclaimed without notice, which fits background and batch work better than anything latency-sensitive.

Reserve capacity if:

  • You know your GPU-hours for the next several months and want a locked rate and guaranteed availability without buying hardware that could be worth 25-45% less in three years on the H100 and A100 curves above.
  • You're budgeting multi-month spend and want price certainty without the clearing-account overhead of something like the compute futures contracts CME is bringing to market, which hedge price risk but still don't guarantee you a physical GPU.

Own if:

  • Your utilization is high and steady enough, multi-year, specialized hardware access, or a regulatory requirement that genuinely needs owned infrastructure, that the GPU pays for a meaningful share of its own cost before the resale curve steepens past year two or three.
  • You have a concrete plan to resell or redeploy the hardware before it reaches the back half of its generation's curve, not an assumption that you'll figure that out later.

Two honest limits worth stating plainly. First, renting protects you from depreciation risk going forward; it does nothing for hardware you've already bought and need to recover value on today, beyond the separate option of listing idle capacity to demand rather than taking a resale-broker discount. Second, a rented rate isn't insulated from the same cloud-price compression this post documents either, on Spheron or anywhere else, so "rent instead of buy" solves the depreciation problem, not the "will this rate still look good in a year" problem. And if your refresh plan is actually pointed at next-generation systems like GB200, GB300, or Rubin's R100, note that those ship on Spheron as reservation and pre-order capacity rather than on-demand, so the "switch to the new generation next month" argument that applies cleanly to H100, H200, and B200 doesn't carry over as neatly to the newest systems yet. If you're weighing the generational jump itself rather than just the rent-vs-own question, the RTX 5090 vs H100 vs B200 comparison is the place to work through what you'd actually be refreshing into.

Planning a GPU Refresh Cycle Rental Strategy

How long should a GPU last before you refresh it? The honest answer, using the curves above: plan your primary workload window around the first 24-36 months, where H100 and A100 both still hold 60% or more of their value, and treat anything past that as a card you're running for the marginal cost of keeping it, not for its resale upside.

A practical sequence for timing a refresh:

  1. Map your workload's real horizon. A 6-month fine-tuning project and a 3-year inference service shouldn't use the same generation decision. Short horizons favor renting through the window; long, steady-utilization horizons are where owning starts to pencil out.
  2. Check where your current generation sits on its curve. Use the table above as a baseline: an H100 at 30 months is already in the 60-75% retained-value band and closing on the steeper drop past year three.
  3. Watch NVIDIA's own cadence, not just your budget cycle. With an annual rhythm confirmed since GTC 2025 and Rubin following on schedule at CES 2026, assume a named successor to whatever you're running lands within about 12 months, which is what compresses the resale window on the generation you're holding.
  4. Separate the book question from the market question before you decide. If you're optimizing reported expense, a 5-6 year useful-life schedule is what your finance team will use. If you're deciding whether to sell, rent, or hold a specific card, use the tracked resale curve instead, since that's the number a real buyer will actually pay.
  5. If you own idle capacity, list it before it ages further down the curve. A half-empty rack of H100s six months from its next value step-down is worth more listed to demand today than it is worth waiting for a liquidation sale later.
  6. If you're renting, match the tier to the commitment. On-demand for workloads still finding their shape, spot for background jobs that can tolerate reclamation, reserved for the multi-month window where you want a locked rate without buying hardware at all.

A GPU refresh cycle rental plan that treats depreciation as four separate curves, each with its own book number and market number, gives you a real answer to when to move, not just a sense that GPUs lose value fast in general.

Renting through a generation's useful life means you never have to guess which depreciation curve your own hardware is actually on. Compare current rates on Spheron's H200 GPU rental or B200 GPU pricing before you plan your next refresh.

Get started on Spheron →

FAQ / 04

Frequently Asked Questions

Tracked H100 resale transactions retain 95-100% of value in the first 12 months, slide to 60-75% by the 24-36 month mark, and settle to a 25-35% floor past 60 months, according to Mercatus's H100 depreciation tracker. That is the market curve, not the accounting schedule: a hyperscaler's books depreciate the same card on a straight line over 5-6 years, which runs well ahead of what the card is actually worth by year three.

Book depreciation is an accounting estimate: Microsoft, Alphabet, and Meta all extended server useful life to 5.5-6 years between 2023 and 2025, which lowers the depreciation expense they report each quarter. Market resale value is what a buyer will actually pay for that specific card today, set by auction and broker transactions. The two numbers are built for different purposes and routinely disagree, especially once a GPU generation passes its second birthday.

Rent through on-demand or spot capacity if you want this generation's compute without carrying resale risk when the next SKU ships, since you are never holding depreciating hardware on a balance sheet. Reserve capacity if you want rate certainty over several months without buying. Own if you have a multi-year, steady-utilization workload where the hardware pays for itself well before year three, and you have a plan to sell or redeploy it before the value curve steepens.

Yes. NVIDIA confirmed at GTC 2025 that it has moved its data center GPU lineup to an annual release rhythm starting with the Blackwell generation, replacing the roughly two-year cadence that preceded it. Rubin followed within the same yearly rhythm at CES 2026, which means a GPU generation now has a named successor on the roadmap within about 12 months of its own launch instead of 24.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min