A neocloud GPU provider is built from the ground up to rent out GPUs, not a general-purpose cloud that bolted GPU instances onto an existing fleet of CPU servers. That distinction decides what you're actually buying or selling: bare-metal access to a specific GPU SKU on a dedicated high-speed fabric, versus a slice of a much larger, more abstracted compute layer. This post defines the term precisely, draws the lines against colocation and the hyperscalers using the frameworks analysts actually use to split up the market, and ends on the practical question a GPU owner has to answer: build a direct sales motion, lease into colocation, or list capacity through a marketplace.
TL;DR: What Is a Neocloud GPU and How Does GPU-as-a-Service Work?
- Definition: a neocloud is a cloud provider built to rent GPU compute; SemiAnalysis says it coined the term and it now names an entire industry segment.
- Four-tier market: SemiAnalysis splits GPU clouds into Traditional Hyperscalers, Neocloud Giants (CoreWeave, Crusoe, Nebius, Lambda), Emerging Neoclouds, and Brokers/Platforms/Aggregators.
- Scale: Synergy Research Group projects neocloud revenue near $180 billion by 2030, a 69% compound annual growth rate.
- Price gap: Hashrate Index reports neoclouds pricing up to 85% below comparable hyperscaler GPU rates, while hyperscale capacity has reportedly sold out into 2028-2029.
- Live example: Spheron's H100 on-demand rate is $2.64/hr as of 09 Oct 2026; see the full cloud GPU provider lineup.
What Is a Neocloud GPU Provider? GPU-as-a-Service vs Traditional Cloud
A neocloud is a cloud provider built specifically to rent GPU compute, rather than a general-purpose cloud that added GPUs later as one product line among many. Runpod's own explainer on the category puts it plainly: the term is a contraction of "new cloud," and it picked up widespread adoption through 2025 as the number of GPU-first providers multiplied.
What makes GPU-as-a-Service a different product from a hyperscaler's GPU instance, not just a cheaper version of it, comes down to three architectural choices neoclouds tend to make consistently. First, minimal virtualization: renters get bare-metal or lightly virtualized access to the GPU rather than a hypervisor-managed slice of it. Second, a dedicated high-performance backend fabric, usually InfiniBand or RDMA over Converged Ethernet (RoCE), kept physically separate from the front-end network that handles ordinary traffic. Third, exposed hardware topology: the renter can see and reason about which GPUs sit on which NVLink domain or InfiniBand switch, instead of that detail being abstracted away behind an instance type name. These three choices are what make multi-node training and tightly coupled inference workloads behave predictably on a neocloud in a way they often don't on a shared, general-purpose cloud fabric.
Where the Term Came From
SemiAnalysis, the research firm best known for its GPU supply-chain analysis, says it coined "neocloud." The firm's own framing is that the word has outgrown its origin:
"Neocloud is now an entire industry. The word was coined by a SemiAnalysis researcher. Some of the companies using it still have no idea."
That's worth sitting with for a second, because it explains why the term gets used loosely. A word coined to describe one specific kind of provider has since been stretched to cover everything from billion-dollar, multi-gigawatt operators to a five-person team reselling a handful of racked GPUs through a web dashboard. Precision matters here, which is why the next section draws the boundaries analysts actually use.
How Neoclouds Differ from Colocation and from Hyperscalers
"Neocloud," "colocation," and "hyperscaler" get used interchangeably in casual conversation about GPU supply, but they sit at different layers of the same stack. Knowing which layer you're buying from, or selling into, changes what you're actually negotiating.
Neocloud vs Colocation: Who Owns the GPUs vs Who Owns the Floor Space
Colocation providers, Equinix, CoreSite, Digital Realty among them, sell physical space, power, and cooling. They do not run the GPU service layer. CoreSite's own description of the category puts the GPU side of it this way: "neoclouds focus on GPUaaS optimized for AI workloads." A colocation provider's own contribution stops at the physical infrastructure, space, power, and cooling, that hosts those operations; it isn't the GPU service layer running on top of it.
The Uptime Institute, cited on that same CoreSite post, frames the category from the other direction, describing neoclouds as "a new category of cloud providers...specializing in AI infrastructure as a service." A neocloud can own its own data centers outright, or it can lease colocation space and run its GPU fleet, networking, and service layer on top of someone else's floor, power, and cooling contract. Either way, the colocation provider's product ends at the rack; the neocloud's product is the GPU hour the renter actually books.
Neocloud vs Hyperscaler: GPU-First vs GPU-as-One-of-Hundreds
A hyperscaler is a diversified cloud business where GPU instances are one line item among hundreds of services, storage, databases, serverless functions, managed Kubernetes, and so on, typically running on a heavily virtualized, shared network fabric and priced at a premium that reflects that breadth and the enterprise support wrapped around it.
A neocloud is built around GPUs first. Gartner analyst Enrique Castera has pointed out that neocloud providers are emerging partly because the US hyperscalers are themselves launching sovereign cloud services, and that neoclouds differentiate through AI-optimized infrastructure and high-performance workloads rather than breadth of service. In other words: hyperscalers compete on everything; neoclouds compete on one thing done deliberately, which is why the architectural choices in the section above (bare-metal access, dedicated fabric, exposed topology) show up so consistently across the category.
The Neocloud GPU Business Model: Four Ways to Sell GPU Capacity
SemiAnalysis's own framework for the GPU cloud market splits providers into four categories, and it's a useful map for seeing where any given name on a pricing comparison actually sits:
| Category | What it means | Examples named by SemiAnalysis |
|---|---|---|
| Traditional Hyperscalers | Diversified cloud businesses, GPUs are one product among many, premium pricing | AWS, Azure, GCP, Oracle, and others |
| Neocloud Giants | GPU-first operators with 100k+ H100-equivalent planned capacity | CoreWeave, Crusoe, Nebius, Lambda |
| Emerging Neoclouds | Smaller operators, often under 10k GPUs, including regional and sovereign-AI providers | Not individually named |
| Brokers/Platforms/Aggregators | Capital-light intermediaries, including marketplace models | Not individually named |
That fourth category, capital-light aggregators, is where a lot of the newer activity sits, and it only exists because of a real gap between what GPU owners need and what GPU renters want. GPU shortage conditions through 2026 have kept hyperscaler reserved capacity booked out, which pushes renters toward whichever neocloud or aggregator actually has the SKU available this week, not whichever one they'd prefer to use on principle.
Own Infrastructure, Colocation-Leased, and Asset-Light Marketplace Models
Underneath that four-category market map, Vast.ai's framework describes three distinct ways a neocloud actually builds and operates, and it maps cleanly onto the "Neocloud Giants," "Emerging Neoclouds," and "Brokers/Platforms/Aggregators" rows above:
- Own infrastructure. Build dedicated data centers with full-stack control over power, cooling, networking, and the GPUs themselves. This is the path the Neocloud Giants have taken, and it requires billions of dollars in capital expenditure before the first GPU-hour is ever sold. NVIDIA's own backstop financing arrangements with neoclouds exist largely to make this capital-intensive path possible for operators who don't have the balance sheet to self-fund it, and the utilization math behind why that financing structure gets built the way it does is worth reading in full there.
- Colocation-leased. Lease space in existing data centers and scale the GPU fleet without taking on construction risk or multi-year build timelines. This is a common path for Emerging Neoclouds that want to compete on specific regions or GPU SKUs without the capital outlay of the first model.
- Asset-light marketplace. Aggregate distributed GPU capacity without owning the hardware at all, acting as a coordination layer between providers that have GPUs and customers that want them. This is the Brokers/Platforms/Aggregators row, and it's structurally a different business: revenue comes from matching and transaction flow, not from GPU depreciation schedules. Fluidstack, Runpod, CoreWeave, and Lambda are among the named providers worth comparing on pricing and billing model if you're evaluating where a given GPU-as-a-Service provider actually sits in this landscape.
The Economics at a Glance: Utilization, Contracts, and Why Marketplaces Exist
The category is growing fast by every analyst's count, though the exact numbers differ depending on what's being measured. Synergy Research Group forecasts neocloud revenues reaching approximately $180 billion by 2030, a 69% compound annual growth rate from current levels. Jeremy Duke, founder and Chief Analyst at Synergy Research Group, frames the pace of the underlying market this way: "GPUaaS and GenAI platform services are currently growing at around 165% per year and neoclouds are gaining share" of that spend.
Gartner's own forecast puts the broader AI cloud market at $267 billion by 2030, with specialized neocloud providers capturing roughly 20% of it. Vast.ai cites a similar trajectory from a different starting point: the neocloud market expanding from roughly $42 billion to over $250 billion by 2030. The spread between these numbers is a reminder that "neocloud market size" isn't a single agreed-upon figure yet, because different firms draw the category boundary slightly differently, which is itself evidence of how new the term still is.
Hashrate Index puts today's split at roughly one-third of AI workloads running on neoclouds versus two-thirds on hyperscalers, with the overall GPUaaS market currently estimated at $4 to $6 billion and growing at double-digit rates. There are more than 100 neoclouds globally by its count, but only 10 to 15 operating at meaningful scale in the US, which tells you the "Emerging Neoclouds" row in the table above is both crowded and thin: lots of entrants, few with real scale.
On price, the gap is the whole reason renters go looking for a neocloud in the first place. Hashrate Index reports neoclouds pricing as much as 85% below comparable hyperscaler GPU rates, and part of why that gap persists is that hyperscale capacity has been reported sold out into 2028-2029. The Uptime Institute independently documented 66% cost savings for certain GPU instances bought from neoclouds versus hyperscalers. On Spheron, for example, H100 SXM5 on-demand runs $2.64/hr as of 09 Oct 2026, a live illustration of how far GPU-as-a-Service pricing can sit below the rate card most teams still default to budgeting against.
Pricing fluctuates based on GPU availability. Spheron rates above are live as of 09 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.
None of this explains utilization math or break-even thresholds on its own, those depend on capex, power costs, financing terms, and contract mix, which is a deeper rabbit hole than a definitional post can responsibly cover. The full utilization-to-profit model for a representative H100 cluster, including the point at which a financed cluster flips from loss to profit, is in the backstop-financing post linked above, and it's the right place to go next if the business model section raised the obvious follow-up question of "at what utilization does any of this actually pay for itself."
Build-Your-Own Go-to-Market vs Listing Through a Marketplace
For a GPU owner, data center, or neocloud with capacity to sell, the three build models from the section above collapse into a simpler practical choice once the infrastructure already exists: go direct, lease into colocation and sell through your own channel, or list through an asset-light marketplace and let someone else bring the demand.
What Each Path Demands of an Operator
Going direct means building the thing a hyperscaler or Neocloud Giant already has: a sales team, a self-serve signup flow, billing and metering infrastructure, and enough brand recognition that a renter searching for GPU capacity finds you instead of the three other names on the same search results page. Spheron's own comparison against SF Compute is a useful concrete example of two different shapes this can take even within the marketplace-adjacent world: one platform built around instant, per-GPU self-serve access, the other built around cluster-scale CLI purchasing for research-grade workloads. Neither approach is free to stand up; both required years of engineering and go-to-market investment before either had enough renter traffic to fill a fleet reliably.
Listing through a marketplace instead trades that build cost for a cut of the margin and a loss of full control over the buyer relationship, in exchange for demand that already exists. The honest tradeoff an operator has to weigh is capital versus control: building a direct channel keeps 100% of the revenue and the full customer relationship but costs years and real money before it generates reliable utilization; listing into an existing marketplace gets capacity in front of qualified demand faster, at the cost of sharing economics with the platform that aggregated that demand.
The constraint isn't only commercial, either. An operator deciding how to sell capacity is also managing physical limits that don't move with a pricing decision. Power, not GPU count, is the binding bottleneck for a growing share of AI data center capacity in 2026, and that constraint shapes how much capacity an operator even has to sell in the first place, regardless of which go-to-market path it chooses for selling what it's got.
Getting Your Capacity in Front of Renters via Spheron
If the direct-build path is the one without a sales team or existing brand recognition to lean on, the asset-light marketplace model from earlier in this post is the one designed to solve exactly that gap. Spheron's GPU Supplier Program is built for data centers and neoclouds with idle H100, H200, B200, or B300 capacity: list your inventory, set your own on-demand and reserved pricing, and sign your own contracts directly with the buyer. Spheron aggregates demand from AI teams looking for that specific hardware and handles billing, metering, and monthly settlement on top, without setting your prices, inserting itself into the signed contract, or reselling your capacity under another brand.
The practical case for listing, in Spheron's own framing, comes down to a few things a direct build has to earn the hard way: buyers come looking for specific GPUs, so qualified requests land without cold outreach, paid ads, or a dedicated sales hire; the operator sets the floor and approves every deal rather than competing in a reverse auction; and training clusters, single-node fine-tuning jobs, and per-minute inference demand all route through the same intake pipeline instead of requiring separate sales motions for each.
That said, this is a demand-routing and billing layer, not a substitute for a managed enterprise sales relationship. An operator with near-full utilization already locked in through its own direct contracts has less obvious upside from listing. The program's own intake form is structured in GPU-count brackets starting at 1-64 GPUs and running up to 1024+, which signals it's built around operators with at least a small fleet to place, not a single idle card. And because the operator still sets pricing and approves every deal, listing doesn't remove the work of actually defining a sensible price floor and SLA terms; it only removes the work of finding the buyer to quote them to.
If you're a data center or neocloud operator with idle H100, H200, B200, or B300 capacity, listing it where qualified demand already exists is usually faster than building that demand channel yourself.
Frequently Asked Questions
A neocloud is a cloud provider built specifically to rent out GPU compute, rather than a general-purpose cloud that added GPU instances to an existing CPU fleet. [SemiAnalysis says it coined the term](https://x.com/SemiAnalysis_/status/2101830640773525612), a contraction of 'new cloud,' and that it now names an entire industry segment, from giants like CoreWeave and Nebius down to small regional operators and asset-light marketplaces.
A hyperscaler (AWS, Azure, GCP) is a diversified cloud business where GPU instances are one product line among hundreds, typically running on a heavily virtualized, shared network fabric and priced at a premium that reflects that breadth. A neocloud is built around GPUs first: minimal virtualization, a dedicated high-performance backend fabric (InfiniBand or RDMA-based Ethernet), and exposed hardware topology instead of an abstracted instance type. [Gartner analyst Enrique Castera notes](https://www.e4ds.com/sub_view.asp?idx=22913) neoclouds are emerging partly because hyperscalers are launching their own sovereign cloud offerings, and neoclouds differentiate through AI-optimized infrastructure rather than service breadth.
Colocation providers (Equinix, CoreSite, Digital Realty) sell physical space, power, and cooling. They do not manage the GPUs or the service layer on top of them. A neocloud is the managed GPU-compute layer, which may run on its own data centers or lease colocation space to scale faster than it could build. [CoreSite itself describes](https://www.coresite.com/blog/neoclouds-are-gathering-infrastructure-for-the-ai-age) neoclouds as focused on GPUaaS optimized for AI workloads, while its own contribution as a colocation provider is the physical infrastructure that hosts those operations.
[Hashrate Index reports](https://hashrateindex.com/blog/what-is-a-neocloud-gpu-cloud-providers-ai) neoclouds pricing as much as 85% below comparable hyperscaler GPU rates, and the [Uptime Institute has documented](https://www.coresite.com/blog/neoclouds-are-gathering-infrastructure-for-the-ai-age) 66% cost savings for certain GPU instances bought from neoclouds instead of hyperscalers. As one live data point, Spheron's H100 SXM5 on-demand rate is $2.64/hr as of 09 Oct 2026; check current rates before budgeting, since marketplace pricing moves with GPU availability.
The asset-light marketplace model is built for exactly this: list idle capacity, set your own on-demand and reserved pricing, and let the platform aggregate demand instead of running outbound sales. Spheron's GPU Supplier Program works this way for data centers and neoclouds with H100, H200, B200, or B300 capacity to list, handling demand aggregation, billing, metering, and monthly settlement while the operator keeps pricing and contract control.






