Engineering

GPU Colocation Explained: How Idle Racks Earn Revenue (2026)

Back to BlogWritten by Published Oct 8, 2026
GPU ColocationGPU Colocation PricingGPU Colocation ServicesGPU Rack Power DensityNeocloud EconomicsInfiniBandGPU CloudAI Infrastructure
GPU Colocation Explained: How Idle Racks Earn Revenue (2026)

A colocation deal built for standard enterprise servers tops out around 3-5 kW per rack. A single GB200 NVL72 rack draws roughly 125-130 kW, and NVIDIA's next generation after that is specified even higher. That gap, not the hardware cost, is why GPU colocation has become its own category with its own pricing, its own facility requirements, and its own failure modes, and why a GPU owner who gets it checked off still has an idle rack if nobody finds out it exists.

This post covers what makes a rack GPU-ready at 2026 NVIDIA densities, what that colocation deal actually costs and why the number swings so widely between owners, and where a demand-aggregating marketplace fits once the facility side of the decision is already made.

TL;DR: What GPU Colocation Actually Means

GPU Colocation vs Traditional Server Colocation

The two look like the same product from a sales sheet: rack space, power, cooling, a network handoff. At the density NVIDIA's current GPUs run at, they stop being comparable on almost every spec that determines whether a facility can actually host the rack.

DimensionStandard server colocationGPU colocation (2026 NVIDIA racks)
Power per rack3-5 kW baseline10-15 kW baseline, scaling past 100 kW for NVL72-class racks
CoolingAir-cooled, standard CRAC/CRAHLiquid-cooled above ~50 kW/rack, CDU-fed coolant loops
Floor loadingStandard raised floor ratingReinforced flooring; standard raised-floor ratings are insufficient at GB200-class density
Power deliverySingle 20 kW rack PDU, often dual-redundantMultiple PDUs or a dedicated busway retrofit per rack
Networking1-25 Gbps cross-connectsInfiniBand fabric sized for multi-Tbps east-west GPU-to-GPU traffic
Typical tenantGeneral enterprise compute, storageTraining clusters, inference fleets, neoclouds

Power Density, Rack by Rack: Standard Colo vs. GPU Colo

The 3-5 kW figure for standard colocation is not a marketing floor, it is what most existing data center electrical and cooling plant was engineered around for the last two decades of enterprise IT. A rack running general-purpose servers, storage arrays, or network gear rarely needs more than that, and facilities built to that spec make up the bulk of the colocation market a GPU owner might otherwise assume is available to them.

GPU infrastructure resets that baseline. Netrality's analysis puts GPU colocation at 10-15 kW per rack as the floor, scaling to 25-30 kW or more for dense AI training racks, before a single NVL72-class rack enters the picture. NVIDIA's GB200 NVL72 rack carries a total TDP of approximately 125-130 kW. That is not a difference of degree from a standard colo rack, it is a different category of electrical and mechanical plant, and a facility that has never retrofitted for it cannot simply sell the rack space and hope the power follows.

Cooling, Floor Loading, and Networking Are Different Problems at This Density

Three problems show up together once a rack crosses into GPU-colocation density, and all three have to be solved before power is even turned on.

Cooling stops being passive. At a rack running near 130 kW, the facility needs a working coolant distribution loop, not just more CRAC units, before the rack can run at spec. Air cooling becomes impractical above roughly 50 kW per rack, which a GB200-class rack clears by a factor of two or more.

Floor loading becomes a structural question, not an electrical one. Standard raised-floor ratings in most colocation builds were never engineered for GB200-class density, and reinforcing the floor is a line item retrofit projects frequently miss until the rack is already on-site.

Networking has to move multi-terabit east-west traffic, not just north-south traffic to the internet. A standard colocation cross-connect handles the second kind of traffic fine; it was never sized for GPUs talking to each other.

Why GPU Owners Colocate Instead of Running Their Own Facility End-to-End

Building a purpose-built facility from the ground up means owning the utility interconnect negotiation, the mechanical and electrical engineering, the permitting timeline, and the standing cost of a facility that depreciates whether or not it's full. Colocation lets a GPU owner buy into shared power infrastructure, an existing utility interconnect, and a facility that's already past permitting, and put capital into GPUs instead of concrete.

It also shortens time to revenue. A retrofit or a purpose-built GPU suite inside an existing colocation facility can go from signed deal to powered rack in months; a ground-up build is a multi-year undertaking even before the first server ships. For an owner trying to get GPUs earning money against a depreciation clock, that difference is the whole argument for colocating rather than building.

This is also where NVIDIA's own involvement in facility design shows up, and where it's worth separating two things that look similar but aren't. NVIDIA is not paying colocation operators to host racks. What it has done is co-design reference architectures with cooling vendors and certify facility designs through its DGX-Ready Data Center program, so an owner can point to a facility that's already validated against Blackwell and Vera Rubin specs rather than negotiating thermal requirements from scratch. Vertiv and NVIDIA jointly engineered a liquid cooling reference architecture that supports up to 132 kW per rack, and Vertiv has since announced gigawatt-scale reference architectures for NVIDIA's Omniverse DSX Blueprint aimed at AI factory deployments across platforms including Vera Rubin. That's a design and certification relationship, separate from the equity and financing arrangements NVIDIA has entered into with some neoclouds, which is a different story covered further down in this post.

What a GPU Colocation Deal Actually Requires

Three things have to be true of a facility before a GPU rack can go live in it at 2026 densities: enough power delivered to the rack, a cooling loop that can remove the heat that power turns into, and a network fabric fast enough that the GPUs can actually talk to each other. Get any one wrong and the other two don't matter.

Power Density by NVIDIA Rack Generation (H100 Through Rubin Ultra)

The jump in per-rack power draw across NVIDIA's last few generations is the single biggest reason colocation retrofits keep happening instead of being a one-time project.

Rack generationApproximate power per rackCoolingTarget timing
H100-era HGX (air-cooled, general GPU colo)10-15 kW baseline, up to 25-30 kW+ denseAirShipping since 2023
GB200 NVL72~125-130 kWLiquidShipping
GB300 NVL72up to 142 kWLiquidShipping
Vera Rubin NVL72190-230 kW (trade-press estimate; NVIDIA hasn't published an official figure)LiquidFull production targeted June 2026
Rubin Ultra NVL576 ("Kyber")~600 kWLiquid, beyond what current CDU loops are built forTargeted H2 2027

A rack-mount power distribution unit is typically rated around 20 kW with double redundancy. A single GB200-class rack at over 125 kW needs several of those PDUs, or more commonly a dedicated busway retrofit, just to deliver power to one rack. That's before the facility has solved cooling or networking for it. As Omkar Nimbalkar, IBM's vice president of multi-vendor support services, told Data Center Knowledge: "The design question is never how many GPUs you can buy, but how many you can safely run if a power supply fails."

Cooling: Where Air Stops Working and Liquid Cooling Takes Over

Air cooling becomes impractical above roughly 50 kW per rack, which is the ceiling most existing colocation facilities were designed below. Every rack generation past H100-era HGX systems in the table above sits well over that line, which is why liquid cooling isn't an upgrade option for a GPU colocation facility at current NVIDIA densities, it's a prerequisite.

The cooling vendors building for this are explicit about the ceiling they're targeting today. Vertiv co-designed its liquid cooling reference architecture directly with NVIDIA to support up to 132 kW per rack, and the company's newer gigawatt-scale Omniverse DSX Blueprint work extends that partnership toward the Vera Rubin generation. Rubin Ultra's roughly 600 kW spec sits well beyond what a 132 kW-class coolant loop services today, which is the honest version of where the ceiling is heading next, and why facilities are already planning a second retrofit rather than treating the current generation as the finish line.

A GPU cluster's InfiniBand network exists because training a model is a synchronized operation: AI training pipelines spanning multiple nodes need synchronized, collective GPU communication and GPUDirect RDMA with low jitter. NVIDIA's RDMA InfiniBand support lets one machine access another's memory directly, without involving the operating system, which cuts CPU overhead and latency, and that's what keeps the synchronization fast enough to be worth doing at all.

The bandwidth that requires is not close to what a standard colocation cross-connect is provisioned for. Each GPU on an 8-GPU node connects through a 400 Gbps network interface, putting total node bandwidth at 3.2 Tbps, a figure far beyond what standard colocation cross-connects are sized for. A facility selling GPU colocation without rethinking its cross-connect capacity from the ground up is selling a product it cannot actually deliver once the rack is under real training load.

GPU Colocation Pricing: What It Actually Costs

There's no single number here, and anyone quoting one flat rate per rack is skipping the variables that actually set the price. What's consistent is the shape of the bill.

The Cost Components That Make Up a Colocation Bill

A GPU colocation invoice breaks down into a handful of recurring components, and power is usually the largest by a wide margin once a rack crosses into liquid-cooled territory:

  • Power: metered or committed, usually the single biggest line item at GB200-class density and above
  • Cooling infrastructure: CDU capacity, coolant loop maintenance, and the facility-level chiller plant backing it
  • Floor space and structural reinforcement: denser racks need reinforced flooring and often a smaller physical footprint per kW than legacy colocation pricing assumes
  • Networking and cross-connects: InfiniBand-capable fabric and the cross-connect fees that come with it
  • Redundancy and SLAs: N+1 or 2N power and cooling redundancy, uptime guarantees, remote-hands support

For a sense of what that adds up to at scale: a cost model for a 1,024-GPU H100 cluster in our piece on NVIDIA's neocloud backstop financing itemizes colocation and networking at roughly $75,000/month, separate from a $38,000/month power line calculated at $0.08/kWh, alongside GPU capex amortization and financing overhead as the other major cost lines. Divide that $75,000 across the cluster and colocation plus networking works out to roughly $73 per GPU per month, before power, capex, or financing enter the picture. That per-GPU figure moves with region, contract length, and power rates, but it's a real anchor for what the facility side of a large deployment costs next to the hardware itself, and it's the kind of number worth re-running against your own quote before signing a multi-year colocation contract.

Why GPU Colocation Pricing Varies So Much Between Owners

The same rack generation can price out very differently between two facilities, and the gap usually traces back to one of these:

  • Utility power rates, which vary several-fold by region and directly scale the largest line item on the bill
  • Retrofit vs. purpose-built, since a facility retrofitting an existing floor for liquid cooling amortizes that capex differently than one built GPU-ready from day one
  • Redundancy tier, where Tier III and Tier IV power/cooling redundancy carry materially different standing costs
  • Contract structure, since multi-year committed capacity prices differently than month-to-month colocation
  • Whether cooling is metered or bundled, which shifts risk between the facility and the tenant when coolant loop maintenance costs spike

None of this is published pricing in the way GPU rental rates are. A colocation quote is closer to a real estate lease than a cloud price list, negotiated per deal, per region, per power contract. That's also why a figure like "$X per kW per month" from one facility tells you little about what another facility will quote for the same rack.

Colocation figures above reflect estimates and industry cost models current as of early October 2026, and vary materially by region, power contract, and facility. Check with the facility directly for a current quote, and see current GPU rental pricing → for how colocation costs compare to renting GPU time outright.

Colocation vs Listing on a GPU Marketplace

This is the part that's easy to conflate and genuinely isn't the same decision. Colocation is a facility contract: it gets a GPU owner the power, cooling, and connectivity to run the hardware. Listing on a marketplace is a demand problem: it gets renters to a rack that already has all of that solved. An owner can have a flawless colocation deal, every PDU, CDU, and InfiniBand link provisioned correctly, and still have an idle rack if nobody outside the owner's own sales team knows the capacity exists.

The market backdrop makes that gap more expensive to ignore, not less. The global data center colocation market was valued at $88.91B in 2025 and is projected to reach $216.37B by 2031, a 15.98% CAGR, with AI workload demand cited as the primary driver. AI represented about a quarter of all data center workloads in 2025, and JLL projects that share could reach half of all workloads by 2030, with inference projected to overtake training as the dominant AI workload type in 2027. That's demand growing faster than most individual owners' ability to go find it themselves.

Two genuinely different paths handle that demand problem:

Direct colocation only, sourcing renters yourself. The owner keeps full control over pricing and every contract, but also owns outbound sales, buyer vetting, billing and collections, and the risk that a half-empty rack sits half-empty because the right buyer never found it. This fits an owner who already has an established book of customers and doesn't need a new demand channel.

Colocation plus a demand-aggregating marketplace. The facility relationship stays exactly the same, the marketplace layer sits on top of it and routes buyers who are already looking for a specific GPU SKU. This is where Spheron's partner program fits: it brings vetted buyers with verified budgets and use cases, handles billing, metering, and collection, and aggregates demand across SKUs so a half-empty rack still finds a buyer. It explicitly does not set the owner's prices, does not sit between the owner and the signed contract once a deal closes, and does not resell the owner's capacity under a different brand, because Spheron doesn't own GPUs and isn't competing with the owners listing on it.

The honest limit on the second path: a marketplace is a demand channel, not a colocation broker. It doesn't negotiate the power, cooling, or facility contract between a GPU owner and the data center, so that relationship, whether an existing deal or a new one, has to already be in place before a listing makes sense. And listing isn't instant or automatic: it goes through a vetting review on uptime history, networking, and reference deployments before capacity goes live.

The owners best served by adding a marketplace layer are the ones who've already cleared the facility side of this post, power, cooling, and InfiniBand-capable networking sorted, and are specifically short on demand, not short on infrastructure. An owner who hasn't secured colocation yet needs to work through the facility requirements covered above first, or look at a DGX-Ready certified facility, before a demand channel is the relevant question at all.

For context on what buyers actually screen for before renting from a given owner's fleet, our guide to 2026 GPU export controls covers the KYC and compliance checks that increasingly sit alongside the usual InfiniBand and uptime questions, and our breakdown of multi-GPU cluster networking covers the NVLink and InfiniBand tradeoffs that determine what a renter actually needs from your fabric.

How Spheron's Partner Program Fits Into a Colocation Strategy

Once the colocation deal is signed and the rack is powered, cooled, and networked, the partner program is the layer that gets it rented. The published process runs: apply with fleet details, SKUs, regions, and data center tier; get vetted on uptime history, networking, and reference deployments, typically under a week; list inventory and set your own on-demand and reserved pricing; receive demand routed by specification; approve and provision the requests you choose to accept; get paid on a monthly settlement cycle. Average onboarding runs 1-2 weeks end to end.

The program currently supports H100, H200, B200, and B300 capacity. It's built around a specific division of labor: the owner keeps pricing control and keeps the contract once a deal signs, and Spheron's side is demand aggregation, buyer vetting, and billing, not a layer that inserts itself into the relationship between the owner and the renter. For a GPU owner who's already solved power, cooling, and networking the way this post describes and is looking for the demand side of the equation, that's the gap the program is built to close.

A colocated rack that clears every power, cooling, and networking spec in this post is still just a power bill without renters finding it. Spheron's partner program routes vetted demand to capacity GPU owners have already colocated.

See the Spheron partner program →

FAQ / 05

Frequently Asked Questions

GPU colocation is renting rack space, power, and cooling in a third-party data center for GPU servers you own, rather than building your own facility or renting GPU time from someone else's. The owner keeps the hardware and the customer relationships; the facility provides power, cooling, floor space, and network cross-connects. It differs from standard server colocation mainly in density: a single GB200 NVL72 rack draws approximately 125-130 kW of total TDP, far above what a standard enterprise rack is provisioned for.

No, not directly. NVIDIA does not pay colocation facilities to host racks. What it does is certify facility designs through its DGX-Ready Data Center program and work with cooling vendors like Vertiv on reference architectures, so a facility can prove it meets Blackwell and Vera Rubin power and thermal specs before a GPU owner signs a lease. Separately, NVIDIA has taken equity and provided financing to some neoclouds that then lease colocation space; that financing relationship is distinct from the colocation deal itself.

Vertiv is the cooling vendor that has publicly co-designed a liquid cooling reference architecture with NVIDIA, built to support up to 132 kW per rack. Vertiv has also announced gigawatt-scale reference architectures for NVIDIA's Omniverse DSX Blueprint, aimed at AI factory deployments across platforms including NVIDIA's upcoming Vera Rubin generation.

Because the facility is selling a different product. Standard colocation bills mostly for floor space and a modest power allotment, typically 3-5 kW per rack. A GPU colocation deal bills for 10 to over 100 kW of continuous power per rack depending on the NVIDIA generation, plus the coolant distribution units, reinforced flooring, and high-bandwidth cross-connects that density requires. Power alone usually becomes the largest line item once a rack crosses the roughly 50 kW threshold where air cooling stops working and liquid cooling becomes mandatory.

They solve different problems and most owners need both. Colocation is the facility contract: power, cooling, floor space, and network connectivity for hardware you own. A GPU marketplace is the demand channel: it routes renters to your already-colocated capacity, handles billing and metering, and lets a half-empty rack find a buyer without you running outbound sales. Neither replaces the other; a colocated rack with no renters is just a power bill, and a marketplace listing with no facility behind it has nothing to sell.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min