How Wafer runs Wafer Pass and Wafer Serverless on B300 On-demand capacity from Spheron
On-demand access to NVIDIA B300 capacity through the Spheron marketplace on both Spot and dedicated tiers, powering production inference for Wafer Pass and Wafer Serverless on the newest Blackwell Ultra hardware.
About Wafer
Wafer is a Y Combinator-backed startup with seed backing from Fifty Years and Liquid2. The pitch is direct: ship the fastest inference in the world. Autonomous agents profile, diagnose, and optimize inference across the full stack, then deliver the result through two products: Wafer Pass, a flat-rate subscription, and Wafer Serverless, pay-per-token.
Wafer Pass is a flat-rate subscription with three tiers: Lite, Starter, and Privacy. Each tier sets a request cap inside a rolling five-hour window and opens access to every model Wafer hosts. Billing is weekly, monthly, or yearly with 20 percent off the yearly plan. Cancel anytime. The Privacy tier adds Zero Data Retention, the configuration production agents and teams handling sensitive data pick by default.
Wafer Serverless prices the same lineup per million tokens, with cache-read tokens billed at 10 percent of input and no minimums or commitments. Qwen3.5-397B-Turbo and GLM 5.1-Turbo are the open-source models live today. More are on the way.
Wafer publishes 2.8x faster than base SGLang on Qwen3.5-397B, with throughput hitting 408 tokens per second against 145 on the SGLang baseline. Numbers like that need the newest NVIDIA silicon, all the time. B300 sits at the top of the Blackwell Ultra line and is not on every cloud's catalog yet. Wafer cannot build a serving stack on hardware they cannot get, and they cannot scale it on hardware they cannot afford to keep running.
The challenge
Wafer Pass and Wafer Serverless need to run on the fastest available NVIDIA silicon, and right now that is B300. The promise to customers is the 2.8x speedup on Qwen3.5-397B and the rest of the open-source lineup. The only way to keep that promise is to serve traffic on Blackwell Ultra hardware, all the time. The problem is sourcing it.
Most GPU clouds either do not list B300 yet or gate access behind reservation forms. Reservations come with minimums, lead times, and capacity meetings. For a flat-rate subscription product and a per-token API, reservation-only B300 is the wrong shape. Inference traffic moves with users. Reserved blocks bill the same whether requests are flowing or not, and the moment traffic spikes past the reserved window there is no extra B300 to fall back on.
Cost was the other half of the problem. Wafer Pass starts at $3 a week. Wafer Serverless prices Qwen3.5-397B-Turbo at $0.60 per million input tokens and $3.60 per million output tokens. Those numbers only work if the underlying GPU cost lines up. The few providers offering B300 priced it as a long-tail reservation: annual commits and on-demand sticker rates that left no room in the unit economics for a flat-rate subscription.
A Spot tier on B300, the kind of preemptible capacity that fits stateless inference cleanly, was not on the menu anywhere. Spot only exists when a provider has excess B300 sitting idle. On the newest Blackwell Ultra hardware, that excess does not exist on hyperscaler fleets yet.
Spot on B300 is the price tier nobody else was offering. Every other cloud quoted us reserved blocks at on-demand sticker rates, on hardware we needed to scale with user traffic. That math does not work for a flat-rate subscription where the margin lives or dies on per-hour GPU cost.
Why Wafer chose Spheron
Spheron is a marketplace, not a single provider. Behind one API are H100, H200, B200, B300, A100, GH200, L40S, and other NVIDIA capacity from data center partners across multiple regions, on both Spot and dedicated tiers. The marketplace structure is how B300 shows up before it lands in mainstream catalogs: new partners coming online with new hardware plug into one platform instead of building distribution from scratch. It is also how Spot on B300 becomes possible. When a partner has excess Blackwell Ultra capacity at a given hour, the marketplace surfaces it as a preemptible Spot tier instead of letting the node sit idle. Dedicated B300 sits on the same marketplace through the same API, so the tier choice is a workload decision, not a procurement one.
For Wafer Pass and Wafer Serverless, that translates directly into the speed and cost combination the products promise. Inference traffic runs on B300 across a mix of Spot and dedicated nodes, picked per workload. Stateless, elastic load rides Spot for the price advantage. Steadier baselines run on dedicated for predictability. The 2.8x speedup customers buy on Qwen3.5-397B is delivered on the same Blackwell Ultra silicon in either case. When a Spot node gets preempted, requests route to the next available B300 in the marketplace and end users see no change.
Per-minute marketplace pricing means cost ties to actual GPU hours, not to a quarterly reserved block paid up front. No reserved blocks. No capacity meetings. No minimum spend. Spot pulls the per-hour cost down for the parts of the fleet that can take preemptions, dedicated holds the steady baseline, and the combination of the two is what keeps Wafer Pass flat-rate at $3 a week and Wafer Serverless competitive on per-token pricing.
We run B300 from Spheron on a mix of Spot and dedicated tiers, all from the same marketplace, and pick whichever fits each part of the workload. Nowhere else can we do that on the newest Blackwell Ultra hardware. Everywhere else, B300 is a reservation form.
The setup
Wafer's inference fleet runs on B300 from Spheron. Spot and dedicated nodes side by side, picked per workload, all on per-minute billing. The serving stack absorbs Spot preemptions: when a node gets reclaimed, requests route to the next available B300 and user-facing latency stays inside the speed targets the products are sold on. Bare metal where it matters for raw throughput, VMs where it does not. Networking and storage are pre-configured for inference.
Deployment is fully self-serve. When the team needs a B300 node, they bring one up directly from Spheron's infrastructure the moment availability lines up, on whichever tier fits the workload. No procurement, no contracts, no sales call, no waiting on a ticket. The same API covers H100, H200, B200, and the rest of the NVIDIA catalog, so a different GPU is one deploy call away. As Spheron adds new B300 supply across data center partners, that capacity shows up in the same marketplace and Wafer's stack picks it up automatically based on real-time availability and price.
The outcome
Wafer Pass and Wafer Serverless run on the latest Blackwell Ultra hardware before most teams have access to it, at a per-hour GPU cost no other provider offers. Subscribers get the 2.8x speedup on Qwen3.5-397B on the same B300 silicon other teams are still waiting on. Flat-rate Pass tiers and per-token Serverless prices both stay sustainable because the underlying mix is Spot for elastic capacity and dedicated for steady baselines, never long-term reservations at hyperscaler rates.
Treating B300 as a marketplace commodity beats treating it as a reserved asset at hyperscaler rates. That is what keeps Wafer's products fast, cheap, and elastic. As Spheron adds more B300 partners, more capacity flows to Wafer on both Spot and dedicated tiers, no integration changes required.
Wafer Pass and Wafer Serverless depend on getting time on whatever NVIDIA shipped most recently, at a price that lets us run it constantly. Spheron is the only marketplace where we can run B300 on a mix of Spot and dedicated tiers from the same API, and we deploy directly the moment availability lines up. We ship the fastest open-source inference on the market at a cost the rest of the market cannot match.
H200 rental at the best market price, on long-term commitment terms we negotiated, from a vetted Tier 3+ data center partner.
PremAI builds confidential AI for regulated industries. Spheron found their H200 rental at the best price in the market, negotiated long-term commitment terms, validated the data center, and stays on as PremAI's GPU rental sourcing partner.
In-region H200 from a Tier 3+ neo cloud, with bandwidth past the India DC norm.
Compute Desk needed H200 in India for one of its buyers during the GPU shortage. Spheron sourced the rental direct from a Tier 3+ neo cloud, with bandwidth past the 100 Mbps cap most India DCs hand customers and a rate below channel-partner quotes.
Need GPU capacity sourced for you?
Tell us what you're building, the GPU type you need, and the region. We source it from one of our vetted data center partners and get you on-demand access with per-minute billing.