Set temperature to 0 in vLLM, send the same prompt twice, and you can still get two different completions back. Not because of GPU floating-point non-associativity alone, the explanation that's circulated for years, but because the kernels underneath vLLM aren't batch invariant: the same request produces different numeric results depending on what else happens to be sitting in the batch alongside it. Thinking Machines Lab traced this to three specific kernels in September 2025, vLLM shipped a VLLM_BATCH_INVARIANT flag to fix it, and turning that flag on has a throughput cost worth pricing before you flip it on for a production RL loop, an eval suite, or an audit trail.
TL;DR: vLLM Batch Invariance and Deterministic Inference
- Cause: Temperature-0 nondeterminism in vLLM traces to batch-variant kernels (RMSNorm, matmul, attention), not floating-point non-associativity alone, per Thinking Machines Lab's September 2025 analysis.
- Fix: Setting
VLLM_BATCH_INVARIANT=1pins vLLM's kernels so a request's output no longer depends on batch size or batch composition, documented in vLLM's batch invariance docs. - Cost: SGLang's comparable deterministic-inference mode measured an average 34.35% throughput slowdown (24.4% to 46.0% range) against Thinking Machines Lab's original 61.5% overhead, per the SGLang project's own benchmarks.
- Hardware: Requires NVIDIA GPUs with compute capability 8.0+, or Intel XPUs on the Triton attention backend.
- Scheduling gotcha: vLLM's async scheduling is on by default and can still break batch invariance even with the flag set, since batch membership depends on wall-clock timing, not just kernel math.
- Who needs it: RL rollout and training consistency, deterministic evals, and audit trails, not general-purpose serving.
Why Greedy Decoding Still Gives You Different Outputs on Different Days
Greedy decoding, temperature 0, is supposed to be the deterministic case: pick the highest-probability token at every step, no sampling. In practice, teams running eval suites or regression tests against a live vLLM endpoint see completions drift run to run anyway, and the usual explanation reaches for floating-point non-associativity combined with GPU parallelism: the fact that (a + b) + c isn't guaranteed to equal a + (b + c) once you sum across thousands of concurrent threads.
Thinking Machines Lab's analysis of the problem calls that explanation incomplete. Non-associativity is real, but on its own it doesn't explain run-to-run drift, because a fixed sequence of floating-point operations on the same inputs produces the same output every time, deterministically, associative or not. What actually varies is which sequence of operations runs, because vLLM's kernels change their execution strategy depending on the shape of the batch around your request. As the team put it, "the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies."
That load variance isn't a bug, it's the point of continuous batching. vLLM re-forms the running batch at every decode step, evicting finished sequences and admitting queued ones the instant a slot frees up, which is exactly what keeps GPU utilization high. It's also exactly what makes your batch neighbors, and therefore your kernel execution path, a moving target on every retry, every traffic spike, and every server restart.
What "Batch Invariant" Actually Means at the Kernel Level
A kernel is batch invariant when a single row's output depends only on that row's own data, never on which other rows share the batch with it or how many there are. Most of vLLM's kernels aren't: RMSNorm changes its reduction order by batch size, matmul dispatches to different tiling configurations by batch shape, and attention picks a different split-KV strategy by batch and cache length, so the same request can produce different logits purely from what else is queued next to it.
RMSNorm: Reduction Order Changes When Batch Size Changes
RMSNorm normalizes each token's hidden-state vector by the root-mean-square of its own values, a sum-then-normalize reduction repeated at every layer. The GPU kernel computing that sum has to decide how to split the reduction work across threads, and the efficient split depends on how many rows (tokens across the batch) it's reducing at once. A larger batch parallelizes the reduction differently than a smaller one, which changes the order floating-point additions happen in. Floating-point addition isn't associative, so a different order produces a slightly different sum, a tiny numeric difference that then compounds across every layer of the model.
Matrix Multiplication: Shape-Adaptive Kernels vs Fixed Configs
The matmuls in a transformer's forward pass, the projections in every attention and MLP block, are dispatched to GEMM (general matrix multiply) kernels that select a tiling and accumulation strategy based on the shape of the operation. That shape is partly determined by batch size: more rows in flight changes which tile configuration the kernel picks as fastest. Pick a different configuration, and the order partial products get summed in changes too, again a rounding difference, again compounding across layers. Whether a given request's matmul is compute bound or memory bound at all depends on this same batch-shape sensitivity, a relationship covered in more depth in our roofline model breakdown for LLM inference.
Attention: Split-KV Strategy Size Adapts to the Batch, and That's the Bug
FlashAttention-style kernels split the KV cache into chunks and reduce partial attention outputs across those chunks in parallel, a strategy usually called split-K or split-KV. How many chunks the kernel picks, and how large each one is, adapts to batch size and cache length, because that adaptation is what makes the kernel fast across a wide range of workloads. It also means the order partial attention outputs get combined in changes with the batch, which is the third and last of the three kernels Thinking Machines Lab names as the source of the nondeterminism: RMSNorm's reduction order, matmul's shape-adaptive tiling, and attention's adaptive split-KV sizing.
Where This Sits in LLM Inference Optimization
LLM inference optimization decisions get made along one axis almost universally: how much throughput a GPU-hour buys. That axis has three real levers. Batching and scheduling strategy decides how many requests share a forward pass and how quickly a finished slot gets refilled, and it scales further than most teams assume before batch size stops paying off linearly. Kernel choice decides how efficiently the GPU executes each forward pass, whether that's a fused RMSNorm, a tuned GEMM, or a FlashAttention variant. And serving engine choice, vLLM versus TensorRT-LLM versus SGLang, bundles a specific combination of the first two plus a scheduler and a memory manager into one deployable stack. Our decision framework for choosing between engines breaks down where each one wins on that throughput axis.
Batch invariance sits on a different axis entirely: not tokens per second, but whether the same request produces the same tokens twice. It doesn't compete with the three levers above, it trades against them directly. The same adaptive-kernel-selection logic that batch invariance disables exists specifically because it makes the throughput axis better. Pin it, and the throughput axis gets worse by design, which is the subject of the next section.
Turning It On: VLLM_BATCH_INVARIANT and What It Costs in Tokens/sec
Setting the VLLM_BATCH_INVARIANT environment variable to 1, applied before importing vLLM or passed through vllm serve, pins vLLM's kernels so a given row's result no longer depends on batch size or batch composition. That's the entire mechanism: a fixed execution strategy in place of an adaptive one, per kernel, for RMSNorm, matmul, and attention alike.
vLLM doesn't publish a single head-to-head throughput number for the flag on its own, but two adjacent measurements put a real range on what pinning these kernels costs. SGLang built its own deterministic-inference mode on top of Thinking Machines Lab's batch-invariant operators, adding deterministic attention with fixed split-KV sizes across its FlashInfer, FlashAttention 3, and Triton backends. The SGLang team measured that mode's overhead at an average 34.35% slowdown, ranging from 24.4% to 46.0% depending on workload, against a reported 61.5% overhead for Thinking Machines Lab's original approach; CUDA graphs recovered roughly 2.8x of the throughput that deterministic mode otherwise gave up.
The more direct vLLM number comes from a harder test: bitwise-exact consistency between a training engine and an inference engine, which requires batch invariance on both sides. vLLM and TorchTitan combined batch-invariant kernels to get bitwise-consistent numerics for on-policy reinforcement learning, auditing every kernel invocation across both frameworks' forward passes. In their test case, Qwen3 1.7B on GSM8K, the bitwise-exact training run measured 2.4x slower than the non-bitwise baseline, though it reached a higher total reward in fewer steps, while the non-bitwise run's KL divergence drifted away from 0.0 over the course of training. A 24% to 46% inference-only overhead and a 2.4x training-loop overhead are two different workloads, not the same number twice, but they bracket the range a team should expect: real, and large enough to budget for, not a rounding error on the GPU bill.
Hardware and Backend Requirements
Batch invariance in vLLM requires NVIDIA GPUs with compute capability 8.0 or higher, or Intel XPUs running the Triton attention backend. It's tested against a wide model list: DeepSeek, Qwen3 in its dense, MoE, and vision variants, Qwen2.5, Llama 3, GPT-OSS, Mistral, Phi, Granite 3.1, EXAONE 4.0, and OLMo 2. Enabling it also disables custom all-reduce in tensor-parallel mode, along with async tensor parallelism and sequence parallelism, because their reduce-scatter communication paths aren't themselves batch invariant. That's a real multi-GPU cost on top of the per-kernel one: turning the flag on gives up two of vLLM's own communication optimizations, not just kernel-level speed.
What Async Scheduling Silently Undoes
The flag only pins the kernels. It does nothing about which requests the scheduler puts in the same batch, and vLLM's async scheduling, on by default, composes the next step's batch before the current one retires. That means batch membership still depends on wall-clock timing rather than a fixed order, so async scheduling can undo batch invariance even with VLLM_BATCH_INVARIANT=1 set. Anyone who has tuned continuous batching, PagedAttention, and chunked prefill together on vLLM has already seen how much the scheduler's own timing decisions shape what lands in a given step, and that's precisely the mechanism at issue here: the kernel-level fix and the scheduler-level cause are two separate settings, and only fixing one of them doesn't get you determinism.
vLLM's own guarantee isn't universal yet, either. Open issue tracking shows fused-MoE kernels, and at least one pooling model path, reported as non-invariant even with the flag set, meaning MoE-heavy deployments running VLLM_BATCH_INVARIANT=1 today shouldn't assume full coverage without checking their specific model against the current issue list.
When You Actually Need This: RL Rollouts, Evals, and Audit Trails
The clearest case is on-policy reinforcement learning, where the model that generates rollouts and the model being trained on them need to agree on what happened. As the vLLM project put it, "reinforcement learning has been shown to amplify tiny numerical mismatches between trainer and sampler, leading to non-deterministic and unstable training behavior." The SGLang team frames the same failure mode from the sampling side: "the indeterminism of inference results can implicitly transform on-policy reinforcement learning (RL) into off-policy RL," which is a correctness problem, not just a noisy metric, since the whole point of on-policy RL is that the policy being trained matches the policy that generated the data.
Two other cases share the same logic without the RL machinery. Deterministic evals matter whenever a benchmark score needs to be reproducible run to run, because a shifting completion on an exact-match or code-execution eval looks like model regression when it's actually kernel noise. Audit trails matter in regulated or safety-sensitive settings where a specific output needs to be reproducible after the fact, for review, dispute resolution, or compliance, and "the model probably would have said something like this" isn't good enough.
When You Don't: The Case for Leaving It Off
Most production serving doesn't need bitwise determinism, and paying a 24% to 46% inference-time throughput cost, or something closer to 2.4x in a training loop, for reproducibility nobody downstream is checking is a real cost with no matching benefit. If your traffic is customer-facing chat, retrieval-augmented generation, or anything where two runs producing slightly different phrasing of the same answer is invisible to the end user, the adaptive kernels vLLM ships by default are doing their job, and turning them off just spends GPU-hours on a guarantee nobody consumes.
This is also where renting rather than buying makes the decision cheaper to get right. Spheron bills GPU rentals per minute past a 20-minute minimum, with both on-demand and spot instances and full root access to a bare-metal or dedicated VM, which is what actually lets a team set VLLM_BATCH_INVARIANT=1, swap attention backends, and measure the real throughput hit on their own model and traffic shape, rather than reasoning about it from someone else's benchmark. A team pricing that test against an H100 running SGLang's measured 24% to 46% overhead range can compare $2.65/hr on-demand against $2.20/hr spot, with the caveat that spot capacity can be reclaimed without notice, which is its own kind of reproducibility problem for an RL rollout that needs to run uninterrupted. Spheron aggregates live rates across 5+ providers, so that comparison is visible before committing a training loop's GPU-hour budget to the overhead; Spheron's docs cover provisioning the instance itself.
None of that changes the arithmetic, though. Renting the hardware doesn't reduce what VLLM_BATCH_INVARIANT=1 costs in GPU-hours for the same token volume, and it doesn't close vLLM's own open gaps, like the fused-MoE non-invariance still tracked in the project's issue history. It just makes the decision of whether to pay that cost a short, cheap experiment instead of a guess. For sustained RL training, where the overhead compounds over many GPU-hours, that's worth pricing carefully; for the occasional flaky eval you need to debug once, a single short-lived instance is enough to confirm the mechanism, without changing anything about how the rest of the fleet serves traffic. GPU pricing fluctuates by availability and provider, so check current rates before committing a training loop's budget to either tier.
Pricing fluctuates based on GPU availability. Spheron rates above are live as of 01 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.
Batch invariance is a reproducibility trade you make deliberately, one experiment at a time, not a default worth serving traffic on. Rent the instance, run the comparison against your own model, then decide.
Frequently Asked Questions
Because greedy decoding removes sampling randomness but not kernel nondeterminism. vLLM's kernels for RMSNorm, matrix multiplication, and attention pick different execution strategies depending on the size and composition of the batch a request lands in, and continuous batching means that composition changes on every decode step. The same request can produce different logits, and therefore a different top-1 token, purely from what else is queued alongside it that step.
A kernel is batch invariant when a single request's output depends only on that request's own data, never on which other requests share its batch or how many there are. vLLM ships most kernels as batch-variant by default because adaptive tiling and adaptive split-KV strategies improve throughput. The VLLM_BATCH_INVARIANT environment variable pins them to a fixed strategy instead, trading some of that throughput for reproducibility.
Set VLLM_BATCH_INVARIANT=1 before importing vLLM, or pass it through vllm serve, on an NVIDIA GPU with compute capability 8.0+ or an Intel XPU running the Triton attention backend. That pins RMSNorm, matmul, and attention kernels to fixed, batch-independent execution strategies. It doesn't by itself stop vLLM's async scheduling from assembling different batches on different runs, so you also need to account for that, and even then, documented gaps remain, including fused-MoE kernels.
Batch size doesn't change what the model has learned, but it can change the exact tokens a specific request produces at temperature 0, because it changes which kernel execution strategy runs. That's a difference in exact wording, not a difference in overall quality. For tasks where exact-match matters, like graded evals or RL reward signals, that token-level difference is the whole problem.






