The vLLM OOM preemption fix almost never lives in one flag. Every ticket we've traced comes down to gpu_memory_utilization and max_num_seqs tuned as if they were unrelated dials instead of one system. gpu_memory_utilization decides how much memory the engine is allowed to claim in total; max_num_seqs decides how fast the engine spends that allowance against incoming requests. Set the first too conservatively and you leave throughput on the table with VRAM sitting idle. Set the second too aggressively for what the first left behind and you get either a startup crash or a runtime throughput cliff that looks like a bug and isn't. This guide treats them as a pair from the first section to the last, and ends with a script for finding the right pair on your own hardware instead of guessing at one.
TL;DR: What's the vLLM OOM Preemption Fix for gpu_memory_utilization and max_num_seqs?
- The fix is a pair, not a flag: raise
gpu_memory_utilizationwhen VRAM sits unused; lowermax_num_seqs, thenmax_num_batched_tokens, under OOM orRECOMPUTEwarnings. - Default:
gpu_memory_utilizationis 0.9 in current vLLM, pulled back from 0.95 because startup profiling can misjudge peak usage. - Preemption mode: vLLM V1 defaults to
RECOMPUTE, notSWAP, since recompute costs less than swap in the V1 scheduler. - The gap most guides miss: CUDA graph capture memory scales with
max_num_seqsoutside the KV cache pool, so it can OOM even when cache-block math checks out. - Watch:
vllm:kv_cache_usage_percandvllm:num_preemptions_totalcatch this before a user reports it. - Spheron bills H100 GPU rentals per minute, cheap enough for this post's tuning sweep.
What gpu_memory_utilization Actually Allocates
gpu_memory_utilization is not a KV-cache setting with a misleading name. It's the fraction of total GPU memory vLLM profiles and reserves for the model executor as a whole, and that includes model weights, activation memory during the forward pass, CUDA graph buffers if graph capture is enabled, and the KV cache pool, all drawing from the same budget. vLLM's own engine argument documentation is explicit that this is a per-instance limit that doesn't account for other processes sharing the GPU, which matters the moment you're colocating vLLM with a display server, a second inference process, or anything else that also touches VRAM.
Weights come out first and are fixed for a given model and precision. Activation memory and CUDA graph buffers come out next and scale with batch size and sequence length, which is the part most people underestimate. Whatever is left over becomes the KV cache pool, and the KV cache pool is what determines how many concurrent sequences vLLM can actually hold in flight. On an 80GB H100 at the default 0.9, vLLM's own optimization guide notes the usable pool works out to roughly 72GB for weights and KV cache combined, not the full 80GB, because the remaining 10% is deliberately left as a safety margin against profiling error.
That margin exists for a documented reason. gpu_memory_utilization's default was originally 0.95, and it was lowered to 0.9 mainly because memory profiling at initialization time can be inaccurate, which is a polite way of saying the estimate used to be wrong often enough to crash instances that should have fit. vLLM maintainer Woosuk Kwon put it directly when the change was discussed: "we should make the memory profiling more accurate and set gpu-memory-utilization back to 0.95." The number moved down as a stopgap for a measurement problem, not because 90% is some universal safe threshold for every model and workload.
If you want an absolute number instead of a fraction, vLLM's newer kv_cache_memory_bytes argument lets you set an exact KV cache byte budget per GPU. When you set it to anything other than None, it ignores `gpu_memory_utilization` entirely and bypasses the fraction-based profiling step altogether. That's useful when you're running vLLM alongside something else on the same card and need a hard, predictable ceiling rather than a percentage of whatever VRAM happens to be free at profiling time.
The Two Failure Modes: Startup OOM vs Runtime Preemption
These are different failures with different fixes, and conflating them is why "just lower gpu_memory_utilization" advice sometimes makes things worse instead of better.
Startup crash: "No available memory for the cache blocks"
This one happens before you ever serve a request. vLLM profiles memory, subtracts weights and activation overhead from your gpu_memory_utilization budget, and finds there isn't enough left to allocate even the minimum number of KV cache blocks the engine needs to start. The fix here is almost always to raise gpu_memory_utilization toward its ceiling (if you have headroom to give), reduce max_model_len so each sequence's KV cache slot is smaller, or move to a GPU with more VRAM. Lowering max_num_seqs does nothing for this failure, because the crash happens before any sequences are scheduled.
Runtime collapse: the PreemptionMode.RECOMPUTE warning and the throughput cliff
This one happens under load, after the server has been serving requests successfully. The KV cache pool fills up as concurrent sequences accumulate, and vLLM has to make room for new work by evicting an in-flight sequence. In vLLM V1, the default preemption mode is `RECOMPUTE` rather than `SWAP`, a change the project made because recomputation has lower overhead than swapping to host memory in the V1 architecture. When it fires, you'll see a log line close to: "Sequence group 0 is preempted by PreemptionMode.RECOMPUTE mode because there is not enough KV cache space." That sequence's cached state is discarded, and the engine has to reprocess its entire prompt from scratch the next time it's scheduled, which is strictly more total compute than if the sequence had never been preempted at all. This is why preemption shows up as a throughput cliff rather than a clean error: the server stays up, requests still complete, and the failure is only visible as p99 latency and tokens/sec quietly getting worse under the exact load levels that used to work fine.
The two failures point in different directions. Startup OOM says your memory budget is too tight for what you're trying to run at all. Runtime preemption says your memory budget is fine on average but max_num_seqs lets you promise more concurrent capacity than the KV cache pool can actually back.
Symptom: Low Throughput With VRAM to Spare, Raise gpu_memory_utilization First
If nvidia-smi shows meaningful unused VRAM on a dedicated GPU and throughput is lower than you'd expect for the hardware, the KV cache pool is probably the bottleneck, not compute. A smaller KV cache pool means fewer sequences can be held concurrently, which means the batch scheduler has less work available to fill each forward pass, which shows up as underutilized tensor cores even though the GPU "looks" busy in coarse utilization metrics.
The fix here is straightforward: raise gpu_memory_utilization in increments (0.85 to 0.9, then toward 0.92-0.95 if monitoring confirms it's safe) and re-measure. This only works if the GPU is dedicated to this vLLM instance. On a shared card, raising the fraction eats into memory something else needs, and you'll trade a throughput problem for an OOM problem in whatever's colocated with you.
This is also the point at which it's worth checking whether you're actually compute-bound instead. Telling memory-bound and compute-bound stalls apart with nvidia-smi and DCGM matters here, because raising gpu_memory_utilization does nothing if the real constraint is SM occupancy rather than KV cache headroom.
Symptom: OOM or Preemption Under Load, Lower max_num_seqs (Then max_num_batched_tokens)
If you're seeing OOM errors under load, or PreemptionMode.RECOMPUTE warnings climbing in your logs, the KV cache pool is oversubscribed relative to what max_num_seqs is promising the scheduler. vLLM's own tuning guidance for this exact failure lays out the fix hierarchy in order: increase `gpu_memory_utilization`, or decrease `max_num_seqs` / `max_num_batched_tokens`, or increase `tensor_parallel_size` / `pipeline_parallel_size` to spread the model (and its memory footprint) across more GPUs.
max_num_seqs is a ceiling on concurrency, not a target you're entitled to hit. Setting it to 256 doesn't mean vLLM will comfortably run 256 concurrent sequences; it means vLLM is allowed to try, and if the KV cache pool can only back 180 of them under your current context-length distribution, the other 76 slots are where preemption comes from. Lower the ceiling first, watch num_preemptions_total go flat, and only then consider whether you actually need more concurrent capacity badly enough to raise gpu_memory_utilization or add parallelism instead.
If lowering max_num_seqs alone doesn't fully clear the problem, max_num_batched_tokens is the next lever, and it's a genuine latency tradeoff rather than a free win. vLLM's guidance is concrete on both directions: values around 2048 favor inter-token latency, because fewer prefill chunks interrupt decode steps for other sequences, while values above 8192 favor throughput and time-to-first-token, particularly for smaller models running on large GPUs where there's headroom to batch aggressively. Chunked prefill, which interleaves prefill and decode work, is enabled by default whenever possible in vLLM V1 and batches all pending decode requests ahead of any prefill operation, which is part of why max_num_batched_tokens has more effect on the prefill-decode balance than max_num_seqs does once chunked prefill is active.
For deterministic testing rather than waiting for preemption to show up under real traffic, num_gpu_blocks_override lets you force a specific KV cache block count regardless of what profiling would have picked, and it exists primarily to test preemption scenarios deterministically. Set it artificially low in a staging environment and you can reproduce a preemption cascade on demand instead of hoping it recurs in production.
The Trap the Math Doesn't Catch: CUDA Graph Capture Memory Scales With max_num_seqs Too
Here's the part the KV-cache arithmetic won't warn you about. A vLLM forum thread describes a case where raising --max-num-seqs from 90 to 92, a two-sequence bump, triggered a runtime CUDA OOM even though model loading and KV cache allocation both succeeded at startup. The KV cache pool had room. The OOM happened anyway. The unexplained gap is CUDA graph capture: vLLM captures graphs for a range of batch sizes to avoid Python-level kernel launch overhead at inference time, and the memory those captured graphs consume scales with the batch sizes you've told it to expect, which is a direct function of max_num_seqs. That memory sits outside the KV cache accounting entirely, so a max_num_seqs increase that looks safe by cache-block math can still blow the graph capture budget on its own.
vLLM gives you two ways to control this independently of the KV cache tuning above. You can disable CUDA graph capture entirely with enforce_eager=True, which trades some steady-state throughput for a lower and more predictable fixed memory overhead, or you can shrink the `cudagraph_capture_sizes` list so vLLM only captures graphs for the batch sizes you actually expect to hit rather than a wide default range. If you've lowered max_num_seqs and are still seeing an OOM that the KV cache math says shouldn't happen, this is the next place to look, not a sign that the KV cache tuning was wrong.
This is also where the deployment layer matters. vLLM Model Runner V2 changes how the model execution layer handles kernel dispatch, which shifts some of these fixed-overhead tradeoffs; if you're running MRV2, re-verify capture memory behavior separately rather than assuming the legacy runner's numbers still apply.
A Repeatable Tuning Script: Sweep Both Flags and Log Tokens/Sec on a Rented GPU
Guessing at a single gpu_memory_utilization and max_num_seqs pair and shipping it is how most of these tickets get created in the first place. A small grid sweep, run once against a representative request pattern, replaces the guess with data. Here's a minimal version that launches vLLM at each combination, fires a fixed load pattern, and logs throughput alongside the preemption counter so you can see the tradeoff directly instead of inferring it:
#!/bin/bash
# sweep_vllm_memory.sh - grid search gpu_memory_utilization x max_num_seqs
MODEL="meta-llama/Llama-3.1-8B-Instruct"
UTILS=(0.80 0.85 0.90 0.95)
SEQS=(64 128 192 256)
RESULTS="sweep_results.csv"
echo "gpu_memory_utilization,max_num_seqs,tokens_per_sec,preemptions" > "$RESULTS"
for util in "${UTILS[@]}"; do
for seqs in "${SEQS[@]}"; do
echo "Testing util=$util seqs=$seqs"
vllm serve "$MODEL" \
--gpu-memory-utilization "$util" \
--max-num-seqs "$seqs" \
--port 8000 &
VLLM_PID=$!
sleep 45 # wait for startup and profiling to finish
# replace with your real load generator; this hits the metrics endpoint
# after a fixed benchmark run against /v1/completions
python3 bench_client.py --endpoint http://localhost:8000 --duration 60 \
--concurrency 200 --output /tmp/bench.json
TOKENS_SEC=$(python3 -c "import json; print(json.load(open('/tmp/bench.json'))['tokens_per_sec'])")
PREEMPT=$(curl -s http://localhost:8000/metrics | grep 'vllm:num_preemptions_total' | tail -1 | awk '{print $2}')
echo "$util,$seqs,$TOKENS_SEC,$PREEMPT" >> "$RESULTS"
kill $VLLM_PID
wait $VLLM_PID 2>/dev/null
sleep 10
done
done
echo "Sweep complete. Sort by tokens_per_sec where preemptions == 0:"
sort -t, -k3 -rn "$RESULTS" | awk -F, '$4 == 0'The point of sorting by throughput only among rows with zero preemptions is that a cell with high tokens/sec and nonzero preemptions is lying to you: some of that throughput is recompute work spent on sequences that got evicted, not net-new completed work. The best cell in this grid is the highest-throughput row that never preempted, not the highest-throughput row on the sheet.
Running this grid is a short-lived job, sixteen combinations at roughly two minutes each is well under an hour end to end, which is exactly the kind of workload where paying for reserved capacity doesn't make sense. Spheron bills GPU rentals per minute with no long-term commitment, on-demand or spot, so you can spin up an H100 SXM5 for the duration of the sweep and shut it down when the grid finishes rather than holding a reservation for a one-off benchmark. It's worth being clear about the limits of that fit: nothing on Spheron's pricing page addresses preemption or scheduler behavior specifically, because that's entirely a vLLM-side configuration problem, and it behaves identically on Spheron, a hyperscaler, or your own hardware. What per-minute billing buys you here is a cheap place to run the sweep, not a shortcut around doing it.
Watch the Two Metrics That Predict This Before It Pages You
vLLM exposes both signals you need through its Prometheus endpoint, and neither requires custom instrumentation. vllm:num_preemptions_total is a Prometheus counter tracking cumulative preemptions from the engine; vllm:kv_cache_usage_perc is a gauge, scaled 0 to 1, for the fraction of KV cache blocks currently in use. Together they turn "throughput mysteriously dropped" into a specific, attributable cause before a user files a ticket about it.
A healthy instance shows kv_cache_usage_perc fluctuating with load but rarely pinned at or near 1.0, and num_preemptions_total staying flat between scrapes. If usage is consistently pegged near the ceiling and the preemption counter is climbing steadily rather than staying flat, that's the pair confirming the diagnosis from the sections above: max_num_seqs is set higher than the KV cache pool gpu_memory_utilization leaves behind can actually support, well before it shows up as a user-visible latency spike. For the alerting and dashboard setup to wire these into Grafana alongside the rest of your GPU fleet metrics, GPU monitoring for ML covers the DCGM and Prometheus stack this plugs into.
Preemption thrash is also worth framing at the fleet level, not just the single-instance level. An instance that's technically "up" and reporting high GPU utilization while quietly recomputing a growing share of its work is a goodput problem: the utilization number looks fine and the completed-request rate is what's actually degrading. GPU goodput engineering covers that gap in more detail if you're diagnosing this across a cluster rather than one node.
Quick Reference: The vLLM OOM Preemption Fix, by Symptom
| Symptom | Move this flag first | Why |
|---|---|---|
| "No available memory for the cache blocks" at startup | Raise gpu_memory_utilization, or lower max_model_len | Crash happens before any sequences are scheduled, so max_num_seqs isn't involved yet |
OOM or PreemptionMode.RECOMPUTE warnings under load | Lower max_num_seqs | Concurrency ceiling is oversubscribing the KV cache pool that gpu_memory_utilization actually leaves behind |
Preemption persists after lowering max_num_seqs | Lower max_num_batched_tokens | Controls the prefill/decode token budget per forward pass independently of sequence count |
OOM despite max_num_seqs reduction and cache math checking out | Shrink cudagraph_capture_sizes or set enforce_eager=True | CUDA graph capture memory scales with max_num_seqs outside the KV cache accounting entirely |
| Low throughput with VRAM visibly unused | Raise gpu_memory_utilization | Small KV cache pool limits concurrent sequences the scheduler has available to batch |
| Need to reproduce preemption on demand for testing | Set num_gpu_blocks_override low | Forces a fixed cache block count so the failure is deterministic instead of load-dependent |
Model still won't fit after maxing gpu_memory_utilization | Raise tensor_parallel_size or pipeline_parallel_size | Spreads weights and KV cache across more GPUs instead of squeezing one |
No specific Spheron rate is quoted above. GPU pricing fluctuates based on availability and varies by region, so check current GPU pricing → for live rates as of your read date, not this post's publish date of 29 Sep 2026, before running the sweep script above.
If you're still deciding between vLLM and another serving stack before investing time in tuning it this closely, our decision framework for vLLM, TensorRT-LLM, and SGLang is the right starting point. And if gpu_memory_utilization and max_num_seqs are maxed out and you still need more concurrent capacity than one GPU's memory can back, prefix cache sharing across nodes is the next lever, not a bigger max_num_seqs value on a single instance. The deeper reason both flags fight over the same scarce resource is covered in why GPU memory, not compute, is usually the binding constraint on inference for LLM serving generally.
Running the sweep in this post takes minutes, not hours, which is exactly what per-minute billing on Spheron is built for.
Quick Setup Guide
Set num_gpu_blocks_override to a small value to force preemption on demand instead of waiting for it to happen under real load. This exists specifically so preemption scenarios can be tested deterministically rather than chased in production logs.
Check vllm:kv_cache_usage_perc and vllm:num_preemptions_total in your Prometheus metrics. Usage pinned near 1.0 with a climbing preemption counter means max_num_seqs is set too high for the KV cache pool gpu_memory_utilization actually leaves behind.
Low throughput with VRAM headroom to spare: raise gpu_memory_utilization first. OOM or preemption warnings under load: lower max_num_seqs, then max_num_batched_tokens if the problem persists. Do not move both at once, or you lose the ability to tell which change fixed it.
If lowering max_num_seqs doesn't fully resolve a startup or early-request OOM, the cudagraph_capture_sizes list scales with max_num_seqs and can still be claiming more fixed memory than expected. Shrink the capture size list or set enforce_eager=True to isolate whether graph capture, not the KV cache, is the actual ceiling.
Run a small grid across gpu_memory_utilization and max_num_seqs on a rented GPU, logging tokens/sec and preemption count per cell. The combination that maximizes throughput while keeping num_preemptions_total at zero is the one to ship, not the combination that maximizes either flag in isolation.
Frequently Asked Questions
It sets the fraction of total GPU memory vLLM is allowed to profile and claim for the model executor as a whole: weights, activation memory, CUDA graph buffers, and the KV cache pool, combined into one number. It is not a KV-cache-only setting, and vLLM does not know or care about other processes sharing the same GPU, so the true ceiling is lower than the raw VRAM figure whenever anything else (a display service, another framework, a second vLLM instance) is also resident.
Two common causes. First, gpu_memory_utilization is a fraction vLLM commits to at startup based on a profiling pass, not a live measurement, so if that pass under- or overestimates activation memory, the actual peak can exceed what was reserved. Second, raising max_num_seqs increases the size of the CUDA graphs vLLM captures for larger batch sizes, and that capture memory sits outside the KV cache accounting entirely, so a request that looks fine on paper (KV blocks available) can still fail during graph replay.
0.9 is vLLM's current default and a reasonable starting point on a dedicated GPU. The default was 0.95 for a period before the project pulled it back to 0.9, specifically because startup memory profiling can be inaccurate and 0.95 left too little margin for that error. Push it toward 0.95 only after you've confirmed with monitoring that the instance isn't OOMing at peak load, and never on a GPU shared with another process.
No, and past a certain point it does the opposite. max_num_seqs is an upper bound on concurrent sequences, not a promise of concurrency. Once the KV cache pool fills, vLLM has to preempt in-flight sequences to make room for new ones, and each preemption throws away that sequence's cached state and forces a full recompute from scratch. A higher ceiling that triggers preemption costs more total compute per completed request than a lower one that never does.
max_num_seqs caps how many sequences can be scheduled concurrently. max_num_batched_tokens caps how many tokens across all of those sequences can be processed in a single forward pass, which is the lever that actually governs prefill-versus-decode balance. A smaller value (around 2048) favors inter-token latency because fewer prefill chunks interrupt decode steps; a larger value (vLLM recommends above 8192 for throughput, especially with smaller models on large GPUs) favors time-to-first-token by letting prefills complete faster.






