vLLM

11 guides in this topic

vLLM is the serving engine most self-hosted LLM deployments on Spheron run on, and this cluster is where the actual configuration lives: Docker commands, tensor-parallel setup, FP8 quantization flags, and the KV-cache tuning that separates a working deployment from a fast one.

Some posts here are model-specific (deploying Llama 4 Scout, Devstral, Magistral, or DeepSeek V3.2 Speciale with vLLM), and some cover vLLM's own architecture as it evolves: Model Runner V2's throughput gains, disaggregated serving with Mooncake's KV-cache transfer engine, or vLLM-Omni's split encoder, prefill, and decode pools for multimodal models. One post compares vLLM against Ollama directly for teams deciding whether they need a production serving engine at all or something simpler for local development.

The pillar guide, vLLM Production Deployment 2026, has the working multi-GPU tensor-parallel Docker command most of the other posts assume you already have, tested on Spheron H100 SXM5 with FP8 quantization and Prometheus monitoring wired in. Start there for the baseline setup, then branch into whichever model or serving pattern matches what you're deploying.

Start Here

All vLLM Guides

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute with no contracts and no minimum. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min