Engineering

GPU Requirements Cheat Sheet 2026: VRAM + Cost, 18 Models

Back to BlogWritten by Published May 14, 2026Updated
GPU RequirementsVRAMOpen Source AILlama 4DeepSeekQwenMistralLLM DeploymentAI Infrastructure
GPU Requirements Cheat Sheet 2026: VRAM + Cost, 18 Models

The Quick Reference Table

This table covers the most popular open-source models in production today. VRAM numbers assume inference (not training) with the most common quantization level for each model size. For any model not listed here, run it through the live VRAM calculator: it reads the exact parameter count from any HuggingFace model and matches it to the cheapest GPU with current pricing. This table is the static snapshot; the calculator is the live version. For a model well past this table's range, the 2.8T Kimi K3 VRAM breakdown shows what multi-node sizing looks like at the extreme.

ModelParametersQuantizationVRAM NeededMin GPU ConfigSpheron Cost/hr
Llama 4 Scout109B (17B active)INT4~67 GB1x H100 80GBFrom $2.20/hr
Llama 4 Scout109B (17B active)FP16~218 GB4x H100 80GBFrom $8.80/hr
Llama 4 Maverick400B (17B active)INT4~240 GB4x H100 80GBFrom $8.80/hr
Llama 4 Maverick400B (17B active)FP16~800 GB8x H200 141GBFrom $26.48/hr
DeepSeek V3.2671B (37B active)FP8~700 GB8x H200 141GB (8x H100 80GB is insufficient at FP8)From $26.48/hr
DeepSeek V4-Pro1.6T (49B active)FP4+FP8~865 GB8x H200 141GBFrom $26.48/hr
DeepSeek V4-Flash284B (13B active)FP4+FP8~160 GB2x H200 141GBFrom $6.62/hr
GLM 5.1754B (40B active)FP8~754 GB8x H200 141GBFrom $26.48/hr
Qwen 3.6 35B-A3B35B (3B active)FP8~35 GB1x L40S 48GBFrom $0.96/hr
Qwen 3.5 397B397B (17B active)FP8~397 GB8x H100 80GBFrom $17.60/hr
Qwen 3.5 397B397B (17B active)INT4~240 GB4x H100 80GBFrom $8.80/hr
Qwen 3.5 27B27BFP8~27 GB1x H100 80GBFrom $2.20/hr
Qwen 3.5 9B9BFP8~9 GB1x RTX 4090 24GBFrom $0.72/hr
Qwen 2.5 72B72BINT4~43 GB1x H100 80GBFrom $2.20/hr
Qwen 2.5 72B72BFP16~144 GB2x H100 80GBFrom $4.40/hr
Qwen 3 32B32BINT4~20 GB1x RTX 4090 24GB (~20K context at FP16 KV)From $0.72/hr
Qwen 3 32B32BFP16~66 GB1x H100 80GBFrom $2.20/hr
Mistral Large 2123BINT4~74 GB1x H100 80GB (~30K context at FP16 KV)From $2.20/hr
Mistral Large 2123BFP16~246 GB4x H100 80GBFrom $8.80/hr
Kimi K2.51T (32B active)INT4~600 GB8x H100 80GBFrom $17.60/hr
Nemotron Ultra 253B253BINT4~152 GB2x H100 80GB (short context), 2x H200 141GBFrom $4.40/hr
Nemotron Ultra 253B253BFP16~506 GB8x H100 80GBFrom $17.60/hr
Nemotron 3 Super (120B MoE)120B (12B active)NVFP4~80 GB1x B200 192GB (native FP4)From $5.34/hr
Llama 3.3 70B70BINT4~42 GB1x H100 80GBFrom $2.20/hr
Llama 3.3 70B70BFP16~140 GB2x H100 80GBFrom $4.40/hr
Phi-4 14B14BINT4~9 GB1x RTX 4090 24GBFrom $0.72/hr
Phi-4 14B14BFP16~29 GB1x A100From $1.21/hr

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 17 Sep 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

Multi-GPU costs are the per-GPU rate times the GPU count. INT4 figures use ~0.6 bytes per parameter, which is what real GPTQ, AWQ and GGUF Q4_K_M files weigh, not the theoretical 0.5. The context limits in brackets assume an FP16 KV cache; an FP8 KV cache roughly doubles them. For full vLLM setup steps for Qwen 3.5, see the Qwen 3.5 deployment guide.

Nemotron 3 Super's NVFP4 checkpoint is 80.4 GB, larger than 120B × 0.5 bytes suggests, because NVFP4 carries a scale for every 16 weights and several layers stay at higher precision. NVFP4 is a Blackwell-native format, so a B200 runs it at full speed. For a full deployment walkthrough of Nemotron 3 Super including VRAM breakdowns across all precision tiers and vLLM configuration, see the Nemotron 3 Super GPU deployment guide. For a model-first view organized by VRAM tier with Spheron cost-per-tier calculations, see the best open-source LLMs to self-host in 2026 by VRAM tier.

Note on DeepSeek V3.2 VRAM constraint: At FP8, the 671B parameters require roughly 671 GB for weights alone, plus ~30-60 GB for KV cache, activations, and framework overhead. That puts minimum practical VRAM at ~700 GB, which exceeds 8× H100 80GB (640 GB total). Production deployment requires 8× H200 141GB (1,128 GB total, $38.40/hr) or a split across multiple H100 nodes with CPU offload. See our DeepSeek V3.2 deployment guide for setup details.

Note on DeepSeek V4: it ships in two sizes, both in a native mixed format with FP4 MoE experts and FP8 for everything else. V4-Pro (1.6T total, 49B active) is a ~865 GB checkpoint, which fits one 8× H200 node (1,128 GB) with room for KV cache. V4-Flash (284B, 13B active) is ~160 GB and fits 2× H200. An FP8 conversion roughly doubles the expert weights, so don't size from the parameter count at 1 byte each. See the DeepSeek V4-Pro deployment guide and the DeepSeek V4-Flash guide for multi-node and expert parallelism configuration.

Note on GLM 5.1 and Qwen 3.6: GLM 5.1 is a 754B MoE with 40B active parameters, so it needs a full 8× H200 node at FP8; see the GLM-5.1 deployment guide. Qwen 3.6 Plus is not in the table because Alibaba serves it through its API only, with no downloadable weights. The open Qwen 3.6 models are the 35B-A3B MoE listed above and a dense 27B.

Note on Qwen3.7 Flash: it's deliberately absent from the table above. Alibaba never published downloadable weights for it, so there's no VRAM figure to size against. If you came here looking for Qwen3.7 Flash GPU requirements, the real answer is that Qwen3.6-35B-A3B is the closest self-hostable sibling and the one this cheat sheet's Qwen rows actually apply to.

Note on Qwen3.8-Max: also absent from the table above; at 2.4T total parameters it needs a multi-node cluster far outside this table's single- and dual-GPU scope, so see the dedicated Qwen3.8-Max GPU requirements guide for the full FP8 and quantized cluster sizing.


Every time a new model releases, the first question is always the same: "What GPU do I need to run this?"

The answer is scattered across Hugging Face model cards, Reddit threads, and provider-specific docs, none of which agree with each other. This post puts it all in one place: exact VRAM requirements, recommended GPU configurations, and estimated cloud costs for every major open-source model available in May 2026. For deeper technical understanding, see our comprehensive GPU memory requirements guide.

Bookmark this page. We update it as new models release.

For teams planning infrastructure beyond 2026, the NVIDIA Rubin R100 guide covers next-generation GPU specs and cloud availability projections. Rubin ships in H2 2026, so if you want to lock capacity early, you can register for R100 Rubin availability with your GPU count, timeline, and workload, and the team reaches out as allocation opens.

Llama 4 Scout

Llama 4 Scout (109B total, 17B active, MoE) is the most cost-effective entry point in the Llama 4 family. At INT4 quantization, the ~67 GB weight footprint fits on a single H100 80GB from $2.20/hr on Spheron. The 10M token context window is the standout spec: no other model in this comparison comes close at this price, though it's also where a naive VRAM estimate falls apart; our Llama 4 VRAM calculator breakdown covers exactly how much of that headline context you can actually afford to use. Suitable for conversational AI, RAG, and long-document tasks where context length matters. See the Llama 4 deployment guide for vLLM configuration.

DeepSeek V3.2

DeepSeek V3.2 (671B total, 37B active, MoE) leads on math and multi-step reasoning. At FP8, minimum practical VRAM is ~700 GB, requiring 8× H200 141GB ($38.40/hr on Spheron). The high hardware bar is the cost of entry for its benchmark leadership. For coding and general tasks, Qwen3-32B on a single H100 at $2.98/hr is a better dollar-per-quality choice. See our DeepSeek V3.2 deployment guide for multi-node setup and expert offloading strategies.

Qwen 3.5 27B

Qwen 3.5 27B is a dense 27B model that fits on a single H100 80GB at FP8 (~27 GB VRAM, from $2.20/hr). It is the largest dense model in the Qwen 3.5 line-up, a solid choice for teams that need more capability than the 9B but can't justify multi-GPU costs. For full vLLM setup steps, see the Qwen 3.5 deployment guide.

Qwen 2.5 72B

Qwen 3 has no 72B dense model; its largest dense model is 32B, so the 72B slot belongs to Qwen 2.5 72B. At INT4 it fits on a single H100 80GB with ~43 GB of weights at $2.98/hr, leaving plenty of room for KV cache. At FP16, it needs 2× H100s ($5.96/hr). Before you pay for the extra memory, note that the Qwen3 technical report shows Qwen3-32B-Base beating Qwen2.5-72B-Base on 10 of 15 benchmarks at less than half the parameters, so the 32B is usually the better buy. For how Qwen 3 stacks up against Llama 4 Scout and DeepSeek on coding and reasoning, see the DeepSeek vs Llama 4 vs Qwen 3 comparison.

Mistral Large 2

Mistral Large 2 (123B) at INT4 has ~74 GB of weights, which fits on a single H100 80GB at $2.98/hr but leaves room for only about 30K tokens of FP16 KV cache (88 layers × 8 KV heads works out to ~350 KB per token). Switch to an FP8 KV cache or move to an H200 for longer context. At FP16 (~246 GB), it needs 4× H100s at $11.92/hr. It has strong instruction following and function calling. Note the license: Mistral Large 2 ships under the Mistral Research License, so commercial deployment needs a separate license from Mistral. For teams needing a 100B+ class model on a single GPU without MoE complexity, it's a straightforward option. For the deployment guide, see the Mistral deployment guide.

Kimi K2.5

Kimi K2.5 (1T total, 32B active, MoE) at INT4 needs ~600 GB for weights, fitting on 8× H100 80GB ($23.84/hr on Spheron). Its compressed MLA attention keeps the KV cache small, so the roughly 80 GB left across the node's 640 GiB goes further than it would on a GQA model. Despite 1T total parameters, the 32B active parameter count keeps per-token compute manageable. Its strength is long-context reasoning and agentic tasks. Moonshot's newer release follows a similar hardware profile: see the Kimi K2.6 deployment guide for updated VRAM math and agentic-swarm configuration (300-agent swarms, 4,000 coordinated steps).

How to Read This Table

Parameters vs Active Parameters. Mixture-of-Experts models (Llama 4, DeepSeek, Kimi K2.5) have a total parameter count much higher than their active parameter count. The total count determines VRAM needs (all weights must be loaded). The active count determines inference speed: fewer active parameters means faster per-token generation.

VRAM Needed. This is the approximate memory for the model weights alone at the listed quantization level. INT4 rows use ~0.6 bytes per parameter because real 4-bit files carry per-block scales on top of the 4-bit weights. The figure doesn't include the KV cache, and the KV cache can't be estimated as a percentage of this number: it scales with context length, layer count and KV head count, and quantizing the weights does nothing to it. Llama 3.3 70B needs about 10 GB of FP16 KV cache per request at 32K context and about 40 GB at 128K, which is as much as its INT4 weights. Phi-4 needs about 3.4 GB at its 16K maximum context, over a third of its INT4 weights. Use the KV formula below, then add 1-2 GB for CUDA context and the serving framework.

Min GPU Config. The minimum hardware configuration that fits the model. For production inference with concurrent users, you'll often want more VRAM than the minimum, the extra headroom enables higher throughput through larger batch sizes.

VRAM Calculation Rules of Thumb

If a model is not in the table above, you can estimate VRAM requirements using these formulas:

FP16 (half precision): Parameters in billions × 2 = VRAM in GB. A 70B model needs ~140 GB in FP16.

FP8 (8-bit weights): Parameters in billions × 1 = VRAM in GB. A 70B model needs ~70 GB in FP8.

INT4 (4-bit quantization): Parameters in billions × 0.6 = VRAM in GB. A 70B model needs ~42 GB in INT4. The theoretical figure is 0.5, but GPTQ, AWQ and GGUF store a scale (and often a zero-point) for every block of 32-128 weights and usually keep embeddings at 16-bit, so real files land between 0.55 and 0.65. Qwen's own Qwen3-32B Q4_K_M file is 19.8 GB for 32.8B parameters, 0.60 bytes per parameter.

KV cache: Per token, KV cache = 2 × layers × KV heads × head dimension × bytes per element (2 at FP16, 1 at FP8). For common GQA models at FP16 that is 128 KB per token for Llama 3.1 8B, 256 KB for Qwen3-32B and 320 KB for Llama 3.3 70B, or roughly 130-330 MB per 1K tokens of context per concurrent request. At 128K context with 8 concurrent requests, Llama 3.3 70B needs over 300 GB of FP16 KV cache. An FP8 KV cache halves every one of these figures.

GPU Selection Guide by Workload

Development and Experimentation

Best option: 1x RTX 4090 (from $0.72/hr) or 1x A100 (from $1.21/hr)

For prompt engineering, testing model quality, and running small-batch inference, you don't need datacenter GPUs. Most models up to 32B parameters fit on a single RTX 4090 with INT4 quantization. Larger models (70B) fit on a single A100 80GB.

Spot instances suit experimentation because an interruption costs you little while you're iterating. Spot capacity can be reclaimed without notice, so compare both tiers on the pricing page before you pick.

Single-Model Production Inference

Best option: 1x H100 (from $2.20/hr) or 1x H200 (from $3.31/hr)

For serving one model to production traffic, an H100 handles most 70B-class models comfortably. The H200's extra VRAM (141 GB vs 80 GB) gives you headroom for longer contexts and higher concurrency without sharding across multiple GPUs. See how best NVIDIA GPUs for LLMs compares these options for your specific use case.

Large Model Serving (100B+)

Best option: Multi-GPU H100 or B300 cluster

Past roughly 130B parameters, even INT4 weights outgrow a single 80 GB card. For the largest models (DeepSeek V3.2, DeepSeek V4-Pro, GLM 5.1, Kimi K2.5), 8-GPU configurations are the minimum.

On Spheron, you can provision multi-GPU baremetal servers with H100, H200, and B300 GPUs. B300 instances, from $5.40/hr per GPU, give you 288 GB VRAM per GPU, meaning a single B300 can run models that would require 2-4 H100s. NVIDIA isn't the only rack-scale option at this end of the table anymore either: AMD's Helios platform packs 72 MI455X GPUs into a single rack, and our AMD Helios rack-scale deployment guide covers when fractional MI-series capacity is worth considering over a full NVIDIA cluster commit.

Training and Fine-Tuning

Best option: H100 or B300 with NVLink

Fine-tuning requires more VRAM than inference because you need to store the model weights, gradients, and optimizer states simultaneously. A general rule: full fine-tuning with Adam in mixed precision needs about 16 bytes per parameter (weights, gradients, an FP32 master copy and two optimizer moments), roughly 8x the FP16 inference weights. A 70B model needs over 1.1 TB before activations, which is more than one 8x H100 node.

QLoRA dramatically reduces fine-tuning memory requirements. With QLoRA, you can fine-tune a 70B model on a single H100 80GB, or Llama 4 Scout on a single H100 using ~71 GB of VRAM. Learn how to deploy Llama 4 on GPU cloud for production fine-tuning.

The Cost Efficiency Tiers

Not every model needs the most expensive GPU. Here's how to think about cost efficiency:

Tier 1: Consumer and Workstation GPUs

RTX 4090, L40S

The RTX 4090's 24 GB runs models up to about 32B at INT4 with short context: Qwen 3 32B, Phi-4 14B, Llama 3.1 8B, Mistral 7B. The L40S's 48 GB stretches to 70B at INT4, with roughly 20K tokens of FP16 KV cache left over.

See our L40S inference benchmark guide for a detailed performance and pricing analysis.

Tier 2: Single Datacenter GPU

A100 80GB, H100 80GB, H200 141GB

Run models up to 70-100B parameters quantized, or up to 40B in full precision. The sweet spot for most production deployments. Covers: Qwen 2.5 72B, Llama 3.3 70B, Llama 4 Scout (quantized), Mistral Large 2 (quantized).

Tier 3: Multi-GPU Clusters

4-8x H100, 4-8x H200, 4-8x B300

Required for 100B+ models and the largest MoE architectures. Covers: Llama 4 Maverick, DeepSeek V3.2, DeepSeek V4, GLM 5.1, Kimi K2.5, Nemotron Ultra 253B (FP16).

Video Generation GPU Requirements

Video AI models have different requirements from LLMs. Where a 70B LLM at INT4 needs ~42GB VRAM, a 5-second 720p Wan 2.1 clip needs 65–80GB on an H100. The top open-source video models (Wan 2.1/2.2, HunyuanVideo) require datacenter hardware and cannot run on consumer GPUs.

ModelMin VRAMMin GPUNotes
LTX-2.3 (720p)24–32GBRTX 4090 (fp8) or RTX 5090fp8 quant required below 32GB
CogVideoX-1.5-5B (720p)24–32GBRTX 40908-bit reduces to ~16GB
Wan 2.1/2.2 (480p)40–48GBH100 PCIeWith fp8 quantization
Wan 2.1/2.2 (720p)65–80GBH100 SXM5Tight on 80GB
HunyuanVideo (720p)60–80GBH200 SXMH100 carries OOM risk

For model-by-model quality and cost comparisons, see AI video generation GPU guide. For robotics workloads, see Cosmos world foundation model GPU requirements covering VRAM and instance sizing for synthetic training data pipelines.

Time series foundation models (Chronos-T5-Large at ~2.5 GB, Lag-Llama at <0.1 GB) fit comfortably in the L40S tier without H100-class hardware. See the full deployment guide for complete VRAM tables by model variant.

What's Coming Next

The model landscape moves fast. The frontier of open weights already runs from 754B (GLM 5.1) to 1.6T (DeepSeek V4-Pro) total parameters, and each release pushes VRAM requirements higher. GPU improvements (B300's 288 GB, Rubin on the roadmap) and 4-bit formats like NVFP4 and MXFP4 are what keep that hardware within reach.

We'll update this cheat sheet as new models ship. If you're planning GPU infrastructure for the next 6-12 months, size for the largest model you expect to run, add its KV cache at the context length and concurrency you'll actually serve, and leave room for the next model up.

Need GPUs matched to your model? Spheron has H100, H200, B300, A100, and RTX 4090 instances with per-minute billing and no commitments. Pick the GPU that fits your VRAM requirements and deploy in minutes.

Explore GPU rental options →

FAQ / 04

Frequently Asked Questions

Llama 4 Scout (109B total, 17B active) requires ~67 GB for its weights at INT4 quantization, fitting on a single H100 80 GB with room for a moderate KV cache. At FP16 it needs ~218 GB, requiring 4x H100s or a single B300. Llama 4 Maverick (400B total, 17B active) requires ~240 GB at INT4, needing 4x H100s. At FP16 it needs ~800 GB, requiring 8x H200s.

No. DeepSeek V3.2 has 671B total parameters (37B active). At FP8, weights alone require ~671 GB, and including KV cache and activations the minimum practical VRAM lands around 700 GB. That exceeds 8x H100 80GB (640 GB total), so production deployment requires 8x H200 141GB (1,128 GB total) or a multi-node split with CPU offload.

A real 4-bit file for a 70B model is about 42 GB, since GPTQ, AWQ and GGUF land near 0.6 bytes per parameter rather than 0.5. The cheapest card that holds it is an L40S (48 GB) from $0.96/hr on Spheron, which leaves about 7 GB for KV cache, roughly 20K tokens at FP16. An A100 80GB from $1.21/hr on Spheron leaves room for 32K+ context. For INT8, a single H100 80GB from $2.20/hr on Spheron holds the ~70 GB of weights with a modest KV cache. Rates as of 06 Oct 2026.

Multiply the parameter count (in billions) by bytes per parameter: FP16 = 2, INT8 = 1, and about 0.6 for INT4, because real 4-bit files store a scale for every small block of weights and never reach the theoretical 0.5. A 70B model at INT4 needs about 42 GB for weights. Then add the KV cache, which does not scale with weight size: 2 x layers x KV heads x head dimension x bytes per element, per token, times context length, times concurrent requests. Llama 3.3 70B at FP16 KV needs about 10 GB per request at 32K context and 40 GB at 128K. Add 1-2 GB for CUDA context and the serving framework. For MoE models, use total parameter count (not active), since all expert weights must be loaded.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min