Nemotron 3 Super hit 60.47% on SWE-Bench Verified when NVIDIA released it on March 11, 2026, ahead of GTC (March 16-19, 2026). The number that makes it interesting for production deployment is 12B: that's the active parameter count per forward pass, despite a 120B total parameter count. This guide covers the GPU math, vLLM configuration, and cost breakdown for running it yourself on cloud infrastructure.
TL;DR: What GPU Do You Need to Self-Host Nemotron 3 Super?
- Model: Nemotron 3 Super is a 120B-total, 12B-active hybrid Mamba-Transformer MoE, and all 120B weights must sit in VRAM.
- NVFP4: the NVFP4 checkpoint is 80.4 GB, so it does not fit one 80 GB H100. Run it on 1x B200 (192 GB) or 1x B300, which have native FP4 tensor cores.
- FP8 on Hopper: about 120 GB of weights fits 2x H100 or 1x H200 (141 GB) with modest KV cache.
- BF16: about 240 GB, so 8x H100 for production and 4x H100 for short-context evaluation only.
- Price: Spheron B200 runs $8.63/hr and H100 $2.98/hr per GPU on-demand, as of 02 Oct 2026. Compare live rates.
Before diving into deployment, if you're already running vLLM for other models, the vLLM production deployment guide covers the baseline setup you'll want in place first. For NVIDIA's newest open-weight model with full dense-transformer architecture and stronger reasoning benchmarks, see the Nemotron Ultra 253B deployment guide. For the largest model in the Nemotron 3 family, the Nemotron 3 Ultra 550B deployment guide covers multi-node GPU planning and expert-parallel vLLM config for the 550B MoE tier.
Nemotron 3 Super Architecture: What the Hybrid Mamba-Transformer MoE Means for Inference
Nemotron 3 Super is not a standard transformer. Understanding the architecture directly affects how you configure vLLM, what hardware you need, and where performance bottlenecks will appear.
Standard transformer attention layers compute query-key-value products across the entire sequence at every layer. Memory consumption grows quadratically with sequence length because every token attends to every other token. For long context inference, this is expensive.
SSM (Structured State Space Model) layers, as implemented in Mamba, replace the attention computation with a recurrent state that processes tokens sequentially. Memory consumption scales linearly with sequence length rather than quadratically. The trade-off: SSM layers maintain a running state that is fundamentally sequential during decode, which limits certain forms of batching.
Nemotron 3 Super interleaves SSM layers with standard attention layers through the same model stack. You get the KV cache savings of SSM for long sequences while keeping the parallel attention layers that handle complex reasoning. IBM uses the same interleaved SSM+attention pattern in Granite 4.1 30B, which is worth comparing if you need Apache 2.0 licensing or IBM's cryptographic signing alongside Mamba-style memory efficiency. See the IBM Granite 4.1 deployment guide for the Granite-specific vLLM config and sizing tables.
The MoE (Mixture of Experts) component is separate from the SSM/attention distinction. The model has 120B total parameters split across experts, but only roughly 12B activate per forward pass. This is why you need to load all 120B into VRAM even though only 10% compute on any given token.
LatentMoE and Multi-Token Prediction
Two architectural features distinguish Nemotron 3 Super from standard MoE models.
LatentMoE: Rather than routing input tokens directly to experts, the model first projects tokens into a compressed latent space before expert selection. This lets the routing mechanism activate 4x more experts at the same compute cost compared to standard MoE. More experts contributing to each token improves quality without a proportional VRAM or compute increase.
Multi-Token Prediction (MTP): The model is trained to predict multiple future tokens per forward pass, enabling native speculative decoding. Unlike most speculative decoding setups that require a separate smaller draft model, Nemotron 3 Super handles draft generation internally. This reduces inference latency for medium-to-long responses without extra GPU allocation for a draft model.
Memory Access Patterns
SSM layers have linear sequence scaling vs quadratic for attention. The model supports context windows up to 1M tokens. For a 128K context window, the KV cache from attention layers grows at the usual rate, but SSM layers add only a fixed-size recurrent state buffer per layer regardless of context length. At 1M context, this difference becomes substantial: pure transformer KV cache would be enormous, while Nemotron 3 Super's SSM layers add no additional cache pressure.
For production deployments, the single-B200 NVFP4 vLLM command in this guide uses --max-model-len 32768 as a practical default. The 4x H100 BF16 evaluation configuration uses --max-model-len 16384 due to limited headroom after loading weights. The full 1M context is available on configurations with sufficient VRAM headroom. For repository-level code analysis or document processing with very long inputs, you can increase --max-model-len as long as your GPU has headroom.
The practical implication: Nemotron 3 Super's effective VRAM for long context is lower than a comparable parameter-count pure transformer. A 120B dense model at BF16 with 128K context would require substantially more total memory than Nemotron 3 Super at the same precision.
Prefill vs Decode Differences
During decode, SSM layers compute token by token with a recurrent state. This is sequential by design. Standard attention layers can process in parallel with chunked prefill, but SSM layers cannot correctly initialize their recurrent state across chunk boundaries without special handling.
This is a correctness issue, not just a performance issue. vLLM enables chunked prefill by default in recent versions for long-context efficiency. For Nemotron 3 Super, you must pass --no-enable-chunked-prefill until you have validated that your specific vLLM version handles SSM chunk boundaries correctly. Incorrect SSM state initialization produces wrong outputs, not just slower outputs. Note: vLLM 0.17.1 includes SSM-aware chunked prefill improvements that may handle chunk boundaries correctly; test explicitly before relying on this in production.
| Architecture Component | Standard Transformer | Nemotron 3 Super |
|---|---|---|
| Attention layers | All layers | Alternating with SSM |
| KV cache growth | O(layers x seq_len) | Reduced (fewer attention layers) |
| Decode state | Stateless | SSM layers maintain recurrent state |
| Long-context VRAM | High | Lower (fewer KV cache layers) |
| Active params per token | 100% | ~10% (12B of 120B) |
GPU Requirements: VRAM Budgets by Precision Tier
The key insight with MoE models: you load all 120B parameters into VRAM even though only 12B activate per token. This is fundamentally different from a 12B dense model. You need enough VRAM for the full parameter count, not just the active parameters.
For the general formula and a broader VRAM reference across model families, see GPU memory requirements for LLMs.
VRAM Formula for MoE Models
VRAM = (total_params x bytes_per_param) + KV_cache + activation_overheadWorked examples for Nemotron 3 Super (120B total parameters):
- BF16: 120B x 2 bytes = 240 GB minimum. Requires 8x H100 80GB for comfortable production headroom. 4x H100 (320 GB total) technically holds the weights but leaves only 80 GB across 4 cards for KV cache and activations, so treat 4x H100 as an evaluation floor, not a production configuration. A single B200 (192 GB HBM3e) cannot hold the BF16 model (240 GB > 192 GB).
- FP8: 120B x 1 byte = 120 GB minimum. Fits on 2x H100 80GB with reasonable headroom for KV cache, or on a single H200 (141 GB) with modest KV cache.
- NVFP4: 120B x 0.5 bytes = 60 GB for the raw 4-bit weights, but the published NVFP4 checkpoint is 80.4 GB on disk once scale factors and the layers kept at higher precision are included. That does not practically fit a single H100 80GB: nothing is left for KV cache, Mamba state, or the CUDA context. A single B200 (192 GB) holds it with roughly 100 GB to spare.
- GGUF Q4_K_M: Roughly 0.6 bytes per parameter, about 72 GB. It loads on a single H100 80GB only with a small context window, and llama.cpp can offload layers to CPU for environments without an 80 GB GPU.
GPU Tier Recommendations
| Use Case | Precision | GPUs Needed | Spheron Option |
|---|---|---|---|
| Single-GPU serving (Blackwell) | NVFP4 | 1x B200 | $8.63/hr |
| Single-GPU serving (Hopper) | FP8 | 1x H200 | $4.80/hr |
| Staging / small production | FP8 | 2x H100 | $5.96/hr |
| Evaluation only (short context) | BF16 | 4x H100 | $11.92/hr |
| Production (BF16 recommended) | BF16 | 8x H100 | $23.84/hr |
The B200 row is the one to start with: a single B200 (192 GB HBM3e) holds the 80.4 GB NVFP4 checkpoint with roughly 100 GB left for KV cache, and Blackwell runs NVFP4 on native FP4 tensor cores. B200 is also sold as spot at $4.49/hr, but spot capacity can be reclaimed without notice. If you are staying on Hopper, FP8 on one H200 or two H100s is the path; the NVFP4 checkpoint is too large for a single 80 GB card. Note that the full BF16 model requires 240 GB, which exceeds B200 VRAM, so use 8x H100 for BF16 production workloads. See the B200 GPU rental page for availability. If your 8x H100 tensor-parallel cluster is starting to hit interconnect limits at longer context windows, the AMD Helios rack-scale MI455X guide covers a UALink-based alternative built specifically for that kind of multi-node scaling.
Pricing fluctuates based on GPU availability. Spheron rates above are live as of 02 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.
Deploying with vLLM: Configuration for Hybrid Architectures
vLLM is the fastest path to a working Nemotron 3 Super endpoint. TensorRT-LLM gives higher throughput for sustained production but requires NVIDIA's TRT-LLM repo with the nemotron branch and significantly more setup time. See vLLM vs TensorRT-LLM vs SGLang benchmarks for a framework comparison.
Prerequisites
- CUDA 12.4+ for Hopper (H100); CUDA 12.8+ for Blackwell (B200, B300)
- vLLM 0.17.1+ (includes native Mamba kernel support; no separate
causal-conv1dormamba-ssmpackages needed) - Python 3.10+
Earlier vLLM versions do not include the SSM hybrid kernel required for Nemotron 3 Super. If your existing vLLM installation is older, --trust-remote-code alone will not fix the missing kernel; you need to upgrade to 0.17.1 or later.
Installation
pip install "vllm>=0.17.1"
nvcc --version # Verify CUDA 12.4+ (Hopper/H100) or 12.8+ (Blackwell/B200, B300)Download the Model
pip install huggingface-hub
huggingface-cli download nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 \
--local-dir /models/nemotron-superFor the NVFP4 quantized checkpoint:
huggingface-cli download nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 \
--local-dir /models/nemotron-super-nvfp4For FP8:
huggingface-cli download nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 \
--local-dir /models/nemotron-super-fp8Single B200: NVFP4 Launch Command
vLLM picks up the NVFP4 quantization settings from the checkpoint config, so no quantization flag is needed. This needs a Blackwell GPU: at 80.4 GB the checkpoint does not fit a single 80 GB H100.
python -m vllm.entrypoints.openai.api_server \
--model /models/nemotron-super-nvfp4 \
--gpu-memory-utilization 0.92 \
--max-model-len 32768 \
--trust-remote-code \
--no-enable-chunked-prefill \
--port 8000Hopper: FP8 on 2x H100 or 1x H200
The FP8 checkpoint is about 120 GB. Split it across two H100s, or set --tensor-parallel-size 1 on a single H200 (141 GB) and keep --max-model-len modest.
python -m vllm.entrypoints.openai.api_server \
--model /models/nemotron-super-fp8 \
--tensor-parallel-size 2 \
--gpu-memory-utilization 0.90 \
--max-model-len 32768 \
--trust-remote-code \
--no-enable-chunked-prefill \
--port 8000Multi-GPU BF16: 4x H100 with Tensor Parallelism
This configuration is for evaluation only (short context). With 4x H100 (320 GB total), the BF16 weights consume ~240 GB, leaving only ~80 GB across 4 cards for KV cache and activations. Using 128K context at this tier will OOM. Cap --max-model-len at 16384 for evaluation workloads, or use 8x H100 for production with larger context windows.
python -m vllm.entrypoints.openai.api_server \
--model /models/nemotron-super \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.90 \
--max-model-len 16384 \
--trust-remote-code \
--no-enable-chunked-prefill \
--port 8000Mamba-Specific vLLM Flags
| Flag | Value | Why |
|---|---|---|
--trust-remote-code | required | Nemotron uses custom model code |
--no-enable-chunked-prefill | required initially | SSM state initialization across chunk boundaries is a correctness issue; test without chunked prefill first (vLLM 0.17.1 may handle this correctly; validate before enabling) |
--max-num-seqs | 32-64 | SSM decode state limits effective batch parallelism vs pure attention |
--gpu-memory-utilization | 0.90-0.92 | Leave headroom for SSM activation buffers |
Quantization Options: NVFP4 vs FP8 vs GGUF Q4_K_M
NVFP4 (NVIDIA-Native, Blackwell-First)
NVFP4 is the recommended starting point for single-GPU deployment on B200 or B300:
- Native FP4 tensor core support on Blackwell (B200, B300). H100 and H200 (Hopper) have no FP4 tensor cores, so NVFP4 is not a native format there, and on an 80 GB H100 the checkpoint does not fit anyway.
- 0.5 bytes per parameter for the 4-bit weights; the full checkpoint is 80.4 GB with scale factors and higher-precision layers
- Requires CUDA 12.8+ (Blackwell) and vLLM's built-in NVFP4 kernel
- Built for Blackwell hardware; on Hopper, use the FP8 checkpoint instead
- Quality degradation is minor for coding tasks. The SWE-Bench gap vs BF16 is typically less than 2%.
In vLLM, the NVFP4 checkpoint's own config sets the quantization method, so you only point --model at it.
FP8 (Good Middle Ground)
FP8 gives better quality than NVFP4 at the cost of 2x the VRAM:
- 1 byte per parameter weight
- Supported natively on H100 (E4M3 format)
- Enable in vLLM:
--quantization fp8 - The Hopper path for this model: fits 2x H100 or 1x H200, and suits quality-sensitive production where full BF16 is overkill
GGUF Q4_K_M (CPU-Loadable, llama.cpp Path)
If you don't have access to an 80 GB GPU and need to run locally or on smaller hardware:
- Works with llama.cpp if a GGUF conversion exists for the checkpoint
- Can offload layers to CPU for low-QPS workloads
- Higher latency vs vLLM GPU path
- About 72 GB for the full model (roughly 0.6 bytes per parameter); fits a single H100 80GB only with a small context window
| Format | VRAM (120B model) | Throughput | Quality vs BF16 | Hardware |
|---|---|---|---|---|
| BF16 | ~240 GB | Baseline | 100% | 8x H100 (4x minimum for eval) |
| FP8 | ~120 GB | 1.5-1.8x BF16 | ~98% | 2x H100 or 1x H200 |
| NVFP4 | 80.4 GB | 2.5-3x BF16 | ~96% | 1x B200 or 1x B300 |
| GGUF Q4_K_M | ~72 GB | Slower (CPU path) | ~95% | 1x H100 (small context) or CPU |
One clarification on NVIDIA's advertised 5x throughput claim: that figure compares Nemotron 3 Super against the previous Nemotron Super model, reflecting both architectural improvements and efficiency gains across the model family. NVIDIA also publishes a separate 4x throughput claim comparing NVFP4 on Blackwell (B200) against FP8 on Hopper (H100). The 2.5-3x figure in the table above is our rough estimate for NVFP4 vs BF16 on Blackwell hardware, not a published NVIDIA number.
For deeper analysis of FP4 quantization economics, see FP4 quantization on Blackwell GPUs and KV cache optimization guide.
Benchmarks: SWE-Bench, Throughput, and What the 5x Claim Actually Means
SWE-Bench Verified Results
| Model | SWE-Bench Verified | Active Params | Notes |
|---|---|---|---|
| Nemotron 3 Super | 60.47% | 12B | Hybrid Mamba-MoE, released March 2026 |
| DeepSeek V4 | Not publicly benchmarked on SWE-Bench Verified as of Mar 2026 | ||
| GPT-5.4 | Not publicly benchmarked on SWE-Bench Verified as of Mar 2026 |
SWE-Bench scores vary depending on test harness and scaffold. The 60.47% figure is from NVIDIA's announcement using their evaluation setup. Your production results with the same model weight may differ based on prompt engineering and tool scaffolding.
For throughput comparison methodology across frameworks, see vLLM vs TensorRT-LLM vs SGLang benchmarks.
The 5x Throughput Claim
NVIDIA's 5x claim compares Nemotron 3 Super against the previous Nemotron Super model, reflecting architectural and efficiency gains across the model family. NVIDIA separately claims 4x throughput for NVFP4 on Blackwell vs FP8 on Hopper. Neither figure is a direct NVFP4 vs BF16 comparison on the same hardware.
A concrete estimate: at NVFP4 on a single B200, expect roughly 2,000-4,000 tokens/sec for decode depending on batch size. BF16 on 4x H100 runs approximately 800-1,500 tokens/sec. The single-B200 NVFP4 path needs one GPU instead of four, which usually wins on throughput per dollar at moderate batch sizes; compare the live hourly rates before you commit.
These are derived estimates. Actual throughput depends heavily on sequence lengths, batch sizes, and deployment configuration. Run benchmarks on your specific workload before committing to a hardware tier.
Agentic Coding Task Throughput
Coding agents produce longer outputs than chat: multi-file edits, test generation, and code review outputs often run 4,000-12,000 tokens per task. At 3,000 tokens/sec decode on a single B200 NVFP4, a 9,000-token agent response takes about 3 seconds. At 1,000 tokens/sec on BF16 4x H100, the same output takes 9 seconds.
For cost-per-task analysis, see the cost section below.
Cost to Run Nemotron 3 Super: Monthly Estimates for Enterprise Coding Workloads
| Configuration | $/hr (Spheron) | Best For |
|---|---|---|
| 1x B200 NVFP4 | $8.63/hr | Single-GPU production on Blackwell |
| 1x B200 spot | $4.49/hr | Interruptible batch or dev work (can be reclaimed) |
| 1x H200 FP8 | $4.80/hr | Single-GPU serving on Hopper, modest context |
| 2x H100 FP8 | $5.96/hr | Small team production |
| 4x H100 BF16 | $11.92/hr | Evaluation only (short context; use 8x H100 for production) |
| 8x H100 BF16 | $23.84/hr | High-throughput production |
For a monthly budget, multiply the hourly rate by about 730 hours of continuous serving.
Pricing fluctuates based on GPU availability. Spheron rates above are live as of 02 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.
Relevant hardware pages: H100 GPU rental, A100 GPU rental for teams currently on A100 planning a migration path.
For strategies to reduce GPU spend, see GPU cost optimization playbook.
Cost Per Coding Task
This is a worked example with its input rate frozen: it uses the $7.43/hr B200 on-demand rate we recorded on 02 Apr 2026, not today's price. Assume an agentic coding task produces 8,000 output tokens at 3,000 tokens/sec on 1x B200 NVFP4:
- Time per task: ~2.7 seconds
- Cost per task: $7.43/hr / 3,600 sec x 2.7 sec = $0.0056
- At 10,000 tasks/day: ~7.5 busy GPU-hours, about $56/day of compute, or ~$1,670/month. A B200 kept running around the clock bills all 24 hours.
Compare this to per-token API pricing for similarly capable proprietary models. At scale (10,000+ tasks/day), self-hosting on Spheron usually beats per-token API pricing. The breakeven point depends on the specific API you're comparing against, but for coding agents running at high volume, the math favors self-hosting well before you hit 5,000 tasks/day.
Nemotron 3 Super vs DeepSeek V4 vs GPT-5.4: Choosing the Right Coding Model
| Model | SWE-Bench | Total Params | Active Params | Context | Self-Hostable | Approx Cost per 1M tokens |
|---|---|---|---|---|---|---|
| Nemotron 3 Super | 60.47% | 120B | 12B | 1M | Yes | ~$0.69 (NVFP4, 1x B200, 02 Apr 2026 rate) |
| DeepSeek V4 | Not published | TBD | TBD | TBD | Yes | TBD |
| GPT-5.4 | Not published | Closed | Closed | TBD | No | API only |
When to Choose Nemotron 3 Super
- Your coding agents process long context windows (file-level or repo-level code review): SSM layers reduce KV cache pressure at long sequences, giving you more context per dollar
- You need NVIDIA ecosystem tooling: TensorRT-LLM, NIM microservices, NeMo framework
- You want to avoid third-party API dependencies for enterprise compliance or data residency
When DeepSeek V4 Makes Sense
- You already have a DeepSeek V3/V3.2 deployment and want to minimize migration cost
- See the DeepSeek V3.2 deployment guide for setup details
Production Decision Framework
| SWE-Bench threshold needed | VRAM budget | Compliance requirement | Recommendation |
|---|---|---|---|
| 60%+ | 141 GB single card | Data residency required | Nemotron 3 Super FP8 on H200 (modest context) |
| 60%+ | 192 GB single card | Data residency required | Nemotron 3 Super NVFP4 on B200 |
| 60%+ | 4x 80GB | API acceptable | Nemotron 3 Super BF16 on 4x H100 (evaluation only, short context; OOM risk at longer contexts, use 8x H100 for production) |
| Any | Minimal | API acceptable | Proprietary API (no infra overhead) |
| 55%+ | Existing DeepSeek cluster | Flexible | DeepSeek V3.2 (minimize migration) |
Production Checklist
- GPU provisioned with correct VRAM tier for chosen precision
- CUDA version verified (
nvcc --version): 12.4+ for Hopper (H100); 12.8+ for Blackwell (B200, B300) - vLLM 0.17.1+ installed (
pip install "vllm>=0.17.1") - Model checkpoint downloaded and checksum verified
- vLLM server launched with
--trust-remote-codeand correct--tensor-parallel-size - Chunked prefill disabled (
--no-enable-chunked-prefill); test and re-enable only after baseline benchmarks validate correctness - Health endpoint verified:
curl http://localhost:8000/health - GPU utilization monitored via
nvidia-smi dmon -s u - Load balancer or reverse proxy in front of vLLM for production traffic
- Pricing alert set up via Spheron dashboard to avoid unexpected cost overruns
For GPU monitoring tooling, see GPU monitoring for ML and production GPU cloud architecture.
Nemotron 3 Super's hybrid architecture makes it one of the most VRAM-efficient 120B-class coding models available. At NVFP4 it runs on a single B200, and FP8 fits two H100s, all on Spheron with no enterprise contracts.
On-demand H100 → | Check B200 availability → | View all GPU pricing →
Quick Setup Guide
Determine whether you need NVFP4 (80.4 GB checkpoint, runs on a single B200 or B300), FP8 (about 120 GB, fits on 2x H100 or 1x H200), or BF16 (requires 8x H100 for production, 4x H100 as an evaluation minimum). Use the VRAM budget tables in this guide to pick your tier before provisioning.
Log into app.spheron.ai, select your GPU tier from the catalog, and launch a bare-metal or containerized B200 instance for NVFP4, or an H100/H200 instance for FP8. Prefer bare-metal for MIG access or direct nvme storage. Note the public IP and SSH key.
On the GPU instance, install vLLM 0.17.1 or later via pip: pip install "vllm>=0.17.1". Verify CUDA version with nvcc --version: CUDA 12.4+ is required for Hopper (H100); CUDA 12.8+ is required for Blackwell (B200, B300). vLLM 0.17.1+ includes native Mamba kernel support. No additional packages are required.
Use huggingface-cli download nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 --local-dir /models/nemotron-super to pull the BF16 model (no --include filter; you need config.json, tokenizer files, and the shard index in addition to the weight shards). For the NVFP4 quantized version, pull nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 instead. For FP8, use nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8.
On a single B200, run: python -m vllm.entrypoints.openai.api_server --model /models/nemotron-super-nvfp4 --tensor-parallel-size 1 --gpu-memory-utilization 0.92 --trust-remote-code --no-enable-chunked-prefill --port 8000. vLLM reads the NVFP4 quantization settings from the checkpoint config. The NVFP4 checkpoint is 80.4 GB, which leaves roughly 100 GB on a B200 for KV cache, but does not fit a single 80 GB H100. On Hopper, point --model at /models/nemotron-super-fp8 with --tensor-parallel-size 2 on 2x H100, or --tensor-parallel-size 1 on one H200. The --no-enable-chunked-prefill flag disables chunked prefill for correctness: SSM layers cannot correctly initialize their recurrent state across chunk boundaries, producing wrong outputs if omitted. Note: NVIDIA recommends enabling chunked prefill on vLLM 0.17.1+, which includes SSM-aware chunk boundary handling. Start with --no-enable-chunked-prefill and re-enable only after you have validated correct outputs on your workload. For multi-GPU BF16, change --model to /models/nemotron-super and set --tensor-parallel-size 4 (vLLM manages inter-GPU coordination internally).
Send a test request via curl: curl http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"/models/nemotron-super-nvfp4","messages":[{"role":"user","content":"Write a Python function that finds the longest palindrome in a string."}]}'. Compare latency against a baseline model to verify the deployment is serving correctly.
Frequently Asked Questions
All 120B parameters must sit in VRAM even though only 12B are active per token. The NVFP4 checkpoint is 80.4 GB, so it does not fit a single 80 GB H100; run it on 1x B200 (192 GB HBM3e) or 1x B300, which also have native FP4 tensor cores. On Hopper, use FP8 (about 120 GB of weights) on 2x H100 or 1x H200 (141 GB). Full BF16 needs about 240 GB: 8x H100 for production, 4x H100 as a short-context evaluation floor. Spheron B200 runs $8.63/hr per GPU as of 02 Oct 2026.
Yes. vLLM 0.17.1+ includes Mamba kernel support for Nemotron 3 Super. You need to pass --model nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (or the NVFP4/FP8 variant) with --trust-remote-code and set tensor-parallel-size to match your GPU count. The hybrid SSM layers use different memory access patterns than pure attention, so chunked prefill should be disabled initially.
NVFP4 on a single B200 costs $8.63/hr on-demand or $4.49/hr on spot on Spheron. FP8 on 2x H100 costs $5.96/hr, and BF16 evaluation on 4x H100 costs $11.92/hr. Rates are as of 02 Oct 2026; multiply the hourly figure by about 730 for a month of continuous serving. Spot capacity can be reclaimed without notice.
Nemotron 3 Super mixes SSM (Mamba) layers with standard transformer attention layers in the same stack. SSM layers process sequences with O(sequence length) memory instead of O(sequence length^2), which means long-context inference costs far less VRAM than a pure transformer of similar capability. The trade-off is that SSM layers have sequential state that cannot be parallelized the same way as attention, so prefill batching behaves differently.
On SWE-Bench Verified, Nemotron 3 Super scores 60.47%. The hybrid MoE architecture keeps only 12B parameters active per token, which gives competitive throughput vs dense models. For pure coding tasks with long context windows (file-level edits, multi-file PRs), Nemotron 3 Super's SSM layers reduce KV cache pressure substantially.






