Tutorial

Ling 3.0 Flash VL GPU Requirements: 124B MoE Setup (2026)

Back to BlogWritten by Published Oct 1, 2026
Ling 3.0 Flash VL GPU RequirementsLing 3.0 Flash VLVision Language ModelMoE InferencevLLMGPU CloudHybrid AttentionMultimodal AILLM Deployment
Ling 3.0 Flash VL GPU Requirements: 124B MoE Setup (2026)

Ant Group's inclusionAI team pushed Ling-3.0-flash-VL's weights to Hugging Face on September 4, 2026, with an official announcement following five days later, extending the text-only Ling-3.0-flash into a vision-language model built for GUI and agentic tasks rather than static image captioning. It's a 124B-parameter Mixture-of-Experts model that activates only 5.5B parameters per token, runs a hybrid Kimi Delta Attention (KDA) and Gated MLA backbone, and ships with an official vLLM launch command straight from its own model card. The Ling 3.0 Flash VL GPU requirements swing hard depending on which of the four published checkpoints you load, from a four-GPU BF16 cluster down to a single card at INT4.

This guide works through the real VRAM math for BF16, FP8, FP4, and INT4, then deploys Ling-3.0-flash-VL with vLLM on rented GPU hardware from provisioning to a verified multimodal request.

TL;DR: Ling 3.0 Flash VL GPU Requirements

  • Architecture: 124B total parameters, 5.5B active per token, a 42-layer hybrid stack alternating Kimi Delta Attention and Gated MLA layers at a 5:1 ratio, per Ant's model card.
  • VRAM by precision: roughly 248GB at BF16, 124GB at FP8, and 62GB at FP4 or INT4, before KV cache and the vision encoder.
  • GPU fit: BF16 needs tp=4 on H200/B200/B300-class GPUs (tp=8 on H100); FP8 drops to tp=1 on B200/B300 or tp=2 on H200/H100; FP4 and INT4 run at tp=1 on any of them, per SGLang's cookbook recipe.
  • Pricing: Spheron's H200 SXM5 runs $4.79/hr on-demand as of 01 Oct 2026, the tier that fits the FP8 checkpoint at tp=2.
  • License: MIT, so there's no usage-restriction layer beyond what your GPU provider's own terms impose.

Compare live GPU configurations on Spheron's pricing page.

What's New in Ling 3.0 Flash VL: Hybrid KDA/MLA Attention, MTP, and the Visual Feedback Loop

Ling-3.0-flash-VL's headline change isn't image captioning, it's a visual feedback loop built for GUI and agentic tasks: the model observes the outcome of an action, compares it against the stated goal, identifies where the two diverge, and corrects course, rather than producing a one-shot description of what's in a frame, according to Pandaily's coverage of Ant Group's release. That loop is what separates "describe this screenshot" from "click the right button, notice it didn't work, and try again," which is closer to what GUI automation agents actually need.

Underneath that feedback loop sits the hybrid attention backbone that drives this guide's VRAM math. The model runs a 42-layer hybrid design alternating Kimi Delta Attention (KDA) and Gated MLA layers at a 5:1 ratio, per the model card. Work that ratio out across 42 layers and you land on roughly 35 KDA layers interleaved with 7 Gated MLA layers. If KDA sounds familiar, it should: Moonshot AI's Kimi K2.7/K3 lineage originated the mechanism, and our Kimi K3 deployment guide covers what KDA's fixed-size recurrent state means for KV cache planning in more depth than fits here. The short version: KDA layers hold a bounded recurrent state that doesn't grow with context length, while the Gated MLA layers still keep a compressed per-token cache. Mixing the two means your KV cache budget scales sub-linearly with context rather than linearly, which matters more at a 262,144-token window than it would at 8K.

The Ling lineage this model descends from also bakes in Multi-Token Prediction (MTP) as a core design choice rather than a bolt-on. Ant Group's own Ling-V2 architecture notes list it alongside the rest of the backbone: "Guided by Ling Scaling Laws, Ling 2.0 adopts a 1/32 activation ratio MoE architecture, with empirically optimized design choices in expert granularity, shared expert ratio, attention ratio, aux-loss free + sigmoid routing strategy, MTP loss, QK-Norm, half RoPE, and more," per inclusionAI's Ling-V2 repository. That MTP head isn't limited to training scaffolding either: the same repo documents it working as a speculative-decoding draft model at serving time, noting "MTP is supported for base model, and not yet for chat model. You can add parameter --speculative-algorithm NEXTN to start command." Ling-3.0-flash-VL's own card doesn't call MTP out by name, so treat it as an inherited family trait worth testing against your own workload rather than a documented day-0 feature of the VL checkpoint specifically.

Vision input runs through a separate path from that attention backbone: a ViT encoder feeds a two-layer MLP projector into the same transformer stack, and video frames get spatial and temporal position encoding through VideoRoPE rather than being flattened into plain sequence position, per the model card. That's architecturally adjacent to how Qwen 3.5's hybrid attention design handles its own GDN-based backbone, though Qwen 3.5's Gated DeltaNet replaces attention in 75% of layers rather than Ling's lighter 5:1 split.

Ling 3.0 Flash Vision Model Deploy: Where VL Fits in the Ling 3.0 Family

Deploying the Ling 3.0 Flash vision model means picking between two checkpoints released a few weeks apart. Ant Group pushed Ling-3.0-flash-VL's weights to Hugging Face on September 4, 2026, with the official announcement landing five days later on September 9, 2026, according to OrcaRouter's writeup. FP4 and INT4 quantized checkpoints followed as a second drop, alongside two weeks of free access through OpenRouter, per Ant Ling's own announcement on X. If you want to try the model before provisioning anything, Novita AI runs a free, time-limited hosted endpoint capped at a 262K input and 32K output token limit, per AlphaSignal's coverage, which lines up with the 262,144-token context length the model's own SGLang launch example sets.

The model ships under an MIT license, so there's no extra usage-restriction layer on top of whatever your GPU provider's own terms impose.

On the Artificial Analysis Intelligence Index v4.1.1, Ling-3.0-flash-VL scores 42, a 4-point improvement over the text-only Ling-3.0-flash, according to the model card. That's a general-intelligence benchmark, not a vision-specific one, so the gain reflects the model overall rather than image understanding in isolation. We walk through what that score difference means for the deploy-or-skip decision in the comparison section near the end of this guide.

Ling 3.0 Flash VL GPU Requirements: VRAM Across BF16, FP8, FP4, and INT4

Ling-3.0-flash-VL's 124B total parameters set four very different VRAM floors depending on which checkpoint you load: roughly 248GB at BF16, 124GB at FP8, and about 62GB at either FP4 or INT4, worked out directly from parameter count times bytes-per-parameter. None of those figures include KV cache, the vision encoder, or runtime overhead, so treat them as a floor rather than a provisioning target.

SGLang's own cookbook recipe for this model maps tensor-parallel degree to GPU class for each checkpoint, and it's the clearest published guide to how Ant Group expects the model to be served:

PrecisionApprox. weight size (124B params)SGLang TP degreeGPU classes in the recipe
BF16~248 GBtp=4 on 288GB/141GB-class GPUs, tp=8 on H100GB300, B300, B200, H200, H100
FP8~124 GBtp=1 on GB300/B300/B200, tp=2 on H200/H100GB300, B300, B200, H200, H100
FP4~62 GBtp=1 on any of the aboveGB300, B300, B200, H200, H100
INT4~62 GBtp=1 on any of the aboveGB300, B300, B200, H200, H100

That table comes from SGLang's Ling-3.0-flash-VL cookbook page. Two things stand out. First, BF16's tp=4 on H200 or B200-class hardware allocates far more VRAM than the 248GB weight floor needs: 4 GPUs at H200's 141GB each is 564GB. That headroom goes to KV cache for the 262K context window and the vision encoder's activations, not just weight storage. Second, the recipe's own GPU list leads with H20, Nvidia's China-market Hopper variant, for lowest-latency serving. Spheron doesn't carry H20, so substitute H200 or a Blackwell card and don't expect the recipe's exact latency figures to transfer directly.

Mapped onto what's actually rentable: Spheron's catalog carries H100 GPU rental (80GB), H200 GPU rental (141GB), B200 GPU rental (192GB), and B300 GPU rental (288GB), with GB300 listed on the pricing page but not yet available to book, per Spheron's pricing page. A BF16 deployment maps to 4x B300 or 4x H200; on-demand, Spheron prices B300 at $10.57/hr and H200 at $4.79/hr per GPU as of 01 Oct 2026. FP8 fits on a single B200 ($8.47/hr on-demand) or 2x H200. FP4 and INT4 both fit on a single H200, B200, B300, or H100 ($2.64/hr on-demand) per the SGLang table above, once you account for the vision encoder's extra memory: the ViT encoder and its two-layer MLP projector typically add a few GB on top of language-model weights in VLM designs this size, not a step change in the overall footprint.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 01 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

Deploying Ling-3.0-flash-VL with vLLM on Spheron GPU Cloud

Spheron deploys GPU instances in under 2 minutes with per-minute billing after a 20-minute minimum, per Spheron's pricing page, which removes most of the provisioning lag that would otherwise compete with the actual setup time below. The walkthrough targets the FP8 checkpoint at tp=2 on 2x H200, the middle tier between the BF16 cluster and the single-GPU quantized options.

Install vLLM and Pull the Weights

Provision your instance first, matching the tier from the table above, then install vLLM and pull the checkpoint:

bash
pip install -U vllm
python -c "import vllm; print(vllm.__version__)"

huggingface-cli download inclusionAI/Ling-3.0-flash-VL \
  --local-dir /data/models/ling-3.0-flash-vl

Verify the exact repository name on Hugging Face before downloading, since quantized variants typically ship under separate repo names rather than a flag on the BF16 repo. One thing to flag honestly here: whether this model needs --trust-remote-code or a specific parser build is a property of the model release, not something a GPU provider controls. Ling-3.0-flash-VL's own card requires both --trust-remote-code and the ling3-specific tool-call and reasoning parsers below, so a generic vLLM install without those flags won't recognize the architecture.

Launch the Server (BF16 and Quantized Variants)

The model card's own vLLM example, built for the BF16 checkpoint at tp=4, is:

bash
vllm serve "$MODEL_PATH" \
  --trust-remote-code \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.85 \
  --enable-prefix-caching \
  --tool-call-parser ling3 \
  --reasoning-parser ling3

That's sourced directly from the model card. For the FP8 checkpoint on 2x H200, drop --tensor-parallel-size to 2 and point $MODEL_PATH at the FP8 repo, following the same per-checkpoint tensor-parallel pattern SGLang's cookbook documents for this model:

bash
vllm serve /data/models/ling-3.0-flash-vl-fp8 \
  --trust-remote-code \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.85 \
  --enable-prefix-caching \
  --tool-call-parser ling3 \
  --reasoning-parser ling3

For FP4 or INT4 on a single GPU, drop --tensor-parallel-size entirely and point at the quantized repo. Keep --enable-prefix-caching on regardless of precision; with a 262K context window, repeated system prompts and multi-turn GUI sessions benefit from it more than most text-only deployments.

Test a Multimodal Request

Once the server reports healthy, send a test request through the OpenAI-compatible endpoint with an image attached:

bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ling-3.0-flash-vl",
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "text", "text": "What action should I take next in this screenshot?"},
          {"type": "image_url", "image_url": {"url": "https://example.com/screenshot.png"}}
        ]
      }
    ],
    "max_tokens": 256
  }'

Confirm the response references specific elements in the image rather than a generic description, and watch GPU memory with nvidia-smi dmon -s mu while the first few requests land, since image tokens inflate KV cache usage in a way plain text prompts don't.

Vision-Input Handling: What Changes in Serving Config for Multimodal Requests

Serving Ling-3.0-flash-VL's vision input changes two things from a text-only Ling-3.0-flash deployment: how you budget context length, and how you think about concurrency.

On context length, every image or video frame you pass in consumes visual tokens before any text does, the same dynamic that affects any vLLM-served VLM. VideoRoPE's spatial and temporal encoding means video frames carry more positional information than a flat image, so a short video clip can consume a meaningful fraction of your --max-model-len budget on its own. Size --max-model-len around your actual multimodal workload rather than defaulting to the full 262,144-token window, the same guidance our vision language model deployment guide covers for Qwen3-VL, Llama 4 Scout, and InternVL3.

On concurrency, the visual feedback loop this model is built for assumes multiple image turns per session, observe, compare, correct, rather than one image per request. If you're sizing --max-num-seqs for a GUI-automation workload, plan around sessions that each carry several screenshots across a conversation, not single-image requests, or your KV cache budget will run out faster than the per-request token count suggests. The hybrid KDA/MLA split from earlier in this guide helps here: KDA layers' bounded recurrent state means that growth is sub-linear compared to a model running full attention everywhere, but it isn't free, and the Gated MLA layers' compressed cache still grows with every additional image turn in a session.

Ling-3.0-flash-VL vs Ling-3.0-flash: Do You Need the Vision Checkpoint?

Only if your workload actually touches images, video, or GUI automation. Ling-3.0-flash-VL's 4-point Artificial Analysis Intelligence Index gain over the text-only Ling-3.0-flash reflects the model overall, not vision understanding specifically, and it comes paired with the ViT encoder's extra memory and compute on every request, image or not.

The text-only Ling-3.0-flash's own model card highlights a different benchmark set entirely: SWE-Bench Pro 56.6, AIME 2026 93.2, HMMT Feb 2026 87, and SWE-Bench Multilingual 72.4. Those are strong code and math reasoning scores on their own, and if your workload is pure text, coding agents, or math reasoning that never touches an image, the text-only checkpoint skips the vision encoder's overhead entirely while giving up nothing on the tasks it's actually scored against.

The deciding factor comes down to active-parameter economics as much as benchmark scores. Both checkpoints activate a small fraction of their total parameters per token, which is exactly the kind of MoE cost structure our active-parameters-vs-GPU-bill guide breaks down: you pay for VRAM on total parameters regardless of which tokens get routed where, so adding the VL variant's vision path to a workload that never uses it is pure overhead with no corresponding compute saving. If you need the full stack, text, audio, image, and video in one model, Qwen3.5-Omni's deployment guide covers a comparable 30B MoE design that goes further than Ling-3.0-flash-VL into audio input and speech output, at a different hardware tier. For GUI automation and agentic vision work specifically, Ling-3.0-flash-VL's feedback loop is the more targeted tool of the two.

Ling-3.0-flash-VL's checkpoint spread, from a 4-GPU BF16 cluster down to a single INT4 card, means the right rental size depends entirely on which precision you pick, and Spheron's per-minute billing after a 20-minute minimum means testing more than one tier doesn't cost you a day's commitment.

Get started on Spheron →

FAQ / 05

Frequently Asked Questions

Ling-3.0-flash-VL is Ant Group's vision-language extension of its Ling-3.0-flash model: a 124B-parameter Mixture-of-Experts model that activates 5.5B parameters per token, with a 42-layer hybrid backbone alternating Kimi Delta Attention (KDA) and Gated MLA layers at a 5:1 ratio. A ViT encoder plus a two-layer MLP projector handle image input, and video frames use VideoRoPE for spatial and temporal position encoding. Ant Group published the weights on Hugging Face on September 4, 2026 under an MIT license.

The model splits its 42 transformer layers between two attention types at a 5:1 ratio, roughly 35 Kimi Delta Attention (KDA) layers for bounded, linear-cost attention and 7 Gated MLA layers for standard quadratic attention with a compressed KV cache. Vision input goes through a separate ViT encoder and two-layer MLP projector before joining the same token stream the language model processes. Its headline feature is a visual feedback loop built for GUI and agentic tasks: the model observes an action's outcome, compares it to the goal, and corrects course, rather than just describing an image once.

Yes. The model card ships an official vLLM launch command directly: vllm serve "$MODEL_PATH" --trust-remote-code --tensor-parallel-size 4 --gpu-memory-utilization 0.85 --enable-prefix-caching --tool-call-parser ling3 --reasoning-parser ling3. The --trust-remote-code flag and the ling3-specific tool-call and reasoning parsers are both required; a generic vLLM install without them won't recognize the architecture. SGLang also has day-0 support through a dedicated Docker image.

It depends on checkpoint precision. BF16 weights run roughly 248GB, needing tensor-parallel-size 4 on H200 or B200-class GPUs (tp=8 on H100). FP8 drops to about 124GB, fitting tp=1 on a single B200 or B300, or tp=2 on H200 or H100. FP4 and INT4 both land around 62GB and run at tp=1 on any of those GPU classes. On Spheron, that maps to 4x H200 or 4x B300 for BF16, a single B200 or B300 for FP8, and a single H100, H200, B200, or B300 for FP4 or INT4. Spheron's H200 SXM5 runs $4.79/hr on-demand as of 01 Oct 2026; check current rates before provisioning since marketplace pricing moves.

Only if your workload actually touches images, video, or GUI automation. Ling-3.0-flash-VL scores 42 on the Artificial Analysis Intelligence Index v4.1.1, a 4-point improvement over the text-only Ling-3.0-flash, but that gain reflects the model overall, not image understanding specifically, and it comes with the vision encoder's extra memory and compute. If your workload is pure text, code, or math, the text-only Ling-3.0-flash's own benchmarks (SWE-Bench Pro 56.6, AIME 2026 93.2, HMMT Feb 2026 87, SWE-Bench Multilingual 72.4) are strong on their own and skip that overhead.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min