Tutorial

Qwen 3.6 Plus Is API-Only: Self-Host Qwen3.6 Open Models Instead

Back to BlogWritten by Published Apr 6, 2026Updated
Qwen 3.6 PlusQwen 3.6MoELinear AttentionvLLMGPU CloudLLM DeploymentOpen Source AIQwen 3Chain-of-Thought
Qwen 3.6 Plus Is API-Only: Self-Host Qwen3.6 Open Models Instead

Qwen 3.6 Plus cannot be self-hosted. Alibaba never released its weights: it is served only through Alibaba's own API (Qwen Studio and Model Studio) and resold through aggregators like OpenRouter, so there is no model to download, quantize, or load into vLLM. If you landed here wanting to run a Qwen 3.6 model on your own GPUs, the models you can actually deploy are the open-weight releases from the same generation, Qwen3.6-35B-A3B and Qwen3.6-27B. This guide covers what Plus is, then gives real VRAM numbers, GPU picks, and vLLM commands for the two open models.

TL;DR: Can You Deploy Qwen 3.6 Plus on GPU Cloud?

ModelWeightsFP8 sizeSingle-GPU fit (FP8)Spheron rate
Qwen 3.6 PlusClosed, API onlyn/aCannot self-hostn/a
Qwen3.6-35B-A3BOpen, HF download~35 GBL40S 48GB or H100H100 from $2.20/hr
Qwen3.6-27BOpen, HF download~27 GBL40S 48GB or H100L40S from $0.96/hr
Qwen3.6-27BINT4 checkpoint~16 GBRTX 4090 24GBRTX 4090 from $0.72/hr
  • Spheron rates above are as of 15 May 2026. Compare every tier on the GPU pricing page.

What Qwen 3.6 Plus Is, and Why There Is No GPU Config for It

Qwen 3.6 Plus is Alibaba's hosted model in the Qwen 3.6 generation. What Alibaba has confirmed: a 1M-token context window, native multimodal input, and always-on reasoning, meaning it produces a thinking trace before every answer. What Alibaba has not published: a parameter count, an architecture breakdown, or a license for self-hosting, because the model is not distributed.

An earlier version of this post estimated Plus variant sizes and VRAM needs based on prior Qwen releases. Those estimates had nothing to size against and have been removed. There is no Qwen 3.6 Plus repository on Hugging Face, so questions about VRAM, quantization, or tensor parallelism for Plus do not have answers.

The same pattern carried into the next generation: the entire Qwen3.7 line is API-only, Flash, Plus, and Max included. Qwen 3.6 still has an open side, which is where the rest of this guide goes.

When to use the Qwen 3.6 Plus API instead

  • You need the 1M-token context. The open 3.6 models default to 262K for 35B-A3B.
  • Your traffic is bursty or low volume, where paying per token beats keeping a GPU warm.
  • You want Alibaba's hosted behavior exactly, not an approximation.

Self-host an open 3.6 model when you need data residency, predictable cost at sustained load, or freedom from third-party rate limits. Check Alibaba's current per-token rates in Model Studio before comparing, since they change.

The Open Qwen 3.6 Models You Can Self-Host

Qwen3.6-35B-A3B is a multimodal MoE: 35B total parameters, 3B active per token, text, image, and video input, and a 262K-token default context. It mixes Gated DeltaNet and Gated Attention layers, which keeps KV cache growth lower than a pure attention stack at long context. MoE sparsity cuts compute per token, not memory: all 35B parameters still have to sit in VRAM.

Qwen3.6-27B is the dense sibling. Every parameter is active on every token, so per-token compute is higher than 35B-A3B, but memory is smaller and behavior is simpler to reason about when you tune batch size. Check the model card for its context length and thinking-mode controls before you set --max-model-len.

For the previous generation's open models, see the Qwen 3.5 deployment guide, and for the original Qwen 3 family, the Qwen 3 GPU deployment guide.

VRAM Requirements for Qwen3.6-35B-A3B and Qwen3.6-27B

Weight sizes below come from real checkpoint files. INT4 files run about 0.6 bytes per parameter, not 0.5, because quantization scales and a few unquantized layers add overhead.

ModelBF16 weightsFP8 weightsINT4 weights
Qwen3.6-35B-A3B~70 GB~35 GB~21 GB
Qwen3.6-27B~54 GB~27 GB~16 GB

Add KV cache and a few GB of framework overhead on top. With --gpu-memory-utilization 0.9, vLLM caps itself at 90% of the card, so a 48 GB L40S gives you about 43 GB and an 80 GB H100 about 72 GB.

GPU picks by model and precision

Model and precisionGPUHeadroom for KV cache
35B-A3B FP81x L40S 48GB~8 GB, keep context around 32K
35B-A3B FP81x H100 80GB~37 GB, room for long context
35B-A3B BF161x H100 80GB~2 GB, tight, short context only
35B-A3B BF161x H200 141GB~57 GB, the comfortable BF16 option
35B-A3B INT41x RTX 5090 32GB~8 GB, moderate context
27B FP81x L40S 48GB or 1x H100~16 GB on L40S, ~45 GB on H100
27B INT41x RTX 4090 or RTX 5090~5 GB on 4090, ~13 GB on 5090

FP8 on an L40S is the cheapest way to run either model at full quality for most workloads. Step up to a Spheron H100 when you need longer context or higher concurrency, and to an H200 on Spheron if you want BF16 35B-A3B without squeezing KV cache. For single-developer setups, the 27B INT4 checkpoint on an RTX 4090 rental is the lowest-cost path.

What won't work

  • 35B-A3B at FP8 or BF16 on any 24 GB or 32 GB card. 35 GB of weights does not fit.
  • 27B at FP8 on an RTX 4090. 27 GB of weights exceeds 24 GB.
  • FP8 on A100. A100 has no FP8 Tensor Cores. Use --quantization bitsandbytes for INT8 there, or an INT4 checkpoint.
  • Sizing by active parameters. 3B active does not mean 3B worth of VRAM. Plan for all 35B.

For the general method behind these numbers, see the MoE inference optimization guide and the KV cache optimization guide.

Step-by-Step: Deploy Qwen 3.6 Open Models with vLLM on Spheron

Prerequisites

Provision a GPU instance on Spheron that matches the table above. The Spheron quick-start guide covers provisioning, and the vLLM server guide covers configuration.

bash
nvidia-smi
# Verify GPU count, VRAM, and driver version

Install vLLM

bash
pip install "vllm>=0.19.0"
python -c "import vllm; print(vllm.__version__)"

The hybrid Gated DeltaNet and Gated Attention layers need a recent release. If the model class is not recognized, upgrade vLLM before reaching for --trust-remote-code.

Download the weights

bash
# MoE: 35B total, 3B active
huggingface-cli download Qwen/Qwen3.6-35B-A3B \
    --local-dir /data/models/qwen3.6-35b-a3b

# Dense: 27B
huggingface-cli download Qwen/Qwen3.6-27B \
    --local-dir /data/models/qwen3.6-27b

Use persistent storage so an instance restart does not mean a 50-70 GB re-download. For INT4, download a 4-bit (AWQ or GPTQ) checkpoint of the model from Hugging Face instead; vLLM reads the quantization config from the checkpoint, so you do not pass --quantization for it.

Launch the inference server

bash
# Qwen3.6-35B-A3B, FP8 on 1x H100 80GB
vllm serve /data/models/qwen3.6-35b-a3b \
    --served-model-name qwen3.6-35b-a3b \
    --quantization fp8 \
    --gpu-memory-utilization 0.9 \
    --max-model-len 65536 \
    --port 8000

# Qwen3.6-35B-A3B, FP8 on 1x L40S 48GB (shorter context)
vllm serve /data/models/qwen3.6-35b-a3b \
    --served-model-name qwen3.6-35b-a3b \
    --quantization fp8 \
    --gpu-memory-utilization 0.9 \
    --max-model-len 32768 \
    --port 8000

# Qwen3.6-35B-A3B, BF16 on 1x H200 141GB
vllm serve /data/models/qwen3.6-35b-a3b \
    --served-model-name qwen3.6-35b-a3b \
    --gpu-memory-utilization 0.9 \
    --max-model-len 131072 \
    --port 8000

# Qwen3.6-27B, FP8 on 1x L40S or H100
vllm serve /data/models/qwen3.6-27b \
    --served-model-name qwen3.6-27b \
    --quantization fp8 \
    --gpu-memory-utilization 0.9 \
    --max-model-len 32768 \
    --port 8000

--quantization fp8 quantizes the BF16 weights at load time and uses native FP8 Tensor Cores on L40S, H100, and H200. None of these configs need --tensor-parallel-size. If you want more throughput than one card delivers, running several single-GPU replicas behind a load balancer is usually simpler than tensor parallelism for a model this size. If you do split a model across GPUs, --tensor-parallel-size must evenly divide the GPU count.

Test the API

bash
curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{"model": "qwen3.6-35b-a3b", "messages": [{"role": "user", "content": "Explain tensor parallelism in one paragraph."}], "max_tokens": 1024}'

Python OpenAI client:

python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="qwen3.6-35b-a3b",
    messages=[{"role": "user", "content": "Write a binary search in Python."}],
    max_tokens=2048  # leave room for a reasoning trace if thinking is on
)
print(response.choices[0].message.content)

If thinking mode is on, reasoning tokens count toward max_tokens, so a short answer can still use 1,500-3,000 tokens. Set the budget with that in mind, and see the reasoning model inference cost guide for ways to cap it. Watch throughput during testing with nvidia-smi dmon -s pum -d 5 in a second terminal.

Spheron Pricing for Qwen 3.6 Open Model Workloads

SetupGPUOn-demand / hrSpot / hr
27B INT4, dev and testing1x RTX 4090 24GB$0.72/hrcheck live rates
35B-A3B INT4 or 27B INT41x RTX 5090 32GB$0.78/hr$0.89/hr
35B-A3B or 27B FP8, short context1x L40S 48GB$0.96/hr$0.98/hr
35B-A3B or 27B FP8, long context1x H100 80GB$2.65/hr$2.20/hr
35B-A3B BF161x H200 141GB$4.80/hr$3.31/hr

Spot capacity can be reclaimed without notice, so keep production traffic on on-demand instances unless your serving layer handles restarts.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 15 May 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

Troubleshooting

  • Download fails for a "Qwen3.6-Plus" repo: there is no such repository. Use Qwen/Qwen3.6-35B-A3B or Qwen/Qwen3.6-27B.
  • OOM at startup: lower --max-model-len first. vLLM reserves KV cache for the full context length, so halving it frees real memory. On an L40S with 35B-A3B FP8, stay at or below 32K.
  • Model class not found: upgrade vLLM to 0.19.0 or later. Older releases do not know the Qwen 3.6 hybrid layer layout.
  • FP8 errors on A100: A100 has no FP8 support. Switch to --quantization bitsandbytes or an INT4 checkpoint.
  • Truncated answers: raise max_tokens per request. With thinking on, the reasoning trace uses part of the budget before the answer starts.

For agent workloads built on these models, the GPU infrastructure for AI agents guide covers latency and concurrency planning. For the next Qwen generation, the Qwen3.7 Max guide covers a generation that ships API-only as well.


Qwen 3.6 Plus stays behind Alibaba's API, but Qwen3.6-35B-A3B and Qwen3.6-27B run on a single GPU, from an L40S at FP8 to an H200 at full BF16.

Check L40S availability → | H100 GPU pricing → | View all pricing →

Get started on Spheron →

STEPS / 06

Quick Setup Guide

  1. Decide whether you need Qwen 3.6 Plus itself

    Qwen 3.6 Plus is closed-weight. Alibaba serves it only through its own API (Qwen Studio and Model Studio) and through resellers like OpenRouter. If you need Plus specifically, use the API. If you need a Qwen 3.6 model on your own hardware, pick one of the open-weight releases: Qwen3.6-35B-A3B (MoE, 3B active) or Qwen3.6-27B (dense).

  2. Choose a model, precision, and GPU

    Qwen3.6-35B-A3B weighs about 70 GB in BF16, 35 GB in FP8, and 21 GB in INT4. Qwen3.6-27B weighs about 54 GB in BF16, 27 GB in FP8, and 16 GB in INT4. FP8 on a single L40S 48GB or H100 80GB covers both models. BF16 35B-A3B wants an H200. INT4 27B fits an RTX 4090 or RTX 5090.

  3. Provision a GPU instance on Spheron

    Go to app.spheron.ai and provision the GPU matching your precision choice. SSH in and run nvidia-smi to confirm GPU count, VRAM, and driver version.

  4. Install vLLM

    Run pip install 'vllm>=0.19.0' and verify with python -c "import vllm; print(vllm.__version__)". The Qwen 3.6 open models use a hybrid Gated DeltaNet and Gated Attention layout that older vLLM releases do not recognize.

  5. Download the open weights from Hugging Face

    Run huggingface-cli download Qwen/Qwen3.6-35B-A3B --local-dir /data/models/qwen3.6-35b-a3b, or Qwen/Qwen3.6-27B for the dense model. Store weights on persistent storage so a restart does not trigger a re-download. There is no Qwen 3.6 Plus repository to download.

  6. Launch and test the vLLM server

    Run vllm serve /data/models/qwen3.6-35b-a3b --quantization fp8 --gpu-memory-utilization 0.9 --max-model-len 32768 --port 8000, then send a test request to http://localhost:8000/v1/chat/completions. Raise --max-model-len only after checking KV cache headroom with nvidia-smi.

FAQ / 05

Frequently Asked Questions

No. Qwen 3.6 Plus is closed-weight. There is no Hugging Face repository, no GGUF, and no AWQ checkpoint for it, so there is nothing for vLLM or SGLang to load. Access is through Alibaba's API (Qwen Studio and Model Studio) or resellers such as OpenRouter. The self-hostable models in the same generation are Qwen3.6-35B-A3B and Qwen3.6-27B.

Alibaba has not published a parameter count for Qwen 3.6 Plus. Any figure you see quoted for it, or any VRAM estimate derived from one, is a guess. Since the weights are not downloadable, the number would not change your hardware plan anyway.

Qwen3.6-35B-A3B is the usual pick. It is a multimodal MoE with 35B total parameters and 3B active per token, a 262K-token default context, and it accepts text, image, and video input. Qwen3.6-27B is the dense option if you prefer predictable per-token compute over MoE throughput. Neither is a drop-in match for Plus, which offers a 1M-token context through the API.

At FP8 the weights take about 35 GB, so a single L40S 48GB (short context) or H100 80GB (room for longer context) works. At BF16 the weights take about 70 GB, which is tight on an H100 80GB and comfortable on an H200 141GB. An INT4 checkpoint is about 21 GB and fits an RTX 5090 32GB. All 35B parameters must sit in VRAM even though only 3B are active per token.

Yes, at INT4. A 4-bit Qwen3.6-27B checkpoint is about 16 GB, which leaves several GB of the RTX 4090's 24 GB for KV cache at moderate context lengths. FP8 (about 27 GB) and BF16 (about 54 GB) do not fit a 24 GB card.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min