Comparison

DGX Spark vs RTX 5090 for Local LLMs: Tokens Per Second (2026)

Back to BlogWritten by Published Oct 5, 2026
DGX Spark vs RTX 5090DGX SparkRTX 5090GPU BenchmarkLLM InferenceGPU ComparisonBlackwell
DGX Spark vs RTX 5090 for Local LLMs: Tokens Per Second (2026)

The number everyone quotes in the DGX Spark vs RTX 5090 debate is 205 versus roughly 50 tokens per second, the RTX 5090 beating DGX Spark on a single chat session. That number is real, and it's also only half the story. It's a single-stream result: one prompt, one response, nothing else happening on the GPU. Run more than one request at a time, or try to load a model that doesn't fit in 32GB, and the comparison inverts in ways none of the forum threads citing that number mention.

This isn't a gaming benchmark. The RTX 5090 is also NVIDIA's flagship gaming card, but every number below is LLM inference throughput, measured in tokens per second, on real model weights.

TL;DR: DGX Spark vs RTX 5090 for Local LLMs

  • Single-stream decode: On Llama 3.1 8B at Q4_K_M, RTX 5090 hits 150 tok/s against DGX Spark's 38 tok/s with Ollama.
  • Memory: DGX Spark has 128GB unified LPDDR5x at 273 GB/s; RTX 5090 has 32GB GDDR7 at 1,792 GB/s.
  • What fits: RTX 5090 tops out around 32B-class models under quantization; DGX Spark loads 100B+ MoE models like gpt-oss-120b.
  • Concurrency: vLLM serving 8 concurrent requests on a single DGX Spark reaches an aggregate 289-313 tok/s.
  • Price: DGX Spark started at $3,000; RTX 5090 launched at $1,999 MSRP. Spheron rents RTX 5090 on-demand at $0.78/hr as of 06 Oct 2026. Rent an RTX 5090 by the hour.

Spec and Price Snapshot: 128GB Unified Memory vs 32GB GDDR7

DGX Spark and RTX 5090 solve different problems with memory. DGX Spark trades bandwidth for capacity: one unified pool the CPU and GPU share, large enough to hold models no single consumer card can load. RTX 5090 trades capacity for bandwidth: a smaller, much faster pool built for one task at a time.

SpecNVIDIA DGX SparkRTX 5090
Memory128GB unified LPDDR5x (64GB config also exists via OEM partners)32GB GDDR7
Memory bandwidth273 GB/s1,792 GB/s
ComputeUp to 1 petaFLOP FP4 (with sparsity)3,352 AI TOPS FP4 (sparse)
ArchitectureGB10 Grace Blackwell SuperchipGB202 (Blackwell)
Form factorDesktop AI computer, standalonePCIe add-in card, needs a host system
Networking10GbE + ConnectX-7 200Gb/sNone (PCIe only, no NVLink)
Starting price$3,000 (NVIDIA's announced starting price)$1,999 MSRP
PowerDesktop-class draw, no separate PSU sizing needed575W TDP, needs an 850W+ PSU

The DGX Spark's $3,000 figure comes from NVIDIA's original announcement of the platform (then called Project DIGITS), which the company said would be "available in May from NVIDIA and top partners, starting at $3,000." NVIDIA's current DGX Spark product page doesn't post a fixed retail price; it lists the 128GB configuration as the standard offering, with a 64GB option available only through OEM partners, and directs buyers to retailers and partner stores instead. Actual street price depends on which OEM and configuration you pick. The RTX 5090 launched in January 2025 at a $1,999 MSRP.

For the full RTX 5090 datasheet, including CUDA core counts and Tensor Core generation, see our NVIDIA RTX 5090 specs breakdown.

DGX Spark, RTX 5090, and the OEM Clones at a Glance

DGX Spark isn't sold only as an NVIDIA-branded box. NVIDIA's own product page lists its own GB10 Superchip-powered systems built by Acer, ASUS, Dell, Gigabyte, HP, Lenovo, and MSI, under names like the ASUS Ascent GX10, Dell Pro Max (NVIDIA AI Dev), and Gigabyte AI TOP ATOM. PNY sits in a different category on the same page: it's a retailer selling the NVIDIA-branded DGX Spark itself, alongside Amazon, Micro Center, and TD Synnex, not a company building its own GB10 box. Specs are effectively identical across the OEM-built systems, since they're all the same silicon and the same 128GB LPDDR5x pool; the differences show up in chassis design, port selection, included storage, and support terms, not in tokens per second. If you're shopping DGX Spark, treat "DGX Spark" as the chip generation and compare OEM listings on configuration and warranty rather than expecting performance differences between them.

The RTX 5090 has no equivalent OEM-clone market in this sense. It's a standard PCIe card sold by the usual add-in-board partners (ASUS, MSI, Gigabyte, and others) at prices that track GPU market demand rather than a fixed MSRP once retail stock is in play.

Tokens Per Second, Measured: Dense Models From 8B to 70B

Here's what LMSYS, the group behind the Chatbot Arena leaderboard, measured running the same models on both devices. Prefill is the compute-bound phase that generates your time-to-first-token; decode is the bandwidth-bound phase that generates the rest of the response, one token at a time, and is what most single-stream tok/s numbers report.

Model (quantization)DGX Spark prefillDGX Spark decodeRTX 5090 prefillRTX 5090 decode
GPT-OSS 20B (MXFP4)2,053 tok/s49.7 tok/s8,519 tok/s205 tok/s
Llama 3.1 8B (FP8, batch 1)7,991 tok/s20.5 tok/sn/a (SGLang DGX Spark benchmark, no 5090 pairing)n/a
Llama 3.1 70B (FP8)803 tok/s2.7 tok/sDoesn't fit in 32GBDoesn't fit in 32GB

All DGX Spark and GPT-OSS 20B figures above come from LMSYS's DGX Spark review, which ran both devices through the same GPT-OSS 20B MXFP4 workload with Ollama. The Llama 3.1 8B and 70B figures come from the same review's SGLang benchmarks, run on DGX Spark alone.

None of these numbers require you to take LMSYS's word for it. An RTX 5090 on Spheron rents at $0.78/hr on-demand, billed per minute with a 20-minute minimum, which is enough time to pull the same GPT-OSS 20B checkpoint and read your own decode number off the logs. The exact recipe, run against a freshly provisioned Spheron RTX 5090 instance:

bash
# pull the model
ollama pull gpt-oss:20b

# run with verbose timing so Ollama prints eval rate (decode tok/s)
# and prompt eval rate (prefill tok/s) after the response
ollama run gpt-oss:20b --verbose <<< "Explain the difference between prefill and decode in LLM inference."

Ollama's --verbose flag prints eval rate (decode) and prompt eval rate (prefill) directly after the generated response, in the same tok/s units LMSYS reports. Running that against a rented 5090 and comparing the printed eval rate to the table above is a closer check than reading the numbers on faith, and it costs one 20-minute minimum block to do it.

Small Dense Models (8B): Where the RTX 5090's Bandwidth Wins Outright

On an 8B model, there's no ambiguity. DeepResearch Ninja's Ollama benchmarks, run on Llama 3.1 8B at Q4_K_M quantization, measured the RTX 5090 at 150 tokens/sec against DGX Spark's 38 tokens/sec, a result consistent with the roughly 4x gap LMSYS found on GPT-OSS 20B. At this model size, both devices have plenty of headroom (8B at Q4 needs under 8GB), so the RTX 5090's 1,792 GB/s of GDDR7 bandwidth, about 6.6x DGX Spark's 273 GB/s, is simply faster at moving the smaller set of weights required for each decode step. If your workload is a single chat session on an 8B-13B model, the RTX 5090 wins on raw speed without qualification.

Mid-Size Dense Models (20B-27B): the Gap Narrows but Doesn't Close

GPT-OSS 20B is the cleanest mid-size data point LMSYS published, and the 4.1x decode gap (205 vs 49.7 tok/s) holds roughly in line with the 8B result rather than narrowing meaningfully. What does shift at this size is the margin on specialized hardware: LMSYS also benchmarked an RTX Pro 6000 Blackwell on the same GPT-OSS 20B workload and measured 215 tokens/sec decode, essentially matching the consumer RTX 5090 and roughly 4x DGX Spark's rate. The takeaway for the 20B-27B range is that DGX Spark's relative disadvantage doesn't meaningfully improve as dense models grow from 8B into the 20B-class; the bandwidth bottleneck scales with the model, not away from it.

70B Dense: RTX 5090 Can't Load It, DGX Spark Limps at Single Digits

This is where the comparison stops being about speed and starts being about whether the model runs at all. A 70B dense model needs roughly 70GB of memory at FP8 precision, more at BF16, which doesn't fit in the RTX 5090's 32GB under any quantization scheme that preserves usable quality. DGX Spark's 128GB pool can hold it: LMSYS measured Llama 3.1 70B in FP8 on DGX Spark via SGLang at 803 tokens/sec prefill and 2.7 tokens/sec decode. 2.7 tokens/sec is below conversational speed, closer to watching a response assemble word by word than chatting. DGX Spark technically runs 70B models that the RTX 5090 can't touch; "technically runs" and "usable" are different claims at that decode rate.

Prefill Speed and Time-to-First-Token: the Other Half of the Benchmark

Decode tokens per second gets all the attention because it's the number that determines how fast a response streams in. Prefill determines something different: how long you wait before the first token appears at all, which matters more for long prompts, RAG pipelines with retrieved context, or agent calls stuffing tool outputs back into the context window.

Prefill is compute-bound (it processes the whole prompt in parallel), while decode is bandwidth-bound (it generates one token at a time, re-reading the model weights each step). That's why DGX Spark's prefill numbers look much more competitive than its decode numbers: on GPT-OSS 20B it hits 2,053 tok/s prefill against the RTX 5090's 8,519 tok/s, a 4.1x gap that roughly tracks the decode gap. But on Llama 3.1 8B in FP8, DGX Spark's own SGLang benchmark shows 7,991 tokens/sec prefill, a rate that's respectable even without a direct RTX 5090 comparison point from the same run.

The practical read: DGX Spark's time-to-first-token on a short prompt is usually fine. Where it falls behind is the response that follows, token by token, at decode speed. A long RAG-heavy prompt with a short answer plays closer to DGX Spark's strengths than a short prompt with a long generated response.

What 128GB Unified Memory Lets You Load That 32GB Can't

Tokens per second only matters for models that fit. This is the section where DGX Spark's memory advantage stops being a trade-off and becomes the only option.

100B+ MoE Models (gpt-oss-120b): DGX Spark's Actual Home Turf

gpt-oss-120b is a mixture-of-experts model that, even at MXFP4 quantization, needs more memory than any single discrete consumer GPU has. It simply does not run on an RTX 5090; there's no quantization trick that gets a 120B-parameter MoE model into 32GB with usable quality. On DGX Spark, community benchmarks published in llama.cpp's GitHub discussions measured gpt-oss-120b (MXFP4) at roughly 60.6 tokens/sec decode at empty context, dropping to 40.5 tokens/sec at a 32K-token context as KV cache pressure builds. Those numbers aren't fast in absolute terms, but there's no RTX 5090 comparison to make, because the RTX 5090 can't enter the race. This is the workload class DGX Spark exists for: models too large for any single consumer card, running at a pace that's merely moderate instead of impossible.

Batching Changes the Math: DGX Spark at Concurrency vs RTX 5090 Single-Stream

Every number so far has been single-stream: one request, one response, nothing else running. That's the right test for "how fast does my one chat window feel," and it's the test the 205-vs-50 tok/s comparisons online are built on. It's the wrong test for anyone running more than one thing at once, which is most real local-LLM usage once you move past a single interactive chat.

LMSYS's SGLang benchmark shows DGX Spark's decode throughput on Llama 3.1 8B FP8 scaling from 20.5 tokens/sec at batch size 1 to 368 tokens/sec at batch size 32, a roughly 18x increase from batching alone, inside the same 128GB pool and without adding hardware. That's the number nobody citing the single-stream comparison mentions.

A separate benchmark published on NVIDIA's developer forums tested the same harness, Ollama, llama.cpp, and vLLM, on one DGX Spark across four different models. Serving 8 concurrent requests, vLLM's aggregate throughput reached 289-313 tokens/sec depending on the model. Ollama's 8-concurrent aggregate matched its single-stream number exactly, because Ollama processes requests one at a time and doesn't batch. The hardware was identical in both cases; the serving stack was the difference. If you're running DGX Spark with Ollama and comparing it to an RTX 5090's single-stream number, you're comparing DGX Spark's worst case to the RTX 5090's only case. Point DGX Spark at a batching-aware server like vLLM or SGLang instead, and the comparison changes shape entirely for anyone running concurrent agents, multiple users, or parallel evals.

The RTX 5090 can batch too, and its raw per-request speed means it still wins on latency at low concurrency. But at 32GB, its batching headroom runs out fast: KV cache for 32 concurrent 8B requests at reasonable context lengths can exceed what's left after model weights. DGX Spark's 128GB gives batching far more room to work before memory, not compute, becomes the ceiling.

Fine-Tuning and Multi-Model Workflows: Where Spark's Memory Wins

Inference isn't the only local workload. Benjamin Marie, who writes The Kaitchup newsletter on LLM fine-tuning, tested DGX Spark specifically for training workloads and concluded: "Where DGX Spark makes more sense, I think, is fine-tuning and small-scale training, especially for models under ~8B parameters."

His testing documents Unsloth-based fine-tuning on DGX Spark comfortably handling models in the roughly 8B-14B parameter range inside the 128GB unified pool, without the heavy CPU-offload penalties that the same model sizes incur on a 32GB card like the RTX 5090. Fine-tuning needs memory for the model weights, gradients, optimizer states, and activations simultaneously, a much larger footprint than inference alone. On a 32GB card, a 13B-14B fine-tuning run often has to offload layers to system RAM, which is slow enough to change the economics of the whole job. DGX Spark's extra headroom avoids that offload entirely for this size class.

The same memory advantage applies outside fine-tuning, to any workflow running more than one model at once: an embedding model plus a reranker plus a generation model, or multiple quantized variants loaded side by side for A/B comparison. On 32GB, that kind of stacking runs out of room fast. On 128GB, it's routine. For the quantization formats referenced throughout this comparison (MXFP4, Q4_K_M, AWQ), see our GPTQ vs AWQ vs GGUF guide for how the format you pick changes both the memory footprint and the throughput you'll actually see.

DGX Spark vs RTX 5090: Which One to Buy

Neither device wins outright. The right pick depends on whether your bottleneck is speed on one request or capacity across many.

Quick Answer Table by Use Case

Your situationBuy this
One interactive chat session, 8B-20B modelsRTX 5090: 3-4x faster decode, cheaper
Models over 32GB at your target quantizationDGX Spark: only option that loads them
Fine-tuning 8B-14B models without CPU offloadDGX Spark: 128GB avoids offload penalties
Running several concurrent agents or users against one boxDGX Spark with vLLM/SGLang: batching scales inside 128GB
Need the 205 tok/s single-stream speed without buying hardwareRent an RTX 5090 instead of buying one
Want to test both before committing $2,000-3,000+Rent and benchmark your own models first

The RTX 5090 is the better buy for most individual developers doing day-to-day inference on models up to the low 20B range: faster, cheaper, and it doubles as a capable workstation GPU for everything else. DGX Spark earns its price when the question stops being "how fast" and becomes "does it fit," or when you're running enough concurrent load that batching inside 128GB beats a faster GPU serving one request at a time. For a broader shortlist across both consumer and data-center GPUs, see our best NVIDIA GPUs for LLMs guide.

Buying either one is also optional. $0.78/hr gets you an RTX 5090 on Spheron with per-minute billing and a 20-minute minimum, which is enough to run the exact benchmarks in this post against your own models before deciding whether $1,999-3,000+ of hardware is worth it for your workload.

When to Graduate to Rented Cloud GPUs

Both devices hit a wall at roughly the same place: production serving with real concurrency, 24/7 uptime, or models that need multi-GPU tensor parallelism. Neither DGX Spark nor a single RTX 5090 has NVLink to another GPU, so once a workload needs more throughput than one box can batch its way to, the next step is a cloud instance, not a second desktop unit.

Spheron's companion guide to the DGX Spark local-to-cloud pipeline works through the full break-even math for when dev-time cloud GPU spend crosses over DGX Spark's purchase price, and the exact steps for moving a vLLM setup from local development to a production cloud instance. That math favors buying hardware for a single developer under the break-even point; it favors renting once you need uptime guarantees, multi-GPU scale, or more concurrent throughput than one desktop box can deliver.

Renting doesn't replace either device for its core use case. It doesn't give you offline, private, always-on desk-side inference with no meter running, which is the entire point of owning a DGX Spark or an RTX 5090 workstation in the first place. And Spheron's spot instances can be reclaimed without notice, which matters if you're running an unattended overnight fine-tuning job rather than an interactive session. For a single developer still under the break-even threshold the companion guide calculates, buying remains the cheaper option; renting is what makes sense for the production step after.

Once a model outgrows what DGX Spark or an RTX 5090 can serve at usable speed, the next step is a GPU with the bandwidth and NVLink neither desktop box has. Spheron rents H100 and H200 instances by the minute for exactly that handoff.

Compare live GPU pricing on Spheron →

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 06 Oct 2026; other providers and hardware MSRPs reflect their most recently published figures and may have changed. Check current GPU pricing → for live rates.

FAQ / 05

Frequently Asked Questions

It depends on what you're running. For a single interactive chat session on an 8B-20B model, the RTX 5090 is faster and cheaper: it decodes roughly 3-4x more tokens per second and starts around $1,999 MSRP versus DGX Spark's $3,000 NVIDIA-announced starting price (OEM configurations vary). DGX Spark is worth it when you need to load models the RTX 5090 physically can't, such as 100B+ MoE models like gpt-oss-120b, when you're fine-tuning models in the 8B-14B range without CPU offload, or when you run several models or concurrent agent sessions against the same 128GB pool rather than one chat window at a time.

On GPT-OSS 20B (MXFP4) with Ollama, DGX Spark decodes at 49.7 tokens/sec against the RTX 5090's 205 tokens/sec, a roughly 4.1x gap, according to LMSYS's benchmark. On a smaller Llama 3.1 8B model at Q4_K_M quantization, Ollama benchmarks from DeepResearch Ninja measured 150 tokens/sec on the RTX 5090 versus 38 tokens/sec on DGX Spark. Both numbers are single-stream, one request at a time. DGX Spark's decode throughput scales with concurrent requests inside its 128GB pool in a way the RTX 5090's single-stream number doesn't capture.

Not at usable quality. A 70B dense model needs roughly 70GB of memory even at FP8 precision, and significantly more at BF16, which exceeds the RTX 5090's 32GB of GDDR7 regardless of quantization scheme. DGX Spark can load Llama 3.1 70B in FP8 inside its 128GB pool, but LMSYS measured only 2.7 tokens/sec decode throughput there, too slow for interactive use. Neither device is a good fit for 70B-class dense models; that's the point at which both should hand off to a rented multi-GPU cloud instance.

NVIDIA first announced the DGX Spark platform (then Project DIGITS) at a starting price of $3,000, available from NVIDIA and OEM partners that build their own GB10-powered systems, including Acer, ASUS, Dell, Gigabyte, HP, Lenovo, and MSI; exact street pricing varies by retailer and configuration. The RTX 5090 launched at a $1,999 MSRP in January 2025, though street prices fluctuate with GPU market demand. Renting either skips the upfront cost: Spheron lists RTX 5090 on-demand at $0.78/hr as of 06 Oct 2026, billed per minute with a 20-minute minimum.

Yes, and this is where single-stream benchmarks mislead. On Llama 3.1 8B FP8 with SGLang, DGX Spark's decode throughput scales from 20.5 tokens/sec at batch size 1 to 368 tokens/sec at batch size 32, according to LMSYS. Separately, an NVIDIA developer forum post measured vLLM serving 8 concurrent requests on a single DGX Spark at an aggregate 289-313 tokens/sec across four different models, while Ollama's 8-concurrent aggregate matched its single-stream number because Ollama doesn't batch requests. A GPU that only ever serves one chat session at a time never sees this advantage.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min