DGX Spark thermal throttling is not a maybe. NVIDIA's own staff have confirmed on the developer forum that the GB10 chip's power budget is far smaller than the 240W headline number suggests, and several independent owners have logged the exact point where clock speed and temperature cross into a throttled state under real workloads. What's missing from the spec sheet, and from most of the forum threads reporting a crash, is the shape of that decay over time: how fast it happens, how hot the chassis actually runs, and what a clock lock costs you in tokens per second once you apply it. Below is what the handful of people who have logged clock, power, and temperature together, on the same timestamp, over a sustained run have found, read against each other so you can see where DGX Spark's production ceiling actually sits.
TL;DR: Is DGX Spark Thermal Throttling a Real Problem Under Sustained Load?
- Confirmed by NVIDIA. DGX Spark's GPU is capped at 120W inside the unit's 240W system-wide power budget, per NVIDIA staff on the forum.
- Clock decay is steep. An uncapped GB10 unit's SM clock fell from 2450 MHz to 1475 MHz under load, peaking at 96.0C before shutdown.
- Identical units diverge. One GB10 unit ran above 85C for 62 of 73 hours and shut down at 87C; an identical twin held steady at 68C.
- The fix is cheap. Locking clocks with
nvidia-smi -lgccut peak temperature by 20C, with throughput loss as low as 5-15%. - Spheron: rents H100 GPUs by the minute for workloads past a desk unit's reliability ceiling. Check H100 GPU rental availability.
Why a 100W-Capped GB10 Causes DGX Spark Thermal Throttling
A desktop GPU card has its own power connector, its own fan or radiator, and a power target set for that card alone. DGX Spark's GB10 does not work that way, and the gap between what people expect and what the unit actually allows is the root of most of the "overheating" and "half performance" threads.
NVIDIA staff member NVES laid out the real numbers on the developer forum, clarifying a question that comes up constantly: "Spark GPU can draw a max of 120W and the rest of the system including CPU, CX7 and peripherals can draw 100W. The power supply is rated for 240 Watt of peak draw." That means the 240W figure on the spec sheet was never a GPU budget. It is a whole-system ceiling shared between the GPU, the Arm CPU cores, the ConnectX-7 networking, the SSD, and USB-C peripherals.
A separate thread on the same forum breaks the budget down further: GB10's SoC TDP, meaning CPU and GPU combined, is 140W, while the GPU on its own is capped at 120W, with roughly 100W of the system's total headroom going to everything that isn't the GPU. John Carmack reported hitting this ceiling directly: his unit maxed out around 100W of the rated 240W and delivered roughly half the quoted performance. That is not a defective unit behaving badly. It is a 120W-capped GPU doing exactly what its budget allows, inside a system whose headline number was always describing something larger than the chip doing the inference work.
This is the structural reason DGX Spark throttles differently than a discrete desktop card sitting in a tower with its own 450W connector and open airflow. A desktop GPU's thermal ceiling is set by its own cooler and its own power delivery. GB10's is set by a shared budget across a compact, fanless-leaning chassis the size of a small desktop mini PC, where the CPU, networking silicon, and storage are all drawing from the same pool the GPU needs for sustained clock speed.
Test Setup: What to Run, How Long, and How to Log It
The method that actually answers whether a given DGX Spark or GB10 box throttles is simple to describe and tedious to skip: run a continuous inference workload for two-plus hours and log clock speed, power draw, and temperature together, on the same timestamp, for the entire run. A short benchmark pass does not do this. It measures the first few minutes, which, as the data below shows, is often the part of the curve before the real ceiling has been reached.
The command that produces this log is the same nvidia-smi CSV query used for sustained-load testing on any NVIDIA GPU:
nvidia-smi --query-gpu=timestamp,clocks.sm,clocks.max.sm,power.draw,power.limit,temperature.gpu,utilization.gpu \
--format=csv -l 2Run that against a genuinely continuous workload, not a bursty one with idle gaps, and the log shows whether clock speed, power, and temperature move together or independently, which is the whole key to diagnosing what's actually limiting you. The tests referenced through the rest of this piece used exactly this kind of continuous load: sustained Ollama serving, continuous vLLM inference against Qwen3.5-35B-A3B-FP8, and uncapped native stress jobs run specifically to find the ceiling. The same nvidia-smi logging approach and the same thermal-vs-power diagnostic method applies just as directly to a rented cloud GPU; our sustained-load throttling test on a rented H100 runs the cloud version of this exact check.
One caveat worth logging alongside the GPU numbers: nvidia-smi's own reported GPU temperature is not the full picture on a GB10 box. One test found nvidia-smi's reported GPU temperature running 8-15C lower than the system's actual thermal-zone sensor (acpitz) on the same unit, with acpitz reading 92.2C moments before an undiagnosed power-off while nvidia-smi showed only 84C at the same instant. If you're logging only nvidia-smi's GPU temperature field, you may be reading a number that's already several degrees optimistic relative to what the chassis is actually experiencing.
Tokens Per Second, Clock Speed, and Power Draw Over a 2-Hour Sustained Run
No single published test holds a GB10 unit under one continuous two-hour inference job while logging all four signals end to end. What exists instead is a set of overlapping windows from different sustained-load tests, run on different units, at different lengths, that together draw the same decay curve a full two-hour run would show. Stitched together on elapsed time, here is what they report:
| Elapsed time / test | SM clock | Temperature | Power draw | What it shows |
|---|---|---|---|---|
| Idle, before load | ~2411 MHz default | ~45-46C | ~10W | Baseline before a job starts |
| Uncapped stress job, 0-156s | Falls toward 1475 MHz (from 2450 MHz) | Peaks at 96.0C (12% of run at or above 95C) | Peaks at 93.2W | Job hard-powers off before completing |
| Same job, clock capped at 2200 MHz | Held near 2200 MHz | Peaks at 91.1C (0% of run at or above 95C) | Peaks at 74.0W | Job completes in 251s instead of crashing |
| Healthy unit, vLLM serving Qwen3.5-35B-A3B-FP8 | 2522 MHz | not separately logged | 35.65W | 96% GPU utilization, ~50 tok/s sustained |
| Hours 0-73, GB10 unit A | not separately logged | At or above 85C for 62 of 73 hours | not separately logged | Shuts down at 87C, needs to cool to 39C to recover |
| Hours 0-73, GB10 unit B (identical hardware) | not separately logged | Steady 68C throughout | not separately logged | No shutdown, same workload |
The decay in the uncapped row happened inside roughly two and a half minutes, not two hours. That matters more than it first looks: GB10's small, shared-budget chassis can reach a throttled or failed state far faster than a cloud GPU with its own dedicated cooling, which is exactly why a two-hour test window, long enough to cover both the fast initial ramp and any slower multi-hour drift, is the length worth running before you trust a unit with a production job. The 73-hour rows, from two identical GB10 units running the same workload, show the other half of the picture: a unit that survives the first few minutes fine can still diverge sharply from an identical neighbor over the following hours, for reasons that aren't visible in a short benchmark at all.
Thermal-Limited or Power-Limited? How to Tell on DGX Spark
A GB10 unit sitting at a lower-than-expected clock could be thermal-limited, power-limited, or in rare cases suffering an outright power-delivery fault that looks like throttling at a glance. Telling these apart means reading temperature and power draw together, not either one alone.
The general method comes from NVIDIA's own performance documentation: correlate temperature and power draw on the same timestamp. NVIDIA's TensorRT performance documentation puts the thermal throttle point near 85C for most GPUs, and high temperature paired with power sitting below its cap points to a thermal limit, while power pinned flat at its ceiling with temperature still comfortable points to a power limit instead. NVIDIA's Grace performance tuning guide frames the broader mechanism: throttling is usually a continuous governor response to average power, not a single hard cutoff, and "the maximum frequency usually corresponds to the maximum possible performance and is higher than the frequency at which nominal (sustained) performance can be achieved."
On an apparently healthier GB10 unit, one test found power draw was the signal that actually moved under load while clock speed barely did: SM clock sat at 2379-2392 MHz under sustained inference against roughly 2411 MHz idle, while power spiked from idle to a peak between 57W and 82W. On that unit, power draw was the real tell, not clock speed, which stayed nearly flat the entire time.
There's a third, trickier failure mode worth checking for before you assume a low clock means thermal or power throttling at all: a genuine power-delivery fault. One diagnostic case compared a healthy unit running Qwen3.5-35B-A3B-FP8 at a 2522 MHz SM clock, 35.65W power draw, and 96% GPU utilization against a defective unit stuck in a "30W safety mode," showing a similar 2411 MHz clock but only ~4.80W draw and 2% utilization. The clock numbers alone look almost identical between the two units. Only the power draw and utilization expose that the second unit isn't throttling from heat at all, it's a power-delivery fault that happens to produce a clock reading that resembles a healthy idle state. Log all three signals together, or a fault like that reads as "fine" on a glance at clock speed alone.
Separately, the same diagnostic work documents what a genuine unexplained power-off looks like from the inside: no kernel logs, no pstore data, no watchdog reboot record. As the author put it, "It is not a crash. There is no power." That's a useful distinction to carry into any incident review: a GB10 hitting its thermal or power ceiling hard enough doesn't crash in a way you can debug from logs after the fact. It just stops, and the only evidence you have is whatever telemetry you were already logging externally at the time.
Does the OEM Chassis Change the Throttle Curve? Identical GB10 Units, Different Results
The most uncomfortable finding in the sustained-load data isn't about any single unit's peak temperature. It's that two identical units, same chip, same chassis, same workload, can produce entirely different thermal outcomes. On a 73-hour sustained test, one GB10 unit spent 62 of the 73 hours at or above 85C and shut down once it hit 87C, recovering only after cooling back down to 39C, while a second, identical unit held a steady 68C for the entire 73 hours. Same hardware, same software, same job, a roughly 20C gap in steady-state temperature between the two boxes.
That gap is consistent with what's already documented about build variance across the GB10 lineup. Every certified GB10 system, DGX Spark included, runs the identical 140W SoC inside the same 240W system power budget, so the chip itself isn't the variable. Chassis-level thermal design is. ASUS reportedly redesigned its Ascent GX10's chassis between its GTC preview and its final market release specifically to improve cooling, which our comparison of every GB10 box covers in more detail. A unit-to-unit gap this wide inside the same model line suggests cooling headroom in these compact chassis designs is tighter, and more sensitive to manufacturing and assembly variance, than the identical spec sheets across OEM partners would suggest.
The practical takeaway isn't that any particular box is defective. It's that a spec sheet, or even a review of one unit, doesn't tell you where your specific unit sits on that curve. Running the sustained-load test above on your own hardware is the only way to know whether you landed closer to the 68C unit or the 87C one.
The Fix That Actually Works: Locking Clocks With nvidia-smi -lgc
The fix with the most consistent evidence behind it is locking the GPU clock below its default ceiling, rather than letting the driver chase a boost clock the chassis can't sustain. sudo nvidia-smi -lgc <min>,<max> pins the SM clock inside a specified range instead of letting it float up to the card's rated ceiling and then fall back hard once temperature or power limits engage.
The evidence on what this actually costs is better than the worst-case math suggests, and it depends heavily on what kind of workload you're running. On one GB10 unit, capping the clock at 2200 MHz cut peak power draw 21% (from 93.2W to 74.0W) and peak temperature from 96.0C to 91.1C, while the job that previously crashed at 156 seconds uncapped completed in 251 seconds capped, a roughly 9% time cost in exchange for the job actually finishing. A separate test locking clocks to sudo nvidia-smi -lgc 300,2200 against a stock ceiling near 3003 MHz cut peak temperature by roughly 20C with no measurable throughput loss on memory-bandwidth-bound MoE inference, because that workload's real ceiling was GB10's 273 GB/s memory bandwidth, not SM clock speed at all. A third test, locking clocks into a wider 1800-3000 MHz hysteresis band under sustained Ollama load, dropped sustained temperature from 82-84C to 72C; the theoretical worst-case throughput loss was around 27%, but the measured median impact was only 5-15% tok/s.
The pattern across all three: the throughput cost of locking clocks is smallest, sometimes effectively zero, on workloads that are already bandwidth-bound rather than compute-bound, since those workloads were never going to use the full boost clock's worth of compute anyway. The cost grows on more compute-bound jobs, but even there, the measured median losses ran well under the theoretical worst case in every test that actually measured it rather than just calculating it. And on a job that was crashing outright without the lock, finishing slower is a strict improvement over not finishing at all.
Two things worth noting if you apply this yourself: nvidia-smi -lgc requires sudo, and the lock does not persist across a reboot by default, so it needs to be reapplied or scripted into startup if you want it to hold on a box that restarts unattended.
What This Means for Production Use vs Prototyping
Two separate findings push against running DGX Spark unattended for long production jobs, beyond the thermal variance documented above. First, a six-plus day DGX Spark fine-tuning stress test hit a hard system freeze 7.5 hours into one run, at 88% completion, caused by memory fragmentation, a separate reliability failure mode from thermal throttling entirely, and one that, like the thermal variance above, only shows up on genuinely multi-hour jobs rather than short benchmarks. Second, serving-stack choice turns out to matter as much as hardware for anything resembling concurrent production traffic: on one DGX Spark, Ollama's aggregate throughput at 8 concurrent requests matched its single-stream figure across all four tested models, meaning no real batching, while vLLM scaled from a similar single-stream baseline up to 289-313 tok/s under the same concurrent load. A unit that's thermally fine can still underperform badly in production if the serving stack on top of it was never built to batch concurrent requests.
Warranty terms add a formal line under all of this. NVIDIA's standard DGX Spark warranty is explicitly voided by what it calls "Enterprise Use," defined as large-scale datacenter or GPU cluster commercial deployment. A unit bought for a developer's desk is not warranted to be redeployed as an unattended production node, independent of whether it happens to run reliably there.
None of this makes DGX Spark a bad buy for what it's actually built for. A workload that fits inside one prototyping session, a fine-tuning run that checkpoints and pauses, or a memory-bandwidth-bound MoE inference job where the clock-lock throughput cost is near zero are all cases where this hardware does its job well, and renting a cloud GPU for that same short session would be a net cost increase, not a fix. It's also worth sizing that break-even honestly before switching anything: our DGX Spark to GPU cloud pipeline guide works through the hours-of-use math where a desk unit's one-time cost stops beating an hourly rental, and our 3-year rent-vs-buy TCO breakdown covers the same tradeoff at cluster scale. The electricity side of that comparison, what a GPU actually draws against what it costs to run, is covered separately in our AI inference power and electricity cost guide.
Where the data above actually changes the calculus is 24/7 unattended production serving, multi-hour jobs with no human watching for a silent power-off, or deployment at a scale the warranty's own "Enterprise Use" clause explicitly excludes. That's the point where Spheron's on-demand GPU rental is the honest alternative, not a universal upgrade. Spheron bills per minute after a 20-minute minimum, with no long-term commitment, so a reader who hits DGX Spark's reliability ceiling on one job can burst to an $2.65/hr H100 for exactly the hours needed rather than buying a second desk unit. Spheron's bare-metal option gives the same kind of direct, unabstracted nvidia-smi access this piece's entire methodology depends on, and its 99.9% uptime SLA is a guarantee about reachability that a single desk unit, by definition, can't offer. What it doesn't do is fix the GB10 firmware and chassis behavior documented above; that's a different machine, not a patch for this one, and it also gives up DGX Spark's actual selling point for solo, offline prototyping: no per-hour meter and no network dependency. For a workload that genuinely fits inside one local session, renting hardware for it is a cost increase looking for a problem.
Pricing fluctuates based on GPU availability. Spheron rates above are live as of 07 Oct 2026; other figures in this piece reflect their sources' most recently published data and may have changed since. Check current GPU pricing → for live rates.
If your DGX Spark workload has outgrown one box's reliability ceiling, an H100 or H200 on Spheron gives you the same bare-metal
nvidia-smiaccess this piece's method depends on, without the per-unit variance documented above.
Frequently Asked Questions
Yes, confirmed from two directions. NVIDIA staff on the developer forum clarified that DGX Spark's 240W rating covers the whole system, not the GPU alone, which caps the GPU itself at 120W and explains reports of owners getting roughly half the advertised performance. Separately, independent sustained-load tests on GB10 units have logged SM clock falling from a 2450 MHz default to as low as 1475 MHz under continuous load, with one unit spending 12% of its run at or above 95C before a hard power-off.
It varies by unit and workload, which is itself part of the problem. One GB10 box hit a peak of 96.0C on an uncapped job before powering off. On a 73-hour sustained test, one unit spent 62 of the 73 hours at or above 85C and shut down at 87C, while a second, identical unit held a steady 68C for the entire run on the same workload. A separate report on what looked like a healthier unit saw sustained inference peak at 75C with idle around 45-46C.
Lock the GPU clock with nvidia-smi instead of letting it float to its default ceiling. `sudo nvidia-smi -lgc 300,2200` cut one GB10 unit's peak temperature by roughly 20C with no measurable throughput loss on memory-bandwidth-bound MoE inference. A separate test locking clocks into a 1800-3000 MHz band dropped sustained temperature from 82-84C to 72C, with a measured median throughput cost of only 5-15% despite a theoretical worst case near 27%. The setting does not persist across reboots by default, so it needs to be reapplied or scripted into startup.
Because 240W was never the GPU's own budget. NVIDIA staff clarified on the developer forum that GB10's combined CPU+GPU SoC TDP is 140W, the GPU alone is capped at 120W, and DGX Spark's 240W figure is the peak draw for the entire system, with the remaining roughly 100W covering the ConnectX-7 networking, SSD, and USB-C peripherals, not spare GPU headroom. Reports of a unit capping near 100W and delivering about half the quoted performance reflect that shared budget, not a malfunction, though a genuine power-delivery fault can also stall a unit near 30W and look similar at a glance.
Less than the worst-case math suggests, and it depends entirely on whether your workload is bandwidth-bound or compute-bound. On memory-bandwidth-bound MoE inference, one test measured no measurable throughput loss from a clock lock, because GB10's 273 GB/s memory bandwidth was already the real ceiling, not SM clock. On a broader sustained Ollama workload, the theoretical worst case was a 27% throughput cut, but the measured median impact was only 5-15% tok/s. A clock lock that finishes a job in 251 seconds instead of crashing it at 156 seconds uncapped is a net throughput win, not a loss.






