What does the
DGX Spark cost?

Updated 14 September 2026

A DGX Spark comes with monthly costs. Hardware amortisation plus power under load. Adjust the assumptions to estimate your own usage.

Hardware / month€103amortisation of the Spark
Power / month€8170W × 22 working days
Total / month€111hardware plus power

The formula

cost/month = (hardware ÷ amortisation) + (W ÷ 1000 × hours × 22 working days × kWh price)

The calculation adds depreciation and electricity. Hardware is depreciated linearly; electricity only counts during the stated hours under load, on 22 working days per month. Idle hours and weekends are excluded.

Assumptions

  • Power assumes 170W under vLLM load. Maximum TDP is around 240W; actual use is lower.
  • Hardware is the Founders Edition NL excl. VAT. Adjust if you bought it second-hand, get a tax benefit or sourced it another way.
  • Utilisation assumption: all monthly calculations use 22 working days. Electricity counts only during the stated hours, excluding idle time, evenings and weekends.
  • Not included: internet, cooling, space and your time maintaining it. This estimate therefore does not include every cost.
What this number doesn't tell you: whether local is cheaper than a cloud API. For that you also need tokens-per-second, and those differ per model. We work that out below, with the arenathroughput. The wider trade-off is in the Local AI decision guide.

Cost per million output tokens

The Spark costs nearly the same each month whether you run it flat out or not. This comparison explicitly assumes 8 hours per working day and 22 working days per month. The two local columns use the exact Arena office-work and Monday-peak profiles.

Result under these assumptions: the Mistral Small 4 API is cheaper than the same model on your own Spark. Against GPT-5 mini, the result depends on the load profile: the peak run costs slightly less, the office profile more. Base an on-prem decision on actual utilisation, data sovereignty and predictability rather than a general price claim.
Model Precision €/1M · office work (c=10) €/1M · Monday peak
Mistral Small 4 APImistral-small-2603 US, CLOUD Act €0,51 any volume
GPT-5 mini APIgpt-5-mini-2025-08-07 US, CLOUD Act €1,70 any volume
Qwen-3.5 0.8B BF16 €0,23 €0,24
Qwen-3.5 2B BF16 €0,46 €0,29
Gemma-4 E2B BF16 €0,47 €0,35
LFM2.5 2.6B BF16 €0,51 €0,36
Nemotron-3-Nano 4B FP8 €0,58 €0,50
gpt-oss-20b MXFP4 €1,24 €0,57
Nemotron-3-Nano 4B BF16 €0,83 €0,61
Gemma-4 E4B BF16 €1,20 €0,66
Nemotron-3 Nano Omni 30B-A3B NVFP4 €0,84 €0,67
Nemotron-3-Nano 30B-A3B NVFP4 €0,89 €0,67
Qwen-3.5 4B BF16 €0,88 €0,67
Qwen-3.6 35B-A3B NVFP4 €0,95 €0,78
Gemma-4 26B-A4B NVFP4 €1,01 €0,81
Gemma-4 26B-A4B NVFP4 €1,00 €0,82
Ministral-3 8B BF16 €0,85 €0,84
Nemotron-3-Nano 30B-A3B FP8 €1,35 €0,93
Qwen-3.6 35B-A3B FP8 €1,24 €1,01
Qwen-3.5 9B BF16 €1,54 €1,10
Nemotron-3 Nano Omni 30B-A3B FP8 €1,27 -
Granite-4.1 8B BF16 €1,75 €1,29
Nemotron-3 Nano Omni 30B-A3B BF16 €2,35 €1,38
Gemma-4 26B-A4B BF16 €1,95 €1,41
Gemma-4 26B-A4B BF16 €1,97 €1,41
Nemotron-3-Nano 30B-A3B BF16 €2,35 €1,44
Qwen-3.6 35B-A3B BF16 €2,12 €1,54
Nemotron-Cascade-2 30B-A3B BF16 €2,36 €1,54
KAT-Coder V2.5 BF16 €2,24 €1,56
Mistral-Small 4 119B NVFP4 €1,97 €1,65
Nemotron-3-Super 120B-A12B NVFP4 €2,74 €2,47
Muse-Glimmer 30B BF16 €4,39 €3,07
Gemma-4 31B NVFP4 €4,02 €3,43
Qwen-3.8 27B FP8 €3,32 €3,51
Qwen-3.8 27B BF16 €4,80 €3,51
Qwen-3.6 27B FP8 €3,32 €3,54
Qwen-3.6 27B BF16 €4,88 €3,56
Muse-Glimmer 30B BF16 €2,73 €4,33
Gemma-4 31B BF16 €8,03 €4,34
  • Assumption: 8 hours of actual load per working day, 22 working days per month. Bursty usage produces fewer tokens per month and therefore a higher price per token.
  • We compare output tokens because the Arena measures decode. Office work uses test 04, multi-turn at c=10. Monday peak uses test 10; output throughput is total throughput times the scenario's output share. A missing profile remains empty and never falls back to another test.
  • Only equivalent models make a fair comparison. The cleanest point is Mistral-Small 4 119B NVFP4 locally versus the fixed Mistral Small 4 API snapshot: the same model generation in a different location.
  • The runs ran with prefix caching off. On would improve local throughput, and so lower the price.
  • Cloud prices cover output tokens: Mistral Small 4 $0.60 and GPT-5 mini $2.00 per 1M. Both use $1 = €0.85 (2026-08-21). Sources: Mistral, OpenAI and the ECB reference rate.
  • Energy: at 170W, electricity is only about 7% of the monthly cost; the rest is hardware depreciation. Under these assumptions, depreciation accounts for most of the cost. One million Gemma-4 NVFP4 output tokens uses roughly 274 Wh, a few cents. This assumes a flat 170W, since we do not measure power per model.