What does the
DGX Spark cost?
Updated 14 September 2026
A DGX Spark comes with monthly costs. Hardware amortisation plus power under load. Adjust the assumptions to estimate your own usage.
02Assumptions for your own situation
Hardware / month€103amortisation of the Spark
Power / month€8170W × 22 working days
Total / month€111hardware plus power
03How the numbers come together
The formula
cost/month = (hardware ÷ amortisation) + (W ÷ 1000 × hours × 22 working days × kWh price)
The calculation adds depreciation and electricity. Hardware is depreciated linearly; electricity only counts during the stated hours under load, on 22 working days per month. Idle hours and weekends are excluded.
Assumptions
- Power assumes 170W under vLLM load. Maximum TDP is around 240W; actual use is lower.
- Hardware is the Founders Edition NL excl. VAT. Adjust if you bought it second-hand, get a tax benefit or sourced it another way.
- Utilisation assumption: all monthly calculations use 22 working days. Electricity counts only during the stated hours, excluding idle time, evenings and weekends.
- Not included: internet, cooling, space and your time maintaining it. This estimate therefore does not include every cost.
What this number doesn't tell you: whether local is cheaper than a cloud API. For that you also need tokens-per-second, and those differ per model. We work that out below, with the arenathroughput. The wider trade-off is in the Local AI decision guide.
04Local versus cloud
Cost per million output tokens
The Spark costs nearly the same each month whether you run it flat out or not. This comparison explicitly assumes 8 hours per working day and 22 working days per month. The two local columns use the exact Arena office-work and Monday-peak profiles.
Result under these assumptions: the Mistral Small 4 API is cheaper than the same model on your own Spark. Against GPT-5 mini, the result depends on the load profile: the peak run costs slightly less, the office profile more. Base an on-prem decision on actual utilisation, data sovereignty and predictability rather than a general price claim.
| Model | Precision | €/1M · office work (c=10) | €/1M · Monday peak |
|---|---|---|---|
| Mistral Small 4 APImistral-small-2603 | US, CLOUD Act | €0,51 any volume | |
| GPT-5 mini APIgpt-5-mini-2025-08-07 | US, CLOUD Act | €1,70 any volume | |
| Qwen-3.5 0.8B | BF16 | €0,23 | €0,24 |
| Qwen-3.5 2B | BF16 | €0,46 | €0,29 |
| Gemma-4 E2B | BF16 | €0,47 | €0,35 |
| LFM2.5 2.6B | BF16 | €0,51 | €0,36 |
| Nemotron-3-Nano 4B | FP8 | €0,58 | €0,50 |
| gpt-oss-20b | MXFP4 | €1,24 | €0,57 |
| Nemotron-3-Nano 4B | BF16 | €0,83 | €0,61 |
| Gemma-4 E4B | BF16 | €1,20 | €0,66 |
| Nemotron-3 Nano Omni 30B-A3B | NVFP4 | €0,84 | €0,67 |
| Nemotron-3-Nano 30B-A3B | NVFP4 | €0,89 | €0,67 |
| Qwen-3.5 4B | BF16 | €0,88 | €0,67 |
| Qwen-3.6 35B-A3B | NVFP4 | €0,95 | €0,78 |
| Gemma-4 26B-A4B | NVFP4 | €1,01 | €0,81 |
| Gemma-4 26B-A4B | NVFP4 | €1,00 | €0,82 |
| Ministral-3 8B | BF16 | €0,85 | €0,84 |
| Nemotron-3-Nano 30B-A3B | FP8 | €1,35 | €0,93 |
| Qwen-3.6 35B-A3B | FP8 | €1,24 | €1,01 |
| Qwen-3.5 9B | BF16 | €1,54 | €1,10 |
| Nemotron-3 Nano Omni 30B-A3B | FP8 | €1,27 | - |
| Granite-4.1 8B | BF16 | €1,75 | €1,29 |
| Nemotron-3 Nano Omni 30B-A3B | BF16 | €2,35 | €1,38 |
| Gemma-4 26B-A4B | BF16 | €1,95 | €1,41 |
| Gemma-4 26B-A4B | BF16 | €1,97 | €1,41 |
| Nemotron-3-Nano 30B-A3B | BF16 | €2,35 | €1,44 |
| Qwen-3.6 35B-A3B | BF16 | €2,12 | €1,54 |
| Nemotron-Cascade-2 30B-A3B | BF16 | €2,36 | €1,54 |
| KAT-Coder V2.5 | BF16 | €2,24 | €1,56 |
| Mistral-Small 4 119B | NVFP4 | €1,97 | €1,65 |
| Nemotron-3-Super 120B-A12B | NVFP4 | €2,74 | €2,47 |
| Muse-Glimmer 30B | BF16 | €4,39 | €3,07 |
| Gemma-4 31B | NVFP4 | €4,02 | €3,43 |
| Qwen-3.8 27B | FP8 | €3,32 | €3,51 |
| Qwen-3.8 27B | BF16 | €4,80 | €3,51 |
| Qwen-3.6 27B | FP8 | €3,32 | €3,54 |
| Qwen-3.6 27B | BF16 | €4,88 | €3,56 |
| Muse-Glimmer 30B | BF16 | €2,73 | €4,33 |
| Gemma-4 31B | BF16 | €8,03 | €4,34 |
- Assumption: 8 hours of actual load per working day, 22 working days per month. Bursty usage produces fewer tokens per month and therefore a higher price per token.
- We compare output tokens because the Arena measures decode. Office work uses test 04, multi-turn at c=10. Monday peak uses test 10; output throughput is total throughput times the scenario's output share. A missing profile remains empty and never falls back to another test.
- Only equivalent models make a fair comparison. The cleanest point is Mistral-Small 4 119B NVFP4 locally versus the fixed Mistral Small 4 API snapshot: the same model generation in a different location.
- The runs ran with prefix caching off. On would improve local throughput, and so lower the price.
- Cloud prices cover output tokens: Mistral Small 4 $0.60 and GPT-5 mini $2.00 per 1M. Both use $1 = €0.85 (2026-08-21). Sources: Mistral, OpenAI and the ECB reference rate.
- Energy: at 170W, electricity is only about 7% of the monthly cost; the rest is hardware depreciation. Under these assumptions, depreciation accounts for most of the cost. One million Gemma-4 NVFP4 output tokens uses roughly 274 Wh, a few cents. This assumes a flat 170W, since we do not measure power per model.