The Arena.
38 models.
1 DGX Spark.

Compare open-weight models on DGX Spark using an indicative score: 70% external quality, 20% throughput and 10% memory efficiency. Choose a workload and inspect the measurements. Incomplete quality data yields no total score.

Size class

Active comparison: Suitability score (indicative) · Aggregate

The indicative suitability score gives quality the most weight. Small models can be fast without doing your task well. External tests are not interchangeable and do not validate your local precision.

Select # Model Size Context VRAM Knowledge Reasoning Coding Throughput Suitability score (indicative) Details
11 Nemotron-Cascade-2 30B-A3B nvidia · MoE · BF16 30B 1000K 60 GB 79.8MMLU-Pro 76.1GPQA-Diamond 87.2LiveCodeBench v6 62t/s 62.8
12 Qwen-3.6 27B alibaba · Hybrid · BF16 27B 256K 54 GB 86.2MMLU-Pro 87.8GPQA-Diamond 83.9LiveCodeBench v6 26t/s 62.7
13 Qwen-3.5 9B alibaba · Dense · BF16 9B 256K 18 GB 82.5MMLU-Pro 81.7GPQA-Diamond 65.6LiveCodeBench v6 84t/s 62.6
14 Gemma-4 31B google · Dense · NVFP4 31B 256K 22 GB 85.2MMLU-Pro 84.3GPQA-Diamond 80LiveCodeBench v6 26t/s 60.9
15 Gemma-4 31B google · Dense · BF16 31B 256K 62 GB 85.2MMLU-Pro 84.3GPQA-Diamond 80LiveCodeBench v6 15t/s 59.7
16 Gemma-4 E4B google · Dense · BF16 4.5B 128K 9 GB 69.4MMLU-Pro 58.6GPQA-Diamond 52LiveCodeBench v6 101t/s 54.1
gpt-oss-20b openai · MoE · MXFP4 21B 128K 14 GB 85.3MMLU 71.5GPQA-Diamond 130t/s Not available
Granite-4.1 8B ibm · Dense · BF16 8B 128K 16 GB 56MMLU-Pro 42GPQA 85.4HumanEval 59t/s Not available
KAT-Coder V2.5 kwaipilot · MoE · BF16 35B 256K 70 GB 60t/s Not available
LFM2.5 2.6B liquidai · Hybrid · BF16 2.7B 32K 5 GB 59.4LiveCodeBench v6 235t/s Not available

The horizontal axis shows total throughput for the selected workload. The vertical axis shows mean external quality from MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on a fixed 0–100 scale. Only models with the required quality tests, a workload measurement and valid memory data are shown. The blue frontier contains models for which no other model within your filter is at least as fast and at least as high quality, with a strict improvement on one axis. This is not the total score: memory efficiency also contributes to that score.

Shown: 16 of 38 models.

On the Pareto frontier Dominated VRAM (small → large)

Each bubble represents one model profile. Blue marks the Pareto frontier. For a grey model, another model in your filter is at least as good in speed and external quality, and better on one of those axes. Larger bubbles mean more memory for model weights. Open the legend explanations for the precise definitions.

Chart values as a table
Model t/s External quality GB Suitability score
Qwen-3.6 35B-A3B · NVFP4 130.8 83.9 24 72.3
Qwen-3.6 35B-A3B · FP8 100.2 83.9 38 68.8
Gemma-4 26B-A4B · NVFP4 110.5 81.5 24 68.5
Gemma-4 26B-A4B · NVFP4 111.2 80.7 18 68.4
Qwen-3.5 4B · BF16 146 70.4 8 67.3
Qwen-3.6 35B-A3B · BF16 62.3 83.9 70 64.8
Gemma-4 E2B · BF16 212.3 49.1 5 64.4
Qwen-3.6 27B · FP8 35.5 86 31 63.8
Gemma-4 26B-A4B · BF16 65.7 80.7 52 62.9
Gemma-4 26B-A4B · BF16 64.5 80.7 52 62.8
Nemotron-Cascade-2 30B-A3B · BF16 62.3 81 60 62.8
Qwen-3.6 27B · BF16 25.7 86 54 62.7
Qwen-3.5 9B · BF16 84.2 76.6 18 62.6
Gemma-4 31B · NVFP4 25.7 83.2 22 60.9
Gemma-4 31B · BF16 15 83.2 62 59.7
Gemma-4 E4B · BF16 100.5 60 9 54.1

Suitability score (indicative)

70% quality + 20% throughput + 10% throughput per GB. Quality is the unweighted mean of MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on their original 0–100 scales. All three are required; missing or different tests yield no total score. Speed and efficiency are normalised against all models with matching quality data and a measurement for this workload, not your search filter. Aggregate requires all six closed-loop tests. These are editorial weights, not our own quality evaluation: vendor settings may differ and these figures do not validate your precision. Not available means insufficient comparable data, not a poor model.

Hardware score explained

The optional hardware score uses only throughput and throughput per GB, with preset-specific weights. This older technical metric does not assess answer quality and uses different normalisation from the suitability score.

Explanation