The Arena.
38 models.
1 DGX Spark.

Compare open-weight models on DGX Spark using an indicative score: 70% external quality, 20% throughput and 10% memory efficiency. Choose a workload and inspect the measurements. Incomplete quality data yields no total score.

Size class

Active comparison: Suitability score (indicative) · Aggregate

The indicative suitability score gives quality the most weight. Small models can be fast without doing your task well. External tests are not interchangeable and do not validate your local precision.

Select # Model Size Context VRAM Knowledge Reasoning Coding Throughput Suitability score (indicative) Details
01 Qwen-3.6 35B-A3B alibaba · MoE · NVFP4 35B 256K 24 GB 85.2MMLU-Pro 86GPQA-Diamond 80.4LiveCodeBench v6 131t/s 72.3
02 Qwen-3.6 35B-A3B alibaba · MoE · FP8 35B 256K 38 GB 85.2MMLU-Pro 86GPQA-Diamond 80.4LiveCodeBench v6 100t/s 68.8
03 Gemma-4 26B-A4B google · MoE · NVFP4 26B 256K 24 GB 84.8MMLU-Pro 79.9GPQA-Diamond 79.8LiveCodeBench v6 111t/s 68.5
04 Gemma-4 26B-A4B google · MoE · NVFP4 26B 256K 18 GB 82.6MMLU-Pro 82.3GPQA-Diamond 77.1LiveCodeBench v6 111t/s 68.4
05 Qwen-3.5 4B alibaba · Hybrid · BF16 4B 256K 8 GB 79.1MMLU-Pro 76.2GPQA-Diamond 55.8LiveCodeBench v6 146t/s 67.3
06 Qwen-3.6 35B-A3B alibaba · MoE · BF16 35B 256K 70 GB 85.2MMLU-Pro 86GPQA-Diamond 80.4LiveCodeBench v6 62t/s 64.8
07 Gemma-4 E2B google · Dense · BF16 2.3B 128K 5 GB 60MMLU-Pro 43.4GPQA-Diamond 44LiveCodeBench v6 212t/s 64.4
08 Qwen-3.6 27B alibaba · Hybrid · FP8 27B 256K 31 GB 86.2MMLU-Pro 87.8GPQA-Diamond 83.9LiveCodeBench v6 36t/s 63.8
09 Gemma-4 26B-A4B google · MoE · BF16 26B 256K 52 GB 82.6MMLU-Pro 82.3GPQA-Diamond 77.1LiveCodeBench v6 66t/s 62.9
10 Gemma-4 26B-A4B google · MoE · BF16 26B 256K 52 GB 82.6MMLU-Pro 82.3GPQA-Diamond 77.1LiveCodeBench v6 65t/s 62.8

The horizontal axis shows total throughput for the selected workload. The vertical axis shows mean external quality from MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on a fixed 0–100 scale. Only models with the required quality tests, a workload measurement and valid memory data are shown. The blue frontier contains models for which no other model within your filter is at least as fast and at least as high quality, with a strict improvement on one axis. This is not the total score: memory efficiency also contributes to that score.

Shown: 16 of 38 models.

On the Pareto frontier Dominated VRAM (small → large)

Each bubble represents one model profile. Blue marks the Pareto frontier. For a grey model, another model in your filter is at least as good in speed and external quality, and better on one of those axes. Larger bubbles mean more memory for model weights. Open the legend explanations for the precise definitions.

Chart values as a table
Model t/s External quality GB Suitability score
Qwen-3.6 35B-A3B · NVFP4 130.8 83.9 24 72.3
Qwen-3.6 35B-A3B · FP8 100.2 83.9 38 68.8
Gemma-4 26B-A4B · NVFP4 110.5 81.5 24 68.5
Gemma-4 26B-A4B · NVFP4 111.2 80.7 18 68.4
Qwen-3.5 4B · BF16 146 70.4 8 67.3
Qwen-3.6 35B-A3B · BF16 62.3 83.9 70 64.8
Gemma-4 E2B · BF16 212.3 49.1 5 64.4
Qwen-3.6 27B · FP8 35.5 86 31 63.8
Gemma-4 26B-A4B · BF16 65.7 80.7 52 62.9
Gemma-4 26B-A4B · BF16 64.5 80.7 52 62.8
Nemotron-Cascade-2 30B-A3B · BF16 62.3 81 60 62.8
Qwen-3.6 27B · BF16 25.7 86 54 62.7
Qwen-3.5 9B · BF16 84.2 76.6 18 62.6
Gemma-4 31B · NVFP4 25.7 83.2 22 60.9
Gemma-4 31B · BF16 15 83.2 62 59.7
Gemma-4 E4B · BF16 100.5 60 9 54.1

Suitability score (indicative)

70% quality + 20% throughput + 10% throughput per GB. Quality is the unweighted mean of MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on their original 0–100 scales. All three are required; missing or different tests yield no total score. Speed and efficiency are normalised against all models with matching quality data and a measurement for this workload, not your search filter. Aggregate requires all six closed-loop tests. These are editorial weights, not our own quality evaluation: vendor settings may differ and these figures do not validate your precision. Not available means insufficient comparable data, not a poor model.

Hardware score explained

The optional hardware score uses only throughput and throughput per GB, with preset-specific weights. This older technical metric does not assess answer quality and uses different normalisation from the suitability score.

Explanation