The Arena.
38 models.
1 DGX Spark.
Compare open-weight models on DGX Spark using an indicative score: 70% external quality, 20% throughput and 10% memory efficiency. Choose a workload and inspect the measurements. Incomplete quality data yields no total score.
Active comparison: Suitability score (indicative) · Aggregate
The indicative suitability score gives quality the most weight. Small models can be fast without doing your task well. External tests are not interchangeable and do not validate your local precision.
| Select | # | Model | Size | Context | VRAM | Knowledge | Reasoning | Coding | Throughput | Suitability score (indicative) | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 01 | Qwen-3.6 35B-A3B alibaba · MoE · NVFP4 | 35B | 256K | 24 GB | 85.2MMLU-Pro | 86GPQA-Diamond | 80.4LiveCodeBench v6 | 131t/s | 72.3 | ||
| 02 | Qwen-3.6 35B-A3B alibaba · MoE · FP8 | 35B | 256K | 38 GB | 85.2MMLU-Pro | 86GPQA-Diamond | 80.4LiveCodeBench v6 | 100t/s | 68.8 | ||
| 03 | Gemma-4 26B-A4B google · MoE · NVFP4 | 26B | 256K | 24 GB | 84.8MMLU-Pro | 79.9GPQA-Diamond | 79.8LiveCodeBench v6 | 111t/s | 68.5 | ||
| 04 | Gemma-4 26B-A4B google · MoE · NVFP4 | 26B | 256K | 18 GB | 82.6MMLU-Pro | 82.3GPQA-Diamond | 77.1LiveCodeBench v6 | 111t/s | 68.4 | ||
| 05 | Qwen-3.5 4B alibaba · Hybrid · BF16 | 4B | 256K | 8 GB | 79.1MMLU-Pro | 76.2GPQA-Diamond | 55.8LiveCodeBench v6 | 146t/s | 67.3 | ||
| 06 | Qwen-3.6 35B-A3B alibaba · MoE · BF16 | 35B | 256K | 70 GB | 85.2MMLU-Pro | 86GPQA-Diamond | 80.4LiveCodeBench v6 | 62t/s | 64.8 | ||
| 07 | Gemma-4 E2B google · Dense · BF16 | 2.3B | 128K | 5 GB | 60MMLU-Pro | 43.4GPQA-Diamond | 44LiveCodeBench v6 | 212t/s | 64.4 | ||
| 08 | Qwen-3.6 27B alibaba · Hybrid · FP8 | 27B | 256K | 31 GB | 86.2MMLU-Pro | 87.8GPQA-Diamond | 83.9LiveCodeBench v6 | 36t/s | 63.8 | ||
| 09 | Gemma-4 26B-A4B google · MoE · BF16 | 26B | 256K | 52 GB | 82.6MMLU-Pro | 82.3GPQA-Diamond | 77.1LiveCodeBench v6 | 66t/s | 62.9 | ||
| 10 | Gemma-4 26B-A4B google · MoE · BF16 | 26B | 256K | 52 GB | 82.6MMLU-Pro | 82.3GPQA-Diamond | 77.1LiveCodeBench v6 | 65t/s | 62.8 |
The horizontal axis shows total throughput for the selected workload. The vertical axis shows mean external quality from MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on a fixed 0–100 scale. Only models with the required quality tests, a workload measurement and valid memory data are shown. The blue frontier contains models for which no other model within your filter is at least as fast and at least as high quality, with a strict improvement on one axis. This is not the total score: memory efficiency also contributes to that score.
Shown: 16 of 38 models.
Each bubble represents one model profile. Blue marks the Pareto frontier. For a grey model, another model in your filter is at least as good in speed and external quality, and better on one of those axes. Larger bubbles mean more memory for model weights. Open the legend explanations for the precise definitions.
Chart values as a table
| Model | t/s | External quality | GB | Suitability score |
|---|---|---|---|---|
| Qwen-3.6 35B-A3B · NVFP4 | 130.8 | 83.9 | 24 | 72.3 |
| Qwen-3.6 35B-A3B · FP8 | 100.2 | 83.9 | 38 | 68.8 |
| Gemma-4 26B-A4B · NVFP4 | 110.5 | 81.5 | 24 | 68.5 |
| Gemma-4 26B-A4B · NVFP4 | 111.2 | 80.7 | 18 | 68.4 |
| Qwen-3.5 4B · BF16 | 146 | 70.4 | 8 | 67.3 |
| Qwen-3.6 35B-A3B · BF16 | 62.3 | 83.9 | 70 | 64.8 |
| Gemma-4 E2B · BF16 | 212.3 | 49.1 | 5 | 64.4 |
| Qwen-3.6 27B · FP8 | 35.5 | 86 | 31 | 63.8 |
| Gemma-4 26B-A4B · BF16 | 65.7 | 80.7 | 52 | 62.9 |
| Gemma-4 26B-A4B · BF16 | 64.5 | 80.7 | 52 | 62.8 |
| Nemotron-Cascade-2 30B-A3B · BF16 | 62.3 | 81 | 60 | 62.8 |
| Qwen-3.6 27B · BF16 | 25.7 | 86 | 54 | 62.7 |
| Qwen-3.5 9B · BF16 | 84.2 | 76.6 | 18 | 62.6 |
| Gemma-4 31B · NVFP4 | 25.7 | 83.2 | 22 | 60.9 |
| Gemma-4 31B · BF16 | 15 | 83.2 | 62 | 59.7 |
| Gemma-4 E4B · BF16 | 100.5 | 60 | 9 | 54.1 |
Suitability score (indicative)
70% quality + 20% throughput + 10% throughput per GB. Quality is the unweighted mean of MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on their original 0–100 scales. All three are required; missing or different tests yield no total score. Speed and efficiency are normalised against all models with matching quality data and a measurement for this workload, not your search filter. Aggregate requires all six closed-loop tests. These are editorial weights, not our own quality evaluation: vendor settings may differ and these figures do not validate your precision. Not available means insufficient comparable data, not a poor model.
Hardware score explained
The optional hardware score uses only throughput and throughput per GB, with preset-specific weights. This older technical metric does not assess answer quality and uses different normalisation from the suitability score.