The Arena.
38 models.
1 DGX Spark.

Compare open-weight models on DGX Spark using an indicative score: 70% external quality, 20% throughput and 10% memory efficiency. Choose a workload and inspect the measurements. Incomplete quality data yields no total score.

Size class

Active comparison: Suitability score (indicative) · Aggregate

The indicative suitability score gives quality the most weight. Small models can be fast without doing your task well. External tests are not interchangeable and do not validate your local precision.

Select # Model Size Context VRAM Knowledge Reasoning Coding Throughput Suitability score (indicative) Details
Nemotron-3 Nano Omni 30B-A3B nvidia · MoE · FP8 30B 256K 33 GB 109t/s Not available
Nemotron-3 Nano Omni 30B-A3B nvidia · MoE · NVFP4 30B 256K 21 GB 145t/s Not available
Nemotron-3-Nano 30B-A3B nvidia · MoE · NVFP4 30B 256K 21 GB 77.3MMLU-Pro 72.2GPQA-Diamond 63.2LiveCodeBench v5 144t/s Not available
Nemotron-3-Super 120B-A12B nvidia · MoE · NVFP4 120B 256K 60 GB 83.7MMLU-Pro 79.2GPQA 81.2LiveCodeBench v5 45t/s Not available
Qwen-3.8 27B alibaba · Dense · BF16 27B 256K 54 GB 89.2GPQA-Diamond 90.3LiveCodeBench v6 27t/s Not available
Qwen-3.8 27B alibaba · Dense · FP8 27B 256K 30 GB 89.2GPQA-Diamond 90.3LiveCodeBench v6 38t/s Not available
Qwen-3.5 0.8B alibaba · Hybrid · BF16 0.8B 256K 2 GB 29.7MMLU-Pro 11.9GPQA 461t/s Not available
Qwen-3.5 2B alibaba · Hybrid · BF16 2B 256K 4 GB 55.3MMLU-Pro 30.4SuperGPQA 243t/s Not available

The horizontal axis shows total throughput for the selected workload. The vertical axis shows mean external quality from MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on a fixed 0–100 scale. Only models with the required quality tests, a workload measurement and valid memory data are shown. The blue frontier contains models for which no other model within your filter is at least as fast and at least as high quality, with a strict improvement on one axis. This is not the total score: memory efficiency also contributes to that score.

Shown: 16 of 38 models.

On the Pareto frontier Dominated VRAM (small → large)

Each bubble represents one model profile. Blue marks the Pareto frontier. For a grey model, another model in your filter is at least as good in speed and external quality, and better on one of those axes. Larger bubbles mean more memory for model weights. Open the legend explanations for the precise definitions.

Chart values as a table
Model t/s External quality GB Suitability score
Qwen-3.6 35B-A3B · NVFP4 130.8 83.9 24 72.3
Qwen-3.6 35B-A3B · FP8 100.2 83.9 38 68.8
Gemma-4 26B-A4B · NVFP4 110.5 81.5 24 68.5
Gemma-4 26B-A4B · NVFP4 111.2 80.7 18 68.4
Qwen-3.5 4B · BF16 146 70.4 8 67.3
Qwen-3.6 35B-A3B · BF16 62.3 83.9 70 64.8
Gemma-4 E2B · BF16 212.3 49.1 5 64.4
Qwen-3.6 27B · FP8 35.5 86 31 63.8
Gemma-4 26B-A4B · BF16 65.7 80.7 52 62.9
Gemma-4 26B-A4B · BF16 64.5 80.7 52 62.8
Nemotron-Cascade-2 30B-A3B · BF16 62.3 81 60 62.8
Qwen-3.6 27B · BF16 25.7 86 54 62.7
Qwen-3.5 9B · BF16 84.2 76.6 18 62.6
Gemma-4 31B · NVFP4 25.7 83.2 22 60.9
Gemma-4 31B · BF16 15 83.2 62 59.7
Gemma-4 E4B · BF16 100.5 60 9 54.1

Suitability score (indicative)

70% quality + 20% throughput + 10% throughput per GB. Quality is the unweighted mean of MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on their original 0–100 scales. All three are required; missing or different tests yield no total score. Speed and efficiency are normalised against all models with matching quality data and a measurement for this workload, not your search filter. Aggregate requires all six closed-loop tests. These are editorial weights, not our own quality evaluation: vendor settings may differ and these figures do not validate your precision. Not available means insufficient comparable data, not a poor model.

Hardware score explained

The optional hardware score uses only throughput and throughput per GB, with preset-specific weights. This older technical metric does not assess answer quality and uses different normalisation from the suitability score.

Explanation