The Arena.
38 models.
1 DGX Spark.

Compare open-weight models on DGX Spark using an indicative score: 70% external quality, 20% throughput and 10% memory efficiency. Choose a workload and inspect the measurements. Incomplete quality data yields no total score.

Size class

Active comparison: Suitability score (indicative) · Aggregate

The indicative suitability score gives quality the most weight. Small models can be fast without doing your task well. External tests are not interchangeable and do not validate your local precision.

Select # Model Size Context VRAM Knowledge Reasoning Coding Throughput Suitability score (indicative) Details
Ministral-3 3B mistral · Dense · BF16 3B 256K 8 GB 70.7MMLU 53.4GPQA-Diamond 54.8LiveCodeBench t/s Not available
Ministral-3 8B mistral · Dense · BF16 8B 256K 18 GB 76.1MMLU 66.8GPQA-Diamond 61.6LiveCodeBench 119t/s Not available
Mistral-Small 4 119B mistral · MoE · NVFP4 119B 256K 83 GB 71.2GPQA-Diamond 60t/s Not available
Muse-Glimmer 30B meta · Dense · BF16 30B 128K 60 GB 83.5GPQA-Diamond 29t/s Not available
Muse-Glimmer 30B meta · Dense · BF16 30B 128K 60 GB 83.5GPQA-Diamond 52t/s Not available
Nemotron-3-Nano 30B-A3B nvidia · MoE · BF16 30B 256K 62 GB 77.3MMLU-Pro 72.2GPQA-Diamond 63.2LiveCodeBench v5 62t/s Not available
Nemotron-3-Nano 30B-A3B nvidia · MoE · FP8 30B 256K 33 GB 77.3MMLU-Pro 72.2GPQA-Diamond 63.2LiveCodeBench v5 103t/s Not available
Nemotron-3-Nano 4B nvidia · Dense · BF16 4B 256K 8 GB 18.1MMLU-Pro 51.3GPQA-Diamond 51.8LiveCodeBench 147t/s Not available
Nemotron-3-Nano 4B nvidia · Dense · FP8 4B 256K 4 GB 18.1MMLU-Pro 51.3GPQA-Diamond 51.8LiveCodeBench 201t/s Not available
Nemotron-3 Nano Omni 30B-A3B nvidia · MoE · BF16 30B 256K 60 GB 60t/s Not available

The horizontal axis shows total throughput for the selected workload. The vertical axis shows mean external quality from MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on a fixed 0–100 scale. Only models with the required quality tests, a workload measurement and valid memory data are shown. The blue frontier contains models for which no other model within your filter is at least as fast and at least as high quality, with a strict improvement on one axis. This is not the total score: memory efficiency also contributes to that score.

Shown: 16 of 38 models.

On the Pareto frontier Dominated VRAM (small → large)

Each bubble represents one model profile. Blue marks the Pareto frontier. For a grey model, another model in your filter is at least as good in speed and external quality, and better on one of those axes. Larger bubbles mean more memory for model weights. Open the legend explanations for the precise definitions.

Chart values as a table
Model t/s External quality GB Suitability score
Qwen-3.6 35B-A3B · NVFP4 130.8 83.9 24 72.3
Qwen-3.6 35B-A3B · FP8 100.2 83.9 38 68.8
Gemma-4 26B-A4B · NVFP4 110.5 81.5 24 68.5
Gemma-4 26B-A4B · NVFP4 111.2 80.7 18 68.4
Qwen-3.5 4B · BF16 146 70.4 8 67.3
Qwen-3.6 35B-A3B · BF16 62.3 83.9 70 64.8
Gemma-4 E2B · BF16 212.3 49.1 5 64.4
Qwen-3.6 27B · FP8 35.5 86 31 63.8
Gemma-4 26B-A4B · BF16 65.7 80.7 52 62.9
Gemma-4 26B-A4B · BF16 64.5 80.7 52 62.8
Nemotron-Cascade-2 30B-A3B · BF16 62.3 81 60 62.8
Qwen-3.6 27B · BF16 25.7 86 54 62.7
Qwen-3.5 9B · BF16 84.2 76.6 18 62.6
Gemma-4 31B · NVFP4 25.7 83.2 22 60.9
Gemma-4 31B · BF16 15 83.2 62 59.7
Gemma-4 E4B · BF16 100.5 60 9 54.1

Suitability score (indicative)

70% quality + 20% throughput + 10% throughput per GB. Quality is the unweighted mean of MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on their original 0–100 scales. All three are required; missing or different tests yield no total score. Speed and efficiency are normalised against all models with matching quality data and a measurement for this workload, not your search filter. Aggregate requires all six closed-loop tests. These are editorial weights, not our own quality evaluation: vendor settings may differ and these figures do not validate your precision. Not available means insufficient comparable data, not a poor model.

Hardware score explained

The optional hardware score uses only throughput and throughput per GB, with preset-specific weights. This older technical metric does not assess answer quality and uses different normalisation from the suitability score.

Explanation