Compare
side by side.

First select at least 2 models on the Arena.

← Choose other models
Suitability score (indicative)

70% quality + 20% throughput + 10% throughput per GB. Quality is the unweighted mean of MMLU-Pro, GPQA-Diamond and LiveCodeBench v6 on their original 0–100 scales. All three are required; missing or different tests yield no total score. Speed and efficiency are normalised against all models with matching quality data and a measurement for this workload, not your search filter. Aggregate requires all six closed-loop tests. These are editorial weights, not our own quality evaluation: vendor settings may differ and these figures do not validate your precision. Not available means insufficient comparable data, not a poor model.

Change selection or workload

Select at least two profiles to open the comparison.

→ To the Arena
Explanation