Measured on 2026-08-06 with vLLM v0.26.0. In the chat test with ten concurrent requests this profile reaches 24.14 tokens/s per user, with an average of 0.88 seconds to the first token. At 25k context that average wait rises to 17.73 seconds. These runs measure speed and waiting, not the quality of the answers.
—
Suitability score (indicative)
147
Throughput tok/s
8 GB
VRAM
11/11
Reliable benches
Hugging Face →·vLLM v0.26.0·DGX Spark, NVIDIA GB10, 128 GB unified memory·Measured with context 48K·Last measured 6 August 2026
02Vendor quality · reference only
Knowledge: MMLU-Pro; reasoning: GPQA-Diamond; coding: LiveCodeBench v6. Each value names its exact test. Other tests and missing measurements do not qualify for the total score. External figures are not our own evaluation of this precision. Source: vendor model card ↗
3/3
Coverage
18.1
MMLU-Pro
51.3
GPQA-Diamond
51.8
LiveCodeBench
03Performance · BF16
Decode throughput · total t/s · c=10
■ BF16
1k ctx BF16224 t/s
8k ctx BF16161 t/s
4k+turn BF16211 t/s
25k ctx BF1662 t/s
04Test suite · 11 benchmarks
6 closed-loop tests with llama-benchy and 5 open-loop tests with vllm bench serve. Only complete runs without failed sanity checks are eligible; open-loop tests require at least 99% successful requests. Raw results remain visible. Expand “view command” for the reconstructed command. Methodology →
01 · llama-benchyclosed-loop
Chat
Short prompt, long answer. The shape that has to feel like ordinary chat; TTFT decides whether it feels snappy.
pp (prompt)
1024
tg (gen)
1024
depth
0
concurrency
1 · 5 · 10
repeats
3
Tokens/sec · per user
24.1t/s
TTFT · mean
883ms
3 repeats · mean ± stddevview command →hide command ↑
Long chains of thought, 1k in and 4k out, a slow arrival rate of 0.2 because every request costs a lot of decode budget. Tests whether TTFT stays stable.
With 1024 input and 1024 output tokens at ten concurrent requests I measure an average of 24.14 tokens/s per user. The average wait for the first token is 0.88 seconds.
Long context
Wait at 25k context
With 25000 input tokens and 256 output tokens at ten concurrent requests the average wait for the first token is 17.73 seconds. After that the model generates an average of 11.9 tokens/s per user. The measurement alone does not show whether memory, prompt processing or scheduling causes the delay.
Peak load
Requests and wait under peak load
The run handles 0.57 requests/s at a configured arrival rate of 1.5 requests/s. 300 of 300 requests succeed. The p95 wait for the first token is 1.84 seconds.
Measurement coverage
Which measurements count
11 of 11 tests have a measurement that counts under the current Arena rules. No sanity check was recorded for this historical run. That proof is missing, even where the speed measurement counts under the current rules.