Measured on 2026-08-09 with vLLM v0.26.0. In the chat test with ten concurrent requests this profile reaches 42.31 tokens/s per user, with an average of 0.86 seconds to the first token. At 25k context that average wait rises to 15.14 seconds. These runs measure speed and waiting, not the quality of the answers.
—
Suitability score (indicative)
—
Throughput tok/s
8 GB
VRAM
0/11
Reliable benches
Hugging Face →·vLLM v0.26.0·DGX Spark, NVIDIA GB10, 128 GB unified memory·Measured with context 128K·Last measured 9 August 2026
02Vendor quality · reference only
Knowledge: MMLU-Pro; reasoning: GPQA-Diamond; coding: LiveCodeBench v6. Each value names its exact test. Other tests and missing measurements do not qualify for the total score. External figures are not our own evaluation of this precision. Source: vendor model card ↗
3/3
Coverage
70.7
MMLU
53.4
GPQA-Diamond
54.8
LiveCodeBench
03Performance · BF16
Decode throughput · total t/s · c=10
■ BF16
1k ctx BF16—
8k ctx BF16—
4k+turn BF16—
25k ctx BF16—
04Test suite · 11 benchmarks
6 closed-loop tests with llama-benchy and 5 open-loop tests with vllm bench serve. Only complete runs without failed sanity checks are eligible; open-loop tests require at least 99% successful requests. Raw results remain visible. Expand “view command” for the reconstructed command. Methodology →
01 · llama-benchyclosed-loop
Chat
Short prompt, long answer. The shape that has to feel like ordinary chat; TTFT decides whether it feels snappy.
pp (prompt)
1024
tg (gen)
1024
depth
0
concurrency
1 · 5 · 10
repeats
3
Tokens/sec · per user
42.3t/s
TTFT · mean
864ms
Sanity check failed
3 repeats · mean ± stddevview command →hide command ↑
Long chains of thought, 1k in and 4k out, a slow arrival rate of 0.2 because every request costs a lot of decode budget. Tests whether TTFT stays stable.
The chat measurement is excluded because the sanity check failed. The tokens/s present therefore prove no usable model output.
Long context
Wait at 25k context
The test with 25k input tokens is excluded as well. The raw results sit with the test, but they are no valid basis for a performance recommendation.
Peak load
Requests and wait under peak load
The peak measurement is excluded by the failed sanity check. A successful HTTP request is not the same as a usable answer.
Measurement coverage
Which measurements count
0 of 11 tests have a measurement that counts under the current Arena rules. The sanity check failed. This run is therefore excluded from the ranking; the raw figures stay visible for inspection only. Aborted: 11-rate-sweep.