---
title: Qwen-3.5-9B (BF16) · DGX Spark Arena
canonical: https://djangodevreng.nl/en/arena/qwen-3-5-9b-bf16/
license: CC-BY-4.0
source: https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/qwen-3.5/qwen-3.5-9b/bf16/
source_revision: a948cec23b054c62fef595bad3cdb284066922ac
contract_version: 1
suite_version: 2026-08
minimum_request_success_rate: 0.99
sanity_status: captured_unreviewed
attribution: Django de Vreng, https://djangodevreng.nl
---

# Qwen-3.5-9B (BF16)

The best quality per parameter in this arena, and you pay for it in throughput.

Measured on 2026-08-07 with vLLM v0.26.0. In the chat test with ten concurrent requests this profile reaches 12.51 tokens/s per user, with an average of 1.82 seconds to the first token. At 25k context that average wait rises to 37.12 seconds. These runs measure speed and waiting, not the quality of the answers.

## Specs

- Vendor: Alibaba
- Architecture: Dense
- Parameters: 9B
- Precision: BF16
- Context: 256K
- Measured at: 128K
- VRAM: 18.0 GB
- Engine: vLLM v0.26.0
- Hardware: DGX Spark, NVIDIA GB10, 128 GB unified memory
- Model card: https://huggingface.co/Qwen/Qwen3.5-9B
## Quality (model cards)

Knowledge: MMLU-Pro; reasoning: GPQA-Diamond; coding: LiveCodeBench v6. Each value names its exact test. Other tests and missing measurements do not qualify for the total score. External figures are not our own evaluation of this precision.

https://huggingface.co/Qwen/Qwen3.5-9B

| Benchmark | Score |
| --- | --- |
| LiveCodeBench v6 | 65.6 |
| MMLU-Pro | 82.5 |
| GPQA-Diamond | 81.7 |
## Benchmarks on the DGX Spark

Only complete runs without a failed sanity check count. Open-loop tests require at least 99% request success. Raw results remain visible.

| Test | tokens/s per user | tokens/s total | TTFT (ms) | Request success | Ranking |
| --- | --- | --- | --- | --- | --- |
| 01 Translation missing: en.arena_copy.benches.chat.name | 12.51 | 123.0 | 1823.0 | — | Included |
| 02 Translation missing: en.arena_copy.benches.rag-8k.name | 10.25 | 84.0 | 11334.0 | — | Included |
| 03 Translation missing: en.arena_copy.benches.long-output.name | 12.55 | 123.0 | 637.0 | — | Included |
| 04 Translation missing: en.arena_copy.benches.multi-turn.name | 12.07 | 113.0 | 3131.0 | — | Included |
| 05 Translation missing: en.arena_copy.benches.big-context.name | 5.91 | 30.0 | 37117.0 | — | Included |
| 06 Translation missing: en.arena_copy.benches.concurrency-stress.name | 3.82 | 32.0 | 71438.0 | — | Included |
| 07 Translation missing: en.arena_copy.benches.office-baseline.name | 37.56666666666667 | 1239.7 | 1865.02 | 100.0% | Included |
| 08 Translation missing: en.arena_copy.benches.sharegpt.name | 9.161428571428571 | 128.26 | 250.96 | 100.0% | Included |
| 09 Translation missing: en.arena_copy.benches.reasoning.name | 6.092826086956522 | 280.27 | 530.55 | 100.0% | Included |
| 10 Translation missing: en.arena_copy.benches.monday-peak.name | 49.33896551724138 | 1430.83 | 1650.73 | 100.0% | Included |
| 11 Translation missing: en.arena_copy.benches.rate-sweep.name | — | — | — | 100.0% | Included |
## Notes

BF16 without quantisation, dense 9B, roughly 18 GB of weights. KV cache fp8, max_model_len 131072 (model-native 256k). Async scheduling on, prefix caching off, gpu-util 0.95. Runs on vLLM v0.20.1, older than the v23 runs elsewhere in this arena, so compare throughput figures with that in mind.

---

CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Django de Vreng, https://djangodevreng.nl.
Full arena: https://djangodevreng.nl/en/arena/
Raw runs (GitHub): https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/qwen-3.5/qwen-3.5-9b/bf16/
Source JSON: https://djangodevreng.nl/en/arena/qwen-3-5-9b-bf16/source.json
