---
title: Nemotron-3-Nano-4B (BF16) · DGX Spark Arena
canonical: https://djangodevreng.nl/en/arena/nemotron-3-nano-4b-bf16/
license: CC-BY-4.0
source: https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/nemotron-3/nemotron-3-nano-4b/bf16/
source_revision: a948cec23b054c62fef595bad3cdb284066922ac
contract_version: 1
suite_version: 2026-08
minimum_request_success_rate: 0.99
sanity_status: not_recorded
attribution: Django de Vreng, https://djangodevreng.nl
---

# Nemotron-3-Nano-4B (BF16)

Measured on 2026-08-06 with vLLM v0.26.0. In the chat test with ten concurrent requests this profile reaches 24.14 tokens/s per user, with an average of 0.88 seconds to the first token. At 25k context that average wait rises to 17.73 seconds. These runs measure speed and waiting, not the quality of the answers.

## Specs

- Vendor: NVIDIA
- Architecture: Dense
- Parameters: 4B
- Precision: BF16
- Context: 256K
- Measured at: 48K
- VRAM: 8.0 GB
- Engine: vLLM v0.26.0
- Hardware: DGX Spark, NVIDIA GB10, 128 GB unified memory
- Model card: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
## Quality (model cards)

Knowledge: MMLU-Pro; reasoning: GPQA-Diamond; coding: LiveCodeBench v6. Each value names its exact test. Other tests and missing measurements do not qualify for the total score. External figures are not our own evaluation of this precision.

https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16

| Benchmark | Score |
| --- | --- |
| LiveCodeBench | 51.8 |
| MMLU-Pro | 18.1 |
| GPQA-Diamond | 51.3 |
## Benchmarks on the DGX Spark

Only complete runs without a failed sanity check count. Open-loop tests require at least 99% request success. Raw results remain visible.

| Test | tokens/s per user | tokens/s total | TTFT (ms) | Request success | Ranking |
| --- | --- | --- | --- | --- | --- |
| 01 Translation missing: en.arena_copy.benches.chat.name | 24.14 | 224.0 | 883.0 | — | Included |
| 02 Translation missing: en.arena_copy.benches.rag-8k.name | 19.96 | 161.0 | 5546.0 | — | Included |
| 03 Translation missing: en.arena_copy.benches.long-output.name | 25.22 | 158.0 | 345.0 | — | Included |
| 04 Translation missing: en.arena_copy.benches.multi-turn.name | 23.24 | 211.0 | 1458.0 | — | Included |
| 05 Translation missing: en.arena_copy.benches.big-context.name | 11.9 | 62.0 | 17732.0 | — | Included |
| 06 Translation missing: en.arena_copy.benches.concurrency-stress.name | 7.49 | 66.0 | 34588.0 | — | Included |
| 07 Translation missing: en.arena_copy.benches.office-baseline.name | 88.02533333333334 | 1320.38 | 800.46 | 100.0% | Included |
| 08 Translation missing: en.arena_copy.benches.sharegpt.name | 13.109 | 131.09 | 129.09 | 100.0% | Included |
| 09 Translation missing: en.arena_copy.benches.reasoning.name | 10.254999999999999 | 430.71 | 280.71 | 100.0% | Included |
| 10 Translation missing: en.arena_copy.benches.monday-peak.name | 89.3303448275862 | 2590.58 | 823.11 | 100.0% | Included |
| 11 Translation missing: en.arena_copy.benches.rate-sweep.name | — | — | — | 100.0% | Included |
## Notes

Model-ID: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16. Container: vllm/vllm-openai:v0.26.0-aarch64-cu129-ubuntu2404. llama-benchy 0.4.0. Source run: 2026-08-06T05:29:55+02:00. KV-cache fp8_e4m3; max_model_len 49152; gpu_memory_utilization 0,9. Prefix caching off; async scheduling on. Extra server flags: --trust-remote-code. 11 of 11 tests have a measurement that counts under the current Arena rules. No sanity check was recorded for this historical run. That proof is missing, even where the speed measurement counts under the current rules.

---

CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Django de Vreng, https://djangodevreng.nl.
Full arena: https://djangodevreng.nl/en/arena/
Raw runs (GitHub): https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/nemotron-3/nemotron-3-nano-4b/bf16/
Source JSON: https://djangodevreng.nl/en/arena/nemotron-3-nano-4b-bf16/source.json
