---
title: Gemma-4-26B-A4B-it (NVFP4) · DGX Spark Arena
canonical: https://djangodevreng.nl/en/arena/gemma-4-26b-a4b-it-nvfp4-v23/
license: CC-BY-4.0
source: https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gemma-4/gemma-4-26b-a4b-it/nvfp4-v23/
source_revision: a948cec23b054c62fef595bad3cdb284066922ac
contract_version: 1
suite_version: 2026-08
minimum_request_success_rate: 0.99
sanity_status: not_recorded
attribution: Django de Vreng, https://djangodevreng.nl
---

# Gemma-4-26B-A4B-it (NVFP4)

Measured on 2026-08-06 with vLLM v0.26.0. In the chat test with ten concurrent requests this profile reaches 21.1 tokens/s per user, with an average of 1.28 seconds to the first token. At 25k context that average wait rises to 38.8 seconds. These runs measure speed and waiting, not the quality of the answers.

## Specs

- Vendor: NVIDIA (re-quant van Google)
- Architecture: MoE
- Parameters: 26B-A4B
- Precision: NVFP4
- Context: 256K
- Measured at: 128K
- VRAM: 24.0 GB
- Engine: vLLM v0.26.0
- Hardware: DGX Spark, NVIDIA GB10, 128 GB unified memory
- Model card: https://huggingface.co/nvidia/Gemma-4-26B-A4B-NVFP4
## Quality (model cards)

Knowledge: MMLU-Pro; reasoning: GPQA-Diamond; coding: LiveCodeBench v6. Each value names its exact test. Other tests and missing measurements do not qualify for the total score. External figures are not our own evaluation of this precision.

https://huggingface.co/nvidia/Gemma-4-26B-A4B-NVFP4

| Benchmark | Score |
| --- | --- |
| LiveCodeBench v6 | 79.8 |
| MMLU-Pro | 84.8 |
| GPQA-Diamond | 79.9 |
## Benchmarks on the DGX Spark

Only complete runs without a failed sanity check count. Open-loop tests require at least 99% request success. Raw results remain visible.

| Test | tokens/s per user | tokens/s total | TTFT (ms) | Request success | Ranking |
| --- | --- | --- | --- | --- | --- |
| 01 Translation missing: en.arena_copy.benches.chat.name | 21.1 | 141.0 | 1278.0 | — | Included |
| 02 Translation missing: en.arena_copy.benches.rag-8k.name | 16.06 | 123.0 | 8173.0 | — | Included |
| 03 Translation missing: en.arena_copy.benches.long-output.name | 22.94 | 161.0 | 382.0 | — | Included |
| 04 Translation missing: en.arena_copy.benches.multi-turn.name | 19.5 | 172.0 | 2017.0 | — | Included |
| 05 Translation missing: en.arena_copy.benches.big-context.name | 7.24 | 33.0 | 38796.0 | — | Included |
| 06 Translation missing: en.arena_copy.benches.concurrency-stress.name | 3.99 | 33.0 | 73575.0 | — | Included |
| 07 Translation missing: en.arena_copy.benches.office-baseline.name | 81.5325 | 1304.52 | 1056.91 | 100.0% | Included |
| 08 Translation missing: en.arena_copy.benches.sharegpt.name | 13.379 | 133.79 | 153.51 | 100.0% | Included |
| 09 Translation missing: en.arena_copy.benches.reasoning.name | 8.911363636363637 | 392.1 | 338.5 | 100.0% | Included |
| 10 Translation missing: en.arena_copy.benches.monday-peak.name | 68.90214285714286 | 1929.26 | 957.85 | 100.0% | Included |
| 11 Translation missing: en.arena_copy.benches.rate-sweep.name | — | — | — | 100.0% | Included |
## Notes

NVIDIA NVFP4 re-quant on vLLM v0.23.0. KV-cache fp8, prefix caching off, gpu-memory-utilization 0.85. Full suite run on 2026-06-23.

---

CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Django de Vreng, https://djangodevreng.nl.
Full arena: https://djangodevreng.nl/en/arena/
Raw runs (GitHub): https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gemma-4/gemma-4-26b-a4b-it/nvfp4-v23/
Source JSON: https://djangodevreng.nl/en/arena/gemma-4-26b-a4b-it-nvfp4-v23/source.json
