---
title: Gemma-4-26B-A4B-it (NVFP4) · DGX Spark Arena
canonical: https://djangodevreng.nl/fr/arena/gemma-4-26b-a4b-it-nvfp4-v23/
license: CC-BY-4.0
source: https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gemma-4/gemma-4-26b-a4b-it/nvfp4-v23/
source_revision: a948cec23b054c62fef595bad3cdb284066922ac
contract_version: 1
suite_version: 2026-08
minimum_request_success_rate: 0.99
sanity_status: not_recorded
attribution: Django de Vreng, https://djangodevreng.nl
---

# Gemma-4-26B-A4B-it (NVFP4)

Mesuré le 2026-08-06 avec vLLM v0.26.0. Dans le test de chat avec dix requêtes simultanées, ce profil atteint 21,1 tokens/s par utilisateur, avec en moyenne 1,28 secondes jusqu&#39;au premier token. À 25k de contexte, cette attente moyenne monte à 38,8 secondes. Ces runs mesurent la vitesse et l&#39;attente, pas la qualité des réponses.

## Spécifications

- Fournisseur: NVIDIA (re-quant van Google)
- Architecture: MoE
- Paramètres: 26B-A4B
- Précision: NVFP4
- Contexte: 256K
- Contexte testé: 128K
- VRAM: 24.0 GB
- Moteur: vLLM v0.26.0
- Matériel: DGX Spark, NVIDIA GB10, 128 GB unified memory
- Fiche modèle: https://huggingface.co/nvidia/Gemma-4-26B-A4B-NVFP4
## Qualité (fiches modèles)

Connaissances : MMLU-Pro ; raisonnement : GPQA-Diamond ; code : LiveCodeBench v6. Chaque chiffre indique son test exact. Les autres tests et les mesures absentes ne permettent pas de score total. Ces chiffres externes ne sont pas notre propre évaluation de cette précision.

https://huggingface.co/nvidia/Gemma-4-26B-A4B-NVFP4

| Benchmark | Score |
| --- | --- |
| LiveCodeBench v6 | 79.8 |
| MMLU-Pro | 84.8 |
| GPQA-Diamond | 79.9 |
## Benchmarks sur le DGX Spark

Seules les exécutions complètes sans échec de cohérence comptent. Les tests en boucle ouverte exigent au moins 99 % de réussite. Les résultats bruts restent visibles.

| Test | tokens/s par utilisateur | tokens/s au total | TTFT (ms) | Réussite | Classement |
| --- | --- | --- | --- | --- | --- |
| 01 Translation missing: fr.arena_copy.benches.chat.name | 21.1 | 141.0 | 1278.0 | — | Inclus |
| 02 Translation missing: fr.arena_copy.benches.rag-8k.name | 16.06 | 123.0 | 8173.0 | — | Inclus |
| 03 Translation missing: fr.arena_copy.benches.long-output.name | 22.94 | 161.0 | 382.0 | — | Inclus |
| 04 Translation missing: fr.arena_copy.benches.multi-turn.name | 19.5 | 172.0 | 2017.0 | — | Inclus |
| 05 Translation missing: fr.arena_copy.benches.big-context.name | 7.24 | 33.0 | 38796.0 | — | Inclus |
| 06 Translation missing: fr.arena_copy.benches.concurrency-stress.name | 3.99 | 33.0 | 73575.0 | — | Inclus |
| 07 Translation missing: fr.arena_copy.benches.office-baseline.name | 81.5325 | 1304.52 | 1056.91 | 100.0% | Inclus |
| 08 Translation missing: fr.arena_copy.benches.sharegpt.name | 13.379 | 133.79 | 153.51 | 100.0% | Inclus |
| 09 Translation missing: fr.arena_copy.benches.reasoning.name | 8.911363636363637 | 392.1 | 338.5 | 100.0% | Inclus |
| 10 Translation missing: fr.arena_copy.benches.monday-peak.name | 68.90214285714286 | 1929.26 | 957.85 | 100.0% | Inclus |
| 11 Translation missing: fr.arena_copy.benches.rate-sweep.name | — | — | — | 100.0% | Inclus |
## Notes

Re-quant NVIDIA NVFP4 sur vLLM v0.23.0. KV-cache fp8, prefix caching désactivé, gpu-memory-utilization 0.85. Suite complète exécutée le 2026-06-23.

---

CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Django de Vreng, https://djangodevreng.nl.
Arena complète: https://djangodevreng.nl/fr/arena/
Exécutions brutes (GitHub): https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gemma-4/gemma-4-26b-a4b-it/nvfp4-v23/
Source JSON: https://djangodevreng.nl/fr/arena/gemma-4-26b-a4b-it-nvfp4-v23/source.json
