{"name":"gpt-oss-20b","arena":{"id":"gpt-oss-20b-mxfp4","name":"gpt-oss-20b","vram":14,"params":21,"vendor":"openai","quality":{"coding":null,"science":{"bench":"GPQA-Diamond","value":71.5},"knowledge":{"bench":"MMLU","value":85.3}},"contextK":128},"notes":{"en":"Native MXFP4 weights, no separate quantisation step. MoE with 3.6B active parameters out of 21B total. KV cache fp8, max_model_len 131072. Marlin MXFP4 MoE is selected automatically on SM121. The profile deliberately differs from the rest: gpu-util 0.90 instead of 0.95, max-num-seqs 16 and max-cudagraph-capture-size 2048, because the memory peak during CUDA graph capture takes the box down otherwise. Async scheduling on, prefix caching off.","fr":"Poids MXFP4 natifs, pas d'étape de quantisation séparée. MoE avec 3.6B de paramètres actifs sur 21B au total. Cache KV fp8, max_model_len 131072. Le MoE Marlin MXFP4 est choisi automatiquement sur SM121. Le profil diffère volontairement du reste : gpu-util 0.90 au lieu de 0.95, max-num-seqs 16 et max-cudagraph-capture-size 2048, car le pic mémoire pendant la capture du graphe CUDA fait tomber la machine sinon. Async scheduling activé, prefix caching désactivé.","nl":"Model-ID: openai/gpt-oss-20b. Container: vllm/vllm-openai:v0.26.0-aarch64-cu129-ubuntu2404. llama-benchy 0.4.0. Bronrun: 2026-08-14T11:45:36+02:00. KV-cache fp8; max_model_len 131072; gpu_memory_utilization 0,9. Prefix caching uit; async scheduling aan. Extra serverflags: --max-num-batched-tokens 8192 --max-cudagraph-capture-size 2048. Profielvariabelen: TIKTOKEN_ENCODINGS_BASE=/tiktoken_encodings. 11 van de 11 tests hebben een meting die volgens de huidige Arena-regels meetelt. De sanity-check is geslaagd. Dit is een controle op een kort gegenereerd antwoord, geen inhoudelijke evaluatie van de benchmarktaken."},"order":19,"engine":"vLLM v0.26.0","vramGb":14,"verdict":{"en":"Twenty-one billion parameters that behave like a four-billion model on throughput. Chat gives you 25.6 tokens per user with sub-second TTFT, and on the office baseline the queue stays at ten requests while it completes 0.29 of the 0.3 scheduled RPS. No other model above eight billion in this arena keeps that queue so short: Gemma-4 26B NVFP4 sits at sixteen, Qwen-3.6 27B at a hundred and sixty. That comes from combining MoE with 3.6B active parameters and native MXFP4 weights, so there is never an expensive dequantisation step in the path. On the ShareGPT replay the median TTFT lands at 118 milliseconds, faster than models a tenth its size. The brake is where you would expect it: at 25k context with ten concurrent users decode drops to 9.9 tokens per user and TTFT climbs to 27 seconds. For a team that chats, summarises and does tool calls all day, this is one of the few models on this box that is fast and smart at the same time.","fr":"Vingt-et-un milliards de paramètres qui se comportent comme un modèle de quatre milliards côté débit. En chat vous obtenez 25.6 tokens par utilisateur avec un TTFT sous la seconde, et sur la baseline bureau la file reste à dix requêtes pendant qu'il termine 0.29 des 0.3 RPS programmés. Aucun autre modèle au-dessus de huit milliards dans cette arène ne garde cette file aussi courte : Gemma-4 26B NVFP4 est à seize, Qwen-3.6 27B à cent soixante. Cela vient de la combinaison MoE avec 3.6B de paramètres actifs et de poids MXFP4 natifs, donc jamais d'étape de déquantisation coûteuse dans le chemin. Sur le replay ShareGPT la médiane TTFT atteint 118 millisecondes, plus rapide que des modèles dix fois plus petits. Le frein est là où on l'attend : à 25k de contexte avec dix utilisateurs simultanés le decode tombe à 9.9 tokens par utilisateur et le TTFT monte à 27 secondes. Pour une équipe qui chatte, résume et fait des tool calls toute la journée, c'est l'un des rares modèles sur cette machine à être à la fois rapide et intelligent.","nl":"Gemeten op 2026-08-14 met vLLM v0.26.0. In de chattest met tien gelijktijdige verzoeken haalt dit profiel 24,06 tokens/s per gebruiker, met gemiddeld 0,88 seconden tot het eerste token. Bij 25k context stijgt die gemiddelde wachttijd naar 26,28 seconden. Deze runs meten snelheid en wachttijd, niet de kwaliteit van de antwoorden."},"hardware":"DGX Spark, NVIDIA GB10, 128 GB unified memory","modelUrl":"https://huggingface.co/openai/gpt-oss-20b","provider":"OpenAI","precision":"MXFP4","parameters":"21B (3.6B actief)","provenance":{"modelId":"openai/gpt-oss-20b","runPath":"results/gpt-oss/gpt-oss-20b/mxfp4","validity":{"tests":{"01-chat":"complete","02-rag-8k":"complete","08-sharegpt":"complete","09-reasoning":"complete","04-multi-turn":"complete","11-rate-sweep":"complete","03-long-output":"complete","05-big-context":"complete","10-monday-peak":"complete","07-office-baseline":"complete","06-concurrency-stress":"complete"},"sanity":"passed"},"sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","servedName":"gpt-oss-20b-mxfp4","generatedAt":"2026-08-14T11:45:36+02:00","runRevision":"a948cec23b054c62fef595bad3cdb284066922ac","vllmVersion":"0.26.0","suiteVersion":"2026-08","repositoryUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks","contractVersion":1,"contractRevision":"a948cec23b054c62fef595bad3cdb284066922ac","commandProvenance":"reconstructed_from_raw_artifacts","llamaBenchyVersion":"0.4.0"},"architecture":"MoE","measuredContextK":128,"benchmarkResults":[{"id":"gpt-oss-20b-mxfp4-chat-1k","date":"2026-08-14","runs":3,"metrics":{"ttfrSec":{"mean":0.878,"stddev":0.335},"decodeTokensPerSecondTotal":{"mean":180.39,"stddev":44.74},"prefillTokensPerSecondTotal":{"mean":7705.55,"stddev":147.31},"decodeTokensPerSecondPerUser":{"mean":24.06,"stddev":6.36}},"sourceRun":"Test 01","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","workloadId":"chat-1k","sourceLabel":"openai/gpt-oss-20b mxfp4 op de DGX Spark","sourceStatus":"complete","sourceTestId":"01-chat","prefixCaching":false},{"id":"gpt-oss-20b-mxfp4-rag-8k","date":"2026-08-14","runs":3,"metrics":{"ttfrSec":{"mean":6.113,"stddev":3.046},"decodeTokensPerSecondTotal":{"mean":162.14,"stddev":15.13},"prefillTokensPerSecondTotal":{"mean":4920.4,"stddev":1614.68},"decodeTokensPerSecondPerUser":{"mean":22.03,"stddev":1.87}},"sourceRun":"Test 02","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","workloadId":"rag-8k","sourceLabel":"openai/gpt-oss-20b mxfp4 op de DGX Spark","sourceStatus":"complete","sourceTestId":"02-rag-8k","prefixCaching":false},{"id":"gpt-oss-20b-mxfp4-long-output-agents","date":"2026-08-14","runs":3,"metrics":{"ttfrSec":{"mean":0.287,"stddev":0.09},"decodeTokensPerSecondTotal":{"mean":204.02,"stddev":34.13},"prefillTokensPerSecondTotal":{"mean":3001.61,"stddev":2113.85},"decodeTokensPerSecondPerUser":{"mean":23.56,"stddev":4.53}},"sourceRun":"Test 03","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","workloadId":"long-output-agents","sourceLabel":"openai/gpt-oss-20b mxfp4 op de DGX Spark","sourceStatus":"complete","sourceTestId":"03-long-output","prefixCaching":false},{"id":"gpt-oss-20b-mxfp4-multi-turn-office","date":"2026-08-14","runs":3,"metrics":{"ttfrSec":{"mean":1.637,"stddev":0.66},"decodeTokensPerSecondTotal":{"mean":140.97,"stddev":52.06},"prefillTokensPerSecondTotal":{"mean":3222.05,"stddev":3332.08},"decodeTokensPerSecondPerUser":{"mean":24.98,"stddev":1.26}},"sourceRun":"Test 04","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","workloadId":"multi-turn-office","sourceLabel":"openai/gpt-oss-20b mxfp4 op de DGX Spark","sourceStatus":"complete","sourceTestId":"04-multi-turn","prefixCaching":false},{"id":"gpt-oss-20b-mxfp4-large-context-25k","date":"2026-08-14","runs":3,"metrics":{"ttfrSec":{"mean":26.283,"stddev":13.349},"decodeTokensPerSecondTotal":{"mean":44.19,"stddev":2},"prefillTokensPerSecondTotal":{"mean":4043.13,"stddev":472.79},"decodeTokensPerSecondPerUser":{"mean":9.79,"stddev":4.82}},"sourceRun":"Test 05","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","workloadId":"large-context-25k","sourceLabel":"openai/gpt-oss-20b mxfp4 op de DGX Spark","sourceStatus":"complete","sourceTestId":"05-big-context","prefixCaching":false},{"id":"gpt-oss-20b-mxfp4-concurrency-stress-25k","date":"2026-08-14","runs":3,"metrics":{"ttfrSec":{"mean":50.054,"stddev":27.118},"decodeTokensPerSecondTotal":{"mean":47.51,"stddev":0.41},"prefillTokensPerSecondTotal":{"mean":4219.88,"stddev":348.97},"decodeTokensPerSecondPerUser":{"mean":5.76,"stddev":3.47}},"sourceRun":"Test 06","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","workloadId":"concurrency-stress-25k","sourceLabel":"openai/gpt-oss-20b mxfp4 op de DGX Spark","sourceStatus":"complete","sourceTestId":"06-concurrency-stress","prefixCaching":false}],"openLoopResults":[{"id":"gpt-oss-20b-mxfp4-office-random-4k","date":"2026-08-14","metrics":{"e2eMs":{"p50":16080.42,"p90":37446.73,"p95":40713.69,"p99":46763.31,"mean":18124.77},"tpotMs":{"p50":43.52,"p90":57.33,"p95":71.17,"p99":167.91,"mean":48.07},"ttftMs":{"p50":764.7,"p90":1462.92,"p95":1648.75,"p99":2292.97,"mean":826.65}},"sourceRun":"Test 07","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","scenarioId":"office-random-4k","achievedRps":0.291,"sourceStatus":"complete","sourceTestId":"07-office-baseline","configuredRps":0.3,"prefixCaching":false,"totalRequests":200,"successfulRequests":200,"totalTokenThroughput":1315.2,"peakConcurrentRequests":14},{"id":"gpt-oss-20b-mxfp4-sharegpt-replay","date":"2026-08-14","metrics":{"e2eMs":{"p50":3102.33,"p90":13357.68,"p95":15729.19,"p99":23211.64,"mean":5298.94},"tpotMs":{"p50":28.33,"p90":34.2,"p95":36.77,"p99":39.8,"mean":52.41},"ttftMs":{"p50":120.18,"p90":168.12,"p95":182.75,"p99":201.14,"mean":124.47}},"sourceRun":"Test 08","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","scenarioId":"sharegpt-replay","achievedRps":0.299,"sourceStatus":"complete","sourceTestId":"08-sharegpt","configuredRps":0.3,"prefixCaching":false,"totalRequests":250,"successfulRequests":250,"totalTokenThroughput":138.1,"peakConcurrentRequests":9},{"id":"gpt-oss-20b-mxfp4-reasoning","date":"2026-08-14","metrics":{"e2eMs":{"p50":195197.01,"p90":344824.84,"p95":378473.35,"p99":391138.9,"mean":190341.78},"tpotMs":{"p50":54.92,"p90":58.8,"p95":80.12,"p99":168.5,"mean":59.35},"ttftMs":{"p50":256.27,"p90":356.84,"p95":367.02,"p99":400.34,"mean":260.93}},"sourceRun":"Test 09","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","scenarioId":"reasoning","achievedRps":0.093,"sourceStatus":"complete","sourceTestId":"09-reasoning","configuredRps":0.2,"prefixCaching":false,"totalRequests":50,"successfulRequests":50,"totalTokenThroughput":436.3,"peakConcurrentRequests":34},{"id":"gpt-oss-20b-mxfp4-monday-peak-random-4k","date":"2026-08-14","metrics":{"e2eMs":{"p50":37034.27,"p90":73674.3,"p95":76826.33,"p99":81645.71,"mean":38745.73},"tpotMs":{"p50":82.74,"p90":89.59,"p95":94.48,"p99":122.49,"mean":82.56},"ttftMs":{"p50":748.93,"p90":1327.54,"p95":1936.71,"p99":4738.32,"mean":903.65}},"sourceRun":"Test 10","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","scenarioId":"monday-peak-random-4k","achievedRps":0.614,"sourceStatus":"complete","sourceTestId":"10-monday-peak","configuredRps":1.5,"prefixCaching":false,"totalRequests":300,"successfulRequests":300,"totalTokenThroughput":2772.33,"peakConcurrentRequests":28}],"rateSweep":{"date":"2026-08-14","capacity":{"slo2000":{"outputTps":107.9,"ttftP95Ms":1732.3,"achievedRps":0.279,"configuredRps":0.3,"peakConcurrent":10},"slo5000":{"outputTps":324.5,"ttftP95Ms":3289.3,"achievedRps":0.863,"configuredRps":1,"peakConcurrent":74},"slo10000":{"outputTps":324.5,"ttftP95Ms":3289.3,"achievedRps":0.863,"configuredRps":1,"peakConcurrent":74}},"sourceRun":"Test 11","sourceUrl":"https://github.com/djangodevreng/dgx-spark-benchmarks/tree/a948cec23b054c62fef595bad3cdb284066922ac/results/gpt-oss/gpt-oss-20b/mxfp4/","abortedAtRps":null,"queueKneeRps":0.7,"sourceStatus":"complete","sourceTestId":"11-rate-sweep","totalRequests":850,"successfulRequests":850}}