vllm-loop model-improvement progress tracker junk-tokens
models / / #34
▲ MADE POSITIVE PROGRESS

PP8 brought up - 2.75x the KV pool of PP6

Layout-resolver intersection fix makes PP8 work: 9.35M-token KV pool (2.75x PP6), 8.9x concurrency @1M; 512k aggregate still ~60 tok/s @c32 -> step-cost bound. Leads: vLLM #56120 (SM80 Triton knobs), SGLang #38646 (NVFP4 sparse-MLA, 384 B/tok, SM120-only).

When (PT)2026-09-12 08:18 PT
ServingTP1xPP8 (Zanooda 13-patch stack), fp8_ds_mla
Enginezanooda/vllm-sm80-ds41f:v41-sm80 + KV-layout intersection fix
KV

Benchmarks

MetricValueΔ vs previousUnitContextNote
agg_decode_512k_tok_s (agg_decode_512k_tok_s) 60 tok/s 9.9% tok/s PP8 c32 @512k - step-cost bound
kv_pool_tokens (kv_pool_tokens) 9350000 tokens 174.7% tokens PP8 (2.75x PP6)

raw JSON: /api/reports/34