vllm-loop model-improvement progress tracker junk-tokens
models / / #19
▼ MADE NEGATIVE PROGRESS

Hourly #7 - per-stream TPOT is NOT flat with depth

Correction: the earlier 'flat TPOT' note read the aggregate field. True per-stream: 4k 106 ms | 128k 150 ms | 512k c32 294 ms | 512k c64 504 ms (~5x). 500 tok/s @512k needs ~5x lower per-token cost at depth or ~4x concurrency headroom. Upstream: #46994 MERGED (PP MTP + stale topk_indices_buffer…

When (PT)2026-09-11 11:35 PT
ServingTP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs
Enginelocalhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)
KV

What didn't work

  • 'TPOT flat with depth' was an aggregate-field misread - per-stream cost grows ~5x to 512k

raw JSON: /api/reports/19