vllm-loop model-improvement progress tracker junk-tokens
models / / #10
▲ MADE POSITIVE PROGRESS

Benchmark #5 - decode at depth (512k)

Shared-prefix + prefix-caching + ignore_eos, TP4xPP2: per-token decode FLAT with depth (TPOT 8-10 ms @512k, sparse CSA2), but aggregate saturates: 128k c32 214.8, 512k c64 129.6 (26% of the 500 goal). DSpark acceptance ~42% at depth. Cold 512k prefill 326 s (1,608 tok/s).

When (PT)2026-09-10 23:13 PT
ServingTP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs
Enginelocalhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)
KV

Benchmarks

MetricValueΔ vs previousUnitContextNote
agg_decode_128k_tok_s (agg_decode_128k_tok_s) 214.8 tok/s first point tok/s 128k ctx c32
agg_decode_512k_tok_s (agg_decode_512k_tok_s) 129.6 tok/s first point tok/s 512k ctx c64 26% of the 500 tok/s goal
max_num_seqs_lift (128k c16 after raising max-num-seqs) 180 tok/s first point tok/s 128k c16 with max-num-seqs 8 -> 32 (129 -> 180)
prefill_tok_s (prefill_tok_s) 1608 tok/s -29.5% tok/s cold 512k prefill, 326 s

raw JSON: /api/reports/10