vllm-loop model-improvement progress tracker junk-tokens
models / / #17
▲ MADE POSITIVE PROGRESS

PP relay correctness fix adopted - prefill triples

Adopted Schaka's relay fix (topk_indices on the hop + index_k_cache guard): TP1xPP6 prefill 5.6k @32k / 7.2k @128k. TP4xPP2 remains decode-optimal (512k c64 129.6, correct 12/12). TP2xPP4 / TP1xPP8 still boot but needle 0/3 (relay cannot handle TP splits) - rejected.

When (PT)2026-09-11 08:09 PT
ServingTP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs
Enginelocalhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)
KV

Benchmarks

MetricValueΔ vs previousUnitContextNote
prefill_tok_s (prefill_tok_s) 5600 tok/s -2.4% tok/s TP1xPP6 @32k (7.2k @128k) - ~2.5-3x TP4xPP2

What didn't work

  • TP2xPP4 and TP1xPP8: needle 0/3 after relay fix - arbitrary kv-group TP splits still unsupported

raw JSON: /api/reports/17