vllm-loop model-improvement progress tracker junk-tokens
models / / #39
▲ MADE POSITIVE PROGRESS

PP8/c11 live deployment receipt

NVFP4 target + DFlash2 K=7 live on TP1/PP8 (GPUs 1,0,2,3,4,5,7,8; GPU6 excluded): 256K max request, max-num-seqs 11, GPU cache 3,200 blocks = 2,964,172 tokens (~11.3x 256K concurrency); FULL_DECODE_ONLY graphs with DFlash buckets 8-88; healthy LMCache sidecar. Explicitly not a claim that c11…

When (PT)2026-09-03 11:00 PT
ServingTP1/PP8, max-num-seqs 11, scheduler batch cap 4,608
EnginevLLM 0.1.dev20051+g487ecf187, image local/glm53-lmcache-hma:6950fb8ae5a8
KVNVFP4 target + DFlash2 K=7; 2,964,172-token cache

Benchmarks

MetricValueΔ vs previousUnitContextNote
kv_pool_tokens (kv_pool_tokens) 2964172 tokens first point tokens PP8, 3,200 blocks, 256K max request

raw JSON: /api/reports/39