vllm-loop model-improvement progress tracker junk-tokens
models / / #31
▲ MADE POSITIVE PROGRESS

LMCache blocker root-caused AND resolved - serving V4.1

In-process LMCacheConnectorV1 is single-group by construction -> HMA disabled -> vLLM cannot unify V4.1 KV specs (MLA state_content_bytes=584 + SWA compressor group) and startup fails. Correct path: LMCacheMPConnector (group-aware, SupportsHMA) + lmcache MP server WITH GPU access +…

When (PT)2026-09-12 03:56 PT
ServingTP1xPP6 (7,7,7,7,7,5), util 0.95-0.96, DSpark k=5, fp8_ds_mla
Enginelocalhost/vllm-backport-v41:sm80 Schaka v0.13.0 kit (c1b0907b overlay)
KV

What was found

  • Real-OpenCode c1 median ~30 tok/s end-to-end (range 8-48), driven by reasoning tokens + tool-call steps

What was fixed

  • LMCache path: LMCacheMPConnector + GPU-capable MP server + retention knob - cold/hit parity proven

raw JSON: /api/reports/31