models / / #31
▲ MADE POSITIVE PROGRESS
LMCache blocker root-caused AND resolved - serving V4.1
In-process LMCacheConnectorV1 is single-group by construction -> HMA disabled -> vLLM cannot unify V4.1 KV specs (MLA state_content_bytes=584 + SWA compressor group) and startup fails. Correct path: LMCacheMPConnector (group-aware, SupportsHMA) + lmcache MP server WITH GPU access +…
What was found
- Real-OpenCode c1 median ~30 tok/s end-to-end (range 8-48), driven by reasoning tokens + tool-call steps
What was fixed
- LMCache path: LMCacheMPConnector + GPU-capable MP server + retention knob - cold/hit parity proven
raw JSON: /api/reports/31