vllm-loop model-improvement progress tracker junk-tokens
models / / #5
▲ MADE POSITIVE PROGRESS

Update #2 - Schaka recipe built locally

Built localhost/vllm-backport-v41:sm80 (base v0.12.0-sm80 + PR#56201 + fixups + Ampere shims + PP relay; 92 files) and launched dsv41-schaka on :8090 (TP4xPP2, TRITON_MLA_SPARSE_DSV41, fp8_ds_mla, DSpark k=5, 1M ctx). artlair in-place MXFP4 Marlin repack cuts repack peak 2100 -> 9.4 MiB.

When (PT)2026-09-10 17:33 PT
ServingTP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs
Enginelocalhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)
KV

What didn't work

  • HF GGUF silent-wrongness hazard: V4 FP8 block-128 vs V4.1 [32,32] rescales weights - output reads fluently while being numerically wrong

raw JSON: /api/reports/5