vllm-loop model-improvement progress tracker junk-tokens
models / / #2
▲ MADE POSITIVE PROGRESS

MILESTONE — first known SM80/CMP170HX boot

deepseek-v4.1-flash healthy on 8x CMP 170HX (TP8, eager, fp8_ds_mla, Engram in pinned host RAM; init engine 143.5s, ~5 tok/s) - first known V4.1 boot on SM80/CMP170HX anywhere.

When (PT)2026-09-10 11:40 PT
ServingTP8xPP1, DSpark k=5, fp8_ds_mla, 1M ctx
Enginelocal/vllm-dsv41:sm80 @ dsv41-feat@e47aa780
KV

What was found

  • No SM80 attention class existed -> added TRITON_MLA_SPARSE_DSV41 Ampere class (reuses ROCm Triton sparse-MLA)
  • fp8e4nv Triton kernels need uint8 shims on sm80 (indexer / fused_indexer_q / cache / engram)
  • _flashmla_C required by SWA tile-scheduler -> gated on availability

What didn't work

  • Greedy output incoherent and nondeterministic - forward pass numerically wrong (correctness open)

raw JSON: /api/reports/2