vllm-loop model-improvement progress tracker junk-tokens
models / / #41
● AM STUCK

AutoRound W4A16 correctness gate FAIL

Current gate FAIL (correctness): PP4 eager AR loads with the SM80 Triton fallback (BF16 KV, no cache/speculation) but five identical cold 1K-token greedy requests diverge as early as output position 18; forced single-pass sparse-KV also diverges (position 19). No performance claim admitted.

When (PT)2026-09-08 15:30 PT
ServingTP1/PP4 eager AR (GPUs 4,5,7,8; GPU6 never selected)
Enginewtdcode/vllm-backport, DeepGEMM SM80 fix @ ffa38541; Intel/GLM-5.3-Flash-W4A16-AutoRound rev 5eee1846
KV

What didn't work

  • Single-pass sparse-KV (VLLM_TRITON_MLA_FORCE_KV_SPLITS=1): still diverges @token 19
  • Persistent CUDA top-k disabled (portable top_k_per_row_decode): still diverges @tokens 18-40
  • MoE auto->Marlin reproduces the same nondeterministic token family; Triton MoE cannot load AutoGPTQ W4A16 expert bias (rejected before load)

What's next

  • Capture exact per-step SM80 Triton indexer-logit + selected-top-k hashes and PP boundary state to locate the first divergent tensor
  • Scheduler/KV/MTP/PP6/DFlash changes are not justified before that evidence

raw JSON: /api/reports/41