vllm-loop model-improvement progress tracker junk-tokens

glm-5.3-flash

▲ MADE POSITIVE PROGRESS

Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but… (2026-09-08 17:55 PT)

Benchmark progress over time

One chart per metric. Hover a point for date (PT), value, and what happened; click to open the report.

Status log

When (PT)StatusSummary
2026-09-08 17:55 PT ▲ positive Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but… report →
2026-09-08 15:30 PT ● stuck Current gate FAIL (correctness): PP4 eager AR loads with the SM80 Triton fallback (BF16 KV, no cache/speculation) but five identical cold 1K-token greedy requests diverge as early as output position 18; forced single-pass sparse-KV also diverges (position 19). No performance claim admitted. report →
2026-09-08 05:00 PT ▼ negative Scripted depth/concurrency campaign complete with NO PP6 production promotion: K7 passes the correctness + token-accounting matrix, but c3 depth is the limiting production-shaped behavior (c3/c1 retained rate 0.87 @32k -> 0.10 @192k); K3 failed its eager admission gate (branch stops before… report →
2026-09-03 11:00 PT ▲ positive NVFP4 target + DFlash2 K=7 live on TP1/PP8 (GPUs 1,0,2,3,4,5,7,8; GPU6 excluded): 256K max request, max-num-seqs 11, GPU cache 3,200 blocks = 2,964,172 tokens (~11.3x 256K concurrency); FULL_DECODE_ONLY graphs with DFlash buckets 8-88; healthy LMCache sidecar. Explicitly not a claim that c11… report →
2026-09-02 05:00 PT ▲ positive PP7 DFlash2 source-and-KV audit and PP8 native-cache audit + final audit complete (reports: 2026-09-02 series). Date-precision only for this backfill entry. report →
2026-09-01 05:00 PT ▲ positive GLM-5.3-Flash PP4 DFlash2 inference bring-up investigated; investigation + dashboard baselines established (report: 2026-09-01-glm53-flash-pp4-inference-investigation). Date-precision only for this backfill entry. report →

All reports (6)

When (PT)StatusTitleSummary
2026-09-08 17:55 PT Reproducible source/image lane established Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but…
2026-09-08 15:30 PT AutoRound W4A16 correctness gate FAIL Current gate FAIL (correctness): PP4 eager AR loads with the SM80 Triton fallback (BF16 KV, no cache/speculation) but five identical cold 1K-token greedy requests diverge as early as output position 18; forced single-pass sparse-KV also diverges (position 19). No performance claim admitted.
2026-09-08 05:00 PT PP6 depth/concurrency campaign conclusion Scripted depth/concurrency campaign complete with NO PP6 production promotion: K7 passes the correctness + token-accounting matrix, but c3 depth is the limiting production-shaped behavior (c3/c1 retained rate 0.87 @32k -> 0.10 @192k); K3 failed its eager admission gate (branch stops before…
2026-09-03 11:00 PT PP8/c11 live deployment receipt NVFP4 target + DFlash2 K=7 live on TP1/PP8 (GPUs 1,0,2,3,4,5,7,8; GPU6 excluded): 256K max request, max-num-seqs 11, GPU cache 3,200 blocks = 2,964,172 tokens (~11.3x 256K concurrency); FULL_DECODE_ONLY graphs with DFlash buckets 8-88; healthy LMCache sidecar. Explicitly not a claim that c11…
2026-09-02 05:00 PT PP7/PP8 DFlash2 final audits PP7 DFlash2 source-and-KV audit and PP8 native-cache audit + final audit complete (reports: 2026-09-02 series). Date-precision only for this backfill entry.
2026-09-01 05:00 PT PP4 DFlash2 bring-up investigation GLM-5.3-Flash PP4 DFlash2 inference bring-up investigated; investigation + dashboard baselines established (report: 2026-09-01-glm53-flash-pp4-inference-investigation). Date-precision only for this backfill entry.