glm-5.3-flash
▲ MADE POSITIVE PROGRESS
Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but… (2026-09-08 17:55 PT)
Benchmark progress over time
One chart per metric. Hover a point for date (PT), value, and what happened; click to open the report.
Status log
| When (PT) | Status | Summary | |
|---|---|---|---|
| 2026-09-08 17:55 PT | ▲ positive | Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but… | report → |
| 2026-09-08 15:30 PT | ● stuck | Current gate FAIL (correctness): PP4 eager AR loads with the SM80 Triton fallback (BF16 KV, no cache/speculation) but five identical cold 1K-token greedy requests diverge as early as output position 18; forced single-pass sparse-KV also diverges (position 19). No performance claim admitted. | report → |
| 2026-09-08 05:00 PT | ▼ negative | Scripted depth/concurrency campaign complete with NO PP6 production promotion: K7 passes the correctness + token-accounting matrix, but c3 depth is the limiting production-shaped behavior (c3/c1 retained rate 0.87 @32k -> 0.10 @192k); K3 failed its eager admission gate (branch stops before… | report → |
| 2026-09-03 11:00 PT | ▲ positive | NVFP4 target + DFlash2 K=7 live on TP1/PP8 (GPUs 1,0,2,3,4,5,7,8; GPU6 excluded): 256K max request, max-num-seqs 11, GPU cache 3,200 blocks = 2,964,172 tokens (~11.3x 256K concurrency); FULL_DECODE_ONLY graphs with DFlash buckets 8-88; healthy LMCache sidecar. Explicitly not a claim that c11… | report → |
| 2026-09-02 05:00 PT | ▲ positive | PP7 DFlash2 source-and-KV audit and PP8 native-cache audit + final audit complete (reports: 2026-09-02 series). Date-precision only for this backfill entry. | report → |
| 2026-09-01 05:00 PT | ▲ positive | GLM-5.3-Flash PP4 DFlash2 inference bring-up investigated; investigation + dashboard baselines established (report: 2026-09-01-glm53-flash-pp4-inference-investigation). Date-precision only for this backfill entry. | report → |
All reports (6)
| When (PT) | Status | Title | Summary |
|---|---|---|---|
| 2026-09-08 17:55 PT | ▲ | Reproducible source/image lane established | Last ~18h delta: SM80 startup blocker found and fixed (GLM sparse indexer invoked DeepGEMM metadata on unsupported SM80 -> capability-gated Triton fp8 MQA-logits fallback, source ffa38541); pinned NVCR CUDA base digests; KPool 11/11 and AutoRound W4A16 2/2 CPU gates pass repeatedly. PP4 loads but… |
| 2026-09-08 15:30 PT | ● | AutoRound W4A16 correctness gate FAIL | Current gate FAIL (correctness): PP4 eager AR loads with the SM80 Triton fallback (BF16 KV, no cache/speculation) but five identical cold 1K-token greedy requests diverge as early as output position 18; forced single-pass sparse-KV also diverges (position 19). No performance claim admitted. |
| 2026-09-08 05:00 PT | ▼ | PP6 depth/concurrency campaign conclusion | Scripted depth/concurrency campaign complete with NO PP6 production promotion: K7 passes the correctness + token-accounting matrix, but c3 depth is the limiting production-shaped behavior (c3/c1 retained rate 0.87 @32k -> 0.10 @192k); K3 failed its eager admission gate (branch stops before… |
| 2026-09-03 11:00 PT | ▲ | PP8/c11 live deployment receipt | NVFP4 target + DFlash2 K=7 live on TP1/PP8 (GPUs 1,0,2,3,4,5,7,8; GPU6 excluded): 256K max request, max-num-seqs 11, GPU cache 3,200 blocks = 2,964,172 tokens (~11.3x 256K concurrency); FULL_DECODE_ONLY graphs with DFlash buckets 8-88; healthy LMCache sidecar. Explicitly not a claim that c11… |
| 2026-09-02 05:00 PT | ▲ | PP7/PP8 DFlash2 final audits | PP7 DFlash2 source-and-KV audit and PP8 native-cache audit + final audit complete (reports: 2026-09-02 series). Date-precision only for this backfill entry. |
| 2026-09-01 05:00 PT | ▲ | PP4 DFlash2 bring-up investigation | GLM-5.3-Flash PP4 DFlash2 inference bring-up investigated; investigation + dashboard baselines established (report: 2026-09-01-glm53-flash-pp4-inference-investigation). Date-precision only for this backfill entry. |