{"model":"deepseek-4.1-flash","title":"Hourly update #18 — 32K master table online (PP8)","status":"positive","summary":"Built the one-command 32K master table; PP8 baseline c1 20.6 / c4 53.4 / c8 78.9 Gen tok/s, 512K@c16 98.9 tok/s; util 0.96 OOMs 512K so lowered to 0.94","config":{"serving":"TP1xPP8 (5,5,5,5,5,5,5,5), max-num-seqs 32, util 0.94","engine":"vllm-backport-v41:sm80-v13 (v0.13.0 + Schaka DSV41 kit + patches 0007/0008)","kv":"fp8_ds_mla, 7,859,683-token pool (7.50x concurrency @1M), block 128","notes":"TRITON_MLA_SPARSE_DSV41, DSpark k=5, VLLM_SPARSE_DECODE_HEAD_BLOCK_SIZE=32"},"benchmarks":[{"metric":"c1_gen_tok_s","label":"c1 Gen tok/s @ 32K","value":20.6,"unit":"tok/s","context":"32K in / 256 out, 16 real-text prompts, ignore_eos, PP8 util 0.94"},{"metric":"c4_gen_tok_s","label":"c4 Gen tok/s @ 32K","value":53.4,"unit":"tok/s","context":"32K in / 256 out, 16 prompts"},{"metric":"c8_gen_tok_s","label":"c8 Gen tok/s @ 32K","value":78.9,"unit":"tok/s","context":"32K in / 256 out, 16 prompts"},{"metric":"c8_c1_scaling","label":"c8/c1 scaling @ 32K","value":3.82,"unit":"x","context":"derived from the 32K table"},{"metric":"agg_512k_c16_tok_s","label":"512K aggregate decode @ c16","value":98.9,"unit":"tok/s","context":"shared 487K prefix, pure decode, output 256"},{"metric":"accept_pct_c1","label":"DSpark acceptance @ c1 32K","value":23.4,"unit":"%","context":"server /metrics, real text"},{"metric":"kv_pool_tokens","label":"KV token capacity (PP8)","value":7859683,"unit":"tokens","context":"util 0.94"}],"found":["vllm bench `custom` dataset needs pandas (vllm[bench]) and the model has no chat template -\u003e must pass --skip-chat-template","util 0.96 with max-num-seqs 32 OOMs the 512K prefill (torch.OutOfMemoryError in a GEMM, GPU 63.01/63.39 GiB)"],"fixed":["one-command 32K master-table harness (scripts/master_table.py): grid -\u003e per-cell vllm bench serve -\u003e JSON DB -\u003e markdown table","folded in VLLM_SPARSE_DECODE_HEAD_BLOCK_SIZE=32 (patches/0008)","lowered PP8 util 0.96 -\u003e 0.94 so the 512K cell is reliable (pool 7.86M tokens, 7.50x @1M)"],"tried":["random vs sonnet vs custom dataset for the master table (random gives ~1% acceptance -\u003e unrepresentative; custom real-text gives ~23%)","util 0.96 vs 0.94 for the 512K prefill"],"failed":["util 0.96 + SEQS 32: 512K prefill OOM (regression-gate failure, not a small regression)","sonnet dataset: requires a Jinja chat template this tokenizer does not ship","custom dataset without pandas: ImportError"],"next":["attack the c4/c8 ITL p99 tail (~0.6-2.3 s vs ~100 ms median) via chunked-prefill/scheduling (first loop hypothesis)","bake pandas into the serving image so the master table never depends on an ad-hoc install","re-run the 32K table at util 0.94 as the frozen baseline; then one hypothesis per loop"],"notes":"First cycle of the new 32K-master-table loop. PP8 is the serving layout; PP6 remains an A/B reference. 500 tok/s @512k still far off (98.9 at c16 shared-prefix).","links":[{"label":"hourly #18 report","url":"https://serve.context.roboalch.com/a/lnk_8DBA63JSXVGTX67VYP02MCHFW0"}]}