{"model":"deepseek-4.1-flash","title":"Benchmark #5 - decode at depth (512k)","status":"positive","summary":"Shared-prefix + prefix-caching + ignore_eos, TP4xPP2: per-token decode FLAT with depth (TPOT 8-10 ms @512k, sparse CSA2), but aggregate saturates: 128k c32 214.8, 512k c64 129.6 (26% of the 500 goal). DSpark acceptance ~42% at depth. Cold 512k prefill 326 s (1,608 tok/s).","occurred_at":"2026-09-11T06:13:00Z","config":{"serving":"TP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs","engine":"localhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)"},"benchmarks":[{"metric":"agg_decode_128k_tok_s","value":214.8,"unit":"tok/s","context":"128k ctx c32"},{"metric":"agg_decode_512k_tok_s","value":129.6,"unit":"tok/s","context":"512k ctx c64","note":"26% of the 500 tok/s goal"},{"metric":"prefill_tok_s","value":1608,"unit":"tok/s","context":"cold 512k prefill, 326 s"},{"metric":"max_num_seqs_lift","label":"128k c16 after raising max-num-seqs","value":180,"unit":"tok/s","context":"128k c16 with max-num-seqs 8 -\u003e 32 (129 -\u003e 180)"}]}