{"model":"deepseek-4.1-flash","title":"c1 optimization - TP8xPP1 wins at every depth","status":"positive","summary":"Single-stream TP8xPP1 (zero pipeline hops): 98.1/95.0/92.6/88.4 tok/s @8/512/4k/32k vs TP4xPP2 89.4/84.2/87.6/54.1 - decisive at 32k. Accuracy gates pass (needle 4/4, coding 5/5). Tradeoff: 512k aggregate 113 vs 130. TP8xPP1 becomes the served config for c1-first.","occurred_at":"2026-09-11T20:34:00Z","config":{"serving":"TP8xPP1, DSpark k=5, fp8_ds_mla, 1M ctx","engine":"local/vllm-dsv41:sm80 @ dsv41-feat@e47aa780"},"benchmarks":[{"metric":"c1_decode_tok_s","value":98.1,"unit":"tok/s","context":"TP8xPP1, 8-tok prompt (best of 3)"},{"metric":"c1_decode_32k_tok_s","value":88.4,"unit":"tok/s","context":"TP8xPP1 @32k prompt (TP4xPP2 was 54.1)"}],"found":["Under load the GPUs draw ~100 W of 250 W at 1470 MHz - c1 is memory/latency-bound, not power-bound","DSpark acceptance 93.6% on predictable content (len 4.68) vs ~35-45% on open prose"]}