{"model":"deepseek-4.1-flash","title":"Benchmark #1 - baseline (random 1024/256)","status":"positive","summary":"First local benchmark on our box, TP4xPP2, random 1024-in/256-out, server-counted: c1 21.71 out tok/s (TPOT 43.0 ms, DSpark acceptance 6.80% on random tokens), c8 74.60 (peak 80.0). TTFT includes per-shape JIT warmup.","occurred_at":"2026-09-11T03:24:00Z","config":{"serving":"TP4xPP2, DSpark k=5, fp8_ds_mla, 1M ctx, CUDA graphs","engine":"localhost/vllm-backport-v41:sm80 (vLLM 0.12.0-sm80 + PR#56201 + Ampere shims + PP relay)"},"benchmarks":[{"metric":"c1_decode_tok_s","value":21.71,"unit":"tok/s","context":"random 1024-in/256-out, TP4xPP2","note":"first local measurement"},{"metric":"agg_decode_c8_tok_s","value":74.6,"unit":"tok/s","context":"random 1024-in/256-out, TP4xPP2 (peak 80.0)"},{"metric":"dspark_accept_pct","value":6.8,"unit":"%","context":"random tokens - misleadingly low on random data"}]}