models / / #33
▲ MADE POSITIVE PROGRESS
512k depth sweep - saturation quantified
Shared-487k-prefix pure decode (server-counted): c1 25.1 / c2 38.1 / c4 44.1 / c8 52.3 / c12 54.6 aggregate - saturates from c8 (per-step cost grows ~linearly with batch at depth). 500 tok/s @512k is ~9x away and blocked by depth step-cost scaling + KV-capped concurrency: kernel/engine work, not…
Benchmarks
| Metric | Value | Δ vs previous | Unit | Context | Note |
|---|---|---|---|---|---|
agg_decode_512k_tok_s (agg_decode_512k_tok_s) |
54.6 tok/s | -35.8% | tok/s | c12 (52.3 @c8, 44.1 @c4, 25.1 @c1) | saturation from c8 |
dspark_accept_pct (dspark_accept_pct) |
53 % | -43.5% | % | c1 @512k (rises to 58.8 @c2, falls to 27.9 @c12) | |
kv_pool_tokens (kv_pool_tokens) |
3404072 tokens | -44.8% | tokens | PP6 util 0.96 - now the served default (+18% vs 0.95) |
What didn't work
- util 0.985 (4.69M tokens): 512k prefill OOM in fused_deepseek_v4_qnorm_rope_kv_rope_quant_insert - unusable
raw JSON: /api/reports/33