models / / #20
▲ MADE POSITIVE PROGRESS
c1 optimization - TP8xPP1 wins at every depth
Single-stream TP8xPP1 (zero pipeline hops): 98.1/95.0/92.6/88.4 tok/s @8/512/4k/32k vs TP4xPP2 89.4/84.2/87.6/54.1 - decisive at 32k. Accuracy gates pass (needle 4/4, coding 5/5). Tradeoff: 512k aggregate 113 vs 130. TP8xPP1 becomes the served config for c1-first.
Benchmarks
| Metric | Value | Δ vs previous | Unit | Context | Note |
|---|---|---|---|---|---|
c1_decode_32k_tok_s (c1_decode_32k_tok_s) |
88.4 tok/s | first point | tok/s | TP8xPP1 @32k prompt (TP4xPP2 was 54.1) | |
c1_decode_tok_s (c1_decode_tok_s) |
98.1 tok/s | -6.6% | tok/s | TP8xPP1, 8-tok prompt (best of 3) |
What was found
- Under load the GPUs draw ~100 W of 250 W at 1470 MHz - c1 is memory/latency-bound, not power-bound
- DSpark acceptance 93.6% on predictable content (len 4.68) vs ~35-45% on open prose
raw JSON: /api/reports/20