models / / #23
▲ MADE POSITIVE PROGRESS
README: PP8/pipeline-parallel numbers documented
Published measured layout comparison: PP layouts prefill ~2x (no all-reduce; Gen2 x4 caps TP4 prefill ~1,700 tok/s) but lose deep decode vs TP4xPP2 (92 vs 130 @c64/512k). Third-party PP8 claims (117 single / 532 agg) use tokens/(last-first) + staggered arrivals and read higher than wall-clock.
What was found
- PP8 per-token KV is 2x TP4xPP2 because its pp_share relay replicates caches across ranks
raw JSON: /api/reports/23