models / / #3
▲ MADE POSITIVE PROGRESS
Correctness root cause localized
Per-branch probe: attention 'o' absmax ~2.0 (sparse MLA works) but _o_proj output absmax ~0.013 - a ~150x shrink; residual goes FFN-only and explodes (0.6 -> 3.5e29 at KV-source layers). One bug: fused inverse-RoPE + wo_a einsum + wo_b.
What was found
- o absmax ~2.0, _o_proj absmax ~0.013 -> ~150x shrink of the attention contribution
What didn't work
- Ruled out: tokenization/template, Engram (DSV41_ZERO_ENGRAM), wo_a layout, backend selection, NaN
raw JSON: /api/reports/3