spon-extension
Extending the SPON paper (arXiv:2512.12744) with layer-wise allocation and mechanistic interpretability — where to place SPON biases and what they encode.
spon-extension extends Xu/Gao/Weng/Ma’s SPON (resting-neuron bias distribution) paper with two new experiments — where to place SPON biases and what those biases encode mechanistically.
- Attention
o_projSPON recovers 77–80% of damage vs 23% for MLPdown_proj(3–4× better, module-isolated). - Top-50% of layers capture 85% of the benefit — Pareto-optimal at half the parameter cost.
- Different training configs converge to the same biases (0.943 avg cosine similarity across shared layers).
- Fully reproducible on a single A100 in ~37 minutes.