spon-extension

Extending the SPON paper (arXiv:2512.12744) with layer-wise allocation and mechanistic interpretability — where to place SPON biases and what they encode.

spon-extension extends Xu/Gao/Weng/Ma’s SPON (resting-neuron bias distribution) paper with two new experiments — where to place SPON biases and what those biases encode mechanistically.

  • Attention o_proj SPON recovers 77–80% of damage vs 23% for MLP down_proj (3–4× better, module-isolated).
  • Top-50% of layers capture 85% of the benefit — Pareto-optimal at half the parameter cost.
  • Different training configs converge to the same biases (0.943 avg cosine similarity across shared layers).
  • Fully reproducible on a single A100 in ~37 minutes.

View the repository