RESEARCH · RESEARCH · #1255
Representational Simplicity and Circuit Size Dissociate via Adversarial Training (arXiv:2609.35890v1)
This new arXiv paper uses adversarial continual training starting from the same GPT-2 Small checkpoint to test whether representational/attributional simplicity (measured via sparse-autoencoder decomposability and SAE feature engagement) predicts smaller causal circuits for a fixed faithfulness level. Results: adversarially robust models are more SAE-decomposable and use fewer SAE features for task attribution, while circuit size depends on the faithfulness threshold—standard models lead or tie below ~85% faithfulness, but robust models require substantially fewer edges at high faithfulness (90%, 95%); trends were checked across a parametric sweep and a second corpus.
KEY POINTS
- This new arXiv paper uses adversarial continual training starting from the same GPT-2 Small checkpoint to test whether representational/attributional simplicity (measured via sparse-autoencoder decomposability and SAE feature engagement) predicts smaller causal circuits for a fixed faithfulness level.
- Results: adversarially robust models are more SAE-decomposable and use fewer SAE features for task attribution, while circuit size depends on the faithfulness threshold—standard models lead or tie below ~85% faithfulness, but robust models require substantially fewer edges at high faithfulness (90%, 95%); trends were checked across a parametric sweep and a second corpus.
- This matters because it provides the first controlled empirical test linking representational/attributional simplicity to causal circuit size, challenging assumptions used in mechanistic interpretability and robustness evaluation.
WHY IT MATTERS
This matters because it provides the first controlled empirical test linking representational/attributional simplicity to causal circuit size, challenging assumptions used in mechanistic interpretability and robustness evaluation.