Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #10919

adversarial training

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

Representational Simplicity and Circuit Size Dissociate via Adversarial Training (arXiv:2609.35890v1)

This new arXiv paper uses adversarial continual training starting from the same GPT-2 Small checkpoint to test whether representational/attributional simplicity (measured via sparse-autoencoder decomposability and SAE feature engagement) predicts smaller causal circuits for a fixed faithfulness level. Results: adversarially robust models are more SAE-decomposable and use fewer SAE features for task attribution, while circuit size depends on the faithfulness threshold—standard models lead or tie below ~85% faithfulness, but robust models require substantially fewer edges at high faithfulness (90%, 95%); trends were checked across a parametric sweep and a second corpus.

7.0