RESEARCH · RESEARCH · #629
Dynamically Scaled Activation Steering (DSAS) adaptively modulates steering in generative models
DSAS is a method-agnostic framework that decouples when to steer from how to steer by computing context-dependent scaling factors to modulate the strength of existing activation-steering transformations across layers and inputs. The authors report that DSAS improves the trade-off between toxicity mitigation and utility preservation, adds minimal compute overhead, improves interpretability by highlighting tokens that require steering, and can be jointly optimized with steering functions; the paper was accepted to the UniReps workshop at NeurIPS 2025 and the authors say code will be made available on GitHub.
KEY POINTS
- DSAS is a method-agnostic framework that decouples when to steer from how to steer by computing context-dependent scaling factors to modulate the strength of existing activation-steering transformations across layers and inputs.
- The authors report that DSAS improves the trade-off between toxicity mitigation and utility preservation, adds minimal compute overhead, improves interpretability by highlighting tokens that require steering, and can be jointly optimized with steering functions; the paper was accepted to the UniReps workshop at NeurIPS 2025 and the authors say code will be made available on GitHub.
- By applying steering only when and where needed, DSAS can improve safety-utility trade-offs and make steering behavior more interpretable across model types.
WHY IT MATTERS
By applying steering only when and where needed, DSAS can improve safety-utility trade-offs and make steering behavior more interpretable across model types.