Keep It CALM — limits of global unsafety and a local counterfactual safeguard for text-to-image models
This arXiv preprint analyzes the geometric limits of using a single reusable global safety signal (an "unsafe" direction or toxic subspace) for text-to-image generation, showing a coverage–selectivity trade-off: compact unsafe subspaces miss heterogeneous unsafe semantics, while broader removals distort benign prompts. The paper introduces CALM (Counterfactual Adaptive Local Modulation), a training-free, prompt-local method that uses matched unsafe–benign anchors to minimally edit violating token representations and suppress residual unsafe components, and reports improved unsafe-content suppression while preserving benign utility.