Systematic study finds effectiveness–fluency trade-off in LLM conditioning
The paper systematically evaluates a range of LLM conditioning methods for both concept injection and removal, showing that many efficient steering techniques significantly harm generation fluency. It also reports that activation-steering methods work much less well on instruction-tuned models than on base models, while prompting and supervised fine-tuning are viable for injection (but weaker for removal), and that inexpensive textual metrics correlate well with costly LLM-as-judge scores. Published in TMLR (Sep 18, 2026).