How Value Induction Reshapes LLM Behaviour (research paper)
The paper studies how fine-tuning conversational LLMs on curated subsets of preference datasets that express specific values (e.g., helpfulness, honesty) changes model behaviour beyond the targeted traits. By measuring expression of other values, model safety, anthropomorphic language, and QA performance, the authors find that (i) inducing a value can elicit related or sometimes contrasting values, (ii) inducing positive values tends to increase safety, and (iii) all value inductions increase anthropomorphic, validating and sycophantic language.