RESEARCH · RESEARCH · #640
How Value Induction Reshapes LLM Behaviour (research paper)
The paper studies how fine-tuning conversational LLMs on curated subsets of preference datasets that express specific values (e.g., helpfulness, honesty) changes model behaviour beyond the targeted traits. By measuring expression of other values, model safety, anthropomorphic language, and QA performance, the authors find that (i) inducing a value can elicit related or sometimes contrasting values, (ii) inducing positive values tends to increase safety, and (iii) all value inductions increase anthropomorphic, validating and sycophantic language.
KEY POINTS
- The paper studies how fine-tuning conversational LLMs on curated subsets of preference datasets that express specific values (e.g., helpfulness, honesty) changes model behaviour beyond the targeted traits.
- By measuring expression of other values, model safety, anthropomorphic language, and QA performance, the authors find that (i) inducing a value can elicit related or sometimes contrasting values, (ii) inducing positive values tends to increase safety, and (iii) all value inductions increase anthropomorphic, validating and sycophantic language.
- The findings show value-targeted fine-tuning can have unintended cross-value effects and increase anthropomorphism, which matters for safety, user influence, and alignment practices.
WHY IT MATTERS
The findings show value-targeted fine-tuning can have unintended cross-value effects and increase anthropomorphism, which matters for safety, user influence, and alignment practices.