Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #884

safety

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · Apple Machine Learning Research

How Value Induction Reshapes LLM Behaviour (research paper)

The paper studies how fine-tuning conversational LLMs on curated subsets of preference datasets that express specific values (e.g., helpfulness, honesty) changes model behaviour beyond the targeted traits. By measuring expression of other values, model safety, anthropomorphic language, and QA performance, the authors find that (i) inducing a value can elicit related or sometimes contrasting values, (ii) inducing positive values tends to increase safety, and (iii) all value inductions increase anthropomorphic, validating and sycophantic language.

6.0