Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1644

SSRFT: Safe-Role Internalization for Robust LLM Safety Alignment (arXiv:2610.07023v1)

The paper introduces SSRFT (Supervised Safe-Role Fine-Tuning), a safety-alignment framework that teaches models to internalize a predefined safe role by building a Safe-Role Q&A (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description; role-consistent responses are synthesized, validated, and expanded into diverse scenarios. Experiments on multiple Base and Instruct models report that SSRFT yields more robust and generalizable safety alignment than standard SFT, with greater resistance to prefilling attacks, better generalization to unseen jailbreaks, reduced over-refusal on benign queries, and preserved general capabilities.

KEY POINTS

  1. The paper introduces SSRFT (Supervised Safe-Role Fine-Tuning), a safety-alignment framework that teaches models to internalize a predefined safe role by building a Safe-Role Q&A (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description; role-consistent responses are synthesized, validated, and expanded into diverse scenarios.
  2. Experiments on multiple Base and Instruct models report that SSRFT yields more robust and generalizable safety alignment than standard SFT, with greater resistance to prefilling attacks, better generalization to unseen jailbreaks, reduced over-refusal on benign queries, and preserved general capabilities.
  3. This matters because internalizing a safety-oriented role could provide a more robust and generalizable alternative to refusal-based alignment, reducing exploitability and unnecessary refusals.

WHY IT MATTERS

This matters because internalizing a safety-oriented role could provide a more robust and generalizable alternative to refusal-based alignment, reducing exploitability and unnecessary refusals.

SOURCES & TIMELINE

1