Tech Meridian ← LIVE FEED
RU

NEWS · RESEARCH · #355

arXiv:2609.16454v1 — Fine-tuning reduces mode collapse and over-dispersion in LLMs

This arXiv preprint analyzes when large language models exhibit mode collapse (under-diversity) or over-dispersion and shows that sufficient supervised fine-tuning (SFT) data moves model output diversity toward the target distribution. The paper derives a bias–variance decomposition for the gap in collision probability, proves an absolute gap bound proportional to the square root of the KL divergence to the target, and validates these claims on synthetic languages, four LLMs fine-tuned on human survey responses, and the CodeNet code dataset.

KEY POINTS

  1. This arXiv preprint analyzes when large language models exhibit mode collapse (under-diversity) or over-dispersion and shows that sufficient supervised fine-tuning (SFT) data moves model output diversity toward the target distribution.
  2. The paper derives a bias–variance decomposition for the gap in collision probability, proves an absolute gap bound proportional to the square root of the KL divergence to the target, and validates these claims on synthetic languages, four LLMs fine-tuned on human survey responses, and the CodeNet code dataset.
  3. The paper gives a theoretical and empirical account that supervised fine-tuning can correct diversity miscalibration in LLMs, guiding how much SFT data is needed and how to evaluate diversity.

WHY IT MATTERS

The paper gives a theoretical and empirical account that supervised fine-tuning can correct diversity miscalibration in LLMs, guiding how much SFT data is needed and how to evaluate diversity.

SOURCES & TIMELINE

1