Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #967

Compressing streaming neural audio encoders via latent-space distillation

The paper proposes distilling streaming audio tokenizers by training a student encoder to regress the teacher’s pre-quantizer latent representations with a squared-error loss, using a single affine layer to handle width mismatch. At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without fine-tuning, and outperforms an independently trained tokenizer of the same capacity by 3.9% relative; the recipe applies to tokenizers pretrained alone or jointly with a language model.

KEY POINTS

  1. The paper proposes distilling streaming audio tokenizers by training a student encoder to regress the teacher’s pre-quantizer latent representations with a squared-error loss, using a single affine layer to handle width mismatch.
  2. At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without fine-tuning, and outperforms an independently trained tokenizer of the same capacity by 3.9% relative; the recipe applies to tokenizers pretrained alone or jointly with a language model.
  3. Smaller, accurate tokenizers reduce memory, power and latency pressure for always-on on-device dictation systems that compete with sparsely activated language-model experts in DRAM.

WHY IT MATTERS

Smaller, accurate tokenizers reduce memory, power and latency pressure for always-on on-device dictation systems that compete with sparsely activated language-model experts in DRAM.

SOURCES & TIMELINE

1