RESEARCH · RESEARCH · #967
Compressing streaming neural audio encoders via latent-space distillation
The paper proposes distilling streaming audio tokenizers by training a student encoder to regress the teacher’s pre-quantizer latent representations with a squared-error loss, using a single affine layer to handle width mismatch. At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without fine-tuning, and outperforms an independently trained tokenizer of the same capacity by 3.9% relative; the recipe applies to tokenizers pretrained alone or jointly with a language model.
KEY POINTS
- The paper proposes distilling streaming audio tokenizers by training a student encoder to regress the teacher’s pre-quantizer latent representations with a squared-error loss, using a single affine layer to handle width mismatch.
- At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without fine-tuning, and outperforms an independently trained tokenizer of the same capacity by 3.9% relative; the recipe applies to tokenizers pretrained alone or jointly with a language model.
- Smaller, accurate tokenizers reduce memory, power and latency pressure for always-on on-device dictation systems that compete with sparsely activated language-model experts in DRAM.
WHY IT MATTERS
Smaller, accurate tokenizers reduce memory, power and latency pressure for always-on on-device dictation systems that compete with sparsely activated language-model experts in DRAM.