RESEARCH · RESEARCH · #992
Semi-supervised federated ASR: online pseudo-labels with server update stabilization
The paper studies semi-supervised federated learning for automatic speech recognition and shows that two coupled design axes—the teacher that generates pseudo-labels (online per-client vs. broadcast global) and a server-side labeled-data 'anchor' that continues training between rounds—determine stability and final performance. With the recommended stabilization (interleaved server training and tuned augmentation/batch settings), their recipe outperforms the strongest prior method on 9 of 11 test pairs, reducing the gap to fully supervised FL by 20.8% on average in-domain and 10.0% cross-domain.
KEY POINTS
- The paper studies semi-supervised federated learning for automatic speech recognition and shows that two coupled design axes—the teacher that generates pseudo-labels (online per-client vs.
- broadcast global) and a server-side labeled-data 'anchor' that continues training between rounds—determine stability and final performance.
- With the recommended stabilization (interleaved server training and tuned augmentation/batch settings), their recipe outperforms the strongest prior method on 9 of 11 test pairs, reducing the gap to fully supervised FL by 20.8% on average in-domain and 10.0% cross-domain.
WHY IT MATTERS
This provides a practical, empirically validated recipe for stabilizing semi-supervised federated ASR—identifying when online pseudo-labeling helps and how server-side labeled updates are essential to prevent divergence—making SSFL more viable in speech applications.