Pretrained ASR pseudo-labeling for noisy police audio uses LLM judging (arXiv:2609.30469v1)
This arXiv paper evaluates pseudo-labeling to adapt foundation ASR models (Whisper and Qwen3-ASR) to very noisy Broadcast Police Communication (BPC) corpora from Baltimore and Chicago. The authors find internal confidence metrics (log-probabilities, STAR scores) poorly separate good from bad pseudo-labels, propose an external LLM-as-a-judge filtering paradigm that discards contextually implausible transcripts and reduces WER of pseudo-labeled training sets, and introduce a cross-model pseudo-labeling approach where one model is fine-tuned on pseudo-labels from another.