NEWS · MODELS · #976
AWS packages WhisperX DLC for SageMaker to deliver speaker‑labeled, per‑word timestamps
AWS provides a WhisperX Deep Learning Container (DLC) for Amazon SageMaker that bundles OpenAI Whisper with wav2vec2 forced alignment and speaker diarization to produce per-word timestamps and speaker labels. The GPU-ready container conforms to SageMaker’s serving contract, supports real-time and asynchronous endpoints, and outputs json, verbose_json, srt, and vtt formats (example image tag: whisperx:3.8.6-cu128-amzn2023-sagemaker).
KEY POINTS
- AWS provides a WhisperX Deep Learning Container (DLC) for Amazon SageMaker that bundles OpenAI Whisper with wav2vec2 forced alignment and speaker diarization to produce per-word timestamps and speaker labels.
- The GPU-ready container conforms to SageMaker’s serving contract, supports real-time and asynchronous endpoints, and outputs json, verbose_json, srt, and vtt formats (example image tag: whisperx:3.8.6-cu128-amzn2023-sagemaker).
- Packaging WhisperX as a SageMaker DLC makes production-grade, speaker‑labeled ASR with precise per‑word timestamps directly deployable in AWS, easing captioning, compliance, and analytics workflows.
WHY IT MATTERS
Packaging WhisperX as a SageMaker DLC makes production-grade, speaker‑labeled ASR with precise per‑word timestamps directly deployable in AWS, easing captioning, compliance, and analytics workflows.