Jun 13Vibe with Hermes Agent — Bengaluru · RSVP
ToolsMCPBlogResearchCommunityStar on GitHub
All Tools
W
OtherFreeOpen Source

WHISPERX

Word-level timestamps and speaker diarization for Whisper

BSD-2-Clause

ABOUT

Standard Whisper transcription provides segment-level timestamps without word-level precision, speaker identification, or robust voice activity detection — gaps that limit its use in production subtitle generation, meeting transcription, and accessibility workflows. WhisperX solves this by adding phoneme-based forced alignment for millisecond-precise word timestamps, speaker diarization through voice embedding clustering, and integrated VAD for silence detection, delivering production-ready, speaker-attributed transcripts with accurate word-level timing.

INSTALL
pip install whisperx

INTEGRATION GUIDE

1. Generate precise word-level subtitles with speaker identification for video content 2. Transcribe multi-speaker meetings and conversations with accurate speaker diarization 3. Build accessibility tools requiring millisecond-accurate speech-to-text alignment 4. Process call center recordings with speaker attribution and silence-based segmentation 5. Create searchable transcripts with word-level timestamps for media archiving

TAGS

speech-recognitionaudiotranscriptiondiarizationasrwhisperalignmentsubtitles