Troubleshooting of whisperX: Difference between revisions

Tag: wikieditor
 
(3 intermediate revisions by the same user not shown)
Line 20: Line 20:
# Enter <code>export HF_TOKEN='your Hugging Face token'</code>  
# Enter <code>export HF_TOKEN='your Hugging Face token'</code>  
# Run bash command
# Run bash command
=== WhisperX Speaker Diarization Issue ===
'''Problem''': Speaker diarization failed on an interview transcript — over 500 segments were all labeled `SPEAKER_00`, even though it was clearly a two-person Q&A interview. The output looked like a single person talking.
'''Solution'''
Installed the new whisperx 3.8.6 (torch 2.8 cu126) in a venv. Diarization automatically switched to use the new-generation `pyannote/speaker-diarization-community-1` model.
'''Version Info'''
Old environment: whisperX 3.1.1


=== Repeated Same Dialog Issue ===
=== Repeated Same Dialog Issue ===
Line 161: Line 172:
# fr = French
# fr = French


Supported language codes can be found in the Whisper tokenizer documentation at: https://github.com/openai/whisper/blob/main/whisper/tokenizer.py
Supported language codes can be found in the [https://github.com/openai/whisper/blob/main/whisper/tokenizer.py Whisper tokenizer documentation] & [https://github.com/m-bain/whisperX/blob/main/EXAMPLES.md whisperX/EXAMPLES.md at main · m-bain/whisperX]


== whisperX Transcript File Format Guide ==
== whisperX Transcript File Format Guide ==
Line 173: Line 184:
If using for the first time, it's recommended to directly open the srt format
If using for the first time, it's recommended to directly open the srt format


== Further reading ==
* [https://medium.com/@planetoid/how-to-add-punctuation-to-whisper-transcripts-using-ai-619362c9160c How to Add Punctuation to Whisper Transcripts Using AI | Medium]


[[Category: Generative AI]] [[Category: Software]] [[Category: Revised with LLMs]]
[[Category: Generative AI]] [[Category: Software]] [[Category: Revised with LLMs]]