Troubleshooting of whisperX: Difference between revisions
Tag: wikieditor |
|||
| (3 intermediate revisions by the same user not shown) | |||
| Line 20: | Line 20: | ||
# Enter <code>export HF_TOKEN='your Hugging Face token'</code> | # Enter <code>export HF_TOKEN='your Hugging Face token'</code> | ||
# Run bash command | # Run bash command | ||
=== WhisperX Speaker Diarization Issue === | |||
'''Problem''': Speaker diarization failed on an interview transcript — over 500 segments were all labeled `SPEAKER_00`, even though it was clearly a two-person Q&A interview. The output looked like a single person talking. | |||
'''Solution''' | |||
Installed the new whisperx 3.8.6 (torch 2.8 cu126) in a venv. Diarization automatically switched to use the new-generation `pyannote/speaker-diarization-community-1` model. | |||
'''Version Info''' | |||
Old environment: whisperX 3.1.1 | |||
=== Repeated Same Dialog Issue === | === Repeated Same Dialog Issue === | ||
| Line 161: | Line 172: | ||
# fr = French | # fr = French | ||
Supported language codes can be found in the | Supported language codes can be found in the [https://github.com/openai/whisper/blob/main/whisper/tokenizer.py Whisper tokenizer documentation] & [https://github.com/m-bain/whisperX/blob/main/EXAMPLES.md whisperX/EXAMPLES.md at main · m-bain/whisperX] | ||
== whisperX Transcript File Format Guide == | == whisperX Transcript File Format Guide == | ||
| Line 173: | Line 184: | ||
If using for the first time, it's recommended to directly open the srt format | If using for the first time, it's recommended to directly open the srt format | ||
== Further reading == | |||
* [https://medium.com/@planetoid/how-to-add-punctuation-to-whisper-transcripts-using-ai-619362c9160c How to Add Punctuation to Whisper Transcripts Using AI | Medium] | |||
[[Category: Generative AI]] [[Category: Software]] [[Category: Revised with LLMs]] | [[Category: Generative AI]] [[Category: Software]] [[Category: Revised with LLMs]] | ||