Troubleshooting of whisperX: Difference between revisions

No edit summary
Tag: wikieditor
 
(5 intermediate revisions by the same user not shown)
Line 20: Line 20:
# Enter <code>export HF_TOKEN='your Hugging Face token'</code>  
# Enter <code>export HF_TOKEN='your Hugging Face token'</code>  
# Run bash command
# Run bash command
=== WhisperX Speaker Diarization Issue ===
'''Problem''': Speaker diarization failed on an interview transcript — over 500 segments were all labeled `SPEAKER_00`, even though it was clearly a two-person Q&A interview. The output looked like a single person talking.
'''Solution'''
Installed the new whisperx 3.8.6 (torch 2.8 cu126) in a venv. Diarization automatically switched to use the new-generation `pyannote/speaker-diarization-community-1` model.
'''Version Info'''
Old environment: whisperX 3.1.1


=== Repeated Same Dialog Issue ===
=== Repeated Same Dialog Issue ===
Line 106: Line 117:
* '''Compatibility:''' Ensure that the versions of CUDA, cuDNN, and PyTorch are compatible with each other. Refer to the [https://pytorch.org/get-started/previous-versions/ PyTorch documentation] for version compatibility details.
* '''Compatibility:''' Ensure that the versions of CUDA, cuDNN, and PyTorch are compatible with each other. Refer to the [https://pytorch.org/get-started/previous-versions/ PyTorch documentation] for version compatibility details.
* '''Virtual Environments:''' If you’re using a virtual environment, make sure it has access to the system’s CUDA and cuDNN installations. You might need to install CUDA and cuDNN within the virtual environment or ensure that the environment variables are correctly set.
* '''Virtual Environments:''' If you’re using a virtual environment, make sure it has access to the system’s CUDA and cuDNN installations. You might need to install CUDA and cuDNN within the virtual environment or ensure that the environment variables are correctly set.
=== WhisperX Audio Transcription Commands for Multiple Languages ===
Standard Command: English transcription output
<pre>
whisperx /path/to/audio/file.wav \
        --model large-v3 \
        --language en \
        --diarize \
        --batch_size 24 \
        --no_align \
        --chunk_size 10 \
        --hf_token your_huggingface_token \
        --output_dir /path/to/output/directory \
        --output_format all
</pre>
To change output to Thai, modify the following parameters:
Method 1: Set language to Thai {{kbd | key=<nowiki>--language th</nowiki>}}
<pre>
whisperx /path/to/audio/file.wav \
        --model large-v3 \
        --language th \
        --diarize \
        --batch_size 24 \
        --no_align \
        --chunk_size 10 \
        --hf_token your_huggingface_token \
        --output_dir /path/to/output/directory \
        --output_format all
</pre>
Method 2: Auto-detect language (remove the {{kbd | key=<nowiki>--language</nowiki>}} parameter)
<pre>
whisperx /path/to/audio/file.wav \
        --model large-v3 \
        --diarize \
        --batch_size 24 \
        --no_align \
        --chunk_size 10 \
        --hf_token your_huggingface_token \
        --output_dir /path/to/output/directory \
        --output_format all
</pre>
Common Language Code Reference:
# th = Thai
# zh = Chinese
# en = English
# ja = Japanese
# ko = Korean
# es = Spanish
# fr = French
Supported language codes can be found in the [https://github.com/openai/whisper/blob/main/whisper/tokenizer.py Whisper tokenizer documentation] & [https://github.com/m-bain/whisperX/blob/main/EXAMPLES.md whisperX/EXAMPLES.md at main · m-bain/whisperX]


== whisperX Transcript File Format Guide ==
== whisperX Transcript File Format Guide ==
Line 117: Line 184:
If using for the first time, it's recommended to directly open the srt format
If using for the first time, it's recommended to directly open the srt format


== Further reading ==
* [https://medium.com/@planetoid/how-to-add-punctuation-to-whisper-transcripts-using-ai-619362c9160c How to Add Punctuation to Whisper Transcripts Using AI | Medium]


[[Category: Generative AI]] [[Category: Software]] [[Category: Revised with LLMs]]
[[Category: Generative AI]] [[Category: Software]] [[Category: Revised with LLMs]]