home/categories/media/benchflow-ai-skillsbench-tasks-speaker-diarization-subtitles-environment-skills-multimodal-fusion-skill-md
mediacontent-media

multimodal-fusion-for-speaker-diarization

Combine visual features (face detection, lip movement analysis) with audio features to improve speaker diarization accuracy in video files. Use OpenCV for face detection and lip movement tracking, then fuse visual cues with audio-based speaker embeddings. Essential when processing video files with multiple visible speakers or when audio-only diarization needs visual validation.

benchflow-ai
maintainer
benchflow-ai
更新日 1/23/2026
スター
946
フォーク
244
quick start

Installation and usage

Combine visual features (face detection, lip movement analysis) with audio features to improve speaker diarization accuracy in video files. Use OpenCV for face detection and lip movement tracking, then fuse visual cues with audio-based speaker embeddings. Essential when processing video files with multiple visible speakers or when audio-only diarization needs visual validation.

インストール
$ install --globalskills.sh
使い方

インストール後、ターミナルで以下のコマンドを実行してこのスキルを使用できます:

skills use multimodal-fusion-for-speaker-diarization