DiariZen
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0201226" target="_blank" >RIV/00216305:26230/26:0201226 - isvavai.cz</a>
Výsledek na webu
<a href="https://github.com/BUTSpeechFIT/DiariZen" target="_blank" >https://github.com/BUTSpeechFIT/DiariZen</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
DiariZen
Popis výsledku v původním jazyce
DiariZen is a cutting-edge speaker diarization toolkit developed by BUT Speech@FIT, combining end-to-end neural diarization (EEND) based on WavLM and Conformer with VBx clustering for accurate and scalable “who spoke when” analysis. Built on the Pyannote framework, it offers modularity, reproducibility, and seamless integration into speech processing pipelines. Structured pruning ensures efficiency without sacrificing performance.
Název v anglickém jazyce
DiariZen
Popis výsledku anglicky
DiariZen is a cutting-edge speaker diarization toolkit developed by BUT Speech@FIT, combining end-to-end neural diarization (EEND) based on WavLM and Conformer with VBx clustering for accurate and scalable “who spoke when” analysis. Built on the Pyannote framework, it offers modularity, reproducibility, and seamless integration into speech processing pipelines. Structured pruning ensures efficiency without sacrificing performance.
Klasifikace
Druh
R - Software
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Jazykověda, umělá inteligence a jazykové a řečové technologie: od výzkumu k aplikacím</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2024
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Interní identifikační kód produktu
DiariZen
Technické parametry
DiariZen is a speaker diarization toolkit driven by AudioZen and Pyannote 3.1. Languages: Jupyter Notebook 54.8%, Python 44.7%, Shell 0.5% The code in GitHub repository is licensed under the MIT license. The pre-trained model weights are released under the CC BY-NC 4.0 license. https://github.com/BUTSpeechFIT/DiariZen https://huggingface.co/BUT-FIT/diarizen-wavlm-large-s80-md
Ekonomické parametry
DiariZen powers DiCoW, a companion tool that guides Whisper-based ASR for speaker-attributed transcription. The combined system achieved promising results—winning the Jury Prize at CHiME-8 and placing 2nd in the MLC-SLM Challenge. DiariZen has also been successfully adopted by multiple top-performing teams in the MISP 2025 Challenge, underscoring its robustness, generalization, and real-world impact.
IČO vlastníka výsledku
00216305
Název vlastníka
Vysoké učení technické v Brně