All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0201226" target="_blank" >RIV/00216305:26230/26:0201226 - isvavai.cz</a>

  • Result on the web

    <a href="https://github.com/BUTSpeechFIT/DiariZen" target="_blank" >https://github.com/BUTSpeechFIT/DiariZen</a>

  • DOI - Digital Object Identifier

Alternative languages

  • Result language

    angličtina

  • Original language name

    DiariZen

  • Original language description

    DiariZen is a cutting-edge speaker diarization toolkit developed by BUT Speech@FIT, combining end-to-end neural diarization (EEND) based on WavLM and Conformer with VBx clustering for accurate and scalable “who spoke when” analysis. Built on the Pyannote framework, it offers modularity, reproducibility, and seamless integration into speech processing pipelines. Structured pruning ensures efficiency without sacrificing performance.

  • Czech name

  • Czech description

Classification

  • Type

    R - Software

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

    <a href="/en/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications</a><br>

  • Continuities

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

  • Publication year

    2024

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Internal product ID

    DiariZen

  • Technical parameters

    DiariZen is a speaker diarization toolkit driven by AudioZen and Pyannote 3.1. Languages: Jupyter Notebook 54.8%, Python 44.7%, Shell 0.5% The code in GitHub repository is licensed under the MIT license. The pre-trained model weights are released under the CC BY-NC 4.0 license. https://github.com/BUTSpeechFIT/DiariZen https://huggingface.co/BUT-FIT/diarizen-wavlm-large-s80-md

  • Economical parameters

    DiariZen powers DiCoW, a companion tool that guides Whisper-based ASR for speaker-attributed transcription. The combined system achieved promising results—winning the Jury Prize at CHiME-8 and placing 2nd in the MLC-SLM Challenge. DiariZen has also been successfully adopted by multiple top-performing teams in the MISP 2025 Challenge, underscoring its robustness, generalization, and real-world impact.

  • Owner IČO

    00216305

  • Owner name

    Vysoké učení technické v Brně