DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0198052" target="_blank" >RIV/00216305:26230/26:0198052 - isvavai.cz</a>
Alternative codes found
RIV/00216208:11320/26:EE5AGRCD
Result on the web
<a href="https://www.sciencedirect.com/science/article/pii/S088523082500066X" target="_blank" >https://www.sciencedirect.com/science/article/pii/S088523082500066X</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.csl.2025.101841" target="_blank" >10.1016/j.csl.2025.101841</a>
Alternative languages
Result language
angličtina
Original language name
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
Original language description
Speaker-attributed automatic speech recognition (ASR) in multi-speakerenvironments remains a significant challenge, particularly when systems conditioned on speaker embeddings fail to generalize to unseen speakers. In this work, we propose Diarization-Conditioned Whisper (DiCoW), a novel approach to target-speaker ASR that leverages speaker diarization outputs as conditioning information. DiCoW extends the pre-trained Whisper model by integrating diarization labels directly, eliminating reliance on speaker embeddings and reducing the need for extensive speaker-specific training data. Our method introduces frame-level diarization-dependent transformations (FDDT) and query-key biasing (QKb) techniques to refine the model's focus on target speakers while effectively handling overlapping speech. By leveraging diarization outputs as conditioning signals, DiCoW simplifies the workflow for multi-speaker ASR, improves generalization to unseen speakers and enables more reliable transcription in real-world multi-speaker recordings. Additionally, we explore the integration of a connectionist temporal classification (CTC) head to Whisper and demonstrate its ability to improvetranscription efficiency through hybrid decoding. Notably, we show that our approach is not limited to Whisper; it also provides similar benefits when applied to the Branchformer model. We validate DiCoW on real-world datasets, including AMI and NOTSOFAR-1 from CHiME-8 challenge, as well as synthetic benchmarks such as Libri2Mix and LibriCSS, enabling direct comparisons with previous methods. Results demonstrate that DiCoW enhances the model's target-speaker ASR capabilities while maintaining Whisper's accuracy and robustness on single-speaker data.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2026
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
COMPUTER SPEECH AND LANGUAGE
ISSN
0885-2308
e-ISSN
1095-8363
Volume of the periodical
95
Issue of the periodical within the volume
1
Country of publishing house
GB - UNITED KINGDOM
Number of pages
19
Pages from-to
1-19
UT code for WoS article
001518605900001
EID of the result in the Scopus database
2-s2.0-105008798895