Target Speaker ASR with Whisper
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0198049" target="_blank" >RIV/00216305:26230/26:0198049 - isvavai.cz</a>
Result on the web
<a href="https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10887683" target="_blank" >https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10887683</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICASSP49660.2025.10887683" target="_blank" >10.1109/ICASSP49660.2025.10887683</a>
Alternative languages
Result language
angličtina
Original language name
Target Speaker ASR with Whisper
Original language description
We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn the space of all speaker embeddings. We find that adding even a single bias term per diarization output type before the first transformer block can transform single-speaker ASR models into target-speaker ASR models. Our approach also supports speaker-attributed ASR by sequentially generating transcripts for each speaker in a diarization output. This simplified method outperforms baseline speech separation and diarization cascade by 12.9% absolute ORC-WER on the NOTSOFAR-1 dataset.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
ISBN
979-8-3503-6874-1
ISSN
—
e-ISSN
—
Number of pages
5
Pages from-to
1-5
Publisher name
IEEE Signal Processing Society
Place of publication
Hyderabad
Event location
Hyderabad
Event date
Apr 6, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—