Combining multilingual resources to enhance end-to-end speech recognition systems for Scandinavian languages
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F46747885%3A24220%2F25%3A00014526" target="_blank" >RIV/46747885:24220/25:00014526 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.sciencedirect.com/science/article/pii/S0167639325000366?pes=vor&utm_source=clarivate&getft_integrator=clarivate" target="_blank" >https://www.sciencedirect.com/science/article/pii/S0167639325000366?pes=vor&utm_source=clarivate&getft_integrator=clarivate</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.specom.2025.103221" target="_blank" >10.1016/j.specom.2025.103221</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Combining multilingual resources to enhance end-to-end speech recognition systems for Scandinavian languages
Popis výsledku v původním jazyce
Languages with limited training resources, such as Danish, Swedish, and Norwegian, pose a challenge to the development of modern end-to-end (E2E) automatic speech recognition (ASR) systems. We tackle this issue by exploring different ways of exploiting existing multilingual resources. Our approaches combine speech data of closely related languages and/or their already trained models. From several proposed options, the most efficient one is based on initializing the E2E encoder parameters by those from other available models, which we call donors. This approach performs well not only for smaller amounts of target language data but also when thousands of hours are available and even when the donor comes from a distant language. We study several aspects of these donor-based models, namely the choice of the donor language, the impact of the data size (both for target and donor models), or the option of using different donor-based models simultaneously. This allows us to implement an efficient data collection process in which multiple donor-based models run in parallel and serve as complementary data checkers. This greatly helps to eliminate annotation errors in training sets and during automated data harvesting. The latter is utilized for efficient processing of diverse public sources (TV, parliament, YouTube, podcasts, or audiobooks) and training models based on thousands of hours. We have also prepared large test sets (link provided) to evaluate all experiments and ultimately compare the performance of our ASR system with that of major ASR service providers for Scandinavian languages.
Název v anglickém jazyce
Combining multilingual resources to enhance end-to-end speech recognition systems for Scandinavian languages
Popis výsledku anglicky
Languages with limited training resources, such as Danish, Swedish, and Norwegian, pose a challenge to the development of modern end-to-end (E2E) automatic speech recognition (ASR) systems. We tackle this issue by exploring different ways of exploiting existing multilingual resources. Our approaches combine speech data of closely related languages and/or their already trained models. From several proposed options, the most efficient one is based on initializing the E2E encoder parameters by those from other available models, which we call donors. This approach performs well not only for smaller amounts of target language data but also when thousands of hours are available and even when the donor comes from a distant language. We study several aspects of these donor-based models, namely the choice of the donor language, the impact of the data size (both for target and donor models), or the option of using different donor-based models simultaneously. This allows us to implement an efficient data collection process in which multiple donor-based models run in parallel and serve as complementary data checkers. This greatly helps to eliminate annotation errors in training sets and during automated data harvesting. The latter is utilized for efficient processing of diverse public sources (TV, parliament, YouTube, podcasts, or audiobooks) and training models based on thousands of hours. We have also prepared large test sets (link provided) to evaluate all experiments and ultimately compare the performance of our ASR system with that of major ASR service providers for Scandinavian languages.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10307 - Acoustics
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
SPEECH COMMUNICATION>
ISSN
0167-6393
e-ISSN
—
Svazek periodika
170
Číslo periodika v rámci svazku
MAY
Stát vydavatele periodika
NL - Nizozemsko
Počet stran výsledku
13
Strana od-do
—
Kód UT WoS článku
001446647800001
EID výsledku v databázi Scopus
2-s2.0-86000574196