Combining multilingual resources to enhance end-to-end speech recognition systems for Scandinavian languages
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F46747885%3A24220%2F25%3A00014526" target="_blank" >RIV/46747885:24220/25:00014526 - isvavai.cz</a>
Result on the web
<a href="https://www.sciencedirect.com/science/article/pii/S0167639325000366?pes=vor&utm_source=clarivate&getft_integrator=clarivate" target="_blank" >https://www.sciencedirect.com/science/article/pii/S0167639325000366?pes=vor&utm_source=clarivate&getft_integrator=clarivate</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.specom.2025.103221" target="_blank" >10.1016/j.specom.2025.103221</a>
Alternative languages
Result language
angličtina
Original language name
Combining multilingual resources to enhance end-to-end speech recognition systems for Scandinavian languages
Original language description
Languages with limited training resources, such as Danish, Swedish, and Norwegian, pose a challenge to the development of modern end-to-end (E2E) automatic speech recognition (ASR) systems. We tackle this issue by exploring different ways of exploiting existing multilingual resources. Our approaches combine speech data of closely related languages and/or their already trained models. From several proposed options, the most efficient one is based on initializing the E2E encoder parameters by those from other available models, which we call donors. This approach performs well not only for smaller amounts of target language data but also when thousands of hours are available and even when the donor comes from a distant language. We study several aspects of these donor-based models, namely the choice of the donor language, the impact of the data size (both for target and donor models), or the option of using different donor-based models simultaneously. This allows us to implement an efficient data collection process in which multiple donor-based models run in parallel and serve as complementary data checkers. This greatly helps to eliminate annotation errors in training sets and during automated data harvesting. The latter is utilized for efficient processing of diverse public sources (TV, parliament, YouTube, podcasts, or audiobooks) and training models based on thousands of hours. We have also prepared large test sets (link provided) to evaluate all experiments and ultimately compare the performance of our ASR system with that of major ASR service providers for Scandinavian languages.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10307 - Acoustics
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
SPEECH COMMUNICATION>
ISSN
0167-6393
e-ISSN
—
Volume of the periodical
170
Issue of the periodical within the volume
MAY
Country of publishing house
NL - THE KINGDOM OF THE NETHERLANDS
Number of pages
13
Pages from-to
—
UT code for WoS article
001446647800001
EID of the result in the Scopus database
2-s2.0-86000574196