Efficient Enhancement of Norwegian ASR Model
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F46747885%3A24220%2F26%3A00014130" target="_blank" >RIV/46747885:24220/26:00014130 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/chapter/10.1007/978-3-032-02548-7_4" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-02548-7_4</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Efficient Enhancement of Norwegian ASR Model
Popis výsledku v původním jazyce
In this contribution, we present two approaches to efficiently enhance end-to-end (E2E) automatic speech recognition (ASR) models for the Norwegian language. Both utilize multilingual models. First, we demonstrate that model performance can be significantly improved if trained with encoder parameters initialized from models created for other languages. This is true not only for closely related languages, like Swedish or Danish, but also for more or less distant ones, like German, English, Italian, or even Ukrainian. This type of model enhancement is achieved without any additional training data, so it requires no extra computation time. Second, having multiple and differently initialized models for Norwegian offers another advantage. They can be used in data harvesting as multiple checkers to validate correct annotations in parallel. This allows us to acquire a large amount of additional training data automatically from various public sources, such as YouTube, parliament, or government archives. We evaluate our final Norwegian model (trained on 2,599 h) on a diverse 18-h test set and compare its performance to major ASR service providers (Google, Microsoft, Speechmatics) and two fine-tuned Whisper models.
Název v anglickém jazyce
Efficient Enhancement of Norwegian ASR Model
Popis výsledku anglicky
In this contribution, we present two approaches to efficiently enhance end-to-end (E2E) automatic speech recognition (ASR) models for the Norwegian language. Both utilize multilingual models. First, we demonstrate that model performance can be significantly improved if trained with encoder parameters initialized from models created for other languages. This is true not only for closely related languages, like Swedish or Danish, but also for more or less distant ones, like German, English, Italian, or even Ukrainian. This type of model enhancement is achieved without any additional training data, so it requires no extra computation time. Second, having multiple and differently initialized models for Norwegian offers another advantage. They can be used in data harvesting as multiple checkers to validate correct annotations in parallel. This allows us to acquire a large amount of additional training data automatically from various public sources, such as YouTube, parliament, or government archives. We evaluate our final Norwegian model (trained on 2,599 h) on a diverse 18-h test set and compare its performance to major ASR service providers (Google, Microsoft, Speechmatics) and two fine-tuned Whisper models.
Klasifikace
Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2026
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů