Minimum effort adaptation of automatic speech recognition system in air traffic management
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0201384" target="_blank" >RIV/00216305:26230/26:0201384 - isvavai.cz</a>
Výsledek na webu
<a href="https://journals.open.tudelft.nl/ejtir/article/view/7531" target="_blank" >https://journals.open.tudelft.nl/ejtir/article/view/7531</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.59490/ejtir.2024.24.4.7531" target="_blank" >10.59490/ejtir.2024.24.4.7531</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Minimum effort adaptation of automatic speech recognition system in air traffic management
Popis výsledku v původním jazyce
Advancements in Automatic Speech Recognition (ASR) technology is exemplified by ubiquitous voice assistants such as Siri and Alexa. Researchers have been exploring the application of ASR for Air Traffic Management (ATM) systems. Initial prototypes utilized ASR to pre-fill aircraft radar labels and achieved a technological readiness level before industrialization (TRL6). However, accurately recognizing infrequently used but highly informative domain-specific vocabulary is still an issue. This includes waypoint names specific to each airspace region and unique airline designators, e.g., "DEXON" or "POBEDA". Traditionally, open-source ASR toolkits or large pre-trained models require substantial domain-specific transcribed speech data to adapt to specialized vocabularies. However, typically, a "universal" ASR engine capable of reliably recognizing a core dictionary of several hundreds of frequently used words suffices for ATM applications. The challenge lies in dynamically integrating the additional region-specific words used less frequently. These uncommon words are crucial for maintaining clear communication within the ATM environment. This paper proposes a novel approach that facilitates the dynamic integration of these new and specific word entities into the existing universal ASR system. This paves the way for "plug-and-play" customization with minimal expert intervention and eliminates the need for extensive fine-tuning of the universal ASR model. The proposed approach demonstrably improves the accuracy of these region-specific words by a factor of approximate to 7 (from 10% F1-score to 70%) for all rare words and approximate to 5 (from 13% F1-score to 64%) for waypoints.
Název v anglickém jazyce
Minimum effort adaptation of automatic speech recognition system in air traffic management
Popis výsledku anglicky
Advancements in Automatic Speech Recognition (ASR) technology is exemplified by ubiquitous voice assistants such as Siri and Alexa. Researchers have been exploring the application of ASR for Air Traffic Management (ATM) systems. Initial prototypes utilized ASR to pre-fill aircraft radar labels and achieved a technological readiness level before industrialization (TRL6). However, accurately recognizing infrequently used but highly informative domain-specific vocabulary is still an issue. This includes waypoint names specific to each airspace region and unique airline designators, e.g., "DEXON" or "POBEDA". Traditionally, open-source ASR toolkits or large pre-trained models require substantial domain-specific transcribed speech data to adapt to specialized vocabularies. However, typically, a "universal" ASR engine capable of reliably recognizing a core dictionary of several hundreds of frequently used words suffices for ATM applications. The challenge lies in dynamically integrating the additional region-specific words used less frequently. These uncommon words are crucial for maintaining clear communication within the ATM environment. This paper proposes a novel approach that facilitates the dynamic integration of these new and specific word entities into the existing universal ASR system. This paves the way for "plug-and-play" customization with minimal expert intervention and eliminates the need for extensive fine-tuning of the universal ASR model. The proposed approach demonstrably improves the accuracy of these region-specific words by a factor of approximate to 7 (from 10% F1-score to 70%) for all rare words and approximate to 5 (from 13% F1-score to 64%) for waypoints.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
50703 - Transport planning and social aspects of transport (transport engineering to be 2.1)
Návaznosti výsledku
Projekt
—
Návaznosti
S - Specificky vyzkum na vysokych skolach
Ostatní
Rok uplatnění
2024
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
European Journal of Transport and Infrastructure Research
ISSN
1567-7133
e-ISSN
1567-7141
Svazek periodika
24
Číslo periodika v rámci svazku
4
Stát vydavatele periodika
NL - Nizozemsko
Počet stran výsledku
21
Strana od-do
133-153
Kód UT WoS článku
001447236400001
EID výsledku v databázi Scopus
2-s2.0-85215400025