Study on Incorporating Tone into Speech Recognition of Vietnamese
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F46747885%3A24220%2F15%3A%230003410" target="_blank" >RIV/46747885:24220/15:#0003410 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1109/ECMSM.2015.7208688" target="_blank" >http://dx.doi.org/10.1109/ECMSM.2015.7208688</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ECMSM.2015.7208688" target="_blank" >10.1109/ECMSM.2015.7208688</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Study on Incorporating Tone into Speech Recognition of Vietnamese
Popis výsledku v původním jazyce
Vietnamese is a syllable-based tonal language where the tone used in syllable pronunciation carries important information about the meaning. In this paper, we investigate several approaches how to incorporate the tone into an acoustic model. We propose 3basic strategies: a) a phoneme-based, b) a vowel-based, and c) a rhyme-based one. Each can be modified so that we obtain 15 different schemes that are described and compared in experiments performed within the framework of large-vocabulary continuous speech recognition of Vietnamese. We show that the phoneme-based context dependent model performs best, particularly when information about the tone is linked to the syllable end. On the test set, made of 85 minutes of mostly broadcast speech, we achieve 74% syllable accuracy rate. The accuracy is further improved to 78% when the pronunciation lexicon and the language model takes into account also 40,000 most frequent syllable pairs.
Název v anglickém jazyce
Study on Incorporating Tone into Speech Recognition of Vietnamese
Popis výsledku anglicky
Vietnamese is a syllable-based tonal language where the tone used in syllable pronunciation carries important information about the meaning. In this paper, we investigate several approaches how to incorporate the tone into an acoustic model. We propose 3basic strategies: a) a phoneme-based, b) a vowel-based, and c) a rhyme-based one. Each can be modified so that we obtain 15 different schemes that are described and compared in experiments performed within the framework of large-vocabulary continuous speech recognition of Vietnamese. We show that the phoneme-based context dependent model performs best, particularly when information about the tone is linked to the syllable end. On the test set, made of 85 minutes of mostly broadcast speech, we achieve 74% syllable accuracy rate. The accuracy is further improved to 78% when the pronunciation lexicon and the language model takes into account also 40,000 most frequent syllable pairs.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
JC - Počítačový hardware a software
OECD FORD obor
—
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2015
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
2015 IEEE International Workshop of Electronics, Control, Measurement, Signals and their application to Mechatronics
ISBN
978-1-4799-6972-2
ISSN
—
e-ISSN
—
Počet stran výsledku
6
Strana od-do
42-47
Název nakladatele
IEEE
Místo vydání
Česká Republika
Místo konání akce
Česká Republika, Liberec
Datum konání akce
1. 1. 2015
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
000363814500013