End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0199932" target="_blank" >RIV/00216305:26230/26:0199932 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.isca-archive.org/interspeech_2025/pothula25_interspeech.pdf" target="_blank" >https://www.isca-archive.org/interspeech_2025/pothula25_interspeech.pdf</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.21437/interspeech.2025-2525" target="_blank" >10.21437/interspeech.2025-2525</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
Popis výsledku v původním jazyce
The scarcity of high-quality annotated data presents a significant challenge in developing effective end-to-end speech-to-text translation (ST) systems, particularly for low-resource languages. This paper explores the hypothesis that weakly labeled data can be used to build ST models for low-resource language pairs. We constructed speech-to-text translation datasets with the help of bitext mining using state-of-the-art sentence encoders. We mined the multilingual Shrutilipi corpus to build Shrutilipi-anuvaad, a dataset comprising ST data for language pairs Bengali-Hindi, Malayalam-Hindi, Odia-Hindi, and Telugu-Hindi. We created multiple versions of training data with varying degrees of quality and quantity to investigate the effect of quality versus quantity of weakly labeled data on ST model performance. Results demonstrate that ST systems can be built using weakly labeled data, with performance comparable to massive multi-modal multilingual baselines such as SONAR and SeamlessM4T.
Název v anglickém jazyce
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
Popis výsledku anglicky
The scarcity of high-quality annotated data presents a significant challenge in developing effective end-to-end speech-to-text translation (ST) systems, particularly for low-resource languages. This paper explores the hypothesis that weakly labeled data can be used to build ST models for low-resource language pairs. We constructed speech-to-text translation datasets with the help of bitext mining using state-of-the-art sentence encoders. We mined the multilingual Shrutilipi corpus to build Shrutilipi-anuvaad, a dataset comprising ST data for language pairs Bengali-Hindi, Malayalam-Hindi, Odia-Hindi, and Telugu-Hindi. We created multiple versions of training data with varying degrees of quality and quantity to investigate the effect of quality versus quantity of weakly labeled data on ST model performance. Results demonstrate that ST systems can be built using weakly labeled data, with performance comparable to massive multi-modal multilingual baselines such as SONAR and SeamlessM4T.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Jazykověda, umělá inteligence a jazykové a řečové technologie: od výzkumu k aplikacím</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Interspeech 2025
ISBN
—
ISSN
—
e-ISSN
2958-1796
Počet stran výsledku
5
Strana od-do
41-45
Název nakladatele
ISCA
Místo vydání
Rotterdam
Místo konání akce
Rotterdam
Datum konání akce
17. 8. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—