Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511625" target="_blank" >RIV/00216208:11320/25:10511625 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/chapter/10.1007/978-3-032-02551-7_11" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-02551-7_11</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-032-02551-7_11" target="_blank" >10.1007/978-3-032-02551-7_11</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
Popis výsledku v původním jazyce
Most pre-trained Vision-Language (VL) models and training data for the downstream tasks are only available in English. Therefore, multilingual VL tasks are solved using cross-lingual transfer: fine-tune a multilingual pre-trained model or transfer the text encoder using parallel data. We study the alternative approach: transferring an already trained encoder using parallel data. We investigate the effect of parallel data: domain and the number of languages, which were out of focus in previous work. Our results show that even machine-translated task data are the best on average, caption-like authentic parallel data outperformed it in some languages. Further, we show that most languages benefit from multilingual training.
Název v anglickém jazyce
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
Popis výsledku anglicky
Most pre-trained Vision-Language (VL) models and training data for the downstream tasks are only available in English. Therefore, multilingual VL tasks are solved using cross-lingual transfer: fine-tune a multilingual pre-trained model or transfer the text encoder using parallel data. We study the alternative approach: transferring an already trained encoder using parallel data. We investigate the effect of parallel data: domain and the number of languages, which were out of focus in previous work. Our results show that even machine-translated task data are the best on average, caption-like authentic parallel data outperformed it in some languages. Further, we show that most languages benefit from multilingual training.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
S - Specificky vyzkum na vysokych skolach<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
28th International Conference on Text, Speech and Dialogue (Part II)
ISBN
978-3-032-02551-7
ISSN
—
e-ISSN
—
Počet stran výsledku
13
Strana od-do
115-127
Název nakladatele
Springer
Místo vydání
Cham, Switzerland
Místo konání akce
Erlangen, Germany
Datum konání akce
25. 8. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001576349100010