Mono- and cross-lingual evaluation of representation language models on less-resourced languages
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AH8HCMKZH" target="_blank" >RIV/00216208:11320/26:H8HCMKZH - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1016/j.csl.2025.101852" target="_blank" >http://dx.doi.org/10.1016/j.csl.2025.101852</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.csl.2025.101852" target="_blank" >10.1016/j.csl.2025.101852</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Mono- and cross-lingual evaluation of representation language models on less-resourced languages
Popis výsledku v původním jazyce
The current dominance of large language models in natural language processing is based on their contextual awareness. For text classification, text representation models, such as ELMo, BERT, and BERT derivatives, are typically fine-tuned for a specific problem. Most existing work focuses on English; in contrast, we present a large-scale multilingual empirical comparison of several monolingual and multilingual ELMo and BERT models using 14 classification tasks in nine languages. The results show, that the choice of best model largely depends on the task and language used, especially in a cross-lingual setting. In monolingual settings, monolingual BERT models tend to perform the best among BERT models. Among ELMo models, the ones trained on large corpora dominate. Cross-lingual knowledge transfer is feasible on most tasks already in a zero-shot setting without losing much performance. © 2025 The Authors
Název v anglickém jazyce
Mono- and cross-lingual evaluation of representation language models on less-resourced languages
Popis výsledku anglicky
The current dominance of large language models in natural language processing is based on their contextual awareness. For text classification, text representation models, such as ELMo, BERT, and BERT derivatives, are typically fine-tuned for a specific problem. Most existing work focuses on English; in contrast, we present a large-scale multilingual empirical comparison of several monolingual and multilingual ELMo and BERT models using 14 classification tasks in nine languages. The results show, that the choice of best model largely depends on the task and language used, especially in a cross-lingual setting. In monolingual settings, monolingual BERT models tend to perform the best among BERT models. Among ELMo models, the ones trained on large corpora dominate. Cross-lingual knowledge transfer is feasible on most tasks already in a zero-shot setting without losing much performance. © 2025 The Authors
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2026
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Computer Speech and Language
ISSN
0885-2308
e-ISSN
—
Svazek periodika
95
Číslo periodika v rámci svazku
2026
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
25
Strana od-do
101852
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105009701388