Morphological Segmentation with Neural Networks: Performance Effects of Architecture, Data Size, and Cross-Lingual Transfer in Seven Languages
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511638" target="_blank" >RIV/00216208:11320/25:10511638 - isvavai.cz</a>
Výsledek na webu
—
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Morphological Segmentation with Neural Networks: Performance Effects of Architecture, Data Size, and Cross-Lingual Transfer in Seven Languages
Popis výsledku v původním jazyce
We present a comparison of neural network-based morphological segmenters trained on morphologically segmented datasets from seven European languages: Czech, English, French, German, Italian, Dutch, and Slovak. Our aim is to investigate how different model architectures and dataset sizes influence segmentation quality, and how performance varies across languages. To this end, we evaluate recurrent and convolutional neural network models and compare them to widely used unsupervised baseline methods. In selecting the datasets, we prioritized linguistic accuracy and segmentation completeness. We also explore the impact of cross-lingual transfer learning on model performance. Our results show that neural models trained on as few as 125 words outperform unsupervised methods. Moreover, for closely related languages, zero-shot cross-lingual transfer learning can also surpass unsupervised baselines. Overall, we observe consistent performance patterns across languages.
Název v anglickém jazyce
Morphological Segmentation with Neural Networks: Performance Effects of Architecture, Data Size, and Cross-Lingual Transfer in Seven Languages
Popis výsledku anglicky
We present a comparison of neural network-based morphological segmenters trained on morphologically segmented datasets from seven European languages: Czech, English, French, German, Italian, Dutch, and Slovak. Our aim is to investigate how different model architectures and dataset sizes influence segmentation quality, and how performance varies across languages. To this end, we evaluate recurrent and convolutional neural network models and compare them to widely used unsupervised baseline methods. In selecting the datasets, we prioritized linguistic accuracy and segmentation completeness. We also explore the impact of cross-lingual transfer learning on model performance. Our results show that neural models trained on as few as 125 words outperform unsupervised methods. Moreover, for closely related languages, zero-shot cross-lingual transfer learning can also surpass unsupervised baselines. Overall, we observe consistent performance patterns across languages.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
28th International Conference on Text, Speech and Dialogue (Part II)
ISBN
978-3-032-02551-7
ISSN
—
e-ISSN
—
Počet stran výsledku
12
Strana od-do
275-286
Název nakladatele
Springer
Místo vydání
Cham, Switzerland
Místo konání akce
Erlangen, Germany
Datum konání akce
25. 8. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—