Derivational Morphemes as Markers of Borrowed Words in Czech
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511658" target="_blank" >RIV/00216208:11320/25:10511658 - isvavai.cz</a>
Výsledek na webu
<a href="https://events.unifr.ch/derimo2025/en/assets/public/files/derimo_2025_proceedings.pdf" target="_blank" >https://events.unifr.ch/derimo2025/en/assets/public/files/derimo_2025_proceedings.pdf</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Derivational Morphemes as Markers of Borrowed Words in Czech
Popis výsledku v původním jazyce
In Czech, borrowed words are often marked by the presence of morphemes of foreign origin. Non- native derivational affixes such as un-, dys-, and anti- frequently appear in these words, typically alongside foreign stems. However, it is unclear to what extent these morphemes independently influence the classification of a word as borrowed or act as a marker for borrowed words. This study investigates the role of derivational morphemes in marking borrowed words in four parts-of- speech (POS) categories: nouns, verbs, adjectives, and adverbs. We extract native and borrowed words from DeriNet, which categorizes words based on their POS tags and loanword status. Using a multinomial Naive Bayes classifier, we perform binary classification to distinguish native and borrowed words, extracting classification probabilities for the morphemes that make up these words. We compare these results with attention-based binary LSTM classifier and RobeCzech, a pre-trained RoBERTa model for Czech that we finetune for our
Název v anglickém jazyce
Derivational Morphemes as Markers of Borrowed Words in Czech
Popis výsledku anglicky
In Czech, borrowed words are often marked by the presence of morphemes of foreign origin. Non- native derivational affixes such as un-, dys-, and anti- frequently appear in these words, typically alongside foreign stems. However, it is unclear to what extent these morphemes independently influence the classification of a word as borrowed or act as a marker for borrowed words. This study investigates the role of derivational morphemes in marking borrowed words in four parts-of- speech (POS) categories: nouns, verbs, adjectives, and adverbs. We extract native and borrowed words from DeriNet, which categorizes words based on their POS tags and loanword status. Using a multinomial Naive Bayes classifier, we perform binary classification to distinguish native and borrowed words, extracting classification probabilities for the morphemes that make up these words. We compare these results with attention-based binary LSTM classifier and RobeCzech, a pre-trained RoBERTa model for Czech that we finetune for our
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
S - Specificky vyzkum na vysokych skolach<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of Fifth International Workshop on Resources and Tools for Derivational Morphology
ISBN
978-2-8399-4786-2
ISSN
—
e-ISSN
—
Počet stran výsledku
10
Strana od-do
109-118
Název nakladatele
University of Fribourg
Místo vydání
Fribourg, Switzerland
Místo konání akce
Fribourg, Switzerland
Datum konání akce
4. 9. 2025
Typ akce podle státní příslušnosti
EUR - Evropská akce
Kód UT WoS článku
—