Towards Development of New Language Resource for Urdu: The Large Vocabulary Word Embeddings
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ARH4EL44J" target="_blank" >RIV/00216208:11320/26:RH4EL44J - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1145/3748308" target="_blank" >http://dx.doi.org/10.1145/3748308</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1145/3748308" target="_blank" >10.1145/3748308</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Towards Development of New Language Resource for Urdu: The Large Vocabulary Word Embeddings
Popis výsledku v původním jazyce
Urdu is a resource-poor language as it lacks natural language processing (NLP) resources. NLP resources include word embeddings, treebanks, parsers, part-of-speech taggers, tokenizers, stemmers, morphological analyzers, and text visualization tools. We propose word embeddings for Urdu language. The vocabulary size of the word embeddings affects the accuracy of the NLP systems. The larger the vocabulary size is, the less the out of vocabulary words the NLP system will encounter. We also show that if word embeddings are trained on a larger amount of text, then they encode more semantic information as compared to those trained on smaller text. We propose word embeddings that have a vocabulary size of 456,905 which is higher than the vocabulary sizes of two state-of-the-art Urdu word embeddings which have vocabulary sizes of 160,413 and 102,214. We have compared the proposed Urdu word embeddings with the state-of-the-art word embeddings for word similarity and transition-based dependency parsing tasks. The datasets used for the word similarity experiments are WordSim353 and SimLex999. The dataset used for the dependency parsing experiments is the Urdu dependencies’ treebank from the Universal Dependencies (UD) version 2.11. We have trained our proposed word embeddings on three corpora which contain Urdu-news-dataset-1M, the CLE-Urdu-Digest corpus, and the news corpus that we have compiled by scrapping Urdu news from Urdupoint website. The intrinsic evaluation for word similarity task shows that our proposed word embeddings have achieved the highest reported and significant Spearman ranked correlation co-efficient value of 0.693 for WordSim353 dataset and the highest and significant Spearman ranked correlation co-efficient value of 0.426 and Pearson correlation co-efficient value of 0.453 for SimLex999 dataset, in comparison with the two state-of-the-art word embeddings. The extrinsic evaluation of the proposed word embeddings for the dependency parsing task results in an increase in the Labelled Attachment Score (LAS) by ≈ 5.16% and ≈ 10.17% from the LAS achieved by the state-of-the-art embeddings. Similarly, the proposed large vocabulary word embeddings have increased the Unlabelled Attachment Score (UAS) of Urdu dependency parsing by ≈ 5.44% and ≈ 5.08% as compared to the UAS of the two state-of-the-art word embeddings. It is also shown that the proposed solution results in less number of out-of-vocabulary words as compared to the state-of-the-art. © 2025 Copyright held by the owner/author(s).
Název v anglickém jazyce
Towards Development of New Language Resource for Urdu: The Large Vocabulary Word Embeddings
Popis výsledku anglicky
Urdu is a resource-poor language as it lacks natural language processing (NLP) resources. NLP resources include word embeddings, treebanks, parsers, part-of-speech taggers, tokenizers, stemmers, morphological analyzers, and text visualization tools. We propose word embeddings for Urdu language. The vocabulary size of the word embeddings affects the accuracy of the NLP systems. The larger the vocabulary size is, the less the out of vocabulary words the NLP system will encounter. We also show that if word embeddings are trained on a larger amount of text, then they encode more semantic information as compared to those trained on smaller text. We propose word embeddings that have a vocabulary size of 456,905 which is higher than the vocabulary sizes of two state-of-the-art Urdu word embeddings which have vocabulary sizes of 160,413 and 102,214. We have compared the proposed Urdu word embeddings with the state-of-the-art word embeddings for word similarity and transition-based dependency parsing tasks. The datasets used for the word similarity experiments are WordSim353 and SimLex999. The dataset used for the dependency parsing experiments is the Urdu dependencies’ treebank from the Universal Dependencies (UD) version 2.11. We have trained our proposed word embeddings on three corpora which contain Urdu-news-dataset-1M, the CLE-Urdu-Digest corpus, and the news corpus that we have compiled by scrapping Urdu news from Urdupoint website. The intrinsic evaluation for word similarity task shows that our proposed word embeddings have achieved the highest reported and significant Spearman ranked correlation co-efficient value of 0.693 for WordSim353 dataset and the highest and significant Spearman ranked correlation co-efficient value of 0.426 and Pearson correlation co-efficient value of 0.453 for SimLex999 dataset, in comparison with the two state-of-the-art word embeddings. The extrinsic evaluation of the proposed word embeddings for the dependency parsing task results in an increase in the Labelled Attachment Score (LAS) by ≈ 5.16% and ≈ 10.17% from the LAS achieved by the state-of-the-art embeddings. Similarly, the proposed large vocabulary word embeddings have increased the Unlabelled Attachment Score (UAS) of Urdu dependency parsing by ≈ 5.44% and ≈ 5.08% as compared to the UAS of the two state-of-the-art word embeddings. It is also shown that the proposed solution results in less number of out-of-vocabulary words as compared to the state-of-the-art. © 2025 Copyright held by the owner/author(s).
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
ACM Transactions on Asian and Low-Resource Language Information Processing
ISSN
2375-4699
e-ISSN
—
Svazek periodika
24
Číslo periodika v rámci svazku
8
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
14
Strana od-do
1-14
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105018452057