All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Towards Development of New Language Resource for Urdu: The Large Vocabulary Word Embeddings

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ARH4EL44J" target="_blank" >RIV/00216208:11320/26:RH4EL44J - isvavai.cz</a>

  • Result on the web

    <a href="http://dx.doi.org/10.1145/3748308" target="_blank" >http://dx.doi.org/10.1145/3748308</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1145/3748308" target="_blank" >10.1145/3748308</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Towards Development of New Language Resource for Urdu: The Large Vocabulary Word Embeddings

  • Original language description

    Urdu is a resource-poor language as it lacks natural language processing (NLP) resources. NLP resources include word embeddings, treebanks, parsers, part-of-speech taggers, tokenizers, stemmers, morphological analyzers, and text visualization tools. We propose word embeddings for Urdu language. The vocabulary size of the word embeddings affects the accuracy of the NLP systems. The larger the vocabulary size is, the less the out of vocabulary words the NLP system will encounter. We also show that if word embeddings are trained on a larger amount of text, then they encode more semantic information as compared to those trained on smaller text. We propose word embeddings that have a vocabulary size of 456,905 which is higher than the vocabulary sizes of two state-of-the-art Urdu word embeddings which have vocabulary sizes of 160,413 and 102,214. We have compared the proposed Urdu word embeddings with the state-of-the-art word embeddings for word similarity and transition-based dependency parsing tasks. The datasets used for the word similarity experiments are WordSim353 and SimLex999. The dataset used for the dependency parsing experiments is the Urdu dependencies’ treebank from the Universal Dependencies (UD) version 2.11. We have trained our proposed word embeddings on three corpora which contain Urdu-news-dataset-1M, the CLE-Urdu-Digest corpus, and the news corpus that we have compiled by scrapping Urdu news from Urdupoint website. The intrinsic evaluation for word similarity task shows that our proposed word embeddings have achieved the highest reported and significant Spearman ranked correlation co-efficient value of 0.693 for WordSim353 dataset and the highest and significant Spearman ranked correlation co-efficient value of 0.426 and Pearson correlation co-efficient value of 0.453 for SimLex999 dataset, in comparison with the two state-of-the-art word embeddings. The extrinsic evaluation of the proposed word embeddings for the dependency parsing task results in an increase in the Labelled Attachment Score (LAS) by ≈ 5.16% and ≈ 10.17% from the LAS achieved by the state-of-the-art embeddings. Similarly, the proposed large vocabulary word embeddings have increased the Unlabelled Attachment Score (UAS) of Urdu dependency parsing by ≈ 5.44% and ≈ 5.08% as compared to the UAS of the two state-of-the-art word embeddings. It is also shown that the proposed solution results in less number of out-of-vocabulary words as compared to the state-of-the-art. © 2025 Copyright held by the owner/author(s).

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>SC</sub> - Article in a specialist periodical, which is included in the SCOPUS database

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    ACM Transactions on Asian and Low-Resource Language Information Processing

  • ISSN

    2375-4699

  • e-ISSN

  • Volume of the periodical

    24

  • Issue of the periodical within the volume

    8

  • Country of publishing house

    US - UNITED STATES

  • Number of pages

    14

  • Pages from-to

    1-14

  • UT code for WoS article

  • EID of the result in the Scopus database

    2-s2.0-105018452057