Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Statistical analysis of Hindi multi-word expressions using multiple threshold method

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AS3CRRDDV" target="_blank" >RIV/00216208:11320/26:S3CRRDDV - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.1007/s44163-025-00291-z" target="_blank" >http://dx.doi.org/10.1007/s44163-025-00291-z</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/s44163-025-00291-z" target="_blank" >10.1007/s44163-025-00291-z</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Statistical analysis of Hindi multi-word expressions using multiple threshold method

  • Popis výsledku v původním jazyce

    Multiword Expressions (MWEs) extraction is one of the important aspects of text processing, which is used to find the correct meaning of a text phrase. MWEs are lexical phrases consisting of two or more words exhibiting semantic property. MWEs play a vital role in many Natural Language Processing (NLP) applications like machine translation, information retrieval, text processing, and other practical applications. Much of the research in this area has focused on the extraction and analysis of MWEs in English and other natural languages. The MWEs in Hindi have not gained much attention from earlier researchers. In the proposed work the statistical aspects of Hindi MWEs are explored using various statistical measures. Different classes of functional classification of Hindi MWEs are considered for the statistical analysis and experiments. This paper mainly focuses on the evaluation of the following statistical measures, Pointwise Mutual Information (PMI), Dice Coefficient (DC), Modified Dice Coefficient (MDC), Lexical Fixedness (LF), Syntactic Fixedness (SF), and Relevance Measure (RM). The dataset used for the evaluation of MWEs is the benchmark dataset collected from Hindi novels written by “Munshi Premchand Ji”. The best statistical measures are also identified for each functional category of Hindi MWEs. Different threshold values have been obtained for the evaluation of the functional categories. The threshold value represents the maximum limit of the corpus size that one can select for efficient evaluation of a specific category of Hindi MWEs. This approach has been applied to two different datasets to compare and justify the obtained results. © The Author(s) 2025.

  • Název v anglickém jazyce

    Statistical analysis of Hindi multi-word expressions using multiple threshold method

  • Popis výsledku anglicky

    Multiword Expressions (MWEs) extraction is one of the important aspects of text processing, which is used to find the correct meaning of a text phrase. MWEs are lexical phrases consisting of two or more words exhibiting semantic property. MWEs play a vital role in many Natural Language Processing (NLP) applications like machine translation, information retrieval, text processing, and other practical applications. Much of the research in this area has focused on the extraction and analysis of MWEs in English and other natural languages. The MWEs in Hindi have not gained much attention from earlier researchers. In the proposed work the statistical aspects of Hindi MWEs are explored using various statistical measures. Different classes of functional classification of Hindi MWEs are considered for the statistical analysis and experiments. This paper mainly focuses on the evaluation of the following statistical measures, Pointwise Mutual Information (PMI), Dice Coefficient (DC), Modified Dice Coefficient (MDC), Lexical Fixedness (LF), Syntactic Fixedness (SF), and Relevance Measure (RM). The dataset used for the evaluation of MWEs is the benchmark dataset collected from Hindi novels written by “Munshi Premchand Ji”. The best statistical measures are also identified for each functional category of Hindi MWEs. Different threshold values have been obtained for the evaluation of the functional categories. The threshold value represents the maximum limit of the corpus size that one can select for efficient evaluation of a specific category of Hindi MWEs. This approach has been applied to two different datasets to compare and justify the obtained results. © The Author(s) 2025.

Klasifikace

  • Druh

    J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    Discover Artificial Intelligence

  • ISSN

    2731-0809

  • e-ISSN

  • Svazek periodika

    5

  • Číslo periodika v rámci svazku

    1

  • Stát vydavatele periodika

    US - Spojené státy americké

  • Počet stran výsledku

    27

  • Strana od-do

    46

  • Kód UT WoS článku

  • EID výsledku v databázi Scopus

    2-s2.0-105004425390