Employing Natural Language Processing Techniques for the Development of a Voting- Based POS Tagger in the Urdu Language
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AWN5VHV5U" target="_blank" >RIV/00216208:11320/26:WN5VHV5U - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.4018/979-8-3693-5231-1.ch002" target="_blank" >http://dx.doi.org/10.4018/979-8-3693-5231-1.ch002</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.4018/979-8-3693-5231-1.ch002" target="_blank" >10.4018/979-8-3693-5231-1.ch002</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Employing Natural Language Processing Techniques for the Development of a Voting- Based POS Tagger in the Urdu Language
Popis výsledku v původním jazyce
The process of sequence labeling (POS) by assigning syntactic tags to words in the given context is an important role in various NLP applications. The core motive of this work is to tackle the morpho- syntactic category of words in Urdu language. This language has lots of computational challenges because of its dual nature. The work comprises different tasks as initially the authors tracked the best combination of feature sets in terms of CRF to entitle the previous results on two stable and wellknown datasets Bushra Jawaid dataset and CLE dataset. Due to syntactic ambiguity, a state- of- the- art voting method has been introduced which is being implemented to overcome the contradictory results of the different machine learning classifiers. The results show significant improvement in the baseline results as the F1- score on a primary dataset is 94.8% and 95.7% on the succeeding dataset. Long short- term memory (LSTM) is used for one of the most diverse and inflectional tasks like part of speech tagging for the Urdu language by achieving an F1- score of 86.7% and 96.1% respectively for both datasets. © 2025 by IGI Global Scientific Publishing.
Název v anglickém jazyce
Employing Natural Language Processing Techniques for the Development of a Voting- Based POS Tagger in the Urdu Language
Popis výsledku anglicky
The process of sequence labeling (POS) by assigning syntactic tags to words in the given context is an important role in various NLP applications. The core motive of this work is to tackle the morpho- syntactic category of words in Urdu language. This language has lots of computational challenges because of its dual nature. The work comprises different tasks as initially the authors tracked the best combination of feature sets in terms of CRF to entitle the previous results on two stable and wellknown datasets Bushra Jawaid dataset and CLE dataset. Due to syntactic ambiguity, a state- of- the- art voting method has been introduced which is being implemented to overcome the contradictory results of the different machine learning classifiers. The results show significant improvement in the baseline results as the F1- score on a primary dataset is 94.8% and 95.7% on the succeeding dataset. Long short- term memory (LSTM) is used for one of the most diverse and inflectional tasks like part of speech tagging for the Urdu language by achieving an F1- score of 86.7% and 96.1% respectively for both datasets. © 2025 by IGI Global Scientific Publishing.
Klasifikace
Druh
C - Kapitola v odborné knize
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název knihy nebo sborníku
Innovations in Optimization and Machine Learning
ISBN
979-8-3693-5233-5
Počet stran výsledku
23
Strana od-do
23-45
Počet stran knihy
504
Název nakladatele
IGI Global
Místo vydání
—
Kód UT WoS kapitoly
—