Probing a pretrained RoBERTa on Khasi language for POS tagging
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F60076658%3A12520%2F25%3A43909662" target="_blank" >RIV/60076658:12520/25:43909662 - isvavai.cz</a>
Result on the web
<a href="https://doi.org/10.1017/nlp.2024.24" target="_blank" >https://doi.org/10.1017/nlp.2024.24</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1017/nlp.2024.24" target="_blank" >10.1017/nlp.2024.24</a>
Alternative languages
Result language
angličtina
Original language name
Probing a pretrained RoBERTa on Khasi language for POS tagging
Original language description
Part of speech (POS) tagging, though considered to be preliminary to any Natural Language Processing (NLP) task, is crucial to account for, especially in low resource language like Khasi that lacks any form of formal corpus. POS tagging is context sensitive. Therefore, the task is challenging. In this paper, we attempt to investigate a deep learning approach to the POS tagging problem in Khasi. A deep learning model called Robustly Optimized BERT Pretraining Approach (RoBERTa) is pretrained for language modelling task. We then create RoBERTa for POS (RoPOS) tagging, a model that performs POS tagging by fine-tuning the pretrained RoBERTa and leveraging its embeddings for downstream POS tagging. The existing tagset that has been designed, customarily, for the Khasi language is employed for this work, and the corresponding tagged dataset is taken as our base corpus. Further, we propose additional tags to this existing tagset to meet the requirements of the language and have increased the size of the existing Khasi POS corpus. Other machine learning and deep learning models have also been tried and tested for the same task, and a comparative analysis is made on the various models employed. Two different setups have been used for the RoPOS model, and the best testing accuracy achieved is 92 per cent. Comparative analysis of RoPOS with the other models indicates that RoPOS outperforms the others when used for inferencing on texts that are outside the domain of the POS tagged training dataset.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Natural Language Processing
ISSN
—
e-ISSN
2977-0424
Volume of the periodical
31
Issue of the periodical within the volume
2
Country of publishing house
GB - UNITED KINGDOM
Number of pages
20
Pages from-to
230-249
UT code for WoS article
001327418500001
EID of the result in the Scopus database
—