BanglaLem: A Transformer-based Bangla Lemmatizer with an Enhanced Dataset
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A429ANMLZ" target="_blank" >RIV/00216208:11320/26:429ANMLZ - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.1016/j.sasc.2025.200244" target="_blank" >http://dx.doi.org/10.1016/j.sasc.2025.200244</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.sasc.2025.200244" target="_blank" >10.1016/j.sasc.2025.200244</a>
Alternative languages
Result language
angličtina
Original language name
BanglaLem: A Transformer-based Bangla Lemmatizer with an Enhanced Dataset
Original language description
Lemmatization plays a crucial role in various natural language processing (NLP) tasks, such as information retrieval, sentiment analysis, text summarization, and text classification. However, Bangla lemmatization remains particularly challenging due to the language's rich morphology and high inflectional complexity. Existing open-access datasets for Bangla lemmatization are limited in size, with the largest containing only 22353 unique inflected words, which constrains the effectiveness of data-driven neural models. To address this limitation, we introduce a novel dataset, BanglaLem, comprising 96040 frequently used inflected words. This dataset has been carefully curated and annotated through a rigorous selection process to enhance the accuracy and efficiency of Bangla lemmatization. Furthermore, we propose a transformer-based approach to lemmatization and evaluate the performance of various pre-trained and trained from-scratch transformer models on this dataset. Among these, the BanglaT5 model achieved the highest exact match accuracy of 94.42% on the test set. The BanglaLem dataset is publicly accessible via the following link. © 2025 The Authors
Czech name
—
Czech description
—
Classification
Type
J<sub>SC</sub> - Article in a specialist periodical, which is included in the SCOPUS database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Systems and Soft Computing
ISSN
2772-9419
e-ISSN
—
Volume of the periodical
7
Issue of the periodical within the volume
2025
Country of publishing house
US - UNITED STATES
Number of pages
27
Pages from-to
200244
UT code for WoS article
—
EID of the result in the Scopus database
2-s2.0-105003572025