Character-level inclusive transformer architecture for information gain in low resource code-mixed language
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A53BDDJYW" target="_blank" >RIV/00216208:11320/26:53BDDJYW - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1007/s00521-022-06983-2" target="_blank" >http://dx.doi.org/10.1007/s00521-022-06983-2</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/s00521-022-06983-2" target="_blank" >10.1007/s00521-022-06983-2</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Character-level inclusive transformer architecture for information gain in low resource code-mixed language
Popis výsledku v původním jazyce
The use of code-mixed languages in social media platforms is very common to communicate in an informal way and has immense importance in a multilingual society, like India. Implementing various NLP tasks on code-mixed language for machine comprehension and NLP applications is the need of the hour. The implementation of complex learning models is difficult due to the scarcity of available code-mixed resources. Designing more effective architectures to perform learning from low resource dataset along with transfer learning settings are the possible solutions. We propose an improvised transformer network (Character Inclusion Transformer) that utilizes and learns character-level information available in the words of code-mixed sentences. The proposed model improves the performance of the transformer model when trained from scratch using low resource code-mixed datasets. We also propose two more architecture settings, useful for transfer learning strategy using the mBERT pre-trained model. Three basic word-level tagging NLP tasks, i.e., NER, POS Tagging, and Language Identification (LID) are considered in the paper where Language Identification is specific to code-mixed language. Six separate datasets, namely IIITH NER, LID FIRE, LID ICON, LID UD, POS ICON, POS UD, have been tested, and results are reported using weighted and macro-average while evaluating precision, recall and F1 score © The Author(s), under exclusive licence to Springer-Verlag London Ltd., part of Springer Nature 2022.
Název v anglickém jazyce
Character-level inclusive transformer architecture for information gain in low resource code-mixed language
Popis výsledku anglicky
The use of code-mixed languages in social media platforms is very common to communicate in an informal way and has immense importance in a multilingual society, like India. Implementing various NLP tasks on code-mixed language for machine comprehension and NLP applications is the need of the hour. The implementation of complex learning models is difficult due to the scarcity of available code-mixed resources. Designing more effective architectures to perform learning from low resource dataset along with transfer learning settings are the possible solutions. We propose an improvised transformer network (Character Inclusion Transformer) that utilizes and learns character-level information available in the words of code-mixed sentences. The proposed model improves the performance of the transformer model when trained from scratch using low resource code-mixed datasets. We also propose two more architecture settings, useful for transfer learning strategy using the mBERT pre-trained model. Three basic word-level tagging NLP tasks, i.e., NER, POS Tagging, and Language Identification (LID) are considered in the paper where Language Identification is specific to code-mixed language. Six separate datasets, namely IIITH NER, LID FIRE, LID ICON, LID UD, POS ICON, POS UD, have been tested, and results are reported using weighted and macro-average while evaluating precision, recall and F1 score © The Author(s), under exclusive licence to Springer-Verlag London Ltd., part of Springer Nature 2022.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Neural Computing and Applications
ISSN
0941-0643
e-ISSN
—
Svazek periodika
37
Číslo periodika v rámci svazku
2
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
19
Strana od-do
559-577
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-85125946949