Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AWR5XKD8T" target="_blank" >RIV/00216208:11320/26:WR5XKD8T - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/abstract/document/11145067" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11145067</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ACCESS.2025.3604068" target="_blank" >10.1109/ACCESS.2025.3604068</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts
Popis výsledku v původním jazyce
Twitter texts related to crisis events often contain vital information about the impact of disasters. However, irregular language structures, slang, and character limitations pose challenges for automated information extraction related to disaster impact. The Semantic Role Labeling (SRL) task can assist in extracting this information by identifying semantic roles in a sentence, such as who was involved, what happened, when, and where. This task provides a more structured understanding of the information related to disaster management. However, traditional based SRL tends to use predicates as the main anchor, which is less flexible in capturing semantic role relationships in unstructured texts such as Twitter. In addition, SRL research in low-resource languages such as Indonesian is still limited due to the lack of labeled datasets. This study aims to develop a more flexible SRL without relying on predicates as the main anchor and applying a semi-supervised learning strategy with pseudo-labeling. This strategy lets the model gradually learn from unlabeled data, enhancing generalization without solely depending on labeled data. A filtering function is implemented to minimize noise from inaccurate pseudo-labels, eliminating low-confidence predictions and ensuring that only high-quality data is used for retraining. Transformer-based models were chosen because of their self-attention mechanisms that enable an adequate understanding of contextual semantic relationships. Experiments demonstrate that the transformer model with Bidirectional Encoder Representations from Transformers (BERT) based architecture adapted for a specific language achieves the highest performance, with an F1-score of 0.863 at a threshold of 0.9. This study contributes to the advancement of SRL for low-resource language (especially for Indonesian) and paves the way for using SRL results in disaster severity analysis to enhance social media-based emergency response.
Název v anglickém jazyce
Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts
Popis výsledku anglicky
Twitter texts related to crisis events often contain vital information about the impact of disasters. However, irregular language structures, slang, and character limitations pose challenges for automated information extraction related to disaster impact. The Semantic Role Labeling (SRL) task can assist in extracting this information by identifying semantic roles in a sentence, such as who was involved, what happened, when, and where. This task provides a more structured understanding of the information related to disaster management. However, traditional based SRL tends to use predicates as the main anchor, which is less flexible in capturing semantic role relationships in unstructured texts such as Twitter. In addition, SRL research in low-resource languages such as Indonesian is still limited due to the lack of labeled datasets. This study aims to develop a more flexible SRL without relying on predicates as the main anchor and applying a semi-supervised learning strategy with pseudo-labeling. This strategy lets the model gradually learn from unlabeled data, enhancing generalization without solely depending on labeled data. A filtering function is implemented to minimize noise from inaccurate pseudo-labels, eliminating low-confidence predictions and ensuring that only high-quality data is used for retraining. Transformer-based models were chosen because of their self-attention mechanisms that enable an adequate understanding of contextual semantic relationships. Experiments demonstrate that the transformer model with Bidirectional Encoder Representations from Transformers (BERT) based architecture adapted for a specific language achieves the highest performance, with an F1-score of 0.863 at a threshold of 0.9. This study contributes to the advancement of SRL for low-resource language (especially for Indonesian) and paves the way for using SRL results in disaster severity analysis to enhance social media-based emergency response.
Klasifikace
Druh
J<sub>ost</sub> - Ostatní články v recenzovaných periodicích
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
IEEE Access
ISSN
2169-3536
e-ISSN
—
Svazek periodika
13
Číslo periodika v rámci svazku
2025
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
29
Strana od-do
158938-158966
Kód UT WoS článku
—
EID výsledku v databázi Scopus
—