Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AWR5XKD8T" target="_blank" >RIV/00216208:11320/26:WR5XKD8T - isvavai.cz</a>
Result on the web
<a href="https://ieeexplore.ieee.org/abstract/document/11145067" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11145067</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ACCESS.2025.3604068" target="_blank" >10.1109/ACCESS.2025.3604068</a>
Alternative languages
Result language
angličtina
Original language name
Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts
Original language description
Twitter texts related to crisis events often contain vital information about the impact of disasters. However, irregular language structures, slang, and character limitations pose challenges for automated information extraction related to disaster impact. The Semantic Role Labeling (SRL) task can assist in extracting this information by identifying semantic roles in a sentence, such as who was involved, what happened, when, and where. This task provides a more structured understanding of the information related to disaster management. However, traditional based SRL tends to use predicates as the main anchor, which is less flexible in capturing semantic role relationships in unstructured texts such as Twitter. In addition, SRL research in low-resource languages such as Indonesian is still limited due to the lack of labeled datasets. This study aims to develop a more flexible SRL without relying on predicates as the main anchor and applying a semi-supervised learning strategy with pseudo-labeling. This strategy lets the model gradually learn from unlabeled data, enhancing generalization without solely depending on labeled data. A filtering function is implemented to minimize noise from inaccurate pseudo-labels, eliminating low-confidence predictions and ensuring that only high-quality data is used for retraining. Transformer-based models were chosen because of their self-attention mechanisms that enable an adequate understanding of contextual semantic relationships. Experiments demonstrate that the transformer model with Bidirectional Encoder Representations from Transformers (BERT) based architecture adapted for a specific language achieves the highest performance, with an F1-score of 0.863 at a threshold of 0.9. This study contributes to the advancement of SRL for low-resource language (especially for Indonesian) and paves the way for using SRL results in disaster severity analysis to enhance social media-based emergency response.
Czech name
—
Czech description
—
Classification
Type
J<sub>ost</sub> - Miscellaneous article in a specialist periodical
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
IEEE Access
ISSN
2169-3536
e-ISSN
—
Volume of the periodical
13
Issue of the periodical within the volume
2025
Country of publishing house
US - UNITED STATES
Number of pages
29
Pages from-to
158938-158966
UT code for WoS article
—
EID of the result in the Scopus database
—