Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AWR5XKD8T" target="_blank" >RIV/00216208:11320/26:WR5XKD8T - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://ieeexplore.ieee.org/abstract/document/11145067" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11145067</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1109/ACCESS.2025.3604068" target="_blank" >10.1109/ACCESS.2025.3604068</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts

  • Popis výsledku v původním jazyce

    Twitter texts related to crisis events often contain vital information about the impact of disasters. However, irregular language structures, slang, and character limitations pose challenges for automated information extraction related to disaster impact. The Semantic Role Labeling (SRL) task can assist in extracting this information by identifying semantic roles in a sentence, such as who was involved, what happened, when, and where. This task provides a more structured understanding of the information related to disaster management. However, traditional based SRL tends to use predicates as the main anchor, which is less flexible in capturing semantic role relationships in unstructured texts such as Twitter. In addition, SRL research in low-resource languages such as Indonesian is still limited due to the lack of labeled datasets. This study aims to develop a more flexible SRL without relying on predicates as the main anchor and applying a semi-supervised learning strategy with pseudo-labeling. This strategy lets the model gradually learn from unlabeled data, enhancing generalization without solely depending on labeled data. A filtering function is implemented to minimize noise from inaccurate pseudo-labels, eliminating low-confidence predictions and ensuring that only high-quality data is used for retraining. Transformer-based models were chosen because of their self-attention mechanisms that enable an adequate understanding of contextual semantic relationships. Experiments demonstrate that the transformer model with Bidirectional Encoder Representations from Transformers (BERT) based architecture adapted for a specific language achieves the highest performance, with an F1-score of 0.863 at a threshold of 0.9. This study contributes to the advancement of SRL for low-resource language (especially for Indonesian) and paves the way for using SRL results in disaster severity analysis to enhance social media-based emergency response.

  • Název v anglickém jazyce

    Transformer-Based Semantic Role Labeling for Crisis Events Using Semi-Supervised Learning on Low-Resource Language Twitter Texts

  • Popis výsledku anglicky

    Twitter texts related to crisis events often contain vital information about the impact of disasters. However, irregular language structures, slang, and character limitations pose challenges for automated information extraction related to disaster impact. The Semantic Role Labeling (SRL) task can assist in extracting this information by identifying semantic roles in a sentence, such as who was involved, what happened, when, and where. This task provides a more structured understanding of the information related to disaster management. However, traditional based SRL tends to use predicates as the main anchor, which is less flexible in capturing semantic role relationships in unstructured texts such as Twitter. In addition, SRL research in low-resource languages such as Indonesian is still limited due to the lack of labeled datasets. This study aims to develop a more flexible SRL without relying on predicates as the main anchor and applying a semi-supervised learning strategy with pseudo-labeling. This strategy lets the model gradually learn from unlabeled data, enhancing generalization without solely depending on labeled data. A filtering function is implemented to minimize noise from inaccurate pseudo-labels, eliminating low-confidence predictions and ensuring that only high-quality data is used for retraining. Transformer-based models were chosen because of their self-attention mechanisms that enable an adequate understanding of contextual semantic relationships. Experiments demonstrate that the transformer model with Bidirectional Encoder Representations from Transformers (BERT) based architecture adapted for a specific language achieves the highest performance, with an F1-score of 0.863 at a threshold of 0.9. This study contributes to the advancement of SRL for low-resource language (especially for Indonesian) and paves the way for using SRL results in disaster severity analysis to enhance social media-based emergency response.

Klasifikace

  • Druh

    J<sub>ost</sub> - Ostatní články v recenzovaných periodicích

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    IEEE Access

  • ISSN

    2169-3536

  • e-ISSN

  • Svazek periodika

    13

  • Číslo periodika v rámci svazku

    2025

  • Stát vydavatele periodika

    US - Spojené státy americké

  • Počet stran výsledku

    29

  • Strana od-do

    158938-158966

  • Kód UT WoS článku

  • EID výsledku v databázi Scopus