Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Domain-Specific Fine-Tuning of IndoBERT for Aspect-Based Sentiment Analysis in Indonesian Travel User-Generated Content

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A5D8FWCGB" target="_blank" >RIV/00216208:11320/26:5D8FWCGB - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://e-journal.unair.ac.id/JISEBI/article/view/63660" target="_blank" >https://e-journal.unair.ac.id/JISEBI/article/view/63660</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.20473/jisebi.11.1.30-40" target="_blank" >10.20473/jisebi.11.1.30-40</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Domain-Specific Fine-Tuning of IndoBERT for Aspect-Based Sentiment Analysis in Indonesian Travel User-Generated Content

  • Popis výsledku v původním jazyce

    Background: Aspect-based sentiment analysis (ABSA) is essential in extracting meaningful insights from user-generated content (UGC) in various domains. In tourism, UGC such as Google Reviews offers essential feedback, but the challenges associated with processing in Indonesian language, including the unique linguistic characteristics, pose difficulties for automatic sentiment, and aspect detection. Recent advancements in transformer-based models, such as BERT, have shown great potential in addressing these challenges by providing context-aware embeddings. Objective: This research aimed to fine-tune IndoBERT, a pre-trained Indonesian language model, to perform information extraction and key aspect detection from tourism-related UGC. The objective was to identify critical aspects of tourism reviews and classify their sentiments. Methods: A dataset of 20,000 Google Reviews, focusing on 20 tourism destinations in DI Yogyakarta and Jawa Tengah, was collected and preprocessed. Multiple fine-tuning experiments were conducted, using a layer-freezing method by adjusting only the top layers of IndoBERT, while freezing others to determine the optimal configuration. The model's performance was evaluated based on validation loss, precision, recall, and F1-score in aspect detection and overall sentiment classification accuracy. Results: The best-performing configuration involved freezing the last six layers and fine-tuning the top six layers of IndoBERT, yielding a validation loss of 0.324. The model achieved precision scores between 0.85 and 0.89 in aspect detection and an overall sentiment classification accuracy of 0.84. Error analysis revealed challenges in distinguishing neutral and negative sentiments and in handling reviews with multiple aspects or mixed sentiments. Conclusion: The fine-tuned IndoBERT model effectively extracted key tourism aspects and classified sentiments from Indonesian UGC. While the model performed well in detecting strong sentiments, improvements are needed to handle neutral and mixed sentiments better. Future work will explore sentiment intensity analysis and aspect segmentation methods to enhance the model's performance. Keywords: Aspect-Based Sentiment Analysis, Fine-tuning, IndoBERT, Sentiment Classification, Tourism Reviews, User-Generated Content

  • Název v anglickém jazyce

    Domain-Specific Fine-Tuning of IndoBERT for Aspect-Based Sentiment Analysis in Indonesian Travel User-Generated Content

  • Popis výsledku anglicky

    Background: Aspect-based sentiment analysis (ABSA) is essential in extracting meaningful insights from user-generated content (UGC) in various domains. In tourism, UGC such as Google Reviews offers essential feedback, but the challenges associated with processing in Indonesian language, including the unique linguistic characteristics, pose difficulties for automatic sentiment, and aspect detection. Recent advancements in transformer-based models, such as BERT, have shown great potential in addressing these challenges by providing context-aware embeddings. Objective: This research aimed to fine-tune IndoBERT, a pre-trained Indonesian language model, to perform information extraction and key aspect detection from tourism-related UGC. The objective was to identify critical aspects of tourism reviews and classify their sentiments. Methods: A dataset of 20,000 Google Reviews, focusing on 20 tourism destinations in DI Yogyakarta and Jawa Tengah, was collected and preprocessed. Multiple fine-tuning experiments were conducted, using a layer-freezing method by adjusting only the top layers of IndoBERT, while freezing others to determine the optimal configuration. The model's performance was evaluated based on validation loss, precision, recall, and F1-score in aspect detection and overall sentiment classification accuracy. Results: The best-performing configuration involved freezing the last six layers and fine-tuning the top six layers of IndoBERT, yielding a validation loss of 0.324. The model achieved precision scores between 0.85 and 0.89 in aspect detection and an overall sentiment classification accuracy of 0.84. Error analysis revealed challenges in distinguishing neutral and negative sentiments and in handling reviews with multiple aspects or mixed sentiments. Conclusion: The fine-tuned IndoBERT model effectively extracted key tourism aspects and classified sentiments from Indonesian UGC. While the model performed well in detecting strong sentiments, improvements are needed to handle neutral and mixed sentiments better. Future work will explore sentiment intensity analysis and aspect segmentation methods to enhance the model's performance. Keywords: Aspect-Based Sentiment Analysis, Fine-tuning, IndoBERT, Sentiment Classification, Tourism Reviews, User-Generated Content

Klasifikace

  • Druh

    J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    Journal of Information Systems Engineering and Business Intelligence

  • ISSN

    2443-2555

  • e-ISSN

  • Svazek periodika

    11

  • Číslo periodika v rámci svazku

    1

  • Stát vydavatele periodika

    US - Spojené státy americké

  • Počet stran výsledku

    11

  • Strana od-do

    30-40

  • Kód UT WoS článku

  • EID výsledku v databázi Scopus

    2-s2.0-105001868603