LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F49777513%3A23520%2F25%3A43976214" target="_blank" >RIV/49777513:23520/25:43976214 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2025.acl-long.41/" target="_blank" >https://aclanthology.org/2025.acl-long.41/</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.18653/v1/2025.acl-long.41" target="_blank" >10.18653/v1/2025.acl-long.41</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation
Popis výsledku v původním jazyce
Cross-lingual aspect-based sentiment analysis (ABSA) involves detailed sentiment analysis in a target language by transferring knowledge from a source language with available annotated data. Most existing methods depend heavily on often unreliable translation tools to bridge the language gap. In this paper, we propose a new approach that leverages a large language model (LLM) to generate high-quality pseudo-labelled data in the target language without the need for translation tools. First, the framework trains an ABSA model to obtain predictions for unlabelled target language data. Next, LLM is prompted to generate natural sentences that better represent these noisy predictions than the original text. The ABSA model is then further fine-tuned on the resulting pseudo-labelled dataset. We demonstrate the effectiveness of this method across six languages and five backbone models, surpassing previous state-of-the-art translation-based approaches. The proposed framework also supports generative models, and we show that fine-tuned LLMs outperform smaller multilingual models.
Název v anglickém jazyce
LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation
Popis výsledku anglicky
Cross-lingual aspect-based sentiment analysis (ABSA) involves detailed sentiment analysis in a target language by transferring knowledge from a source language with available annotated data. Most existing methods depend heavily on often unreliable translation tools to bridge the language gap. In this paper, we propose a new approach that leverages a large language model (LLM) to generate high-quality pseudo-labelled data in the target language without the need for translation tools. First, the framework trains an ABSA model to obtain predictions for unlabelled target language data. Next, LLM is prompted to generate natural sentences that better represent these noisy predictions than the original text. The ABSA model is then further fine-tuned on the resulting pseudo-labelled dataset. We demonstrate the effectiveness of this method across six languages and five backbone models, surpassing previous state-of-the-art translation-based approaches. The proposed framework also supports generative models, and we show that fine-tuned LLMs outperform smaller multilingual models.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_021%2F0008436" target="_blank" >EH23_021/0008436: VaV technologií pro pokročilou digitalizaci v plzeňské metropolitní oblasti (DigiTech)</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
ISBN
979-8-89176-251-0
ISSN
—
e-ISSN
—
Počet stran výsledku
15
Strana od-do
839-853
Název nakladatele
Association for Computational Linguistics
Místo vydání
Kerrville
Místo konání akce
Vídeň
Datum konání akce
27. 7. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001596029800041