LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F49777513%3A23520%2F25%3A43976214" target="_blank" >RIV/49777513:23520/25:43976214 - isvavai.cz</a>
Result on the web
<a href="https://aclanthology.org/2025.acl-long.41/" target="_blank" >https://aclanthology.org/2025.acl-long.41/</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.18653/v1/2025.acl-long.41" target="_blank" >10.18653/v1/2025.acl-long.41</a>
Alternative languages
Result language
angličtina
Original language name
LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation
Original language description
Cross-lingual aspect-based sentiment analysis (ABSA) involves detailed sentiment analysis in a target language by transferring knowledge from a source language with available annotated data. Most existing methods depend heavily on often unreliable translation tools to bridge the language gap. In this paper, we propose a new approach that leverages a large language model (LLM) to generate high-quality pseudo-labelled data in the target language without the need for translation tools. First, the framework trains an ABSA model to obtain predictions for unlabelled target language data. Next, LLM is prompted to generate natural sentences that better represent these noisy predictions than the original text. The ABSA model is then further fine-tuned on the resulting pseudo-labelled dataset. We demonstrate the effectiveness of this method across six languages and five backbone models, surpassing previous state-of-the-art translation-based approaches. The proposed framework also supports generative models, and we show that fine-tuned LLMs outperform smaller multilingual models.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/EH23_021%2F0008436" target="_blank" >EH23_021/0008436: RandD of technologies for advanced digitization in the Pilsen metropolitan area (DigiTech)</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
ISBN
979-8-89176-251-0
ISSN
—
e-ISSN
—
Number of pages
15
Pages from-to
839-853
Publisher name
Association for Computational Linguistics
Place of publication
Kerrville
Event location
Vídeň
Event date
Jul 27, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
001596029800041