LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0193745" target="_blank" >RIV/00216305:26230/26:0193745 - isvavai.cz</a>
Result on the web
<a href="https://aclanthology.org/2025.naacl-long.526/" target="_blank" >https://aclanthology.org/2025.naacl-long.526/</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.18653/v1/2025.naacl-long.526" target="_blank" >10.18653/v1/2025.naacl-long.526</a>
Alternative languages
Result language
angličtina
Original language name
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
Original language description
The generative large language models (LLMs) are increasingly being used for data augmentation tasks, where text samples are LLM-paraphrased and then used for classifier fine-tuning. Previous studies have compared LLM-based augmentations with established augmentation techniques, but the results are contradictory: some report superiority of LLM-based augmentations, while other only marginal increases (and even decreases) in performance of downstream classifiers. A research that would confirm a clear cost-benefit advantage of LLMs over more established augmentation methods is largely missing. To study if (and when) is the LLM-based augmentation advantageous, we compared the effects of recent LLM augmentation methods with established ones on 6 datasets, 3 classifiers and 2 fine-tuning methods. We also varied the number of seeds and collected samples to better explore the downstream model accuracy space. Finally, we performed a cost-benefit analysis and show that LLM-based methods are worthy of deployment only when very small number of seeds is used. Moreover, in many cases, established methods lead to similar or better model accuracies.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
ISBN
979-8-8917-6189-6
ISSN
—
e-ISSN
—
Number of pages
20
Pages from-to
10476-10496
Publisher name
Association for Computational Linguistics
Place of publication
Albuquerque, New Mexico
Event location
Albuquerque, New Mexico
Event date
Apr 29, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
001611654000207