All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0193746" target="_blank" >RIV/00216305:26230/26:0193746 - isvavai.cz</a>

  • Result on the web

    <a href="https://aclanthology.org/2025.findings-emnlp.296/" target="_blank" >https://aclanthology.org/2025.findings-emnlp.296/</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.18653/v1/2025.findings-emnlp.296" target="_blank" >10.18653/v1/2025.findings-emnlp.296</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation

  • Original language description

    The generative large language models (LLMs) are increasingly used for data augmentation tasks, where text samples are paraphrased (or generated anew) and then used for classifier fine-tuning. Existing works on augmentation leverage the few-shot scenarios, where samples are given to LLMs as part of prompts, leading to better augmentations. Yet, the samples are mostly selected randomly and a comprehensive overview of the effects of other (more 'informed') sample selection strategies is lacking. In this work, we compare sample selection strategies existing in few-shot learning literature and investigate their effects in LLM-based textual augmentation. We evaluate this on in-distribution and out-of-distribution classifier performance. Results indicate, that while some 'informed' selection strategies increase the performance of models, especially for out-of-distribution data, it happens only seldom and with marginal performance increases. Unless further advances are made, a default of random sample selection remains a good option for augmentation practitioners.

  • Czech name

  • Czech description

Classification

  • Type

    O - Miscellaneous

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů