All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Iconicity in large language models

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10507824" target="_blank" >RIV/00216208:11210/25:10507824 - isvavai.cz</a>

  • Alternative codes found

    RIV/00216208:11320/26:UUGWL4F6 RIV/61989592:15210/25:73635953

  • Result on the web

    <a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=qqTCHvc.qc" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=qqTCHvc.qc</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1093/llc/fqaf095" target="_blank" >10.1093/llc/fqaf095</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Iconicity in large language models

  • Original language description

    Lexical iconicity, a direct relation between a word&apos;s meaning and its form, is an important aspect of every natural language, most commonly manifesting through sound-meaning associations. Since Large language models&apos; (LLMs&apos;) access to both meaning and sound of text is only mediated (meaning through textual context, sound through written representation, further complicated by tokenization), we might expect that the encoding of iconicity in LLMs would be either insufficient or significantly different from human processing. This study addresses this hypothesis by having GPT-4 generate highly iconic pseudowords in artificial languages. To verify that these words actually carry iconicity, we had their meanings guessed by Czech and German participants (n = 672) and subsequently by LLM-based participants (generated by GPT-4 and Claude 3.5 Sonnet). The results revealed that humans can guess the meanings of pseudowords in the generated iconic language more accurately than words in distant natural languages and that LLM-based participants are even more successful than humans in this task. This core finding is accompanied by several additional analyses concerning the universality of the generated language and the cues that both human and LLM-based participants utilize.

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database

  • CEP classification

  • OECD FORD branch

    60203 - Linguistics

Result continuities

  • Project

    <a href="/en/project/GA24-11725S" target="_blank" >GA24-11725S: Large language models through the prism of corpus linguistics</a><br>

  • Continuities

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    Digital Scholarship in the Humanities

  • ISSN

    2055-7671

  • e-ISSN

    2055-768X

  • Volume of the periodical

    40

  • Issue of the periodical within the volume

    4

  • Country of publishing house

    US - UNITED STATES

  • Number of pages

    22

  • Pages from-to

    1203-1224

  • UT code for WoS article

    001574575700001

  • EID of the result in the Scopus database

    2-s2.0-105020300683