Iconicity in large language models
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10507824" target="_blank" >RIV/00216208:11210/25:10507824 - isvavai.cz</a>
Alternative codes found
RIV/00216208:11320/26:UUGWL4F6 RIV/61989592:15210/25:73635953
Result on the web
<a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=qqTCHvc.qc" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=qqTCHvc.qc</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/llc/fqaf095" target="_blank" >10.1093/llc/fqaf095</a>
Alternative languages
Result language
angličtina
Original language name
Iconicity in large language models
Original language description
Lexical iconicity, a direct relation between a word's meaning and its form, is an important aspect of every natural language, most commonly manifesting through sound-meaning associations. Since Large language models' (LLMs') access to both meaning and sound of text is only mediated (meaning through textual context, sound through written representation, further complicated by tokenization), we might expect that the encoding of iconicity in LLMs would be either insufficient or significantly different from human processing. This study addresses this hypothesis by having GPT-4 generate highly iconic pseudowords in artificial languages. To verify that these words actually carry iconicity, we had their meanings guessed by Czech and German participants (n = 672) and subsequently by LLM-based participants (generated by GPT-4 and Claude 3.5 Sonnet). The results revealed that humans can guess the meanings of pseudowords in the generated iconic language more accurately than words in distant natural languages and that LLM-based participants are even more successful than humans in this task. This core finding is accompanied by several additional analyses concerning the universality of the generated language and the cues that both human and LLM-based participants utilize.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
60203 - Linguistics
Result continuities
Project
<a href="/en/project/GA24-11725S" target="_blank" >GA24-11725S: Large language models through the prism of corpus linguistics</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Digital Scholarship in the Humanities
ISSN
2055-7671
e-ISSN
2055-768X
Volume of the periodical
40
Issue of the periodical within the volume
4
Country of publishing house
US - UNITED STATES
Number of pages
22
Pages from-to
1203-1224
UT code for WoS article
001574575700001
EID of the result in the Scopus database
2-s2.0-105020300683