Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Core vocabulary reveals differences between human word prediction and large language models

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AQQ9DUQJU" target="_blank" >RIV/00216208:11320/26:QQ9DUQJU - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://osf.io/preprints/psyarxiv/dpgac_v2/" target="_blank" >https://osf.io/preprints/psyarxiv/dpgac_v2/</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.31234/osf.io/dpgac_v2" target="_blank" >10.31234/osf.io/dpgac_v2</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Core vocabulary reveals differences between human word prediction and large language models

  • Popis výsledku v původním jazyce

    The question of which words are the most central or important to a language has been explored in various ways. In this study, we propose definitions of core vocabulary that are based on how language is learned, represented, and processed from psychological perspectives, and test these on a word prediction task. We aim to (1) compare core vocabulary based on word frequency in natural language, the content of word associations, and age-of-acquisition in terms of how well they are guessed in word prediction contexts, and (2) investigate the extent to which word prediction in language models aligns with humans, and if there are systematic differences between them, whether these can be captured by core vocabulary measures. Across two experiments, 867 participants completed a task which involved guessing target words that were missing from sentence contexts. Natural language-based core words were generally easier to guess, but when the degree to which these words were naturally predictable in the linguistic environment was taken into account, word association- and acquisition-based core words were easier to predict for reasons that went beyond this. Additionally, language models were able to account for people’s word prediction responses to a considerable extent, but there were also systematic deviations from these predictions which were able to be captured by word association-based coreness. The findings suggest that distributional relationships between words in text is not all there is to human word prediction, but that people may also rely on factors like communicative usefulness and multimodal or extralinguistic information.

  • Název v anglickém jazyce

    Core vocabulary reveals differences between human word prediction and large language models

  • Popis výsledku anglicky

    The question of which words are the most central or important to a language has been explored in various ways. In this study, we propose definitions of core vocabulary that are based on how language is learned, represented, and processed from psychological perspectives, and test these on a word prediction task. We aim to (1) compare core vocabulary based on word frequency in natural language, the content of word associations, and age-of-acquisition in terms of how well they are guessed in word prediction contexts, and (2) investigate the extent to which word prediction in language models aligns with humans, and if there are systematic differences between them, whether these can be captured by core vocabulary measures. Across two experiments, 867 participants completed a task which involved guessing target words that were missing from sentence contexts. Natural language-based core words were generally easier to guess, but when the degree to which these words were naturally predictable in the linguistic environment was taken into account, word association- and acquisition-based core words were easier to predict for reasons that went beyond this. Additionally, language models were able to account for people’s word prediction responses to a considerable extent, but there were also systematic deviations from these predictions which were able to be captured by word association-based coreness. The findings suggest that distributional relationships between words in text is not all there is to human word prediction, but that people may also rely on factors like communicative usefulness and multimodal or extralinguistic information.

Klasifikace

  • Druh

    O - Ostatní výsledky

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů