All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Out-Heroding Herod? — Author-trained GPTs and Original Works in the Perspective of Quantitative Linguistics

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61988987%3A17250%2F25%3AA2603D2P" target="_blank" >RIV/61988987:17250/25:A2603D2P - isvavai.cz</a>

  • Result on the web

    <a href="https://ojs.cuni.cz/pedagogika/article/view/4945" target="_blank" >https://ojs.cuni.cz/pedagogika/article/view/4945</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.14712/23362189.2025.4945" target="_blank" >10.14712/23362189.2025.4945</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Out-Heroding Herod? — Author-trained GPTs and Original Works in the Perspective of Quantitative Linguistics

  • Original language description

    Goals: The paper compares texts created by GPT models trained on the works of prominent Czech authors and the pieces of literature they actually wrote. The goal is to find out (1) whether there are any differences between the two; and if so, (2) in what sphere of language these differences are the most prominent.Methods: The authors used for building GPTs are Karel Čapek, Jaroslav Hašek, Franz Kafka, and Vladislav Vančura. The corpus contains 40 1,000-word text samples per each, 20 of them produced by the respective GPT and 20 taken from the original works. Two investigations are carried out – the first includes calculating 30 morphological, syntactic, and lexical markers for each text; the second  is based on most-frequent-element analyses. The results of the first set are tested on statistical significance via Mann–Whitney U Test.Results: The chatbots do not reflect colloquiality of style and conversation interaction very well, and tend to make texts more narrative. The best results are obtained for Karel Čapek, the worst for Franz Kafka. The stylometric analyses almost always distinguish the AI- and human-generated pieces of language.Conclusions: The texts produced by the author-trained GPTs are still very well distinguishable from those produced by real writers.

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>ost</sub> - Miscellaneous article in a specialist periodical

  • CEP classification

  • OECD FORD branch

    60203 - Linguistics

Result continuities

  • Project

  • Continuities

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    Pedagogika

  • ISSN

    0031-3815

  • e-ISSN

    2336-2189

  • Volume of the periodical

  • Issue of the periodical within the volume

    4

  • Country of publishing house

    CZ - CZECH REPUBLIC

  • Number of pages

    25

  • Pages from-to

    355-379

  • UT code for WoS article

  • EID of the result in the Scopus database