Out-Heroding Herod? — Author-trained GPTs and Original Works in the Perspective of Quantitative Linguistics
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61988987%3A17250%2F25%3AA2603D2P" target="_blank" >RIV/61988987:17250/25:A2603D2P - isvavai.cz</a>
Výsledek na webu
<a href="https://ojs.cuni.cz/pedagogika/article/view/4945" target="_blank" >https://ojs.cuni.cz/pedagogika/article/view/4945</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.14712/23362189.2025.4945" target="_blank" >10.14712/23362189.2025.4945</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Out-Heroding Herod? — Author-trained GPTs and Original Works in the Perspective of Quantitative Linguistics
Popis výsledku v původním jazyce
Goals: The paper compares texts created by GPT models trained on the works of prominent Czech authors and the pieces of literature they actually wrote. The goal is to find out (1) whether there are any differences between the two; and if so, (2) in what sphere of language these differences are the most prominent.Methods: The authors used for building GPTs are Karel Čapek, Jaroslav Hašek, Franz Kafka, and Vladislav Vančura. The corpus contains 40 1,000-word text samples per each, 20 of them produced by the respective GPT and 20 taken from the original works. Two investigations are carried out – the first includes calculating 30 morphological, syntactic, and lexical markers for each text; the second is based on most-frequent-element analyses. The results of the first set are tested on statistical significance via Mann–Whitney U Test.Results: The chatbots do not reflect colloquiality of style and conversation interaction very well, and tend to make texts more narrative. The best results are obtained for Karel Čapek, the worst for Franz Kafka. The stylometric analyses almost always distinguish the AI- and human-generated pieces of language.Conclusions: The texts produced by the author-trained GPTs are still very well distinguishable from those produced by real writers.
Název v anglickém jazyce
Out-Heroding Herod? — Author-trained GPTs and Original Works in the Perspective of Quantitative Linguistics
Popis výsledku anglicky
Goals: The paper compares texts created by GPT models trained on the works of prominent Czech authors and the pieces of literature they actually wrote. The goal is to find out (1) whether there are any differences between the two; and if so, (2) in what sphere of language these differences are the most prominent.Methods: The authors used for building GPTs are Karel Čapek, Jaroslav Hašek, Franz Kafka, and Vladislav Vančura. The corpus contains 40 1,000-word text samples per each, 20 of them produced by the respective GPT and 20 taken from the original works. Two investigations are carried out – the first includes calculating 30 morphological, syntactic, and lexical markers for each text; the second is based on most-frequent-element analyses. The results of the first set are tested on statistical significance via Mann–Whitney U Test.Results: The chatbots do not reflect colloquiality of style and conversation interaction very well, and tend to make texts more narrative. The best results are obtained for Karel Čapek, the worst for Franz Kafka. The stylometric analyses almost always distinguish the AI- and human-generated pieces of language.Conclusions: The texts produced by the author-trained GPTs are still very well distinguishable from those produced by real writers.
Klasifikace
Druh
J<sub>ost</sub> - Ostatní články v recenzovaných periodicích
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Pedagogika
ISSN
0031-3815
e-ISSN
2336-2189
Svazek periodika
—
Číslo periodika v rámci svazku
4
Stát vydavatele periodika
CZ - Česká republika
Počet stran výsledku
25
Strana od-do
355-379
Kód UT WoS článku
—
EID výsledku v databázi Scopus
—