AI Koditex v1
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10507839" target="_blank" >RIV/00216208:11210/25:10507839 - isvavai.cz</a>
Výsledek na webu
<a href="http://hdl.handle.net/11234/1-5991" target="_blank" >http://hdl.handle.net/11234/1-5991</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
AI Koditex v1
Popis výsledku v původním jazyce
AI Koditex is a corpus of Czech texts generated with large language models (LLMs). Its main purpose is to create a resource for comparing human-written texts with LLM-generated text linguistically. The corpus is multi-genre and rich in terms of topics, authors, and text types, and comparabile with existing human-created corpora. The corpus replicates reference human Koditex corpus that follows the Brown Corpus tradition. The new corpus was generated using models from OpenAI, Anthropic, Alphabet, Meta, and DeepSeek, ranging from GPT-3 (davinci-002) to GPT-4.5, and are tagged according to the Universal Dependencies standard (i.e., the texts are tokenized, lemmatized, and morphologically and syntactically annotated). The subcorpus size varies according to the model used (The subcorpus size varies according to the model used (768k tokens per model on average, 21.5M tokens altogether). The raw data and plain texts are freely available for download under the CC BY 4.0 license, the UD annotated data are under CC BY-NC-SA 4.0 licence. The corpus is also accessible through the KonText search interface of the Czech National Corpus (https://www.korpus.cz/kontext/query?corpname=ai_koditex_v1).
Název v anglickém jazyce
AI Koditex v1
Popis výsledku anglicky
AI Koditex is a corpus of Czech texts generated with large language models (LLMs). Its main purpose is to create a resource for comparing human-written texts with LLM-generated text linguistically. The corpus is multi-genre and rich in terms of topics, authors, and text types, and comparabile with existing human-created corpora. The corpus replicates reference human Koditex corpus that follows the Brown Corpus tradition. The new corpus was generated using models from OpenAI, Anthropic, Alphabet, Meta, and DeepSeek, ranging from GPT-3 (davinci-002) to GPT-4.5, and are tagged according to the Universal Dependencies standard (i.e., the texts are tokenized, lemmatized, and morphologically and syntactically annotated). The subcorpus size varies according to the model used (The subcorpus size varies according to the model used (768k tokens per model on average, 21.5M tokens altogether). The raw data and plain texts are freely available for download under the CC BY 4.0 license, the UD annotated data are under CC BY-NC-SA 4.0 licence. The corpus is also accessible through the KonText search interface of the Czech National Corpus (https://www.korpus.cz/kontext/query?corpname=ai_koditex_v1).
Klasifikace
Druh
X - Nezařazeno
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
Výsledek vznikl pri realizaci vícero projektů. Více informací v záložce Projekty.
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů