Comma Distribution In Czech Texts : Variation By Genre And Author, And Error Analysis
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14210%2F25%3A00142284" target="_blank" >RIV/00216224:14210/25:00142284 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.juls.savba.sk/ediela/jc/2025/1/jc25-01.pdf" target="_blank" >https://www.juls.savba.sk/ediela/jc/2025/1/jc25-01.pdf</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.2478/jazcas-2025-0004" target="_blank" >10.2478/jazcas-2025-0004</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Comma Distribution In Czech Texts : Variation By Genre And Author, And Error Analysis
Popis výsledku v původním jazyce
This article investigates the distribution and typology of commas in Czech texts, combining genre-differentiated samples with an annotated error corpus to offer a comprehensive view of punctuation usage and misuse. Building on previous work, we expand the analysis from a small newspaper sample to a broader set of texts, encompassing fiction, blogs, translations, and school dictations. Using a consistent typology of comma usage, we classify 1,000 manually selected instances and identify trends in different textual genres. Furthermore, we examine over 1,000 missing comma errors and more than 200 redundant ones from the self-built error corpus. The results reveal genre-dependent tendencies in comma types, especially in the use of commas preceding connectives and within asyndetic structures. The study offers insights for improving automatic comma insertion systems and deepens our understanding of punctuation norms and deviations in Czech.
Název v anglickém jazyce
Comma Distribution In Czech Texts : Variation By Genre And Author, And Error Analysis
Popis výsledku anglicky
This article investigates the distribution and typology of commas in Czech texts, combining genre-differentiated samples with an annotated error corpus to offer a comprehensive view of punctuation usage and misuse. Building on previous work, we expand the analysis from a small newspaper sample to a broader set of texts, encompassing fiction, blogs, translations, and school dictations. Using a consistent typology of comma usage, we classify 1,000 manually selected instances and identify trends in different textual genres. Furthermore, we examine over 1,000 missing comma errors and more than 200 redundant ones from the self-built error corpus. The results reveal genre-dependent tendencies in comma types, especially in the use of commas preceding connectives and within asyndetic structures. The study offers insights for improving automatic comma insertion systems and deepens our understanding of punctuation norms and deviations in Czech.
Klasifikace
Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
—
Návaznosti
R - Projekt Ramcoveho programu EK
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů