CapekDraCor database and some aspects of quantitative linguistic analysis of the Čapek brothers’ plays
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989592%3A15210%2F25%3A73636249" target="_blank" >RIV/61989592:15210/25:73636249 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1075/cilt.370.15por" target="_blank" >http://dx.doi.org/10.1075/cilt.370.15por</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1075/cilt.370.15por" target="_blank" >10.1075/cilt.370.15por</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
CapekDraCor database and some aspects of quantitative linguistic analysis of the Čapek brothers’ plays
Popis výsledku v původním jazyce
This chapter of a methodological nature presents the newly created CapekDraCor database, which serves as a source of data for quantitative linguistic analysis of the plays of the Čapek brothers. The text focuses on the issue of the specific multi-layered structure of dramas and emphasizes the necessity of separating the dialogues of characters from metatextual elements (stage directions, authorial comments, and character labels) to ensure valid statistical outcomes. By comparing the database with the "capek" corpus in the Czech National Corpus, the analysis illustrates the risks of distorting linguistic parameters, such as lemma/keyword frequencies or part-of-speech distributions, in the absence of adequate text segmentation. The text also offers a comparison of the genre specificities of Karel Čapek's dramatic works using lexicostatistical indexes (ATL, VD, activity/descriptivity index) and demonstrates the application possibilities of the CapekDraCor database using the example of keyword analysis (TF*IDF, specificity score), character co-occurrence networks, and sentiment analysis in the plays The White Disease and R.U.R.
Název v anglickém jazyce
CapekDraCor database and some aspects of quantitative linguistic analysis of the Čapek brothers’ plays
Popis výsledku anglicky
This chapter of a methodological nature presents the newly created CapekDraCor database, which serves as a source of data for quantitative linguistic analysis of the plays of the Čapek brothers. The text focuses on the issue of the specific multi-layered structure of dramas and emphasizes the necessity of separating the dialogues of characters from metatextual elements (stage directions, authorial comments, and character labels) to ensure valid statistical outcomes. By comparing the database with the "capek" corpus in the Czech National Corpus, the analysis illustrates the risks of distorting linguistic parameters, such as lemma/keyword frequencies or part-of-speech distributions, in the absence of adequate text segmentation. The text also offers a comparison of the genre specificities of Karel Čapek's dramatic works using lexicostatistical indexes (ATL, VD, activity/descriptivity index) and demonstrates the application possibilities of the CapekDraCor database using the example of keyword analysis (TF*IDF, specificity score), character co-occurrence networks, and sentiment analysis in the plays The White Disease and R.U.R.
Klasifikace
Druh
C - Kapitola v odborné knize
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název knihy nebo sborníku
Mathematical Modelling in Linguistics and Text Analysis Theory and applications
ISBN
978-90-272-4448-2
Počet stran výsledku
18
Strana od-do
173-190
Počet stran knihy
239
Název nakladatele
John Benjamins Publishing Company
Místo vydání
Amsterdam
Kód UT WoS kapitoly
001609792500016