All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

CapekDraCor database and some aspects of quantitative linguistic analysis of the Čapek brothers’ plays

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989592%3A15210%2F25%3A73636249" target="_blank" >RIV/61989592:15210/25:73636249 - isvavai.cz</a>

  • Result on the web

    <a href="http://dx.doi.org/10.1075/cilt.370.15por" target="_blank" >http://dx.doi.org/10.1075/cilt.370.15por</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1075/cilt.370.15por" target="_blank" >10.1075/cilt.370.15por</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    CapekDraCor database and some aspects of quantitative linguistic analysis of the Čapek brothers’ plays

  • Original language description

    This chapter of a methodological nature presents the newly created CapekDraCor database, which serves as a source of data for quantitative linguistic analysis of the plays of the Čapek brothers. The text focuses on the issue of the specific multi-layered structure of dramas and emphasizes the necessity of separating the dialogues of characters from metatextual elements (stage directions, authorial comments, and character labels) to ensure valid statistical outcomes. By comparing the database with the &quot;capek&quot; corpus in the Czech National Corpus, the analysis illustrates the risks of distorting linguistic parameters, such as lemma/keyword frequencies or part-of-speech distributions, in the absence of adequate text segmentation. The text also offers a comparison of the genre specificities of Karel Čapek&apos;s dramatic works using lexicostatistical indexes (ATL, VD, activity/descriptivity index) and demonstrates the application possibilities of the CapekDraCor database using the example of keyword analysis (TF*IDF, specificity score), character co-occurrence networks, and sentiment analysis in the plays The White Disease and R.U.R.

  • Czech name

  • Czech description

Classification

  • Type

    C - Chapter in a specialist book

  • CEP classification

  • OECD FORD branch

    60203 - Linguistics

Result continuities

  • Project

  • Continuities

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Book/collection name

    Mathematical Modelling in Linguistics and Text Analysis Theory and applications

  • ISBN

    978-90-272-4448-2

  • Number of pages of the result

    18

  • Pages from-to

    173-190

  • Number of pages of the book

    239

  • Publisher name

    John Benjamins Publishing Company

  • Place of publication

    Amsterdam

  • UT code for WoS chapter

    001609792500016