Sibling-texts keyword analysis: exploring topic and register keywords
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10508091" target="_blank" >RIV/00216208:11210/25:10508091 - isvavai.cz</a>
Result on the web
<a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=U5ysKb.edD" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=U5ysKb.edD</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/llc/fqaf037" target="_blank" >10.1093/llc/fqaf037</a>
Alternative languages
Result language
angličtina
Original language name
Sibling-texts keyword analysis: exploring topic and register keywords
Original language description
This article introduces a novel method for refining approaches to distant reading by proposing a procedure that categorizes prominent units (keywords) into two types: those pertinent to the topic and those associated with the genre/register of a text. This differentiation holds significant potential for more accurate modeling of topics and further applications in various domains of digital humanities. For example, register-related keywords may assist in various stylometric tasks, such as functional understanding of text groups (formed through clustering) that exhibit common stylistic traits, thus differentiating themselves from other groups or clusters within the discourse under scrutiny. In the initial step of the procedure, a set of texts with similar genre and register characteristics is identified; we term them 'sibling texts', and their similarity is determined using a model derived from multidimensional analysis. The next step involves conducting multiple parallel keyword analyses on these sibling texts and comparing them for keyword overlaps which mark the register relatedness. The test study on a corpus of Czech parliamentary speeches (Parlcorp) underscores the method's ability to distinguish register-related and topic-related keywords, even in highly homogeneous corpora. The register-related keywords proved to be analytically very useful in substantiating the interpretation of parliamentary subregisters and advancing the understanding of the types of activities (administrative and procedural, deliberating and debating, and government policies) involved in parliamentary discourse.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
60203 - Linguistics
Result continuities
Project
<a href="/en/project/EH22_008%2F0004595" target="_blank" >EH22_008/0004595: Beyond Security: Role of Conflict in Resilience-Building</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Digital Scholarship in the Humanities
ISSN
2055-7671
e-ISSN
2055-768X
Volume of the periodical
40
Issue of the periodical within the volume
3
Country of publishing house
US - UNITED STATES
Number of pages
17
Pages from-to
762-778
UT code for WoS article
001489511400001
EID of the result in the Scopus database
2-s2.0-105015176616