Sibling-texts keyword analysis: exploring topic and register keywords
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10508091" target="_blank" >RIV/00216208:11210/25:10508091 - isvavai.cz</a>
Výsledek na webu
<a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=U5ysKb.edD" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=U5ysKb.edD</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/llc/fqaf037" target="_blank" >10.1093/llc/fqaf037</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Sibling-texts keyword analysis: exploring topic and register keywords
Popis výsledku v původním jazyce
This article introduces a novel method for refining approaches to distant reading by proposing a procedure that categorizes prominent units (keywords) into two types: those pertinent to the topic and those associated with the genre/register of a text. This differentiation holds significant potential for more accurate modeling of topics and further applications in various domains of digital humanities. For example, register-related keywords may assist in various stylometric tasks, such as functional understanding of text groups (formed through clustering) that exhibit common stylistic traits, thus differentiating themselves from other groups or clusters within the discourse under scrutiny. In the initial step of the procedure, a set of texts with similar genre and register characteristics is identified; we term them 'sibling texts', and their similarity is determined using a model derived from multidimensional analysis. The next step involves conducting multiple parallel keyword analyses on these sibling texts and comparing them for keyword overlaps which mark the register relatedness. The test study on a corpus of Czech parliamentary speeches (Parlcorp) underscores the method's ability to distinguish register-related and topic-related keywords, even in highly homogeneous corpora. The register-related keywords proved to be analytically very useful in substantiating the interpretation of parliamentary subregisters and advancing the understanding of the types of activities (administrative and procedural, deliberating and debating, and government policies) involved in parliamentary discourse.
Název v anglickém jazyce
Sibling-texts keyword analysis: exploring topic and register keywords
Popis výsledku anglicky
This article introduces a novel method for refining approaches to distant reading by proposing a procedure that categorizes prominent units (keywords) into two types: those pertinent to the topic and those associated with the genre/register of a text. This differentiation holds significant potential for more accurate modeling of topics and further applications in various domains of digital humanities. For example, register-related keywords may assist in various stylometric tasks, such as functional understanding of text groups (formed through clustering) that exhibit common stylistic traits, thus differentiating themselves from other groups or clusters within the discourse under scrutiny. In the initial step of the procedure, a set of texts with similar genre and register characteristics is identified; we term them 'sibling texts', and their similarity is determined using a model derived from multidimensional analysis. The next step involves conducting multiple parallel keyword analyses on these sibling texts and comparing them for keyword overlaps which mark the register relatedness. The test study on a corpus of Czech parliamentary speeches (Parlcorp) underscores the method's ability to distinguish register-related and topic-related keywords, even in highly homogeneous corpora. The register-related keywords proved to be analytically very useful in substantiating the interpretation of parliamentary subregisters and advancing the understanding of the types of activities (administrative and procedural, deliberating and debating, and government policies) involved in parliamentary discourse.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
<a href="/cs/project/EH22_008%2F0004595" target="_blank" >EH22_008/0004595: Za hranice bezpečnosti: role konfliktu v posilování odolnosti</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Digital Scholarship in the Humanities
ISSN
2055-7671
e-ISSN
2055-768X
Svazek periodika
40
Číslo periodika v rámci svazku
3
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
17
Strana od-do
762-778
Kód UT WoS článku
001489511400001
EID výsledku v databázi Scopus
2-s2.0-105015176616