Fine-grained and Dense Annotation of Czech Propaganda Using Large Language Models
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00142943" target="_blank" >RIV/00216224:14330/25:00142943 - isvavai.cz</a>
Výsledek na webu
<a href="https://nlp.fi.muni.cz/raslan/raslan25.pdf" target="_blank" >https://nlp.fi.muni.cz/raslan/raslan25.pdf</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Fine-grained and Dense Annotation of Czech Propaganda Using Large Language Models
Popis výsledku v původním jazyce
Within the previous project of Czech Propaganda Detection that aimed to recognize manipulative techniques in Czech news articles, the annotation was mostly limited to indicating the presence of a technique in a document. In about 35% of the documents, span-level evidence was also annotated, but only as a support for the document-level labels, resulting in a sparse coverage of the techniques. Thus, the resulting dataset has limitations for training and evaluating more fine-grained propaganda detection models. In this study, we examine the potential of large language models (LLMs) to generate dense span-level annotations of manipulative techniques in Czech news articles. We designed generation prompts tailored to each technique and experimented with several LLMs to produce annotations for a subset of the Czech Propaganda dataset. We present the details of the generation process, including the design of the prompts and the selection of models. We also evaluate the generated annotations both quantitatively and qualitatively, including a manual validation and comparison with human annotations.
Název v anglickém jazyce
Fine-grained and Dense Annotation of Czech Propaganda Using Large Language Models
Popis výsledku anglicky
Within the previous project of Czech Propaganda Detection that aimed to recognize manipulative techniques in Czech news articles, the annotation was mostly limited to indicating the presence of a technique in a document. In about 35% of the documents, span-level evidence was also annotated, but only as a support for the document-level labels, resulting in a sparse coverage of the techniques. Thus, the resulting dataset has limitations for training and evaluating more fine-grained propaganda detection models. In this study, we examine the potential of large language models (LLMs) to generate dense span-level annotations of manipulative techniques in Czech news articles. We designed generation prompts tailored to each technique and experimented with several LLMs to produce annotations for a subset of the Czech Propaganda dataset. We present the details of the generation process, including the design of the prompts and the selection of models. We also evaluate the generated annotations both quantitatively and qualitatively, including a manual validation and comparison with human annotations.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10200 - Computer and information sciences
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_025%2F0008710" target="_blank" >EH23_025/0008710: Na všechno sami: příležitosti a rizika individualizace společnosti</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Recent Advances in Slavonic Natural Language Processing, RASLAN 2025
ISBN
9788026318583
ISSN
2336-4289
e-ISSN
—
Počet stran výsledku
14
Strana od-do
85-98
Název nakladatele
Tribun EU
Místo vydání
Brno, Czech Republic
Místo konání akce
Kouty nad Desnou, Česká Republika
Datum konání akce
1. 1. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—