All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

ParlaCAP: Dataset for tracking political agenda-setting across European parliaments

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10513467" target="_blank" >RIV/00216208:11320/25:10513467 - isvavai.cz</a>

  • Result on the web

    <a href="https://doi.org/10.23669/1ZTELP" target="_blank" >https://doi.org/10.23669/1ZTELP</a>

  • DOI - Digital Object Identifier

Alternative languages

  • Result language

    angličtina

  • Original language name

    ParlaCAP: Dataset for tracking political agenda-setting across European parliaments

  • Original language description

    The ParlaCAP dataset consists of 8 million speeches from 28 European national and regional parliaments, with each speech coded with the sentiment expressed (ParlaSent coding from negative, over neutral, to positive) and the topic discussed (Comparative Agendas Project coding with 22 topics), and rich metadata on the speakers, parties and democracies. The dataset is an extension of the ParlaMint 5.0 dataset, which was primarily focused on the transcripts of parliamentary speeches and their metadata. The ParlaCAP dataset extends the ParlaMint dataset via the &quot;text as data&quot; paradigm by automatically coding topics and sentiment for each speech, simplifying the data to a tabular form, and thereby empowering social science research on agenda setting and negativity in political discourse across a broad set of parliaments. For automatic coding, multilingual transformer models were used, with the ParlaCAP model for topic, and the ParlaSent model for sentiment.

  • Czech name

  • Czech description

Classification

  • Type

    X - Unclassified

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

    <a href="/en/project/LM2023062" target="_blank" >LM2023062: Digital Research Infrastructure for Language Technologies, Arts and Humanities</a><br>

  • Continuities

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů