All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Can LLMs Extract Human-like Fine-grained Evidence for Evidence-based Fact-checking?

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0201605" target="_blank" >RIV/00216305:26230/26:0201605 - isvavai.cz</a>

  • Result on the web

    <a href="https://raslan2025.nlp-consulting.net/" target="_blank" >https://raslan2025.nlp-consulting.net/</a>

  • DOI - Digital Object Identifier

Alternative languages

  • Result language

    angličtina

  • Original language name

    Can LLMs Extract Human-like Fine-grained Evidence for Evidence-based Fact-checking?

  • Original language description

    Misinformation frequently spreads in user comments under online news articles, highlighting the need for effective methods to detect factually incorrect information. To strongly support or refute claims extracted from such comments, it is necessary to identify relevant documents and pinpoint the exact text spans that justify or contradict each claim. This paper focuses on the latter task --- fine-grained evidence extraction for Czech and Slovak claims. We create new dataset, containing two-way annotated fine-grained evidence created by paid annotators. We evaluate large language models (LLMs) on this dataset to assess their alignment with human annotations. The results reveal that LLMs often fail to copy evidence verbatim from the source text, leading to invalid outputs. Error-rate analysis shows that the llama3.1:8b model achieves a high proportion of correct outputs despite its relatively small size, while the gpt-oss-120b model underperforms despite having many more parameters. Furthermore, the models qwen3:14b, deepseek-r1:32b, and gpt-oss:20b demonstrate an effective balance between model size and alignment with human annotations.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

    <a href="/en/project/TQ16000028" target="_blank" >TQ16000028: FactDeMice – Evidence-based Fact-Checking with Fact-consistent translation, Fake Review Detection, and Automatic Misinformative Claim Extraction</a><br>

  • Continuities

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    Proceedings of the Nineteenth Workshop on Recent Advances in Slavonic Natural Languages Processing, RASLAN 2025

  • ISBN

    978-80-263-1858-3

  • ISSN

    2336-4289

  • e-ISSN

  • Number of pages

    11

  • Pages from-to

    25-36

  • Publisher name

  • Place of publication

  • Event location

    Borovets

  • Event date

    Feb 16, 2005

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article