How (un)faithful are explainable LLM-based NLG metrics?
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511664" target="_blank" >RIV/00216208:11320/25:10511664 - isvavai.cz</a>
Result on the web
<a href="https://aclanthology.org/2025.inlg-main.37" target="_blank" >https://aclanthology.org/2025.inlg-main.37</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
How (un)faithful are explainable LLM-based NLG metrics?
Original language description
Explainable NLG metrics are becoming a popular research topic; however, the faithfulness of the explanations they provide is typically not evaluated. In this work, we propose a testbed for assessing the faithfulness of span-based metrics by performing controlled perturbations of their explanations and observing changes in the final score. We show that several popular LLM evaluators do not consistently produce faithful explanations.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
R - Projekt Ramcoveho programu EK
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the 18th International Natural Language Generation Conference
ISBN
979-8-89176-321-0
ISSN
—
e-ISSN
—
Number of pages
42
Pages from-to
617-658
Publisher name
Association for Computational Linguistics
Place of publication
Kerrville, TX, USA
Event location
Hanoi, Vietnam
Event date
Oct 29, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—