Automatic Fact-checking in English and Telugu
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0198540" target="_blank" >RIV/00216305:26230/26:0198540 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2025.lowresnlp-1.15/" target="_blank" >https://aclanthology.org/2025.lowresnlp-1.15/</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Automatic Fact-checking in English and Telugu
Popis výsledku v původním jazyce
Misinformation is a significant problem nowadays, especially in multilingual countries like India, where false claims can be easily spread in multiple languages. Checking claims manually takes a lot of time and resources. To solve this, we use existing large language models (LLMs) that are trained on vast amounts of public data, which can be used to automate the claim verification process. In this project, our objective is to investigate the effectiveness of LLMs in classifying claims and providing justifications in English and Telugu, two widely spoken languages in the southern Indian states of Andhra Pradesh and Telangana. Our experiments demonstrate that LLMs perform better in high-resource languages such as English using baseline approaches, and they achieve improved performance in low-resource languages such as Telugu when provided with supporting documents. A major contribution of this project is the creation of an English and Telugu dataset.
Název v anglickém jazyce
Automatic Fact-checking in English and Telugu
Popis výsledku anglicky
Misinformation is a significant problem nowadays, especially in multilingual countries like India, where false claims can be easily spread in multiple languages. Checking claims manually takes a lot of time and resources. To solve this, we use existing large language models (LLMs) that are trained on vast amounts of public data, which can be used to automate the claim verification process. In this project, our objective is to investigate the effectiveness of LLMs in classifying claims and providing justifications in English and Telugu, two widely spoken languages in the southern Indian states of Andhra Pradesh and Telangana. Our experiments demonstrate that LLMs perform better in high-resource languages such as English using baseline approaches, and they achieve improved performance in low-resource languages such as Telugu when provided with supporting documents. A major contribution of this project is the creation of an English and Telugu dataset.
Klasifikace
Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů