Team LLMinds Submission for ELOQUENT Sensemaking Task
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511648" target="_blank" >RIV/00216208:11320/25:10511648 - isvavai.cz</a>
Výsledek na webu
<a href="https://ceur-ws.org/Vol-4038/paper_114.pdf" target="_blank" >https://ceur-ws.org/Vol-4038/paper_114.pdf</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Team LLMinds Submission for ELOQUENT Sensemaking Task
Popis výsledku v původním jazyce
This paper presents our system submission for the CLEF 2025 ELOQUENT Sensemaking Task, which focuses on generating and answering questions based solely on provided textual materials, such as lecture transcripts and textbooks. Our approach combines open-source large language models (LLMs) in a two-stage pipeline: a Teacher component for question generation, a Student for grounded answering using Retrieval-Augmented Generation (RAG). Ideally, the third stage would be an Evaluator for an assessment of response quality but we leave this part for the future. To ensure transparency, privacy, and accessibility, we prioritized using local models such as LLaMA 7B and DistilQwen 1.5B, avoiding reliance on proprietary APIs. For question generation, we implemented both a simple single-question prompting method and an enhanced generator-discriminator pipeline that scores candidates on answerability, context relevance, and diversity. The Student model leverages a RAG system with FAISS to extract relevant context ch
Název v anglickém jazyce
Team LLMinds Submission for ELOQUENT Sensemaking Task
Popis výsledku anglicky
This paper presents our system submission for the CLEF 2025 ELOQUENT Sensemaking Task, which focuses on generating and answering questions based solely on provided textual materials, such as lecture transcripts and textbooks. Our approach combines open-source large language models (LLMs) in a two-stage pipeline: a Teacher component for question generation, a Student for grounded answering using Retrieval-Augmented Generation (RAG). Ideally, the third stage would be an Evaluator for an assessment of response quality but we leave this part for the future. To ensure transparency, privacy, and accessibility, we prioritized using local models such as LLaMA 7B and DistilQwen 1.5B, avoiding reliance on proprietary APIs. For question generation, we implemented both a simple single-question prompting method and an enhanced generator-discriminator pipeline that scores candidates on answerability, context relevance, and diversity. The Student model leverages a RAG system with FAISS to extract relevant context ch
Klasifikace
Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Jazykověda, umělá inteligence a jazykové a řečové technologie: od výzkumu k aplikacím</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů