Prompting Large Language Models for Church Slavic Translation
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00143188" target="_blank" >RIV/00216224:14330/25:00143188 - isvavai.cz</a>
Výsledek na webu
<a href="https://nlp.fi.muni.cz/raslan/2025/paper3.pdf" target="_blank" >https://nlp.fi.muni.cz/raslan/2025/paper3.pdf</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Prompting Large Language Models for Church Slavic Translation
Popis výsledku v původním jazyce
Church Slavic is a low-resource historical language with limited resources and few experts. We explore the capabilities of off-the-shelf Large Language Models (LLMs) as Church Slavic translators by prompting multiple models in zero-shot and few-shot scenarios. We evaluate four LLMs of varying sizes on 262 sentence pairs translating Church Slavic into English and German, and conduct a second experiment examining the impact of model size using five Qwen2.5-Instruct variants. Our results show that on average EuroLLM-9B-Instruct achieves the best performance, outperforming much larger models. We find minimal benefit from few-shot prompting and performance gaps between English and German as target languages. The automated evaluation metrics suggest that LLMs can produce useful draft translations for Church Slavic, potentially assisting scholars in accessing historical texts.
Název v anglickém jazyce
Prompting Large Language Models for Church Slavic Translation
Popis výsledku anglicky
Church Slavic is a low-resource historical language with limited resources and few experts. We explore the capabilities of off-the-shelf Large Language Models (LLMs) as Church Slavic translators by prompting multiple models in zero-shot and few-shot scenarios. We evaluate four LLMs of varying sizes on 262 sentence pairs translating Church Slavic into English and German, and conduct a second experiment examining the impact of model size using five Qwen2.5-Instruct variants. Our results show that on average EuroLLM-9B-Instruct achieves the best performance, outperforming much larger models. We find minimal benefit from few-shot prompting and performance gaps between English and German as target languages. The automated evaluation metrics suggest that LLMs can produce useful draft translations for Church Slavic, potentially assisting scholars in accessing historical texts.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10200 - Computer and information sciences
Návaznosti výsledku
Projekt
<a href="/cs/project/LM2023062" target="_blank" >LM2023062: Digitální výzkumná infrastruktura pro jazykové technologie, umění a humanitní vědy</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the Nineteenth Workshop on Recent Advances in Slavonic Natural Languages Processing, RASLAN 2025
ISBN
9788026318583
ISSN
2336-4289
e-ISSN
—
Počet stran výsledku
12
Strana od-do
57-68
Název nakladatele
Tribun EU
Místo vydání
Brno
Místo konání akce
Kouty nad Desnou, Czech Republic
Datum konání akce
5. 12. 2025
Typ akce podle státní příslušnosti
CST - Celostátní akce
Kód UT WoS článku
—