Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511623" target="_blank" >RIV/00216208:11320/25:10511623 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2025.iwslt-1.41.pdf" target="_blank" >https://aclanthology.org/2025.iwslt-1.41.pdf</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
Popis výsledku v původním jazyce
This paper describes Charles University sub- mission to the Simultaneous Speech Transla- tion Task of the IWSLT 2025. We cover all four language pairs with a direct or cascade ap- proach. The backbone of our systems is the of- fline Whisper speech model, which we use for both translation and transcription in simultane- ous mode with the state-of-the-art simultaneous policy AlignAtt. We further improve the per- formance by prompting to inject in-domain ter- minology, and we accommodate context. Our cascaded systems further use EuroLLM for un- bounded simultaneous translation. Compared to the Organizers’ baseline, our systems im- prove by 2 BLEU points on Czech to English and 13-22 BLEU points on English to German, Chinese and Japanese on the development sets. Additionally, we also propose a new enhanced measure of speech recognition latency.
Název v anglickém jazyce
Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
Popis výsledku anglicky
This paper describes Charles University sub- mission to the Simultaneous Speech Transla- tion Task of the IWSLT 2025. We cover all four language pairs with a direct or cascade ap- proach. The backbone of our systems is the of- fline Whisper speech model, which we use for both translation and transcription in simultane- ous mode with the state-of-the-art simultaneous policy AlignAtt. We further improve the per- formance by prompting to inject in-domain ter- minology, and we accommodate context. Our cascaded systems further use EuroLLM for un- bounded simultaneous translation. Compared to the Organizers’ baseline, our systems im- prove by 2 BLEU points on Czech to English and 13-22 BLEU points on English to German, Chinese and Japanese on the development sets. Additionally, we also propose a new enhanced measure of speech recognition latency.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Jazykověda, umělá inteligence a jazykové a řečové technologie: od výzkumu k aplikacím</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the 22nd International Conference on Spoken Language Translation (IWSLT 2025)
ISBN
979-8-89176-272-5
ISSN
—
e-ISSN
—
Počet stran výsledku
10
Strana od-do
389-398
Název nakladatele
Association for Computational Linguistics
Místo vydání
Kerrville, TX, USA
Místo konání akce
Wien, Austria
Datum konání akce
31. 7. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—