Corpus of Cross-Lingual Dialogues with Minutes and Detection of Misunderstandings
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511673" target="_blank" >RIV/00216208:11320/25:10511673 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/chapter/10.1007/978-3-032-02551-7_26" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-02551-7_26</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Corpus of Cross-Lingual Dialogues with Minutes and Detection of Misunderstandings
Popis výsledku v původním jazyce
Speech processing and translation technology have the po- tential to facilitate meetings of individuals who do not share any common language. To evaluate automatic systems for such a task, a versatile and realistic evaluation corpus is needed. Therefore, we create and present a corpus of cross-lingual dialogues between individuals without a common language who were facilitated by automatic simultaneous speech transla- tion. The corpus consists of 5 hours of speech recordings with ASR and gold transcripts in 12 original languages and automatic and corrected translations into English. For the purposes of research into cross-lingual summarization, our corpus also includes written summaries (minutes) of the meetings. Moreover, we propose automatic detection of misunderstandings. For an overview of this task and its complexity, we attempt to quantify misun- derstandings in cross-lingual meetings. We annotate misunderstandings manually and also test the ability of current large language models to detect the
Název v anglickém jazyce
Corpus of Cross-Lingual Dialogues with Minutes and Detection of Misunderstandings
Popis výsledku anglicky
Speech processing and translation technology have the po- tential to facilitate meetings of individuals who do not share any common language. To evaluate automatic systems for such a task, a versatile and realistic evaluation corpus is needed. Therefore, we create and present a corpus of cross-lingual dialogues between individuals without a common language who were facilitated by automatic simultaneous speech transla- tion. The corpus consists of 5 hours of speech recordings with ASR and gold transcripts in 12 original languages and automatic and corrected translations into English. For the purposes of research into cross-lingual summarization, our corpus also includes written summaries (minutes) of the meetings. Moreover, we propose automatic detection of misunderstandings. For an overview of this task and its complexity, we attempt to quantify misun- derstandings in cross-lingual meetings. We annotate misunderstandings manually and also test the ability of current large language models to detect the
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Jazykověda, umělá inteligence a jazykové a řečové technologie: od výzkumu k aplikacím</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
28th International Conference on Text, Speech and Dialogue (Part II)
ISBN
978-3-032-02551-7
ISSN
—
e-ISSN
—
Počet stran výsledku
12
Strana od-do
301-312
Název nakladatele
Springer
Místo vydání
Cham, Switzerland
Místo konání akce
Erlangen, Germany
Datum konání akce
25. 8. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—