Corpus of Cross-Lingual Dialogues with Minutes and Detection of Misunderstandings
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511673" target="_blank" >RIV/00216208:11320/25:10511673 - isvavai.cz</a>
Result on the web
<a href="https://link.springer.com/chapter/10.1007/978-3-032-02551-7_26" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-02551-7_26</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
Corpus of Cross-Lingual Dialogues with Minutes and Detection of Misunderstandings
Original language description
Speech processing and translation technology have the po- tential to facilitate meetings of individuals who do not share any common language. To evaluate automatic systems for such a task, a versatile and realistic evaluation corpus is needed. Therefore, we create and present a corpus of cross-lingual dialogues between individuals without a common language who were facilitated by automatic simultaneous speech transla- tion. The corpus consists of 5 hours of speech recordings with ASR and gold transcripts in 12 original languages and automatic and corrected translations into English. For the purposes of research into cross-lingual summarization, our corpus also includes written summaries (minutes) of the meetings. Moreover, we propose automatic detection of misunderstandings. For an overview of this task and its complexity, we attempt to quantify misun- derstandings in cross-lingual meetings. We annotate misunderstandings manually and also test the ability of current large language models to detect the
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
28th International Conference on Text, Speech and Dialogue (Part II)
ISBN
978-3-032-02551-7
ISSN
—
e-ISSN
—
Number of pages
12
Pages from-to
301-312
Publisher name
Springer
Place of publication
Cham, Switzerland
Event location
Erlangen, Germany
Event date
Aug 25, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—