The Development of a Comprehensive Data Set for Systematic Studies of Machine Translation
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A90101%2F21%3A10442378" target="_blank" >RIV/00216208:90101/21:10442378 - isvavai.cz</a>
Result on the web
—
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
The Development of a Comprehensive Data Set for Systematic Studies of Machine Translation
Original language description
This paper presents our on-going efforts to develop a comprehensive data set and benchmark for machine translation beyond highresource languages. The current release includes 500GB of compressed parallel data for almost 3,000 language pairs covering over 500 languages and language variants. We present the structure of the data set and demonstrate its use for systematic studies based on baseline experiments with multilingual neural machine translation between Uralic languages and other language groups. Our initial results show the capabilities of training effective multilingual translation models with skewed training data but also stress the shortcomings with low-resource settings and the difficulties to obtain sufficient information through straightforward transfer from related languages.
Czech name
—
Czech description
—
Classification
Type
C - Chapter in a specialist book
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2021
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Book/collection name
Multilingual Facilitation
ISBN
979-8-7133-6227-0
Number of pages of the result
15
Pages from-to
248-262
Number of pages of the book
298
Publisher name
University of Helsinki
Place of publication
Helsinki
UT code for WoS chapter
—