Mind the Gap: Diverse NMT Models for Resource-Constrained Environments
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511578" target="_blank" >RIV/00216208:11320/25:10511578 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2025.nodalida-1.21/" target="_blank" >https://aclanthology.org/2025.nodalida-1.21/</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Mind the Gap: Diverse NMT Models for Resource-Constrained Environments
Popis výsledku v původním jazyce
We present fast Neural Machine Translation models for 17 diverse languages, developed using Sequence-level Knowledge Distillation. Our selected languages span multiple language families and scripts, including low-resource languages. The distilled models achieve comparable performance while being 10x times faster than transformer-base and 35x times faster than transformer-big architectures. Our experiments reveal that teacher model quality and capacity strongly influence the distillation success, as well as the language script. We also explore the effectiveness of multilingual students. We release publicly our code and models in our Github repository.
Název v anglickém jazyce
Mind the Gap: Diverse NMT Models for Resource-Constrained Environments
Popis výsledku anglicky
We present fast Neural Machine Translation models for 17 diverse languages, developed using Sequence-level Knowledge Distillation. Our selected languages span multiple language families and scripts, including low-resource languages. The distilled models achieve comparable performance while being 10x times faster than transformer-base and 35x times faster than transformer-big architectures. Our experiments reveal that teacher model quality and capacity strongly influence the distillation success, as well as the language script. We also explore the effectiveness of multilingual students. We release publicly our code and models in our Github repository.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
R - Projekt Ramcoveho programu EK
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies
ISBN
978-9908-53-109-0
ISSN
—
e-ISSN
—
Počet stran výsledku
8
Strana od-do
209-216
Název nakladatele
University of Tartu Library
Místo vydání
Tallinn, Estonia
Místo konání akce
Tallinn, Estonia
Datum konání akce
2. 3. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—