Beyond Literal Token Overlap: Token Alignability for Multilinguality
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511589" target="_blank" >RIV/00216208:11320/25:10511589 - isvavai.cz</a>
Result on the web
<a href="https://aclanthology.org/2025.naacl-short.63" target="_blank" >https://aclanthology.org/2025.naacl-short.63</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
Beyond Literal Token Overlap: Token Alignability for Multilinguality
Original language description
Previous work has considered token overlap, or even similarity of token distributions, as predictors for multilinguality and cross-lingual knowledge transfer in language models. However, these very literal metrics assign large distances to language pairs with different scripts, which can nevertheless show good cross-linguality. This limits the explanatory strength of token overlap for knowledge transfer between language pairs that use distinct scripts or follow different orthographic conventions. In this paper, we propose subword token alignability as a new way to understand the impact and quality of multilingual tokenisation. In particular, this metric predicts multilinguality much better when scripts are disparate and the overlap of literal tokens is low. We analyse this metric in the context of both encoder and decoder models, look at data size as a potential distractor, and discuss how this insight may be applied to multilingual tokenisation in future work. We recommend our subword token alignabil
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers)
ISBN
979-8-89176-190-2
ISSN
—
e-ISSN
—
Number of pages
12
Pages from-to
756-767
Publisher name
Association for Computational Linguistics
Place of publication
Kerrville, TX, USA
Event location
Albuquerque, NM, USA
Event date
Apr 29, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—