How to age BERT Well: Continuous Training for Historical Language Adaptation
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AEEMBP847" target="_blank" >RIV/00216208:11320/26:EEMBP847 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.scopus.com/pages/publications/105000195929?origin=resultslist" target="_blank" >https://www.scopus.com/pages/publications/105000195929?origin=resultslist</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
How to age BERT Well: Continuous Training for Historical Language Adaptation
Popis výsledku v původním jazyce
As the application of computational tools increases to digitalize historical archives, automatic annotation challenges persist due to distinct linguistic and morphological features of historical languages like Old English (OE). Existing tools struggle with the historical language varieties due to insufficient training. Previous research has focused on adapting pre-trained language models to new languages or domains but has rarely explored the modeling of language variety across time. Hence, we investigate the effectiveness of continuous language model training for adapting language models to OE on domain-specific data. We compare the continuous training of an English model (EN) and a multilingual model, and use POS tagging for downstream evaluation. Results show that continuous pre-training substantially improves performance. More concretely, EN BERT initially outperformed mBERT with an accuracy of 83% during the language modeling phase. However, on the POS tagging task, mBERT surpassed EN BERT, achieving an accuracy of 94%, which suggests effective performance to the historical language varieties. © 2025 Association for Computational Linguistics.
Název v anglickém jazyce
How to age BERT Well: Continuous Training for Historical Language Adaptation
Popis výsledku anglicky
As the application of computational tools increases to digitalize historical archives, automatic annotation challenges persist due to distinct linguistic and morphological features of historical languages like Old English (OE). Existing tools struggle with the historical language varieties due to insufficient training. Previous research has focused on adapting pre-trained language models to new languages or domains but has rarely explored the modeling of language variety across time. Hence, we investigate the effectiveness of continuous language model training for adapting language models to OE on domain-specific data. We compare the continuous training of an English model (EN) and a multilingual model, and use POS tagging for downstream evaluation. Results show that continuous pre-training substantially improves performance. More concretely, EN BERT initially outperformed mBERT with an accuracy of 83% during the language modeling phase. However, on the POS tagging task, mBERT surpassed EN BERT, achieving an accuracy of 94%, which suggests effective performance to the historical language varieties. © 2025 Association for Computational Linguistics.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proc. Main Conf. Int. Conf. Comput. Linguist., COLING
ISBN
979-8-89176-215-2
ISSN
29512093
e-ISSN
—
Počet stran výsledku
10
Strana od-do
258-267
Název nakladatele
Association for Computational Linguistics (ACL)
Místo vydání
—
Místo konání akce
Abu Dhabi
Datum konání akce
1. 1. 2026
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—