Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0197661" target="_blank" >RIV/00216305:26230/26:0197661 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/chapter/10.1007/978-3-032-09368-4_4" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-09368-4_4</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-032-09368-4_4" target="_blank" >10.1007/978-3-032-09368-4_4</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Popis výsledku v původním jazyce
Self-supervised learning has emerged as a powerful approach for leveraging large-scale unlabeled data to improve model performance in various domains. In this paper, we explore masked self-supervised pre-training for text recognition transformers. Specifically, we propose two modifications to the pre-training phase: progressively increasing the masking probability, and modifying the loss function to incorporate both masked and non-masked patches. We conduct extensive experiments using a dataset of 50M unlabeled text lines for pre-training and four differently sized annotated datasets for fine-tuning. Furthermore, we compare our pre-trained models against those trained with transfer learning, demonstrating the effectiveness of the self-supervised pre-training. In particular, pre-training consistently improves the character error rate of models, in some cases up to 30 % relatively. It is also on par with transfer learning but without relying on extra annotated text lines.
Název v anglickém jazyce
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Popis výsledku anglicky
Self-supervised learning has emerged as a powerful approach for leveraging large-scale unlabeled data to improve model performance in various domains. In this paper, we explore masked self-supervised pre-training for text recognition transformers. Specifically, we propose two modifications to the pre-training phase: progressively increasing the masking probability, and modifying the loss function to incorporate both masked and non-masked patches. We conduct extensive experiments using a dataset of 50M unlabeled text lines for pre-training and four differently sized annotated datasets for fine-tuning. Furthermore, we compare our pre-trained models against those trained with transfer learning, demonstrating the effectiveness of the self-supervised pre-training. In particular, pre-training consistently improves the character error rate of models, in some cases up to 30 % relatively. It is also on par with transfer learning but without relying on extra annotated text lines.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/DH23P03OVV060" target="_blank" >DH23P03OVV060: semANT – Sémantický průzkumník textového kulturního dědictví</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Document Analysis and Recognition – ICDAR 2025 Workshops
ISBN
978-3-032-09367-7
ISSN
—
e-ISSN
—
Počet stran výsledku
18
Strana od-do
53-70
Název nakladatele
Springer Nature Switzerland
Místo vydání
Cham
Místo konání akce
Wuhan, Čína
Datum konání akce
16. 9. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—