TextBite: A Historical Czech Document Dataset for Logical Page Segmentation
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0197678" target="_blank" >RIV/00216305:26230/26:0197678 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/chapter/10.1007/978-3-032-09368-4_8" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-09368-4_8</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-032-09368-4_8" target="_blank" >10.1007/978-3-032-09368-4_8</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
TextBite: A Historical Czech Document Dataset for Logical Page Segmentation
Popis výsledku v původním jazyce
Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches define logical segmenta- tion either through text or geometric objects, relying on OCR or precise geometry. To avoid the need for OCR, we define the task purely as seg- mentation in the image domain. Furthermore, to ensure the evaluation remains unaffected by geometrical variations that do not impact text segmentation, we propose to use only foreground text pixels in the eval- uation metric and disregard all background pixels. To support research in logical document segmentation, we introduce TextBite, a dataset of historical Czech documents spanning the 18th to 20th centuries, fea- turing diverse layouts from newspapers, dictionaries, and handwritten records. The dataset comprises 8,449 page images with 78,863 annotated segments of logically and thematically coherent text. We propose a set of baseline methods combining text region detection and relation predic- tion. The dataset, baselines and evaluation framework can be accessed at https://github.com/DCGM/textbite-dataset.
Název v anglickém jazyce
TextBite: A Historical Czech Document Dataset for Logical Page Segmentation
Popis výsledku anglicky
Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches define logical segmenta- tion either through text or geometric objects, relying on OCR or precise geometry. To avoid the need for OCR, we define the task purely as seg- mentation in the image domain. Furthermore, to ensure the evaluation remains unaffected by geometrical variations that do not impact text segmentation, we propose to use only foreground text pixels in the eval- uation metric and disregard all background pixels. To support research in logical document segmentation, we introduce TextBite, a dataset of historical Czech documents spanning the 18th to 20th centuries, fea- turing diverse layouts from newspapers, dictionaries, and handwritten records. The dataset comprises 8,449 page images with 78,863 annotated segments of logically and thematically coherent text. We propose a set of baseline methods combining text region detection and relation predic- tion. The dataset, baselines and evaluation framework can be accessed at https://github.com/DCGM/textbite-dataset.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/DH23P03OVV060" target="_blank" >DH23P03OVV060: semANT – Sémantický průzkumník textového kulturního dědictví</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Document Analysis and Recognition – ICDAR 2025 Workshops
ISBN
978-3-032-09367-7
ISSN
—
e-ISSN
—
Počet stran výsledku
17
Strana od-do
124-140
Název nakladatele
Springer Nature Switzerland
Místo vydání
Cham
Místo konání akce
Wuhan, Čína
Datum konání akce
16. 9. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—