All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

TextBite: A Historical Czech Document Dataset for Logical Page Segmentation

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0197678" target="_blank" >RIV/00216305:26230/26:0197678 - isvavai.cz</a>

  • Result on the web

    <a href="https://link.springer.com/chapter/10.1007/978-3-032-09368-4_8" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-09368-4_8</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-3-032-09368-4_8" target="_blank" >10.1007/978-3-032-09368-4_8</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    TextBite: A Historical Czech Document Dataset for Logical Page Segmentation

  • Original language description

    Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches define logical segmenta- tion either through text or geometric objects, relying on OCR or precise geometry. To avoid the need for OCR, we define the task purely as seg- mentation in the image domain. Furthermore, to ensure the evaluation remains unaffected by geometrical variations that do not impact text segmentation, we propose to use only foreground text pixels in the eval- uation metric and disregard all background pixels. To support research in logical document segmentation, we introduce TextBite, a dataset of historical Czech documents spanning the 18th to 20th centuries, fea- turing diverse layouts from newspapers, dictionaries, and handwritten records. The dataset comprises 8,449 page images with 78,863 annotated segments of logically and thematically coherent text. We propose a set of baseline methods combining text region detection and relation predic- tion. The dataset, baselines and evaluation framework can be accessed at https://github.com/DCGM/textbite-dataset.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

    <a href="/en/project/DH23P03OVV060" target="_blank" >DH23P03OVV060: semANT - Semantic Document Exploration</a><br>

  • Continuities

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    Document Analysis and Recognition – ICDAR 2025 Workshops

  • ISBN

    978-3-032-09367-7

  • ISSN

  • e-ISSN

  • Number of pages

    17

  • Pages from-to

    124-140

  • Publisher name

    Springer Nature Switzerland

  • Place of publication

    Cham

  • Event location

    Wuhan, Čína

  • Event date

    Sep 16, 2025

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article