Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

TextBite: A Historical Czech Document Dataset for Logical Page Segmentation

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0197678" target="_blank" >RIV/00216305:26230/26:0197678 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://link.springer.com/chapter/10.1007/978-3-032-09368-4_8" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-09368-4_8</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-3-032-09368-4_8" target="_blank" >10.1007/978-3-032-09368-4_8</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    TextBite: A Historical Czech Document Dataset for Logical Page Segmentation

  • Popis výsledku v původním jazyce

    Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches define logical segmenta- tion either through text or geometric objects, relying on OCR or precise geometry. To avoid the need for OCR, we define the task purely as seg- mentation in the image domain. Furthermore, to ensure the evaluation remains unaffected by geometrical variations that do not impact text segmentation, we propose to use only foreground text pixels in the eval- uation metric and disregard all background pixels. To support research in logical document segmentation, we introduce TextBite, a dataset of historical Czech documents spanning the 18th to 20th centuries, fea- turing diverse layouts from newspapers, dictionaries, and handwritten records. The dataset comprises 8,449 page images with 78,863 annotated segments of logically and thematically coherent text. We propose a set of baseline methods combining text region detection and relation predic- tion. The dataset, baselines and evaluation framework can be accessed at https://github.com/DCGM/textbite-dataset.

  • Název v anglickém jazyce

    TextBite: A Historical Czech Document Dataset for Logical Page Segmentation

  • Popis výsledku anglicky

    Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches define logical segmenta- tion either through text or geometric objects, relying on OCR or precise geometry. To avoid the need for OCR, we define the task purely as seg- mentation in the image domain. Furthermore, to ensure the evaluation remains unaffected by geometrical variations that do not impact text segmentation, we propose to use only foreground text pixels in the eval- uation metric and disregard all background pixels. To support research in logical document segmentation, we introduce TextBite, a dataset of historical Czech documents spanning the 18th to 20th centuries, fea- turing diverse layouts from newspapers, dictionaries, and handwritten records. The dataset comprises 8,449 page images with 78,863 annotated segments of logically and thematically coherent text. We propose a set of baseline methods combining text region detection and relation predic- tion. The dataset, baselines and evaluation framework can be accessed at https://github.com/DCGM/textbite-dataset.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

    <a href="/cs/project/DH23P03OVV060" target="_blank" >DH23P03OVV060: semANT – Sémantický průzkumník textového kulturního dědictví</a><br>

  • Návaznosti

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    Document Analysis and Recognition – ICDAR 2025 Workshops

  • ISBN

    978-3-032-09367-7

  • ISSN

  • e-ISSN

  • Počet stran výsledku

    17

  • Strana od-do

    124-140

  • Název nakladatele

    Springer Nature Switzerland

  • Místo vydání

    Cham

  • Místo konání akce

    Wuhan, Čína

  • Datum konání akce

    16. 9. 2025

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku