All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0197672" target="_blank" >RIV/00216305:26230/26:0197672 - isvavai.cz</a>

  • Alternative codes found

    RIV/67985971:_____/26:00645444 RIV/00094943:_____/26:N0000001 RIV/00023221:_____/25:N0000030

  • Result on the web

    <a href="https://link.springer.com/chapter/10.1007/978-3-032-09371-4_4" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-09371-4_4</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-3-032-09371-4_4" target="_blank" >10.1007/978-3-032-09371-4_4</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization

  • Original language description

    We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the late 19th and early 20th centuries. The dataset is designed to support research in document layout analysis and object detection. Each page is annotated with axis-aligned bounding boxes (AABB) representing elements of 25 categories of non-textual elements, such as images, maps, decorative elements, or charts, following the Czech Methodology of image document processing. The annotations were created by expert librarians to ensure accuracy and consistency. The dataset also incorporates pages from multiple, mainly historical, document datasets to enhance variability and maintain continuity. The dataset is divided into development and test subsets, with the test set carefully selected to maintain the category distribution. We provide baseline results using YOLO and DETR object detectors, offering a reference point for future research. The AnnoPage Dataset is publicly available on Zenodo (https://doi.org/10.5281/zenodo.12788419), along with ground-truth annotations in YOLO format.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

    <a href="/en/project/DH23P03OVV033" target="_blank" >DH23P03OVV033: Orbis Pictus – book revival for cultural and creative sectors</a><br>

  • Continuities

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

  • Publication year

    2026

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    Document Analysis and Recognition – ICDAR 2025 Workshops

  • ISBN

    978-3-032-09370-7

  • ISSN

  • e-ISSN

  • Number of pages

    17

  • Pages from-to

    50-66

  • Publisher name

    Springer Nature Switzerland

  • Place of publication

    Cham

  • Event location

    Wuhan, Čína

  • Event date

    Sep 16, 2025

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article