All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Structured Tender Entities Extraction from Complex Tables with Few-short Learning

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AKY6Z7Q25" target="_blank" >RIV/00216208:11320/26:KY6Z7Q25 - isvavai.cz</a>

  • Result on the web

    <a href="https://aclanthology.org/anthology-files/anthology-files/pdf/regnlp/2025.regnlp-1.pdf#page=71" target="_blank" >https://aclanthology.org/anthology-files/anthology-files/pdf/regnlp/2025.regnlp-1.pdf#page=71</a>

  • DOI - Digital Object Identifier

Alternative languages

  • Result language

    angličtina

  • Original language name

    Structured Tender Entities Extraction from Complex Tables with Few-short Learning

  • Original language description

    Extracting structured text from complex tables in PDF tender documents remains a challenging task due to the loss of structural and positional information during the extraction process. AI-based models often require extensive training data, making development from scratch both tedious and time-consuming. Our research focuses on identifying tender entities in complex table formats within PDF documents. To address this, we propose a novel approach utilizing few-shot learning with large language models (LLMs) to restore the structure of extracted text. Additionally, handcrafted rules and regular expressions are employed for precise entity classification. To evaluate the robustness of LLMs with few-shot learning, we employ data-shuffling techniques. Our experiments show that current text extraction tools fail to deliver satisfactory results for complex table structures. However, the few-shot learning approach significantly enhances the structural integrity of extracted data and improves the accuracy of tender entity identification.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    Proceedings of the 1st Regulatory NLP Workshop

  • ISBN

    979-8-89176-217-6

  • ISSN

  • e-ISSN

  • Number of pages

    9

  • Pages from-to

    59-67

  • Publisher name

  • Place of publication

  • Event location

    Abu Dhabi, UAE

  • Event date

    Jan 1, 2026

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article