Structured Tender Entities Extraction from Complex Tables with Few-short Learning
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AKY6Z7Q25" target="_blank" >RIV/00216208:11320/26:KY6Z7Q25 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/anthology-files/anthology-files/pdf/regnlp/2025.regnlp-1.pdf#page=71" target="_blank" >https://aclanthology.org/anthology-files/anthology-files/pdf/regnlp/2025.regnlp-1.pdf#page=71</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Structured Tender Entities Extraction from Complex Tables with Few-short Learning
Popis výsledku v původním jazyce
Extracting structured text from complex tables in PDF tender documents remains a challenging task due to the loss of structural and positional information during the extraction process. AI-based models often require extensive training data, making development from scratch both tedious and time-consuming. Our research focuses on identifying tender entities in complex table formats within PDF documents. To address this, we propose a novel approach utilizing few-shot learning with large language models (LLMs) to restore the structure of extracted text. Additionally, handcrafted rules and regular expressions are employed for precise entity classification. To evaluate the robustness of LLMs with few-shot learning, we employ data-shuffling techniques. Our experiments show that current text extraction tools fail to deliver satisfactory results for complex table structures. However, the few-shot learning approach significantly enhances the structural integrity of extracted data and improves the accuracy of tender entity identification.
Název v anglickém jazyce
Structured Tender Entities Extraction from Complex Tables with Few-short Learning
Popis výsledku anglicky
Extracting structured text from complex tables in PDF tender documents remains a challenging task due to the loss of structural and positional information during the extraction process. AI-based models often require extensive training data, making development from scratch both tedious and time-consuming. Our research focuses on identifying tender entities in complex table formats within PDF documents. To address this, we propose a novel approach utilizing few-shot learning with large language models (LLMs) to restore the structure of extracted text. Additionally, handcrafted rules and regular expressions are employed for precise entity classification. To evaluate the robustness of LLMs with few-shot learning, we employ data-shuffling techniques. Our experiments show that current text extraction tools fail to deliver satisfactory results for complex table structures. However, the few-shot learning approach significantly enhances the structural integrity of extracted data and improves the accuracy of tender entity identification.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the 1st Regulatory NLP Workshop
ISBN
979-8-89176-217-6
ISSN
—
e-ISSN
—
Počet stran výsledku
9
Strana od-do
59-67
Název nakladatele
—
Místo vydání
—
Místo konání akce
Abu Dhabi, UAE
Datum konání akce
1. 1. 2026
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—