Information Extraction in Domain and Generic Documents: Findings from Heuristic-based and Data-driven Approaches

The result's identifiers

Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F23%3AUWYAERTJ" target="_blank" >RIV/00216208:11320/23:UWYAERTJ - isvavai.cz</a>
Result on the web
<a href="http://arxiv.org/abs/2307.00130" target="_blank" >http://arxiv.org/abs/2307.00130</a>
DOI - Digital Object Identifier
—

Alternative languages

Result language
angličtina
Original language name
Information Extraction in Domain and Generic Documents: Findings from Heuristic-based and Data-driven Approaches
Original language description
"Information extraction (IE) plays very important role in natural language processing (NLP) and is fundamental to many NLP applications that used to extract structured information from unstructured text data. Heuristic-based searching and data-driven learning are two main stream implementation approaches. However, no much attention has been paid to document genre and length influence on IE tasks. To fill the gap, in this study, we investigated the accuracy and generalization abilities of heuristic-based searching and data-driven to perform two IE tasks: named entity recognition (NER) and semantic role labeling (SRL) on domain-specific and generic documents with different length. We posited two hypotheses: first, short documents may yield better accuracy results compared to long documents; second, generic documents may exhibit superior extraction outcomes relative to domain-dependent documents due to training document genre limitations. Our findings reveals that no single method demonstrated overwhelming performance in both tasks. For named entity extraction, data-driven approaches outperformed symbolic methods in terms of accuracy, particularly in short texts. In the case of semantic roles extraction, we observed that heuristic-based searching method and data-driven based model with syntax representation surpassed the performance of pure data-driven approach which only consider semantic information. Additionally, we discovered that different semantic roles exhibited varying accuracy levels with the same method. This study offers valuable insights for downstream text mining tasks, such as NER and SRL, when addressing various document features and genres."
Czech name
—
Czech description
—

Classification

Type
O - Miscellaneous
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

Project
—
Continuities
—

Others

Publication year
2023
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Similar results(10)

A Systematic Review on Semantic Role Labeling for Information Extraction in Low-Resource Data Annotating the Tweebank Corpus on Named Entity Recognition and Building NLP Models for Social Media Analysis Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data

What are you looking for?

Quick search

Smart search

Information Extraction in Domain and Generic Documents: Findings from Heuristic-based and Data-driven Approaches

The result's identifiers

Alternative languages

Classification

Result continuities

Others

Similar results(10)

What are you looking for?

Quick search

Smart search

Result description

The result's identifiers

The result's identifiers

Alternative languages

Alternative languages

Classification

Classification

Result continuities

Result continuities

Others

Others

Similar results(10)