Automatic Webpage Content Extraction Based on Structural and Semantic Features
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ABN888AY7" target="_blank" >RIV/00216208:11320/26:BN888AY7 - isvavai.cz</a>
Result on the web
<a href="https://ieeexplore.ieee.org/abstract/document/11006249" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11006249</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICWR65219.2025.11006249" target="_blank" >10.1109/ICWR65219.2025.11006249</a>
Alternative languages
Result language
angličtina
Original language name
Automatic Webpage Content Extraction Based on Structural and Semantic Features
Original language description
The internet is a rich source of textual information, and by purposefully extracting data from its pages, we can obtain vast and suitable datasets for generating language models. However, the current structure of web pages is very diverse and complex. Additionally, web pages often contain irrelevant and sometimes useless information, leading to noise in the data. Developing a tool to extract useful content or eliminate useless content could be an appropriate solution to this problem. Existing methods either extract the main content based on rules, which suffer from decreased performance with changes in page design technologies and require constant updating, or are based on machine learning models. The latter methods also do not perform well across a wide range of pages due to the limited scope of their training datasets or the complexity of the designed models. In this research, we aim to design an intelligent model for extracting main content using the structural, semantic and content features of various page elements. The proposed method divides the page into blocks and by extracting various features, predicts the final label for each block using a machine learning model. To train the proposed model, a dataset of web pages was collected, and the main content of these pages was manually labeled by several volunteers. The advantage of this model is the improved blocking algorithm, which prevents the merging of useful and non-useful texts and avoids generating multiple blocks for a single web page. Another contribution in this research is the design of seven new structural and semantic features According to the results, the proposed method performs 13 % better than other methods.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
2025 11th International Conference on Web Research (ICWR)
ISBN
—
ISSN
2837-8296
e-ISSN
—
Number of pages
5
Pages from-to
276-280
Publisher name
—
Place of publication
—
Event location
Tehran, Iran
Event date
Jan 1, 2026
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—