Automatic Webpage Content Extraction Based on Structural and Semantic Features
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ABN888AY7" target="_blank" >RIV/00216208:11320/26:BN888AY7 - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/abstract/document/11006249" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11006249</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICWR65219.2025.11006249" target="_blank" >10.1109/ICWR65219.2025.11006249</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Automatic Webpage Content Extraction Based on Structural and Semantic Features
Popis výsledku v původním jazyce
The internet is a rich source of textual information, and by purposefully extracting data from its pages, we can obtain vast and suitable datasets for generating language models. However, the current structure of web pages is very diverse and complex. Additionally, web pages often contain irrelevant and sometimes useless information, leading to noise in the data. Developing a tool to extract useful content or eliminate useless content could be an appropriate solution to this problem. Existing methods either extract the main content based on rules, which suffer from decreased performance with changes in page design technologies and require constant updating, or are based on machine learning models. The latter methods also do not perform well across a wide range of pages due to the limited scope of their training datasets or the complexity of the designed models. In this research, we aim to design an intelligent model for extracting main content using the structural, semantic and content features of various page elements. The proposed method divides the page into blocks and by extracting various features, predicts the final label for each block using a machine learning model. To train the proposed model, a dataset of web pages was collected, and the main content of these pages was manually labeled by several volunteers. The advantage of this model is the improved blocking algorithm, which prevents the merging of useful and non-useful texts and avoids generating multiple blocks for a single web page. Another contribution in this research is the design of seven new structural and semantic features According to the results, the proposed method performs 13 % better than other methods.
Název v anglickém jazyce
Automatic Webpage Content Extraction Based on Structural and Semantic Features
Popis výsledku anglicky
The internet is a rich source of textual information, and by purposefully extracting data from its pages, we can obtain vast and suitable datasets for generating language models. However, the current structure of web pages is very diverse and complex. Additionally, web pages often contain irrelevant and sometimes useless information, leading to noise in the data. Developing a tool to extract useful content or eliminate useless content could be an appropriate solution to this problem. Existing methods either extract the main content based on rules, which suffer from decreased performance with changes in page design technologies and require constant updating, or are based on machine learning models. The latter methods also do not perform well across a wide range of pages due to the limited scope of their training datasets or the complexity of the designed models. In this research, we aim to design an intelligent model for extracting main content using the structural, semantic and content features of various page elements. The proposed method divides the page into blocks and by extracting various features, predicts the final label for each block using a machine learning model. To train the proposed model, a dataset of web pages was collected, and the main content of these pages was manually labeled by several volunteers. The advantage of this model is the improved blocking algorithm, which prevents the merging of useful and non-useful texts and avoids generating multiple blocks for a single web page. Another contribution in this research is the design of seven new structural and semantic features According to the results, the proposed method performs 13 % better than other methods.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
2025 11th International Conference on Web Research (ICWR)
ISBN
—
ISSN
2837-8296
e-ISSN
—
Počet stran výsledku
5
Strana od-do
276-280
Název nakladatele
—
Místo vydání
—
Místo konání akce
Tehran, Iran
Datum konání akce
1. 1. 2026
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—