All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Automatic Webpage Content Extraction Based on Structural and Semantic Features

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ABN888AY7" target="_blank" >RIV/00216208:11320/26:BN888AY7 - isvavai.cz</a>

  • Result on the web

    <a href="https://ieeexplore.ieee.org/abstract/document/11006249" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11006249</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1109/ICWR65219.2025.11006249" target="_blank" >10.1109/ICWR65219.2025.11006249</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Automatic Webpage Content Extraction Based on Structural and Semantic Features

  • Original language description

    The internet is a rich source of textual information, and by purposefully extracting data from its pages, we can obtain vast and suitable datasets for generating language models. However, the current structure of web pages is very diverse and complex. Additionally, web pages often contain irrelevant and sometimes useless information, leading to noise in the data. Developing a tool to extract useful content or eliminate useless content could be an appropriate solution to this problem. Existing methods either extract the main content based on rules, which suffer from decreased performance with changes in page design technologies and require constant updating, or are based on machine learning models. The latter methods also do not perform well across a wide range of pages due to the limited scope of their training datasets or the complexity of the designed models. In this research, we aim to design an intelligent model for extracting main content using the structural, semantic and content features of various page elements. The proposed method divides the page into blocks and by extracting various features, predicts the final label for each block using a machine learning model. To train the proposed model, a dataset of web pages was collected, and the main content of these pages was manually labeled by several volunteers. The advantage of this model is the improved blocking algorithm, which prevents the merging of useful and non-useful texts and avoids generating multiple blocks for a single web page. Another contribution in this research is the design of seven new structural and semantic features According to the results, the proposed method performs 13 % better than other methods.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    2025 11th International Conference on Web Research (ICWR)

  • ISBN

  • ISSN

    2837-8296

  • e-ISSN

  • Number of pages

    5

  • Pages from-to

    276-280

  • Publisher name

  • Place of publication

  • Event location

    Tehran, Iran

  • Event date

    Jan 1, 2026

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article