Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Automatic Webpage Content Extraction Based on Structural and Semantic Features

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ABN888AY7" target="_blank" >RIV/00216208:11320/26:BN888AY7 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://ieeexplore.ieee.org/abstract/document/11006249" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11006249</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1109/ICWR65219.2025.11006249" target="_blank" >10.1109/ICWR65219.2025.11006249</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Automatic Webpage Content Extraction Based on Structural and Semantic Features

  • Popis výsledku v původním jazyce

    The internet is a rich source of textual information, and by purposefully extracting data from its pages, we can obtain vast and suitable datasets for generating language models. However, the current structure of web pages is very diverse and complex. Additionally, web pages often contain irrelevant and sometimes useless information, leading to noise in the data. Developing a tool to extract useful content or eliminate useless content could be an appropriate solution to this problem. Existing methods either extract the main content based on rules, which suffer from decreased performance with changes in page design technologies and require constant updating, or are based on machine learning models. The latter methods also do not perform well across a wide range of pages due to the limited scope of their training datasets or the complexity of the designed models. In this research, we aim to design an intelligent model for extracting main content using the structural, semantic and content features of various page elements. The proposed method divides the page into blocks and by extracting various features, predicts the final label for each block using a machine learning model. To train the proposed model, a dataset of web pages was collected, and the main content of these pages was manually labeled by several volunteers. The advantage of this model is the improved blocking algorithm, which prevents the merging of useful and non-useful texts and avoids generating multiple blocks for a single web page. Another contribution in this research is the design of seven new structural and semantic features According to the results, the proposed method performs 13 % better than other methods.

  • Název v anglickém jazyce

    Automatic Webpage Content Extraction Based on Structural and Semantic Features

  • Popis výsledku anglicky

    The internet is a rich source of textual information, and by purposefully extracting data from its pages, we can obtain vast and suitable datasets for generating language models. However, the current structure of web pages is very diverse and complex. Additionally, web pages often contain irrelevant and sometimes useless information, leading to noise in the data. Developing a tool to extract useful content or eliminate useless content could be an appropriate solution to this problem. Existing methods either extract the main content based on rules, which suffer from decreased performance with changes in page design technologies and require constant updating, or are based on machine learning models. The latter methods also do not perform well across a wide range of pages due to the limited scope of their training datasets or the complexity of the designed models. In this research, we aim to design an intelligent model for extracting main content using the structural, semantic and content features of various page elements. The proposed method divides the page into blocks and by extracting various features, predicts the final label for each block using a machine learning model. To train the proposed model, a dataset of web pages was collected, and the main content of these pages was manually labeled by several volunteers. The advantage of this model is the improved blocking algorithm, which prevents the merging of useful and non-useful texts and avoids generating multiple blocks for a single web page. Another contribution in this research is the design of seven new structural and semantic features According to the results, the proposed method performs 13 % better than other methods.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    2025 11th International Conference on Web Research (ICWR)

  • ISBN

  • ISSN

    2837-8296

  • e-ISSN

  • Počet stran výsledku

    5

  • Strana od-do

    276-280

  • Název nakladatele

  • Místo vydání

  • Místo konání akce

    Tehran, Iran

  • Datum konání akce

    1. 1. 2026

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku