Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Evaluating proximity metrics for gene expression data: A hybrid model integrating data mining and machine learning techniques for disease diagnosis systems

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F44555601%3A13440%2F25%3A43899179" target="_blank" >RIV/44555601:13440/25:43899179 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://www.sciencedirect.com/science/article/pii/S1746809425006263" target="_blank" >https://www.sciencedirect.com/science/article/pii/S1746809425006263</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1016/j.bspc.2025.108115" target="_blank" >10.1016/j.bspc.2025.108115</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Evaluating proximity metrics for gene expression data: A hybrid model integrating data mining and machine learning techniques for disease diagnosis systems

  • Popis výsledku v původním jazyce

    This study presents the development and application of a hybrid model for evaluating proximity metrics in high-dimensional gene expression data, integrating data mining and machine learning methods within a comprehensive framework. The research focuses on the comparative analysis of correlation distance, mutual information-based and Wasserstein metrics, assessing their effectiveness for clustering and classification tasks. In the initial modeling stage using gene expression data from over 6,000 patient samples covering 13 cancer types (TCGA dataset), the proposed model achieved classification accuracy exceeding 95.9% and a weighted F1-score above 95.8%. External validation using Alzheimer&apos;s (GSE174367) and Type 2 Diabetes (GSE81608) datasets confirmed the model&apos;s generalizability, with accuracy values reaching 96.28% and 97.43%, and weighted F1-scores of 96.26% and 97.41%, respectively. A stacking model was implemented to enhance classification robustness, compensating for potential clustering errors and delivering consistent performance across varying metrics and cluster structures. The proposed data processing pipeline ensures automated, standardized, and scalable analysis of large-scale gene expression datasets, aligning with the principles of personalized medicine.

  • Název v anglickém jazyce

    Evaluating proximity metrics for gene expression data: A hybrid model integrating data mining and machine learning techniques for disease diagnosis systems

  • Popis výsledku anglicky

    This study presents the development and application of a hybrid model for evaluating proximity metrics in high-dimensional gene expression data, integrating data mining and machine learning methods within a comprehensive framework. The research focuses on the comparative analysis of correlation distance, mutual information-based and Wasserstein metrics, assessing their effectiveness for clustering and classification tasks. In the initial modeling stage using gene expression data from over 6,000 patient samples covering 13 cancer types (TCGA dataset), the proposed model achieved classification accuracy exceeding 95.9% and a weighted F1-score above 95.8%. External validation using Alzheimer&apos;s (GSE174367) and Type 2 Diabetes (GSE81608) datasets confirmed the model&apos;s generalizability, with accuracy values reaching 96.28% and 97.43%, and weighted F1-scores of 96.26% and 97.41%, respectively. A stacking model was implemented to enhance classification robustness, compensating for potential clustering errors and delivering consistent performance across varying metrics and cluster structures. The proposed data processing pipeline ensures automated, standardized, and scalable analysis of large-scale gene expression datasets, aligning with the principles of personalized medicine.

Klasifikace

  • Druh

    J<sub>imp</sub> - Článek v periodiku v databázi Web of Science

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    Biomedical signal processing and control

  • ISSN

    1746-8094

  • e-ISSN

    1746-8108

  • Svazek periodika

    2025

  • Číslo periodika v rámci svazku

    110

  • Stát vydavatele periodika

    GB - Spojené království Velké Británie a Severního Irska

  • Počet stran výsledku

    17

  • Strana od-do

    108-115

  • Kód UT WoS článku

    001529361000002

  • EID výsledku v databázi Scopus

    2-s2.0-105009419762