Ensemble-based clustering and classification pipeline for cancer diagnosis using gene expression data
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F44555601%3A13440%2F25%3A43899385" target="_blank" >RIV/44555601:13440/25:43899385 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.sciencedirect.com/science/article/pii/S1746809425016441?pes=vor&utm_source=scopus&getft_integrator=scopus" target="_blank" >https://www.sciencedirect.com/science/article/pii/S1746809425016441?pes=vor&utm_source=scopus&getft_integrator=scopus</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.bspc.2025.109133" target="_blank" >10.1016/j.bspc.2025.109133</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Ensemble-based clustering and classification pipeline for cancer diagnosis using gene expression data
Popis výsledku v původním jazyce
Objective: Accurate analysis of gene expression data is essential for understanding cancer mechanisms and improving diagnostics. However, the high dimensionality and heterogeneity of transcriptomic profiles often produce unstable clustering and classification outcomes. This study proposes an ensemble-based pipeline designed to improve robustness, interpretability, and clinical utility in cancer diagnostics. Methods: We developed a hybrid framework that integrates the Self-Organizing Tree Algorithm (SOTA) with agglomerative and spectral consensus clustering. Gene expression data from 6310 samples and 18,564 genes across 14 classes were transformed into cluster-based subsets. Classification was performed using Random Forest models with Bayesian hyperparameter optimization and out-of-fold stacking. Clustering quality was evaluated using the Relative Cluster Separation Index (RCSI), Calinski?Harabasz, Silhouette, and PBM indices, while biological interpretation was based on KEGG enrichment and Cytoscape (ClueGO/CluePedia) functional networks. Results: The spectral consensus variant of SOTA achieved the most balanced performance across all metrics. On TCGA data, three- to five-cluster configurations provided high diagnostic accuracy (Accuracy ? 0.975, F1 ? 0.976) with statistically validated improvements (p<0.05) over baseline models. External validation on the independent expO dataset confirmed cross-platform robustness (Accuracy ? 0.811, F1 ? 0.843). Functional enrichment linked each cluster to biologically coherent pathways, including immune regulation, neuroactive signaling, metabolism, and viral response. Conclusion: The proposed ensemble?clustering?classification framework unifies consensus clustering, ensemble learning, and pathway-level validation, offering a reproducible and interpretable approach for cancer gene expression analysis, biomarker discovery, and precision oncology applications.
Název v anglickém jazyce
Ensemble-based clustering and classification pipeline for cancer diagnosis using gene expression data
Popis výsledku anglicky
Objective: Accurate analysis of gene expression data is essential for understanding cancer mechanisms and improving diagnostics. However, the high dimensionality and heterogeneity of transcriptomic profiles often produce unstable clustering and classification outcomes. This study proposes an ensemble-based pipeline designed to improve robustness, interpretability, and clinical utility in cancer diagnostics. Methods: We developed a hybrid framework that integrates the Self-Organizing Tree Algorithm (SOTA) with agglomerative and spectral consensus clustering. Gene expression data from 6310 samples and 18,564 genes across 14 classes were transformed into cluster-based subsets. Classification was performed using Random Forest models with Bayesian hyperparameter optimization and out-of-fold stacking. Clustering quality was evaluated using the Relative Cluster Separation Index (RCSI), Calinski?Harabasz, Silhouette, and PBM indices, while biological interpretation was based on KEGG enrichment and Cytoscape (ClueGO/CluePedia) functional networks. Results: The spectral consensus variant of SOTA achieved the most balanced performance across all metrics. On TCGA data, three- to five-cluster configurations provided high diagnostic accuracy (Accuracy ? 0.975, F1 ? 0.976) with statistically validated improvements (p<0.05) over baseline models. External validation on the independent expO dataset confirmed cross-platform robustness (Accuracy ? 0.811, F1 ? 0.843). Functional enrichment linked each cluster to biologically coherent pathways, including immune regulation, neuroactive signaling, metabolism, and viral response. Conclusion: The proposed ensemble?clustering?classification framework unifies consensus clustering, ensemble learning, and pathway-level validation, offering a reproducible and interpretable approach for cancer gene expression analysis, biomarker discovery, and precision oncology applications.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Biomedical signal processing and control
ISSN
1746-8094
e-ISSN
1746-8108
Svazek periodika
2025
Číslo periodika v rámci svazku
113
Stát vydavatele periodika
GB - Spojené království Velké Británie a Severního Irska
Počet stran výsledku
14
Strana od-do
"nestrankovano"
Kód UT WoS článku
001616477600001
EID výsledku v databázi Scopus
—