Ensemble-based clustering and classification pipeline for cancer diagnosis using gene expression data
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F44555601%3A13440%2F25%3A43899385" target="_blank" >RIV/44555601:13440/25:43899385 - isvavai.cz</a>
Result on the web
<a href="https://www.sciencedirect.com/science/article/pii/S1746809425016441?pes=vor&utm_source=scopus&getft_integrator=scopus" target="_blank" >https://www.sciencedirect.com/science/article/pii/S1746809425016441?pes=vor&utm_source=scopus&getft_integrator=scopus</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.bspc.2025.109133" target="_blank" >10.1016/j.bspc.2025.109133</a>
Alternative languages
Result language
angličtina
Original language name
Ensemble-based clustering and classification pipeline for cancer diagnosis using gene expression data
Original language description
Objective: Accurate analysis of gene expression data is essential for understanding cancer mechanisms and improving diagnostics. However, the high dimensionality and heterogeneity of transcriptomic profiles often produce unstable clustering and classification outcomes. This study proposes an ensemble-based pipeline designed to improve robustness, interpretability, and clinical utility in cancer diagnostics. Methods: We developed a hybrid framework that integrates the Self-Organizing Tree Algorithm (SOTA) with agglomerative and spectral consensus clustering. Gene expression data from 6310 samples and 18,564 genes across 14 classes were transformed into cluster-based subsets. Classification was performed using Random Forest models with Bayesian hyperparameter optimization and out-of-fold stacking. Clustering quality was evaluated using the Relative Cluster Separation Index (RCSI), Calinski?Harabasz, Silhouette, and PBM indices, while biological interpretation was based on KEGG enrichment and Cytoscape (ClueGO/CluePedia) functional networks. Results: The spectral consensus variant of SOTA achieved the most balanced performance across all metrics. On TCGA data, three- to five-cluster configurations provided high diagnostic accuracy (Accuracy ? 0.975, F1 ? 0.976) with statistically validated improvements (p<0.05) over baseline models. External validation on the independent expO dataset confirmed cross-platform robustness (Accuracy ? 0.811, F1 ? 0.843). Functional enrichment linked each cluster to biologically coherent pathways, including immune regulation, neuroactive signaling, metabolism, and viral response. Conclusion: The proposed ensemble?clustering?classification framework unifies consensus clustering, ensemble learning, and pathway-level validation, offering a reproducible and interpretable approach for cancer gene expression analysis, biomarker discovery, and precision oncology applications.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Biomedical signal processing and control
ISSN
1746-8094
e-ISSN
1746-8108
Volume of the periodical
2025
Issue of the periodical within the volume
113
Country of publishing house
GB - UNITED KINGDOM
Number of pages
14
Pages from-to
"nestrankovano"
UT code for WoS article
001616477600001
EID of the result in the Scopus database
—