Examining the Impact of Distance-Based Similarity Metrics on the Performance of Projected Clustering Algorithm for Fingerprint Database Clustering
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F62690094%3A18450%2F25%3A50022579" target="_blank" >RIV/62690094:18450/25:50022579 - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/document/11087686" target="_blank" >https://ieeexplore.ieee.org/document/11087686</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/OJCS.2025.3591192" target="_blank" >10.1109/OJCS.2025.3591192</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Examining the Impact of Distance-Based Similarity Metrics on the Performance of Projected Clustering Algorithm for Fingerprint Database Clustering
Popis výsledku v původním jazyce
The projected clustering (PROCLUS) algorithm is a subspace clustering algorithm based on the k-medoids clustering approach. It is designed to address the challenges of irrelevant received signal strength (RSS) measurements in fingerprint vectors by focusing on meaningful subsets of RSS measurements from wireless APs, known as subspaces. Despite its robustness, its performance heavily depends on the chosen similarity metric, with Euclidean and Manhattan distances being the most common. While many researchers focus on modifying the algorithm to enhance performance, the impact of similarity metrics on clustering performance is often overlooked, despite its critical role in determining accuracy. As such, this article evaluates the clustering performance of the PROCLUS algorithm using five distance-based similarity metrics-Euclidean, Manhattan, Cosine similarity, Canberra, and Chebyshev-across six experimentally generated fingerprint databases. It aims to identify the best similarity metric for maximizing clustering performance for each of the six fingerprint databases, with silhouette scores used as the clustering performance metric. Simulation results show that Cosine similarity is the most effective metric with the PROCLUS algorithm. It consistently produces clusters with the highest silhouette scores, all well above the 0.25 threshold and were 24% to 136% higher than the scores achieved using the other distance-based metrics across the tested fingerprint databases. The Canberra distance performed variably, while Euclidean and Manhattan distances were less reliable. The Chebyshev distance consistently underperformed in all the databases considered. The findings in this article highlight the importance of choosing the appropriate similarity metric to perform clustering operations with the PROCLUS algorithm.
Název v anglickém jazyce
Examining the Impact of Distance-Based Similarity Metrics on the Performance of Projected Clustering Algorithm for Fingerprint Database Clustering
Popis výsledku anglicky
The projected clustering (PROCLUS) algorithm is a subspace clustering algorithm based on the k-medoids clustering approach. It is designed to address the challenges of irrelevant received signal strength (RSS) measurements in fingerprint vectors by focusing on meaningful subsets of RSS measurements from wireless APs, known as subspaces. Despite its robustness, its performance heavily depends on the chosen similarity metric, with Euclidean and Manhattan distances being the most common. While many researchers focus on modifying the algorithm to enhance performance, the impact of similarity metrics on clustering performance is often overlooked, despite its critical role in determining accuracy. As such, this article evaluates the clustering performance of the PROCLUS algorithm using five distance-based similarity metrics-Euclidean, Manhattan, Cosine similarity, Canberra, and Chebyshev-across six experimentally generated fingerprint databases. It aims to identify the best similarity metric for maximizing clustering performance for each of the six fingerprint databases, with silhouette scores used as the clustering performance metric. Simulation results show that Cosine similarity is the most effective metric with the PROCLUS algorithm. It consistently produces clusters with the highest silhouette scores, all well above the 0.25 threshold and were 24% to 136% higher than the scores achieved using the other distance-based metrics across the tested fingerprint databases. The Canberra distance performed variably, while Euclidean and Manhattan distances were less reliable. The Chebyshev distance consistently underperformed in all the databases considered. The findings in this article highlight the importance of choosing the appropriate similarity metric to perform clustering operations with the PROCLUS algorithm.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
20206 - Computer hardware and architecture
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
IEEE OPEN JOURNAL OF THE COMPUTER SOCIETY
ISSN
2644-1268
e-ISSN
2644-1268
Svazek periodika
6
Číslo periodika v rámci svazku
July
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
12
Strana od-do
1248-1259
Kód UT WoS článku
001547282700002
EID výsledku v databázi Scopus
2-s2.0-105011764756