The Effect of Generating Synthetic Data in Smart City Network Systems
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F62690094%3A18450%2F25%3A50022318" target="_blank" >RIV/62690094:18450/25:50022318 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/article/10.1007/s42979-025-03673-3" target="_blank" >https://link.springer.com/article/10.1007/s42979-025-03673-3</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/s42979-025-03673-3" target="_blank" >10.1007/s42979-025-03673-3</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
The Effect of Generating Synthetic Data in Smart City Network Systems
Popis výsledku v původním jazyce
This study examines the effect of synthetic data generation for balancing class distributions on the performance of classification algorithms in smart city network systems. Contrary to the assumption that data balancing improves classification performance, the analysis reveals a more complex impact. Using three publicly available network traffic benchmark datasets and four different balancing techniques, the study evaluates the performance of five classifiers on 65 classification tasks. The findings indicate that, for smaller datasets, classifiers that achieved the highest accuracy on unbalanced data did not benefit from synthetic data generation for minority classes. Although neural network-based classifiers showed improved performance with balanced data, these improvements came at the cost of lower overall classification scores. For larger datasets, balancing through random oversampling of minority classes and undersampling of majority classes helped improve classification. However, these improvements were limited to precision, with no significant gains in recall. The study offers valuable insights into using synthetic data for intrusion detection, emphasizing the challenges of intricate dependencies in network traffic data for generative models. The results align with previous research showing mixed effects of data balancing on classifier performance, contributing to a broader understanding of the limited efficacy of synthetic data in real-world network contexts. This experimental study highlights the need for a systematic benchmarking framework for synthetic data research, ensuring consistency in data balancing and classification processes. This work contributes to the ongoing discourse on the intersection of machine learning and cybersecurity, emphasizing the critical role of data in developing resilient intrusion detection systems. © The Author(s) 2025.
Název v anglickém jazyce
The Effect of Generating Synthetic Data in Smart City Network Systems
Popis výsledku anglicky
This study examines the effect of synthetic data generation for balancing class distributions on the performance of classification algorithms in smart city network systems. Contrary to the assumption that data balancing improves classification performance, the analysis reveals a more complex impact. Using three publicly available network traffic benchmark datasets and four different balancing techniques, the study evaluates the performance of five classifiers on 65 classification tasks. The findings indicate that, for smaller datasets, classifiers that achieved the highest accuracy on unbalanced data did not benefit from synthetic data generation for minority classes. Although neural network-based classifiers showed improved performance with balanced data, these improvements came at the cost of lower overall classification scores. For larger datasets, balancing through random oversampling of minority classes and undersampling of majority classes helped improve classification. However, these improvements were limited to precision, with no significant gains in recall. The study offers valuable insights into using synthetic data for intrusion detection, emphasizing the challenges of intricate dependencies in network traffic data for generative models. The results align with previous research showing mixed effects of data balancing on classifier performance, contributing to a broader understanding of the limited efficacy of synthetic data in real-world network contexts. This experimental study highlights the need for a systematic benchmarking framework for synthetic data research, ensuring consistency in data balancing and classification processes. This work contributes to the ongoing discourse on the intersection of machine learning and cybersecurity, emphasizing the critical role of data in developing resilient intrusion detection systems. © The Author(s) 2025.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/VJ02010016" target="_blank" >VJ02010016: Využití umělé inteligence pro zajištění kybernetické bezpečnosti Smart City</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
SN Computer Science
ISSN
2662-995X
e-ISSN
2661-8907
Svazek periodika
6
Číslo periodika v rámci svazku
2
Stát vydavatele periodika
SG - Singapurská republika
Počet stran výsledku
18
Strana od-do
"Article number: 174"
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-85219672490