Using synthetic data for pretraining partial discharge detection in overhead transmission lines
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989100%3A27240%2F25%3A10259354" target="_blank" >RIV/61989100:27240/25:10259354 - isvavai.cz</a>
Alternative codes found
RIV/61989100:27730/25:10259354
Result on the web
<a href="https://www.nature.com/articles/s41598-025-32642-2" target="_blank" >https://www.nature.com/articles/s41598-025-32642-2</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1038/s41598-025-32642-2" target="_blank" >10.1038/s41598-025-32642-2</a>
Alternative languages
Result language
angličtina
Original language name
Using synthetic data for pretraining partial discharge detection in overhead transmission lines
Original language description
Accurate detection of partial discharges (PDs) in medium-voltage overhead transmission lines is critical for preemptive maintenance and avoiding costly outages, yet it is challenged by scarce labeled data and pervasive electromagnetic interference. This paper investigates a hybrid simulation-and-data-driven framework in which synthetically generated PD signals are used to pretrain deep neural networks and are subsequently fine-tuned on a limited set of real overhead-line measurements. The synthetic pipeline systematically varies PD repetition rates, amplitude distributions, vegetation-contact scenarios, and noise conditions, producing diverse time-series and spectrogram-like representations that approximate real operating environments. We conduct a comprehensive ablation study across multiple architectures—Convolutional Neural Networks (CNNs), a Vision Transformer (ViT), and a Long Short-Term Memory (LSTM) network—and analyze their sensitivity to granular sweeps of synthetic-data parameters. CNN-based models decisively outperform ViT and LSTM counterparts on the spectrogram-based classification task, while ViT and LSTM fail to learn meaningful representation. For the successful CNNs, pretraining on carefully parameterized synthetic datasets—particularly those reflecting higher PD activity, such as our Datasets 3 and 4—consistently improves downstream performance on real data, boosting the Matthews Correlation Coefficient (MCC) on imbalanced, cost-sensitive test sets by roughly 10–20% compared with training from scratch. At the same time, we show that poorly aligned synthetic data can degrade generalization, underscoring the need for accurate noise calibration and domain-aligned simulation. Overall, the results confirm that (i) architectural choice is pivotal for PD detection in overhead lines and (ii) well-designed synthetic data is a powerful, practical lever for achieving reliable and cost-effective PD monitoring when real labeled data are limited. © The Author(s) 2025.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
S - Specificky vyzkum na vysokych skolach
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Scientific Reports
ISSN
2045-2322
e-ISSN
2045-2322
Volume of the periodical
15
Issue of the periodical within the volume
1
Country of publishing house
GB - UNITED KINGDOM
Number of pages
24
Pages from-to
15
UT code for WoS article
001651181400021
EID of the result in the Scopus database
2-s2.0-105026217746