Enhancing CTAO monitoring and alarm subsystems in distributed environments using ServiMon
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68378271%3A_____%2F25%3A00648264" target="_blank" >RIV/68378271:_____/25:00648264 - isvavai.cz</a>
Result on the web
<a href="https://pos.sissa.it/501/775/pdf" target="_blank" >https://pos.sissa.it/501/775/pdf</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.22323/1.501.0775" target="_blank" >10.22323/1.501.0775</a>
Alternative languages
Result language
angličtina
Original language name
Enhancing CTAO monitoring and alarm subsystems in distributed environments using ServiMon
Original language description
ServiMon is a scalable data collection and auditing pipeline designed for service-oriented, cost-efficient quality control in distributed environments, including the CTAO monitoring, logging, and alarm subsystems. Developed within a Docker-based architecture, it leverages cloud-native technologies and distributed computing principles to enhance system observability and reliability. At its core, ServiMon integrates key technologies such as Prometheus, Grafana, Kafka, and Cassandra. Prometheus serves as the primary engine for real-time performance metric collection, enabling efficient monitoring across multiple nodes. Grafana provides interactive, service-oriented data visualization, facilitating system performance analysis. Additionally, Kafka and Cassandra expose system metrics via the JMX Exporter, offering critical insights into infrastructure availability and performance. This contribution exposes how ServiMon could provide an enhancement on scalability, security, and efficiency in a distributed computing environment, such as the CTAO monitoring, logging, and alarm subsystems. This integrated approach not only ensures robust real-time monitoring, but also optimizes operational costs. Furthermore, ServiMon’s ability to generate large volumes of diverse data over time provides a strong foundation for predictive maintenance. By incorporating stochastic and approximate computing techniques, it enables proactive failure detection and system optimization, minimizing downtime and maximizing telescope availability.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10303 - Particles and field physics
Result continuities
Project
Result was created during the realization of more than one project. More information in the Projects tab.
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of Science
ISBN
—
ISSN
1824-8039
e-ISSN
—
Number of pages
8
Pages from-to
775
Publisher name
Sissa Medilab srl
Place of publication
Trieste
Event location
Ženeva
Event date
Jul 15, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—