Augmenting Security Logs with Artificial Intelligence: Are Deep Models the Missing Piece?
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26220%2F26%3A0201114" target="_blank" >RIV/00216305:26220/26:0201114 - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/document/11268672" target="_blank" >https://ieeexplore.ieee.org/document/11268672</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICUMT67815.2025.11268672" target="_blank" >10.1109/ICUMT67815.2025.11268672</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Augmenting Security Logs with Artificial Intelligence: Are Deep Models the Missing Piece?
Popis výsledku v původním jazyce
The analysis of security logs remains a major challenge for modern Security Information and Event Management (SIEM) systems due to insufficient standardization and diversity of log formats. While Artificial Intelligence (AI) offers great potential for automating monitoring, its use is limited by data sensitivity and a lack of annotated datasets. Augmentation can help generate realistic synthetic logs, providing broader opportunities for AI deployment. This article presents a framework for training language models to generate structured log variants, focusing on key metadata fields while maintaining syntactic consistency and semantic relevance. This framework increases data diversity, reduces the need for manual labeling, and facilitates the integration of AI into Security Operations Centers (SOCs), thereby enhancing operational efficiency. A heterogeneous corpus from 49 sources was cleaned, deduplicated, and transformed into semantically distinct entities. Two augmentation strategies were evaluated: Masked Language Modeling (MLM) and Next Word Prediction (NWP). Eight transformer-based models were finetuned and tested on simulated attack scenarios generated using the Atomic Red Team framework and compared with largescale models to assess accuracy and computational efficiency. The results demonstrate the potential of domain-specific language models for context-aware protocol augmentation, contributing to more efficient and automated security systems.
Název v anglickém jazyce
Augmenting Security Logs with Artificial Intelligence: Are Deep Models the Missing Piece?
Popis výsledku anglicky
The analysis of security logs remains a major challenge for modern Security Information and Event Management (SIEM) systems due to insufficient standardization and diversity of log formats. While Artificial Intelligence (AI) offers great potential for automating monitoring, its use is limited by data sensitivity and a lack of annotated datasets. Augmentation can help generate realistic synthetic logs, providing broader opportunities for AI deployment. This article presents a framework for training language models to generate structured log variants, focusing on key metadata fields while maintaining syntactic consistency and semantic relevance. This framework increases data diversity, reduces the need for manual labeling, and facilitates the integration of AI into Security Operations Centers (SOCs), thereby enhancing operational efficiency. A heterogeneous corpus from 49 sources was cleaned, deduplicated, and transformed into semantically distinct entities. Two augmentation strategies were evaluated: Masked Language Modeling (MLM) and Next Word Prediction (NWP). Eight transformer-based models were finetuned and tested on simulated attack scenarios generated using the Atomic Red Team framework and compared with largescale models to assess accuracy and computational efficiency. The results demonstrate the potential of domain-specific language models for context-aware protocol augmentation, contributing to more efficient and automated security systems.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
20203 - Telecommunications
Návaznosti výsledku
Projekt
<a href="/cs/project/VB02000059" target="_blank" >VB02000059: Platforma pro adaptivní dolování znalostí z logových záznamů pomocí technik umělé inteligence</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
2025 17th International Congress on Ultra Modern Telecommunications and Control Systems and Workshops (ICUMT)
ISBN
979-8-3315-7675-2
ISSN
—
e-ISSN
2157-023X
Počet stran výsledku
6
Strana od-do
—
Název nakladatele
IEEE
Místo vydání
Florence, Italy
Místo konání akce
Florencie, Itálie
Datum konání akce
3. 11. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—