Enhancing Privacy While Preserving Context in Text Transformations by Large Language Models
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AQ2EK6NMI" target="_blank" >RIV/00216208:11320/26:Q2EK6NMI - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.3390/info16010049" target="_blank" >http://dx.doi.org/10.3390/info16010049</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.3390/info16010049" target="_blank" >10.3390/info16010049</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Enhancing Privacy While Preserving Context in Text Transformations by Large Language Models
Popis výsledku v původním jazyce
Data security is a critical concern for Internet users, primarily as more people rely on social networks and online tools daily. Despite the convenience, many users are unaware of the risks posed to their sensitive and personal data. This study addresses this issue by presenting a comprehensive solution to prevent personal data leakage using online tools. We developed a conceptual solution that enhances user privacy by identifying and anonymizing named entity classes representing sensitive data while maintaining the original context by swapping source entities for functional data. Our approach utilizes natural language processing methods, combining machine learning tools such as MITIE and spaCy with rule-based text analysis. We employed regular expressions and large language models to anonymize text, preserving its context for further processing or enabling restoration to the original form after transformations. The results demonstrate the effectiveness of our custom-trained models, achieving an F1 score of 0.8292. Additionally, the proposed algorithms successfully preserved context in approximately 93.23% of test cases, indicating a promising solution for secure data handling in online environments. © 2025 by the authors.
Název v anglickém jazyce
Enhancing Privacy While Preserving Context in Text Transformations by Large Language Models
Popis výsledku anglicky
Data security is a critical concern for Internet users, primarily as more people rely on social networks and online tools daily. Despite the convenience, many users are unaware of the risks posed to their sensitive and personal data. This study addresses this issue by presenting a comprehensive solution to prevent personal data leakage using online tools. We developed a conceptual solution that enhances user privacy by identifying and anonymizing named entity classes representing sensitive data while maintaining the original context by swapping source entities for functional data. Our approach utilizes natural language processing methods, combining machine learning tools such as MITIE and spaCy with rule-based text analysis. We employed regular expressions and large language models to anonymize text, preserving its context for further processing or enabling restoration to the original form after transformations. The results demonstrate the effectiveness of our custom-trained models, achieving an F1 score of 0.8292. Additionally, the proposed algorithms successfully preserved context in approximately 93.23% of test cases, indicating a promising solution for secure data handling in online environments. © 2025 by the authors.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Information (Switzerland)
ISSN
2078-2489
e-ISSN
—
Svazek periodika
16
Číslo periodika v rámci svazku
1
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
18
Strana od-do
1-18
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-85215656945