Unlocking the Potential of Federated Learning: The Symphony of Dataset Distillation via Deep Generative Latents
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00377735" target="_blank" >RIV/68407700:21230/25:00377735 - isvavai.cz</a>
Výsledek na webu
<a href="https://doi.org/10.1007/978-3-031-73229-4_2" target="_blank" >https://doi.org/10.1007/978-3-031-73229-4_2</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-031-73229-4_2" target="_blank" >10.1007/978-3-031-73229-4_2</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Unlocking the Potential of Federated Learning: The Symphony of Dataset Distillation via Deep Generative Latents
Popis výsledku v původním jazyce
Data heterogeneity presents significant challenges for federated learning (FL). Recently, dataset distillation techniques have been introduced, and performed at the client level, to attempt to mitigate some of these challenges. In this paper, we propose a highly efficient FL dataset distillation framework on the server side, significantly reducing both the computational and communication demands on local devices while enhancing the clients’ privacy. Unlike previous strategies that perform dataset distillation on local devices and upload synthetic data to the server, our technique enables the server to leverage prior knowledge from pre-trained deep generative models to synthesize essential data representations from a heterogeneous model architecture. This process allows local devices to train smaller surrogate models while enabling the training of a larger global model on the server, effectively minimizing resource utilization. We substantiate our claim with a theoretical analysis, demonstrating the asymptotic resemblance of the process to the hypothetical ideal of completely centralized training on a heterogeneous dataset. Empirical evidence from our comprehensive experiments indicates our method’s superiority, delivering an accuracy enhancement of up to 40% over non-dataset-distillation techniques in highly heterogeneous FL contexts, and surpassing existing dataset-distillation methods by 18%.
Název v anglickém jazyce
Unlocking the Potential of Federated Learning: The Symphony of Dataset Distillation via Deep Generative Latents
Popis výsledku anglicky
Data heterogeneity presents significant challenges for federated learning (FL). Recently, dataset distillation techniques have been introduced, and performed at the client level, to attempt to mitigate some of these challenges. In this paper, we propose a highly efficient FL dataset distillation framework on the server side, significantly reducing both the computational and communication demands on local devices while enhancing the clients’ privacy. Unlike previous strategies that perform dataset distillation on local devices and upload synthetic data to the server, our technique enables the server to leverage prior knowledge from pre-trained deep generative models to synthesize essential data representations from a heterogeneous model architecture. This process allows local devices to train smaller surrogate models while enabling the training of a larger global model on the server, effectively minimizing resource utilization. We substantiate our claim with a theoretical analysis, demonstrating the asymptotic resemblance of the process to the hypothetical ideal of completely centralized training on a heterogeneous dataset. Empirical evidence from our comprehensive experiments indicates our method’s superiority, delivering an accuracy enhancement of up to 40% over non-dataset-distillation techniques in highly heterogeneous FL contexts, and surpassing existing dataset-distillation methods by 18%.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/GA24-11664S" target="_blank" >GA24-11664S: Relační posilované učení pro akceleraci vědy</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Computer Vision – ECCV 2024, Part LXXVIII
ISBN
978-3-031-91569-7
ISSN
0302-9743
e-ISSN
1611-3349
Počet stran výsledku
16
Strana od-do
18-33
Název nakladatele
Springer Nature
Místo vydání
—
Místo konání akce
Milano
Datum konání akce
29. 9. 2024
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001352814300002