Meet-in-Style: Text-Driven Real-Time Video Stylization Using Diffusion Models
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00383170" target="_blank" >RIV/68407700:21230/25:00383170 - isvavai.cz</a>
Výsledek na webu
<a href="https://doi.org/10.1109/MCG.2025.3554312" target="_blank" >https://doi.org/10.1109/MCG.2025.3554312</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/MCG.2025.3554312" target="_blank" >10.1109/MCG.2025.3554312</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Meet-in-Style: Text-Driven Real-Time Video Stylization Using Diffusion Models
Popis výsledku v původním jazyce
We present Meet-in-Style—a new approach to real-time stylization of live video streams using text prompts. In contrast to previous text-based techniques, our system is able to stylize input video at 30 fps on commodity graphics hardware while preserving structural consistency of the stylized sequence and minimizing temporal flicker. A key idea of our approach is to combine diffusion-based image stylization with a few-shot patch-based training strategy that can produce a custom image-to-image stylization network with real-time inference capabilities. Such a combination not only allows for fast stylization, but also greatly improves consistency of individual stylized frames compared to a scenario where diffusion is applied to each video frame separately. We conducted a number of user experiments in which we found our approach to be particularly useful in video conference scenarios enabling participants to interactively apply different visual styles to themselves (or to each other) to enhance the overall chatting experience.
Název v anglickém jazyce
Meet-in-Style: Text-Driven Real-Time Video Stylization Using Diffusion Models
Popis výsledku anglicky
We present Meet-in-Style—a new approach to real-time stylization of live video streams using text prompts. In contrast to previous text-based techniques, our system is able to stylize input video at 30 fps on commodity graphics hardware while preserving structural consistency of the stylized sequence and minimizing temporal flicker. A key idea of our approach is to combine diffusion-based image stylization with a few-shot patch-based training strategy that can produce a custom image-to-image stylization network with real-time inference capabilities. Such a combination not only allows for fast stylization, but also greatly improves consistency of individual stylized frames compared to a scenario where diffusion is applied to each video frame separately. We conducted a number of user experiments in which we found our approach to be particularly useful in video conference scenarios enabling participants to interactively apply different visual styles to themselves (or to each other) to enhance the overall chatting experience.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EF16_019%2F0000765" target="_blank" >EF16_019/0000765: Výzkumné centrum informatiky</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>S - Specificky vyzkum na vysokych skolach
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
IEEE Computer Graphics and Applications
ISSN
0272-1716
e-ISSN
1558-1756
Svazek periodika
45
Číslo periodika v rámci svazku
2
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
10
Strana od-do
47-56
Kód UT WoS článku
001508290000017
EID výsledku v databázi Scopus
2-s2.0-105001518462