Warping: Data-driven Mixture Preprocessing to Boost the Performance of Blind Speech Separation
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F46747885%3A24220%2F25%3A00013595" target="_blank" >RIV/46747885:24220/25:00013595 - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11226076&utm_source=scopus&getft_integrator=scopus" target="_blank" >https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11226076&utm_source=scopus&getft_integrator=scopus</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Warping: Data-driven Mixture Preprocessing to Boost the Performance of Blind Speech Separation
Popis výsledku v původním jazyce
Blind source separation (BSS) can be used to recover speech signals from mixtures recorded by microphones. However, their performances show significant limitations because of the deviations between the instantaneous mixing model and real audio mixtures transformed into the short-time Fourier domain (STFT). This paper presents a data-driven preprocessing technique called mixture warping, which aims to adjust the mixture to obey the instantaneous model as much as possible. As a proof of concept, we demonstrate its effect on a set of reverberant mixtures of two speakers. Warping implemented through a deep neural network is trained to estimate mixtures ideally modified towards the instantaneous model in the least squares sense. By applying it as a preprocessing stage, it boosts the BSS performance by up to 6.4 dB of signal-to-interference (SIR) and 4.7 dB of signal-to-distortion (SDR) on average without requiring any modification to the BSS methods.
Název v anglickém jazyce
Warping: Data-driven Mixture Preprocessing to Boost the Performance of Blind Speech Separation
Popis výsledku anglicky
Blind source separation (BSS) can be used to recover speech signals from mixtures recorded by microphones. However, their performances show significant limitations because of the deviations between the instantaneous mixing model and real audio mixtures transformed into the short-time Fourier domain (STFT). This paper presents a data-driven preprocessing technique called mixture warping, which aims to adjust the mixture to obey the instantaneous model as much as possible. As a proof of concept, we demonstrate its effect on a set of reverberant mixtures of two speakers. Warping implemented through a deep neural network is trained to estimate mixtures ideally modified towards the instantaneous model in the least squares sense. By applying it as a preprocessing stage, it boosts the BSS performance by up to 6.4 dB of signal-to-interference (SIR) and 4.7 dB of signal-to-distortion (SDR) on average without requiring any modification to the BSS methods.
Klasifikace
Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/GA25-18485S" target="_blank" >GA25-18485S: Hybridní extrakce signálů: Synergie fyzikálních, informačně-teoretických a datových znalostí</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů