Unsupervised Estimation of Nonlinear Audio Effects: Comparing Diffusion-Based and Adversarial approaches
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26220%2F26%3A0197846" target="_blank" >RIV/00216305:26220/26:0197846 - isvavai.cz</a>
Result on the web
<a href="https://www.scopus.com/pages/publications/105028968011?origin=resultslist" target="_blank" >https://www.scopus.com/pages/publications/105028968011?origin=resultslist</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
Unsupervised Estimation of Nonlinear Audio Effects: Comparing Diffusion-Based and Adversarial approaches
Original language description
Accurately estimating nonlinear audio effects without access to paired input-output signals remains a challenging problem. This work studies unsupervised probabilistic approaches for solving this task. We introduce a method, novel for this application, based on diffusion generative models for blind system identification, enabling the estimation of unknown nonlinear effects using black- and gray-box models. This study compares this method with a previously proposed adversarial approach, analyzing the performance of both methods under different parameterizations of the effect operator and varying lengths of available effected recordings. Through experiments on guitar distortion effects, we show that the diffusion-based approach provides more stable results and is less sensitive to data availability, while the adversarial approach is superior at estimating more pronounced distortion effects. Our findings contribute to the robust unsupervised blind estimation of audio effects, demonstrating the potential of diffusion models for system identification in music technology.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
20203 - Telecommunications
Result continuities
Project
<a href="/en/project/GA23-07294S" target="_blank" >GA23-07294S: From perceptron to perception: psychoacoustically motivated audio reconstruction using learned components</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>S - Specificky vyzkum na vysokych skolach
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the International Conference on Digital Audio Effects (DAFx)
ISBN
—
ISSN
2413-6689
e-ISSN
—
Number of pages
8
Pages from-to
366-373
Publisher name
Università Politecnica delle Marche
Place of publication
Ancona
Event location
Ancona
Event date
Sep 2, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—