SHDA: Sinkhorn Domain Attention for Cross-Domain Audio Anti-Spoofing
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0199982" target="_blank" >RIV/00216305:26230/26:0199982 - isvavai.cz</a>
Result on the web
<a href="https://ieeexplore.ieee.org/abstract/document/11024052" target="_blank" >https://ieeexplore.ieee.org/abstract/document/11024052</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/TIFS.2025.3576576" target="_blank" >10.1109/TIFS.2025.3576576</a>
Alternative languages
Result language
angličtina
Original language name
SHDA: Sinkhorn Domain Attention for Cross-Domain Audio Anti-Spoofing
Original language description
Audio anti-spoofing algorithms struggle with fake samples from unseen spoofing techniques, even when trained with diverse data sets or data augmentation strategies. Unsupervised domain adaptation (UDA) algorithms have the potential to mitigate this challenge. Typically, UDA assumes that the source and target domains are distinct distributions with clear boundaries and seeks to align model representations between them. However, in anti-spoofing, various spoofing algorithms could cause the distributions of the generated samples to overlap, resulting in unclear domain boundaries. This hinders UDA algorithms from effectively measuring and aligning domain discrepancies. Moreover, forcibly aligning samples with significant discrepancies could diminish the model's discriminative capability. To solve this problem, we propose a domain attention algorithm with optimal transport (OT), termed Sinkhorn Domain Attention (SHDA). Unlike traditional attention mechanisms, SHDA identifies the optimal transfer plan by analyzing the global probability differences among cross-domain samples. Specifically, we first extract audio representations from various domains to compute the overall cost matrix between the source and target domains. Next, we employ Sinkhorn's iteration to calculate the OT coupling matrix, where cross-domain samples with minor differences receive higher transfer weights, while those with substantial differences receive lower weights. Finally, we use the coupling and cost matrices to compute the adaptation loss, effectively transferring the anti-spoofing model from multiple sources to the target domain. We conducted eight cross-domain experiments using eleven well-known anti-spoofing corpora. The results indicate that our label-free SHDA surpassed the state-of-the-art model by 40%.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
S - Specificky vyzkum na vysokych skolach
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
IEEE Transactions on Information Forensics and Security
ISSN
1556-6013
e-ISSN
1556-6021
Volume of the periodical
—
Issue of the periodical within the volume
20
Country of publishing house
US - UNITED STATES
Number of pages
16
Pages from-to
6474-6489
UT code for WoS article
001521429100006
EID of the result in the Scopus database
—