Self-distillation-based domain exploration for source speaker verification under spoofed speech from unknown voice conversion
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0201395" target="_blank" >RIV/00216305:26230/26:0201395 - isvavai.cz</a>
Result on the web
<a href="https://www.sciencedirect.com/science/article/pii/S0167639324001249?pes=vor&utm_source=scopus&getft_integrator=scopus" target="_blank" >https://www.sciencedirect.com/science/article/pii/S0167639324001249?pes=vor&utm_source=scopus&getft_integrator=scopus</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.specom.2024.103153" target="_blank" >10.1016/j.specom.2024.103153</a>
Alternative languages
Result language
angličtina
Original language name
Self-distillation-based domain exploration for source speaker verification under spoofed speech from unknown voice conversion
Original language description
Advancements in voice conversion (VC) technology have made it easier to generate spoofed speech that closely resembles the identity of a target speaker. Meanwhile, verification systems within the realm of speech processing are widely used to identify speakers. However, the misuse of VC algorithms poses significant privacy and security risks by potentially deceiving these systems. To address this issue, source speaker verification (SSV) has been proposed to verify the source speaker's identity of the spoofed speech generated by VCs. Nevertheless, SSV often suffers severe performance degradation when confronted with unknown VC algorithms, which is usually neglected by researchers. To deal with this cross-voice-conversion scenario and enhance the model's performance when facing unknown VC methods, we redefine it as a novel domain adaptation task by treating each VC method as a distinct domain. In this context, we propose an unsupervised domain adaptation (UDA) algorithm termed self-distillation-based domain exploration (SDDE). This algorithm adopts a siamese framework with two branches: one trained on the source (known) domain and the other trained on the target domains (unknown VC methods). The branch trained on the source domain leverages supervised learning to capture the source speaker's intrinsic features. Meanwhile, the branch trained on the target domain employs self-distillation to explore target domain information from multi-scale segments. Additionally, we have constructed a large-scale data set comprising over 7945 h of spoofed speech to evaluate the proposed SDDE. Experimental results on this data set demonstrate that SDDE outperforms traditional UDAs and substantially enhances the performance of the SSV model under unknown VC scenarios. The code for data generation and the trial lists are available at https://github.com/zrtlemontree/cross-domain-source-speaker-verification.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
S - Specificky vyzkum na vysokych skolach
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Speech communication
ISSN
0167-6393
e-ISSN
1872-7182
Volume of the periodical
167
Issue of the periodical within the volume
103153
Country of publishing house
NL - THE KINGDOM OF THE NETHERLANDS
Number of pages
12
Pages from-to
1-12
UT code for WoS article
001391212500001
EID of the result in the Scopus database
2-s2.0-85212173287