Blind Extraction of Target Speech Source: Three ways of Guidance Exploiting Supervised Speaker Embeddings
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F46747885%3A24220%2F22%3A00009854" target="_blank" >RIV/46747885:24220/22:00009854 - isvavai.cz</a>
Result on the web
<a href="https://asap.ite.tul.cz/wp-content/uploads/sites/3/2022/08/paper_16.pdf" target="_blank" >https://asap.ite.tul.cz/wp-content/uploads/sites/3/2022/08/paper_16.pdf</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/IWAENC53105.2022.9914778" target="_blank" >10.1109/IWAENC53105.2022.9914778</a>
Alternative languages
Result language
angličtina
Original language name
Blind Extraction of Target Speech Source: Three ways of Guidance Exploiting Supervised Speaker Embeddings
Original language description
The manuscript deals with the robust extraction of a speaker of interest (SOI) from a mixture of audio sources. A blind algorithm based on independent vector extraction (IVE) is used, which, by definition, extracts an arbitrary source. To focus the extraction towards the SOI, a prior knowledge identifying the target source is required. To this end, the manuscript exploits speaker-identification based on embedding features computed via a pretrained forward sequential memory network (FSMN). We introduce and experimentally validate three ways how this prior knowledge can be employed in a blind algorithm, namely, speaker-specific initialization, pilot signal, and supervised deflation of the mixture. The experiments show that the proposed techniques complement each other and lead to robust identification/extraction of the SOI in difficult mixtures of three speakers.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
Result was created during the realization of more than one project. More information in the Projects tab.
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2022
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
International Workshop on Acoustic Signal Enhancement
ISBN
978-166546867-1
ISSN
—
e-ISSN
—
Number of pages
5
Pages from-to
—
Publisher name
IEEE
Place of publication
—
Event location
Bamberg
Event date
Jan 1, 2022
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
000934046400073