TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0198051" target="_blank" >RIV/00216305:26230/26:0198051 - isvavai.cz</a>
Result on the web
<a href="https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10887574" target="_blank" >https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10887574</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICASSP49660.2025.10887574" target="_blank" >10.1109/ICASSP49660.2025.10887574</a>
Alternative languages
Result language
angličtina
Original language name
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
Original language description
Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been pro- posed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in noisy, multi-talker conditions-a more challenging yet practical case. In this paper, we introduce the Target-Speaker Speech Processing Universal Performance Benchmark (TS-SUPERB), which includes four widely recognized target-speaker processing tasks that require identifying the target speaker and extracting information from the speech mixture. In our benchmark, the speaker embedding extracted from enrollment speech is used as a clue to condition downstream models. The benchmark result reveals the importance of evaluating SSL models in target speaker scenarios, demonstrating that performance cannot be easily inferred from related single-speaker tasks. Moreover, by using a unified SSL-based target speech encoder, consisting of a speaker encoder and an extractor module, we also investigate joint optimization across TS tasks to leverage mutual information and demonstrate its effectiveness.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
Result was created during the realization of more than one project. More information in the Projects tab.
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
ISBN
979-8-3503-6874-1
ISSN
—
e-ISSN
—
Number of pages
5
Pages from-to
1-5
Publisher name
IEEE Signal Processing Society
Place of publication
Hyderabad
Event location
Hyderabad
Event date
Apr 6, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—