TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A2EC93SW8" target="_blank" >RIV/00216208:11320/26:2EC93SW8 - isvavai.cz</a>
Result on the web
<a href="https://ieeexplore.ieee.org/abstract/document/10890256" target="_blank" >https://ieeexplore.ieee.org/abstract/document/10890256</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICASSP49660.2025.10890256" target="_blank" >10.1109/ICASSP49660.2025.10890256</a>
Alternative languages
Result language
angličtina
Original language name
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
Original language description
This work introduces TTS-Transducer – a novel architecture for text-to-speech, leveraging the strengths of audio codec models and neural transducers. Transducers, renowned for their superior quality and robustness in speech recognition, are employed to learn monotonic alignments and allow for avoiding using explicit duration predictors. Neural audio codecs efficiently compress audio into discrete codes, revealing the possibility of applying text modeling approaches to speech generation. However, the complexity of predicting multiple tokens per frame from several codebooks, as necessitated by audio codec models with residual quantizers, poses a significant challenge. The proposed system first uses a transducer architecture to learn monotonic alignments between tokenized text and speech codec tokens for the first codebook. Next, a non-autoregressive Transformer predicts the remaining codes using the alignment extracted from transducer loss. The proposed system is trained end-to-end. We show that TTS-Transducer is a competitive and robust alternative to contemporary TTS systems1.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
ISBN
—
ISSN
15206149
e-ISSN
—
Number of pages
5
Pages from-to
1-5
Publisher name
—
Place of publication
—
Event location
Hyderabad
Event date
Jan 1, 2026
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—