Factorized RVQ-GAN For Disentangled Speech Tokenization
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0199387" target="_blank" >RIV/00216305:26230/26:0199387 - isvavai.cz</a>
Result on the web
<a href="https://www.isca-archive.org/interspeech_2025/khurana25_interspeech.pdf" target="_blank" >https://www.isca-archive.org/interspeech_2025/khurana25_interspeech.pdf</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.21437/Interspeech.2025-2612" target="_blank" >10.21437/Interspeech.2025-2612</a>
Alternative languages
Result language
angličtina
Original language name
Factorized RVQ-GAN For Disentangled Speech Tokenization
Original language description
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives: one from a pre-trained speech encoder (HuBERT) for phoneme-level structure, and another from a text-based encoder (LaBSE) for lexical cues. Experiments on English and multilingual data show that HAC's factorized bottleneck yields disentangled token sets: one aligns with phonemes, while another captures word-level semantics. Quantitative evaluations confirm that HAC tokens preserve naturalness and provide interpretable linguistic information, outperforming single-level baselines in both disentanglement and reconstruction quality. These findings underscore HAC's potential as a unified discrete speech representation, bridging acoustic detail and lexical meaning for downstream speech generation and understanding tasks.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/VK01020132" target="_blank" >VK01020132: Validation of integrating artificial intelligence for receiving emergency calls using a voice chatbot, developed within the research project BV No. VI20192022169, with technology for receiving emergency communications 112 and 150 in the CZE (TCTV 112)</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the Annual Conference of the International Speech Communication Association Interspeech
ISBN
—
ISSN
—
e-ISSN
2958-1796
Number of pages
5
Pages from-to
3514-3518
Publisher name
International Speech Communication Association
Place of publication
Rotterdam, The Netherlands
Event location
Brno
Event date
Aug 30, 2021
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
001613931400123