Learning Optimal Prosody Embedding Codebook based on F0 and Energy
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A29JTPTJH" target="_blank" >RIV/00216208:11320/26:29JTPTJH - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/00216224:14330/25:00141243
Výsledek na webu
<a href="http://dx.doi.org/10.21437/Interspeech.2025-1020" target="_blank" >http://dx.doi.org/10.21437/Interspeech.2025-1020</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.21437/Interspeech.2025-1020" target="_blank" >10.21437/Interspeech.2025-1020</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Learning Optimal Prosody Embedding Codebook based on F0 and Energy
Popis výsledku v původním jazyce
Both the Fundamental frequency (F0) and Energy are prominent features of prosody. Together, they have been used across a wide variety of speech-processing tasks. However, there is a lack of freely available pre-trained vector representations of these features. Therefore, in this paper, we provide the research community with high-quality joint embeddings of the frame-level F0 and Energy features, using the VQ-VAE architecture. By converting the F0 and Energy into a single stream of vector embeddings, we make it possible to seamlessly use prosody in modern architectures, such as multimodal LLMs. In order to ensure maximum embedding quality, we conduct a large-scale hyperparameter search, totaling over 150 experiments on the LibriTTS dataset. We outperform previous works on F0 embeddings, reaching FFE error below 1 percent, while simultaneously embedding the additional feature of Energy. We publish our best-performing models on the HuggingFace website.
Název v anglickém jazyce
Learning Optimal Prosody Embedding Codebook based on F0 and Energy
Popis výsledku anglicky
Both the Fundamental frequency (F0) and Energy are prominent features of prosody. Together, they have been used across a wide variety of speech-processing tasks. However, there is a lack of freely available pre-trained vector representations of these features. Therefore, in this paper, we provide the research community with high-quality joint embeddings of the frame-level F0 and Energy features, using the VQ-VAE architecture. By converting the F0 and Energy into a single stream of vector embeddings, we make it possible to seamlessly use prosody in modern architectures, such as multimodal LLMs. In order to ensure maximum embedding quality, we conduct a large-scale hyperparameter search, totaling over 150 experiments on the LibriTTS dataset. We outperform previous works on F0 embeddings, reaching FFE error below 1 percent, while simultaneously embedding the additional feature of Energy. We publish our best-performing models on the HuggingFace website.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Interspeech 2025
ISBN
—
ISSN
2308-457X
e-ISSN
—
Počet stran výsledku
5
Strana od-do
4728-4732
Název nakladatele
Isca-Int Speech Communication Assoc
Místo vydání
—
Místo konání akce
Baixas
Datum konání akce
1. 1. 2026
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001613931400370