ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511657" target="_blank" >RIV/00216208:11320/25:10511657 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/chapter/10.1007/978-3-032-02548-7_25" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-032-02548-7_25</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-032-02548-7_25" target="_blank" >10.1007/978-3-032-02548-7_25</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Popis výsledku v původním jazyce
We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 h. We combined the sound recordings of the Czech parliamentary speeches with the official transcripts. The recordings were processed with WhisperX and Wav2Vec 2.0 to extract automated audio-text alignment. Our processing pipeline improves upon the ParCzech 3.0 speech recognition version by extracting more data with higher alignment reliability. The dataset is offered in three flexible variants: (1) sentence-segmented for automatic speech recognition and speech synthesis tasks with clean boundaries, (2) unsegmented preserving original utterance flow across sentences, and (3) a raw-alignment for further custom refinement for other possible tasks. All variants maintain the original metadata and are released under a permissive CC-BY license. The dataset is available in the LINDAT repository, with the sentence-segmented and unsegmented variants additi
Název v anglickém jazyce
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Popis výsledku anglicky
We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 h. We combined the sound recordings of the Czech parliamentary speeches with the official transcripts. The recordings were processed with WhisperX and Wav2Vec 2.0 to extract automated audio-text alignment. Our processing pipeline improves upon the ParCzech 3.0 speech recognition version by extracting more data with higher alignment reliability. The dataset is offered in three flexible variants: (1) sentence-segmented for automatic speech recognition and speech synthesis tasks with clean boundaries, (2) unsegmented preserving original utterance flow across sentences, and (3) a raw-alignment for further custom refinement for other possible tasks. All variants maintain the original metadata and are released under a permissive CC-BY license. The dataset is available in the LINDAT repository, with the sentence-segmented and unsegmented variants additi
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EH23_020%2F0008518" target="_blank" >EH23_020/0008518: Jazykověda, umělá inteligence a jazykové a řečové technologie: od výzkumu k aplikacím</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
28th International Conference on Text, Speech and Dialogue (Part I)
ISBN
978-3-032-02548-7
ISSN
—
e-ISSN
—
Počet stran výsledku
10
Strana od-do
299-308
Název nakladatele
Springer
Místo vydání
Cham, Switzerland
Místo konání akce
Erlangen, Germany
Datum konání akce
25. 8. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—