ESPnet-SpeechLM: An Open Speech Language Model Toolkit
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0201388" target="_blank" >RIV/00216305:26230/26:0201388 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2025.naacl-demo.12.pdf" target="_blank" >https://aclanthology.org/2025.naacl-demo.12.pdf</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.18653/v1/2025.naacl-demo.12" target="_blank" >10.18653/v1/2025.naacl-demo.12</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
Popis výsledku v původním jazyce
We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standardizes speech processing tasks by framing them as universal sequential modeling problems, encompassing a cohesive workflow of data preprocessing, pre-training, inference, and task evaluation. With ESPnet-SpeechLM, users can easily define task templates and configure key settings, enabling seamless and streamlined SpeechLM development. The toolkit ensures flexibility, efficiency, and scalability by offering highly configurable modules for every stage of the workflow. To illustrate its capabilities, we provide multiple use cases demonstrating how competitive SpeechLMs can be constructed with ESPnet-SpeechLM, including a 1.7B-parameter model pre-trained on both text and speech tasks, across diverse benchmarks.
Název v anglickém jazyce
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
Popis výsledku anglicky
We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standardizes speech processing tasks by framing them as universal sequential modeling problems, encompassing a cohesive workflow of data preprocessing, pre-training, inference, and task evaluation. With ESPnet-SpeechLM, users can easily define task templates and configure key settings, enabling seamless and streamlined SpeechLM development. The toolkit ensures flexibility, efficiency, and scalability by offering highly configurable modules for every stage of the workflow. To illustrate its capabilities, we provide multiple use cases demonstrating how competitive SpeechLMs can be constructed with ESPnet-SpeechLM, including a 1.7B-parameter model pre-trained on both text and speech tasks, across diverse benchmarks.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
S - Specificky vyzkum na vysokych skolach
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025
ISBN
9798891761919
ISSN
—
e-ISSN
—
Počet stran výsledku
9
Strana od-do
116-124
Název nakladatele
Association for Computational Linguistics (ACL)
Místo vydání
Hybrid, Albuquerque, New Mexico, USA
Místo konání akce
Hybrid, Albuquerque, NM, USA
Datum konání akce
29. 4. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001598875800012