AI-Driven Web-Based Speech Transcription Tool: A Novel Approach for Efficient Evaluation of Verbal Memory Performance
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00159816%3A_____%2F25%3A00082289" target="_blank" >RIV/00159816:_____/25:00082289 - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/document/11284682" target="_blank" >https://ieeexplore.ieee.org/document/11284682</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/JBHI.2025.3620128" target="_blank" >10.1109/JBHI.2025.3620128</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
AI-Driven Web-Based Speech Transcription Tool: A Novel Approach for Efficient Evaluation of Verbal Memory Performance
Popis výsledku v původním jazyce
Memory deficits are prevalent in epilepsy and other brain disorders, significantly affecting quality of life. In particular, patients with mesial temporal lobe epilepsy require ongoing monitoring and repeated memory assessments to track cognitive function, which could benefit from automated transcription tools. We used a classic free recall (FR) verbal memory task, where participants recalled words, to assess their performance by recording, transcribing, and detecting vocalizations of correctly remembered words. Conventional manual speech transcription methods are time-consuming and prone to human error, especially in a noisy hospital environment. To address these limitations, a modified U-Net architecture was employed for noise reduction, resulting in a signal-to-noise ratio (SNR) of 15.8 and a mean squared error (MSE) of 0.0021. We also developed an automated transcription interface utilizing the Whisper speech recognition model, which was fine-tuned for Polish, Czech, and English languages. Dynamic Time Warping (DTW) was applied to provide precise word-level timestamps of vocalization onset and offset. The interface was iteratively refined over eight development cycles, incorporating feedback from target users. Transcription accuracy was evaluated with Word Error Rates (WER) of 10.3% for Czech, 7.1% for Polish, and 5% for English, alongside respective Character Error Rates (CER) of 12%, 10.8%, and 7.5%. Our automated interface outperformed manual transcription, reducing transcription time fourfold and achieving 87.5% user satisfaction. These results demonstrate robust transcription accuracy of the challenging Slavic languages and highlight the potential of automated transcription to streamline speech processing using emerging human-computer interface technologies.
Název v anglickém jazyce
AI-Driven Web-Based Speech Transcription Tool: A Novel Approach for Efficient Evaluation of Verbal Memory Performance
Popis výsledku anglicky
Memory deficits are prevalent in epilepsy and other brain disorders, significantly affecting quality of life. In particular, patients with mesial temporal lobe epilepsy require ongoing monitoring and repeated memory assessments to track cognitive function, which could benefit from automated transcription tools. We used a classic free recall (FR) verbal memory task, where participants recalled words, to assess their performance by recording, transcribing, and detecting vocalizations of correctly remembered words. Conventional manual speech transcription methods are time-consuming and prone to human error, especially in a noisy hospital environment. To address these limitations, a modified U-Net architecture was employed for noise reduction, resulting in a signal-to-noise ratio (SNR) of 15.8 and a mean squared error (MSE) of 0.0021. We also developed an automated transcription interface utilizing the Whisper speech recognition model, which was fine-tuned for Polish, Czech, and English languages. Dynamic Time Warping (DTW) was applied to provide precise word-level timestamps of vocalization onset and offset. The interface was iteratively refined over eight development cycles, incorporating feedback from target users. Transcription accuracy was evaluated with Word Error Rates (WER) of 10.3% for Czech, 7.1% for Polish, and 5% for English, alongside respective Character Error Rates (CER) of 12%, 10.8%, and 7.5%. Our automated interface outperformed manual transcription, reducing transcription time fourfold and achieving 87.5% user satisfaction. These results demonstrate robust transcription accuracy of the challenging Slavic languages and highlight the potential of automated transcription to streamline speech processing using emerging human-computer interface technologies.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
30103 - Neurosciences (including psychophysiology)
Návaznosti výsledku
Projekt
<a href="/cs/project/GF22-28594K" target="_blank" >GF22-28594K: Pokročilé algoritmy pro identifikaci elektrofyziologických příznaků kódování a vybavování lidské paměti v intrakraniálním EEG</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
IEEE Journal of Biomedical and Health Informatics
ISSN
2168-2194
e-ISSN
2168-2208
Svazek periodika
29
Číslo periodika v rámci svazku
12
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
8
Strana od-do
8703-8710
Kód UT WoS článku
001640358300014
EID výsledku v databázi Scopus
—