AI-Driven Web-Based Speech Transcription Tool: A Novel Approach for Efficient Evaluation of Verbal Memory Performance
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00159816%3A_____%2F25%3A00082289" target="_blank" >RIV/00159816:_____/25:00082289 - isvavai.cz</a>
Result on the web
<a href="https://ieeexplore.ieee.org/document/11284682" target="_blank" >https://ieeexplore.ieee.org/document/11284682</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/JBHI.2025.3620128" target="_blank" >10.1109/JBHI.2025.3620128</a>
Alternative languages
Result language
angličtina
Original language name
AI-Driven Web-Based Speech Transcription Tool: A Novel Approach for Efficient Evaluation of Verbal Memory Performance
Original language description
Memory deficits are prevalent in epilepsy and other brain disorders, significantly affecting quality of life. In particular, patients with mesial temporal lobe epilepsy require ongoing monitoring and repeated memory assessments to track cognitive function, which could benefit from automated transcription tools. We used a classic free recall (FR) verbal memory task, where participants recalled words, to assess their performance by recording, transcribing, and detecting vocalizations of correctly remembered words. Conventional manual speech transcription methods are time-consuming and prone to human error, especially in a noisy hospital environment. To address these limitations, a modified U-Net architecture was employed for noise reduction, resulting in a signal-to-noise ratio (SNR) of 15.8 and a mean squared error (MSE) of 0.0021. We also developed an automated transcription interface utilizing the Whisper speech recognition model, which was fine-tuned for Polish, Czech, and English languages. Dynamic Time Warping (DTW) was applied to provide precise word-level timestamps of vocalization onset and offset. The interface was iteratively refined over eight development cycles, incorporating feedback from target users. Transcription accuracy was evaluated with Word Error Rates (WER) of 10.3% for Czech, 7.1% for Polish, and 5% for English, alongside respective Character Error Rates (CER) of 12%, 10.8%, and 7.5%. Our automated interface outperformed manual transcription, reducing transcription time fourfold and achieving 87.5% user satisfaction. These results demonstrate robust transcription accuracy of the challenging Slavic languages and highlight the potential of automated transcription to streamline speech processing using emerging human-computer interface technologies.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
30103 - Neurosciences (including psychophysiology)
Result continuities
Project
<a href="/en/project/GF22-28594K" target="_blank" >GF22-28594K: Advanced algorithms for identification of electrophysiological features underlying encoding and recall of human memory in intracranial EEG</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
IEEE Journal of Biomedical and Health Informatics
ISSN
2168-2194
e-ISSN
2168-2208
Volume of the periodical
29
Issue of the periodical within the volume
12
Country of publishing house
US - UNITED STATES
Number of pages
8
Pages from-to
8703-8710
UT code for WoS article
001640358300014
EID of the result in the Scopus database
—