A Multi-CEFR-Level Learner Corpus Study to Quantify Fluency and Accuracy in Speech
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10507126" target="_blank" >RIV/00216208:11210/25:10507126 - isvavai.cz</a>
Výsledek na webu
<a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=3iV9Xkw8RK" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=3iV9Xkw8RK</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1177/00238309251393170" target="_blank" >10.1177/00238309251393170</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
A Multi-CEFR-Level Learner Corpus Study to Quantify Fluency and Accuracy in Speech
Popis výsledku v původním jazyce
This study strengthens the validation of learner speech assessment in the Common European Framework of Reference (CEFR) by analyzing the quantitative variables related to fluency and accuracy across four CEFR levels (A2, B1, B2, and C1). Drawing on a learner corpus approach, we examine 500,000 tokens from the Louvain International Database of Spoken English Interlanguage (LINDSEI) and its extensions, supplemented by post hoc rater evaluations. Three task types-a semi-monologic topic discussion, a dialogic interaction, and a monologic picture description-are used to elicit variation in speech production. The analysis focuses on speech rates, the frequency of filled and unfilled pauses, and error rates to unveil developmental trends in learner speech. The results reveal strong correlations between these fluency and accuracy metrics and CEFR levels, with speech rate emerging as the most reliable indicator of proficiency. The frequency of unfilled pauses decreases as proficiency increases, while filled pauses, although less critical to fluency assessment, offer insights into speech planning mechanisms. Error rates similarly decline with higher proficiency, reflecting greater accuracy in speech production. Exemplary instances for each CEFR level are presented, offering practical metrics for teaching, assessment, and rater training. While the study's limitations include an overrepresentation of Mandarin Chinese learners and the exclusion of pronunciation errors, these gaps highlight avenues for future research. This study provides empirical, task-sensitive evidence to enrich CEFR can-do descriptors, enhance rater training, and refine speaking assessments, contributing to more effective language teaching, learning, and assessment practices.
Název v anglickém jazyce
A Multi-CEFR-Level Learner Corpus Study to Quantify Fluency and Accuracy in Speech
Popis výsledku anglicky
This study strengthens the validation of learner speech assessment in the Common European Framework of Reference (CEFR) by analyzing the quantitative variables related to fluency and accuracy across four CEFR levels (A2, B1, B2, and C1). Drawing on a learner corpus approach, we examine 500,000 tokens from the Louvain International Database of Spoken English Interlanguage (LINDSEI) and its extensions, supplemented by post hoc rater evaluations. Three task types-a semi-monologic topic discussion, a dialogic interaction, and a monologic picture description-are used to elicit variation in speech production. The analysis focuses on speech rates, the frequency of filled and unfilled pauses, and error rates to unveil developmental trends in learner speech. The results reveal strong correlations between these fluency and accuracy metrics and CEFR levels, with speech rate emerging as the most reliable indicator of proficiency. The frequency of unfilled pauses decreases as proficiency increases, while filled pauses, although less critical to fluency assessment, offer insights into speech planning mechanisms. Error rates similarly decline with higher proficiency, reflecting greater accuracy in speech production. Exemplary instances for each CEFR level are presented, offering practical metrics for teaching, assessment, and rater training. While the study's limitations include an overrepresentation of Mandarin Chinese learners and the exclusion of pronunciation errors, these gaps highlight avenues for future research. This study provides empirical, task-sensitive evidence to enrich CEFR can-do descriptors, enhance rater training, and refine speaking assessments, contributing to more effective language teaching, learning, and assessment practices.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
50301 - Education, general; including training, pedagogy, didactics [and education systems]
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Language and Speech
ISSN
0023-8309
e-ISSN
1756-6053
Svazek periodika
Neuveden
Číslo periodika v rámci svazku
December 26, 2025
Stát vydavatele periodika
GB - Spojené království Velké Británie a Severního Irska
Počet stran výsledku
34
Strana od-do
1-34
Kód UT WoS článku
001648904900001
EID výsledku v databázi Scopus
—