Automated detection of machine translation use in L2 Spanish writing
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ACG2UNPRF" target="_blank" >RIV/00216208:11320/26:CG2UNPRF - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1177/13621688251352263" target="_blank" >http://dx.doi.org/10.1177/13621688251352263</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1177/13621688251352263" target="_blank" >10.1177/13621688251352263</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Automated detection of machine translation use in L2 Spanish writing
Popis výsledku v původním jazyce
Google Translate (GT) has become a popular machine translation (MT) tool among language learners, received by instructors with excitement over its pedagogical potential and concerns about its possible misuse in the classroom, particularly when this misuse goes undetected. This study investigated the suitability of natural language processing (NLP) software for the automated detection of MT use in second language (L2) writing, examining a dataset composed of written samples generated by GT and direct L2 writing produced by intermediate-level postsecondary learners of Spanish. NLP-powered analyses found significant lexical and sentential-level differences, as well as estimated proficiency-level differences across text types. Automated judgments based on lexical diversity and amount of coordination yielded detection accuracy rates of 73.08% each, whereas proficiency estimates informed correct automated judgments with an overall accuracy rate of 86.54%. An automated reverse-translation protocol using probability estimates was capable of differentiating between direct L2 writing and MT-assisted texts 98% of the time, far surpassing human detection rates (73%) found in a previous study for the same dataset. These findings argue strongly for the potential of NLP-driven textual analysis as a reliable tool to assist instructors in detecting unauthorized uses of MT in L2 writing. © The Author(s) 2025
Název v anglickém jazyce
Automated detection of machine translation use in L2 Spanish writing
Popis výsledku anglicky
Google Translate (GT) has become a popular machine translation (MT) tool among language learners, received by instructors with excitement over its pedagogical potential and concerns about its possible misuse in the classroom, particularly when this misuse goes undetected. This study investigated the suitability of natural language processing (NLP) software for the automated detection of MT use in second language (L2) writing, examining a dataset composed of written samples generated by GT and direct L2 writing produced by intermediate-level postsecondary learners of Spanish. NLP-powered analyses found significant lexical and sentential-level differences, as well as estimated proficiency-level differences across text types. Automated judgments based on lexical diversity and amount of coordination yielded detection accuracy rates of 73.08% each, whereas proficiency estimates informed correct automated judgments with an overall accuracy rate of 86.54%. An automated reverse-translation protocol using probability estimates was capable of differentiating between direct L2 writing and MT-assisted texts 98% of the time, far surpassing human detection rates (73%) found in a previous study for the same dataset. These findings argue strongly for the potential of NLP-driven textual analysis as a reliable tool to assist instructors in detecting unauthorized uses of MT in L2 writing. © The Author(s) 2025
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Language Teaching Research
ISSN
1362-1688
e-ISSN
—
Svazek periodika
2025
Číslo periodika v rámci svazku
2025
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
26
Strana od-do
13621688251352263
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105014592650