All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

GermDetect: Verb Placement Error Detection Datasets for Learners of Germanic Languages

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A4WGXVVW8" target="_blank" >RIV/00216208:11320/26:4WGXVVW8 - isvavai.cz</a>

  • Result on the web

    <a href="https://aclanthology.org/2025.bea-1.59/" target="_blank" >https://aclanthology.org/2025.bea-1.59/</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.18653/v1/2025.bea-1.59" target="_blank" >10.18653/v1/2025.bea-1.59</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    GermDetect: Verb Placement Error Detection Datasets for Learners of Germanic Languages

  • Original language description

    Correct verb placement is difficult to acquire for second-language (L2) learners of Germanic languages. However, word order errors and, consequently, verb placement errors, are heavily underrepresented in benchmark datasets of NLP tasks such as grammatical error detection (GED)/correction (GEC) and linguistic acceptability assessment (LA). If they are present, they are most often naively introduced, or classification occurs at the sentence level, preventing the precise identification of individual errors and the provision of appropriate feedback to learners. To remedy this, we present GermDetect: Universal Dependencies-based (UD), linguistically informed verb placement error detection datasets for learners of Germanic languages, designed as a token classification task. As our datasets are UD-based, we are able to provide them in most major Germanic languages: Afrikaans, German, Dutch, Faroese, Icelandic, Danish, Norwegian (Bokmål and Nynorsk), and Swedish. We train multilingual BERT (mBERT) models on GermDetect and show that linguistically informed, UD-based error induction results in more effective models for verb placement error detection than models trained on naively introduced errors. Finally, we conduct ablation studies on multilingual training and find that lower-resource languages benefit from the inclusion of structurally related languages in training.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications

  • ISBN

    979-8-89176-270-1

  • ISSN

  • e-ISSN

  • Number of pages

    12

  • Pages from-to

    818-829

  • Publisher name

    Association for Computational Linguistics

  • Place of publication

  • Event location

    Vienna, Austria

  • Event date

    Jan 1, 2026

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article