All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Towards better language representation in Natural Language Processing A multilingual dataset for text-level Grammatical Error Correction

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10513028" target="_blank" >RIV/00216208:11210/25:10513028 - isvavai.cz</a>

  • Alternative codes found

    RIV/00216208:11320/26:SHY6B8WN

  • Result on the web

    <a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=uVkuQkp2Ul" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=uVkuQkp2Ul</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1075/ijlcr.24033.mas" target="_blank" >10.1075/ijlcr.24033.mas</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Towards better language representation in Natural Language Processing A multilingual dataset for text-level Grammatical Error Correction

  • Original language description

    This paper introduces MultiGEC, a dataset for multilingual Grammatical Error Correction (GEC) in twelve European languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian. MultiGEC distinguishes itself from previous G E C datasets in that it covers several underrepresented languages, which we argue should be included in resources used to train models for Natural Language Processing tasks which, as G E C itself, have implications for Learner Corpus Research and Second Language Acquisition. Aside from multilingualism, the novelty of the MultiGEC dataset is that it consists of full texts - typically learner essays - rather than individual sentences, making it possible to train systems that take a broader context into account. The dataset was built for MultiGEC-2025, the first shared task in multilingual text-level GEC, but it remains accessible after its competitive phase, serving as a resource to train new error correction systems and perform cross-lingual G E C studies.

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database

  • CEP classification

  • OECD FORD branch

    60203 - Linguistics

Result continuities

  • Project

    <a href="/en/project/LM2023044" target="_blank" >LM2023044: Czech National Corpus</a><br>

  • Continuities

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    International Journal of Learner Corpus Research

  • ISSN

    2215-1478

  • e-ISSN

    2215-1486

  • Volume of the periodical

    11

  • Issue of the periodical within the volume

    2

  • Country of publishing house

    NL - THE KINGDOM OF THE NETHERLANDS

  • Number of pages

    27

  • Pages from-to

    309-335

  • UT code for WoS article

    001457603500001

  • EID of the result in the Scopus database

    2-s2.0-105003035015