Natural Language Processing for Dialects of a Language: A Survey
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AKI4ZQ9VX" target="_blank" >RIV/00216208:11320/26:KI4ZQ9VX - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.1145/3712060" target="_blank" >http://dx.doi.org/10.1145/3712060</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1145/3712060" target="_blank" >10.1145/3712060</a>
Alternative languages
Result language
angličtina
Original language name
Natural Language Processing for Dialects of a Language: A Survey
Original language description
State-of-the-art natural language processing (NLP) models are trained on massive training corpora, and report a superlative performance on evaluation datasets. This survey delves into an important attribute of these datasets: the dialect of a language. Motivated by the performance degradation of NLP models for dialectal datasets and its implications for the equity of language technologies, we survey past research in NLP for dialects in terms of datasets, and approaches. We describe a wide range of NLP tasks in terms of two categories: natural language understanding (NLU) (for tasks such as dialect classification, sentiment analysis, parsing, and NLU benchmarks) and natural language generation (NLG) (for summarisation, machine translation, and dialogue systems). The survey is also broad in its coverage of languages which include English, Arabic, German, among others. We observe that past work in NLP concerning dialects goes deeper than mere dialect classification, and extends to several NLU and NLG tasks. For these tasks, we describe classical machine learning using statistical models, along with the recent deep learning-based approaches based on pre-trained language models. We expect that this survey will be useful to NLP researchers interested in building equitable language technologies by rethinking LLM benchmarks and model architectures. © 2025 Copyright held by the owner/author(s)
Czech name
—
Czech description
—
Classification
Type
J<sub>SC</sub> - Article in a specialist periodical, which is included in the SCOPUS database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
ACM Computing Surveys
ISSN
0360-0300
e-ISSN
—
Volume of the periodical
57
Issue of the periodical within the volume
6
Country of publishing house
US - UNITED STATES
Number of pages
37
Pages from-to
1-37
UT code for WoS article
—
EID of the result in the Scopus database
2-s2.0-85219753898