A Rule-Based Parser in Comparison with Statistical Neuronal Approaches in Terms of Grammar Competence
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AJ4MGKWD7" target="_blank" >RIV/00216208:11320/26:J4MGKWD7 - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.3390/app15010087" target="_blank" >http://dx.doi.org/10.3390/app15010087</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.3390/app15010087" target="_blank" >10.3390/app15010087</a>
Alternative languages
Result language
angličtina
Original language name
A Rule-Based Parser in Comparison with Statistical Neuronal Approaches in Terms of Grammar Competence
Original language description
The “Easy Language” standard was created to help individuals with cognitive disabilities understand texts more easily. Typically, text simplification is performed by language experts and is available for limited materials. We introduce a new software tool designed to analyze and simplify any text according to the “Easy Language” rules. This tool uses a rule-based system, conducting a full grammatical analysis of each sentence and then simplifying it into a grammatically correct form. Unlike neuronal approaches, which are based on statistics and are very popular today, our rule-based approach explicitly addresses language ambiguities by examining all possible interpretations and eliminating the incorrect ones. The purpose of the present study is to compare the performance of our rule-base parser with two state-of-the-art statistical parsers, one based on dependencies between words (SpaCy parser) and the other based on linguistic constituents (Stanford parser). Although large language models (LLMs), which are the technical basis of the software ChatGPT, were not designed specifically for grammatical parsing, because of their popularity, users, especially language learners, often ask them grammatical questions as well. Therefore, we use LLMs as supplementary models for comparison. LMMs produce grammatically correct text on any topic; however, their grammar knowledge is implicit within the trained weights. To evaluate how well state-of-the-art methods can perform a grammatical analysis, we parse ten sentences with our tool, the statistical parsers from SpaCy and Stanford, and ask two LLMs equivalent grammar questions. The results show that our rule-based method provides a more informative and reliable grammatical analysis compared to these two parsers and outperforms LLMs in that specific task. © 2024 by the authors.
Czech name
—
Czech description
—
Classification
Type
J<sub>SC</sub> - Article in a specialist periodical, which is included in the SCOPUS database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Applied Sciences (Switzerland)
ISSN
2076-3417
e-ISSN
—
Volume of the periodical
15
Issue of the periodical within the volume
1
Country of publishing house
US - UNITED STATES
Number of pages
37
Pages from-to
1-37
UT code for WoS article
—
EID of the result in the Scopus database
2-s2.0-85214529812