When AI Teaches Geometry Wrong: Systematic Errors in LLM-Generated Explanations and Their Educational Risks
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F44555601%3A13440%2F25%3A43899491" target="_blank" >RIV/44555601:13440/25:43899491 - isvavai.cz</a>
Výsledek na webu
<a href="https://library.iated.org/view/PRIBYL2025WHE" target="_blank" >https://library.iated.org/view/PRIBYL2025WHE</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.21125/iceri.2025.1357" target="_blank" >10.21125/iceri.2025.1357</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
When AI Teaches Geometry Wrong: Systematic Errors in LLM-Generated Explanations and Their Educational Risks
Popis výsledku v původním jazyce
As large language models (LLMs) become increasingly integrated into educational settings, their role in assisting mathematics learning?particularly in geometry?requires critical scrutiny. While LLMs produce fluent and confident responses, these qualities often obscure significant conceptual errors. Geometry, with its reliance on formal definitions, categorical distinctions, and spatial reasoning, offers a robust context in which to evaluate the depth and consistency of AI-generated explanations.This study analyzes the responses of six prominent LLMs (ChatGPT, Claude, Gemini, Mistral Large, Copilot Quick-Nuance, and Copilot Deep-Thinker) to 49 conceptual geometry questions in Czech and 48 in English. The difference in prompt count reflects a translation overlap. In total, 582 responses were examined for correctness, clarity, and internal consistency. To ensure fair comparison within each language, the models were queried in their respective native environments; in particular, English prompts were submitted under simulated English-speaking conditions, including adjustments to system locale and IP address.The results reveal substantial limitations in the models' geometric reasoning. Conceptual errors were inductively categorized into types, including misinterpretation of definitions (e.g., confusing inscribed and circumscribed circles), incorrect geometric properties (e.g., assigning axis symmetry to general parallelograms), and flawed logical inferences (e.g., assuming that all shapes with equal angles must have equal sides). Several models also demonstrated a tendency to "hallucinate" mathematical terminology (e.g., inventing the term ?divergent triangle?) or to provide oversimplified or misleading generalizations. Notably, all six models showed confusion between the terms circle and disk, which led to systematic misclassification of intersection scenarios.Quantitative analysis focused on error frequency per response, error type distribution across models and languages, and variation in model performance. While some models performed marginally better than others, the differences were inconsistent and topic-dependent. The findings emphasize that despite impressive linguistic fluency, current LLMs lack a stable and conceptually grounded understanding of elementary geometry. This has serious implications for their use in education, especially in languages other than English and in tasks requiring formal precision.This paper argues that educational use of LLMs in mathematics should be approached with caution and paired with pedagogical strategies that foster critical evaluation of AI-generated content. The study also underscores the need for benchmarks in mathematical literacy that go beyond symbolic computation and address conceptual depth, especially in multilingual settings.
Název v anglickém jazyce
When AI Teaches Geometry Wrong: Systematic Errors in LLM-Generated Explanations and Their Educational Risks
Popis výsledku anglicky
As large language models (LLMs) become increasingly integrated into educational settings, their role in assisting mathematics learning?particularly in geometry?requires critical scrutiny. While LLMs produce fluent and confident responses, these qualities often obscure significant conceptual errors. Geometry, with its reliance on formal definitions, categorical distinctions, and spatial reasoning, offers a robust context in which to evaluate the depth and consistency of AI-generated explanations.This study analyzes the responses of six prominent LLMs (ChatGPT, Claude, Gemini, Mistral Large, Copilot Quick-Nuance, and Copilot Deep-Thinker) to 49 conceptual geometry questions in Czech and 48 in English. The difference in prompt count reflects a translation overlap. In total, 582 responses were examined for correctness, clarity, and internal consistency. To ensure fair comparison within each language, the models were queried in their respective native environments; in particular, English prompts were submitted under simulated English-speaking conditions, including adjustments to system locale and IP address.The results reveal substantial limitations in the models' geometric reasoning. Conceptual errors were inductively categorized into types, including misinterpretation of definitions (e.g., confusing inscribed and circumscribed circles), incorrect geometric properties (e.g., assigning axis symmetry to general parallelograms), and flawed logical inferences (e.g., assuming that all shapes with equal angles must have equal sides). Several models also demonstrated a tendency to "hallucinate" mathematical terminology (e.g., inventing the term ?divergent triangle?) or to provide oversimplified or misleading generalizations. Notably, all six models showed confusion between the terms circle and disk, which led to systematic misclassification of intersection scenarios.Quantitative analysis focused on error frequency per response, error type distribution across models and languages, and variation in model performance. While some models performed marginally better than others, the differences were inconsistent and topic-dependent. The findings emphasize that despite impressive linguistic fluency, current LLMs lack a stable and conceptually grounded understanding of elementary geometry. This has serious implications for their use in education, especially in languages other than English and in tasks requiring formal precision.This paper argues that educational use of LLMs in mathematics should be approached with caution and paired with pedagogical strategies that foster critical evaluation of AI-generated content. The study also underscores the need for benchmarks in mathematical literacy that go beyond symbolic computation and address conceptual depth, especially in multilingual settings.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
50301 - Education, general; including training, pedagogy, didactics [and education systems]
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
ICERI2025 Proceedings
ISBN
978-84-09-78706-7
ISSN
2340-1095
e-ISSN
2340-1095
Počet stran výsledku
10
Strana od-do
4760-4769
Název nakladatele
IATED
Místo vydání
Sevilla, Spain
Místo konání akce
Sevilla, Spain
Datum konání akce
10. 11. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—