Evaluation of Generative AI Models in Python Code Generation: A Comparative Study
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F62690094%3A18450%2F25%3A50022401" target="_blank" >RIV/62690094:18450/25:50022401 - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/document/10963975" target="_blank" >https://ieeexplore.ieee.org/document/10963975</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ACCESS.2025.3560244" target="_blank" >10.1109/ACCESS.2025.3560244</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Evaluation of Generative AI Models in Python Code Generation: A Comparative Study
Popis výsledku v původním jazyce
This study evaluates leading generative AI models for Python code generation. Evaluation criteria include syntax accuracy, response time, completeness, reliability, and cost. The models tested comprise OpenAI's GPT series (GPT-4 Turbo, GPT-4o, GPT-4o Mini, GPT-3.5 Turbo), Google's Gemini (1.0 Pro, 1.5 Flash, 1.5 Pro), Meta's LLaMA (3.0 8B, 3.1 8B), and Anthropic's Claude models (3.5 Sonnet, 3 Opus, 3 Sonnet, 3 Haiku). Ten coding tasks of varying complexity were tested across three iterations per model to measure performance and consistency. Claude models, especially Claude 3.5 Sonnet, achieved the highest accuracy and reliability. They outperformed all other models in both simple and complex tasks. Gemini models showed limitations in handling complex code. Cost-effective options like Claude 3 Haiku and Gemini 1.5 Flash were budget-friendly and maintained good accuracy on simpler problems. Unlike earlier single-metric studies, this work introduces a multi-dimensional evaluation framework that considers accuracy, reliability, cost, and exception handling. Future work will explore other programming languages and include metrics such as code optimization and security robustness.
Název v anglickém jazyce
Evaluation of Generative AI Models in Python Code Generation: A Comparative Study
Popis výsledku anglicky
This study evaluates leading generative AI models for Python code generation. Evaluation criteria include syntax accuracy, response time, completeness, reliability, and cost. The models tested comprise OpenAI's GPT series (GPT-4 Turbo, GPT-4o, GPT-4o Mini, GPT-3.5 Turbo), Google's Gemini (1.0 Pro, 1.5 Flash, 1.5 Pro), Meta's LLaMA (3.0 8B, 3.1 8B), and Anthropic's Claude models (3.5 Sonnet, 3 Opus, 3 Sonnet, 3 Haiku). Ten coding tasks of varying complexity were tested across three iterations per model to measure performance and consistency. Claude models, especially Claude 3.5 Sonnet, achieved the highest accuracy and reliability. They outperformed all other models in both simple and complex tasks. Gemini models showed limitations in handling complex code. Cost-effective options like Claude 3 Haiku and Gemini 1.5 Flash were budget-friendly and maintained good accuracy on simpler problems. Unlike earlier single-metric studies, this work introduces a multi-dimensional evaluation framework that considers accuracy, reliability, cost, and exception handling. Future work will explore other programming languages and include metrics such as code optimization and security robustness.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
S - Specificky vyzkum na vysokych skolach
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
IEEE Access
ISSN
2169-3536
e-ISSN
2169-3536
Svazek periodika
13
Číslo periodika v rámci svazku
April
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
14
Strana od-do
65334-65347
Kód UT WoS článku
001470367900023
EID výsledku v databázi Scopus
2-s2.0-105003297254