All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Evaluation of Generative AI Models in Python Code Generation: A Comparative Study

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F62690094%3A18450%2F25%3A50022401" target="_blank" >RIV/62690094:18450/25:50022401 - isvavai.cz</a>

  • Result on the web

    <a href="https://ieeexplore.ieee.org/document/10963975" target="_blank" >https://ieeexplore.ieee.org/document/10963975</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1109/ACCESS.2025.3560244" target="_blank" >10.1109/ACCESS.2025.3560244</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Evaluation of Generative AI Models in Python Code Generation: A Comparative Study

  • Original language description

    This study evaluates leading generative AI models for Python code generation. Evaluation criteria include syntax accuracy, response time, completeness, reliability, and cost. The models tested comprise OpenAI&apos;s GPT series (GPT-4 Turbo, GPT-4o, GPT-4o Mini, GPT-3.5 Turbo), Google&apos;s Gemini (1.0 Pro, 1.5 Flash, 1.5 Pro), Meta&apos;s LLaMA (3.0 8B, 3.1 8B), and Anthropic&apos;s Claude models (3.5 Sonnet, 3 Opus, 3 Sonnet, 3 Haiku). Ten coding tasks of varying complexity were tested across three iterations per model to measure performance and consistency. Claude models, especially Claude 3.5 Sonnet, achieved the highest accuracy and reliability. They outperformed all other models in both simple and complex tasks. Gemini models showed limitations in handling complex code. Cost-effective options like Claude 3 Haiku and Gemini 1.5 Flash were budget-friendly and maintained good accuracy on simpler problems. Unlike earlier single-metric studies, this work introduces a multi-dimensional evaluation framework that considers accuracy, reliability, cost, and exception handling. Future work will explore other programming languages and include metrics such as code optimization and security robustness.

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

    S - Specificky vyzkum na vysokych skolach

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    IEEE Access

  • ISSN

    2169-3536

  • e-ISSN

    2169-3536

  • Volume of the periodical

    13

  • Issue of the periodical within the volume

    April

  • Country of publishing house

    US - UNITED STATES

  • Number of pages

    14

  • Pages from-to

    65334-65347

  • UT code for WoS article

    001470367900023

  • EID of the result in the Scopus database

    2-s2.0-105003297254