All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00064203%3A_____%2F25%3A10500720" target="_blank" >RIV/00064203:_____/25:10500720 - isvavai.cz</a>

  • Alternative codes found

    RIV/00216208:11130/25:10500720

  • Result on the web

    <a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=CSQkTTEmVK" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=CSQkTTEmVK</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1186/s41747-025-00591-0" target="_blank" >10.1186/s41747-025-00591-0</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency

  • Original language description

    BACKGROUND: We compared the performance, confidence, and response consistency of five chatbots powered by large language models in solving European Diploma in Radiology (EDiR) text-based multiple-response questions. METHODS: ChatGPT-4o, ChatGPT-4o-mini, Copilot, Gemini, and Claude 3.5 Sonnet were tested using 52 text-based multiple-response questions from two previous EDiR sessions in two iterations. Chatbots were prompted to evaluate each answer as correct or incorrect and grade its confidence level on a scale of 0 (not confident at all) to 10 (most confident). Scores per question were calculated using a weighted formula that accounted for correct and incorrect answers (range 0.0-1.0). RESULTS: Claude 3.5 Sonnet achieved the highest score per question (0.84 +- 0.26, mean +- standard deviation) compared to ChatGPT-4o (0.76 +- 0.31), ChatGPT-4o-mini (0.64 +- 0.35), Copilot (0.62 +- 0.37), and Gemini (0.54 +- 0.39) (p &lt; 0.001). A self-reported confidence in answering the questions was 9.0 +- 0.9 for Claude 3.5 Sonnet followed by ChatGPT-4o (8.7 +- 1.1), compared to ChatGPT-4o-mini (8.2 +- 1.3), Copilot (8.2 +- 2.2), and Gemini (8.2 +- 1.6, p &lt; 0.001). Claude 3.5 Sonnet demonstrated superior consistency, changing responses in 5.4% of cases between the two iterations, compared to ChatGPT-4o (6.5%), ChatGPT-4o-mini (8.8%), Copilot (13.8%), and Gemini (18.5%). All chatbots outperformed human candidates from previous EDiR sessions, achieving a passing grade from this part of the examination. CONCLUSION: Claude 3.5 Sonnet exhibited superior accuracy, confidence, and consistency, with ChatGPT-4o performing nearly as well. The variation in performance among the evaluated models was substantial. RELEVANCE STATEMENT: Variation in performance, consistency, and confidence among chatbots in solving EDiR test-based questions highlights the need for cautious deployment, particularly in high-stakes clinical and educational settings. KEY POINTS: Claude 3.5 Sonnet outperformed other chatbots in accuracy and response consistency. ChatGPT-4o ranked second, showing strong but slightly less reliable performance. All chatbots surpassed EDiR candidates in text-based EDiR questions.

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database

  • CEP classification

  • OECD FORD branch

    30224 - Radiology, nuclear medicine and medical imaging

Result continuities

  • Project

  • Continuities

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    European Radiology Experimental

  • ISSN

    2509-9280

  • e-ISSN

    2509-9280

  • Volume of the periodical

    9

  • Issue of the periodical within the volume

    1

  • Country of publishing house

    CH - SWITZERLAND

  • Number of pages

    8

  • Pages from-to

    79

  • UT code for WoS article

    001553684100001

  • EID of the result in the Scopus database

    2-s2.0-105013688714