Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00064203%3A_____%2F25%3A10500720" target="_blank" >RIV/00064203:_____/25:10500720 - isvavai.cz</a>

  • Nalezeny alternativní kódy

    RIV/00216208:11130/25:10500720

  • Výsledek na webu

    <a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=CSQkTTEmVK" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=CSQkTTEmVK</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1186/s41747-025-00591-0" target="_blank" >10.1186/s41747-025-00591-0</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency

  • Popis výsledku v původním jazyce

    BACKGROUND: We compared the performance, confidence, and response consistency of five chatbots powered by large language models in solving European Diploma in Radiology (EDiR) text-based multiple-response questions. METHODS: ChatGPT-4o, ChatGPT-4o-mini, Copilot, Gemini, and Claude 3.5 Sonnet were tested using 52 text-based multiple-response questions from two previous EDiR sessions in two iterations. Chatbots were prompted to evaluate each answer as correct or incorrect and grade its confidence level on a scale of 0 (not confident at all) to 10 (most confident). Scores per question were calculated using a weighted formula that accounted for correct and incorrect answers (range 0.0-1.0). RESULTS: Claude 3.5 Sonnet achieved the highest score per question (0.84 +- 0.26, mean +- standard deviation) compared to ChatGPT-4o (0.76 +- 0.31), ChatGPT-4o-mini (0.64 +- 0.35), Copilot (0.62 +- 0.37), and Gemini (0.54 +- 0.39) (p &lt; 0.001). A self-reported confidence in answering the questions was 9.0 +- 0.9 for Claude 3.5 Sonnet followed by ChatGPT-4o (8.7 +- 1.1), compared to ChatGPT-4o-mini (8.2 +- 1.3), Copilot (8.2 +- 2.2), and Gemini (8.2 +- 1.6, p &lt; 0.001). Claude 3.5 Sonnet demonstrated superior consistency, changing responses in 5.4% of cases between the two iterations, compared to ChatGPT-4o (6.5%), ChatGPT-4o-mini (8.8%), Copilot (13.8%), and Gemini (18.5%). All chatbots outperformed human candidates from previous EDiR sessions, achieving a passing grade from this part of the examination. CONCLUSION: Claude 3.5 Sonnet exhibited superior accuracy, confidence, and consistency, with ChatGPT-4o performing nearly as well. The variation in performance among the evaluated models was substantial. RELEVANCE STATEMENT: Variation in performance, consistency, and confidence among chatbots in solving EDiR test-based questions highlights the need for cautious deployment, particularly in high-stakes clinical and educational settings. KEY POINTS: Claude 3.5 Sonnet outperformed other chatbots in accuracy and response consistency. ChatGPT-4o ranked second, showing strong but slightly less reliable performance. All chatbots surpassed EDiR candidates in text-based EDiR questions.

  • Název v anglickém jazyce

    Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency

  • Popis výsledku anglicky

    BACKGROUND: We compared the performance, confidence, and response consistency of five chatbots powered by large language models in solving European Diploma in Radiology (EDiR) text-based multiple-response questions. METHODS: ChatGPT-4o, ChatGPT-4o-mini, Copilot, Gemini, and Claude 3.5 Sonnet were tested using 52 text-based multiple-response questions from two previous EDiR sessions in two iterations. Chatbots were prompted to evaluate each answer as correct or incorrect and grade its confidence level on a scale of 0 (not confident at all) to 10 (most confident). Scores per question were calculated using a weighted formula that accounted for correct and incorrect answers (range 0.0-1.0). RESULTS: Claude 3.5 Sonnet achieved the highest score per question (0.84 +- 0.26, mean +- standard deviation) compared to ChatGPT-4o (0.76 +- 0.31), ChatGPT-4o-mini (0.64 +- 0.35), Copilot (0.62 +- 0.37), and Gemini (0.54 +- 0.39) (p &lt; 0.001). A self-reported confidence in answering the questions was 9.0 +- 0.9 for Claude 3.5 Sonnet followed by ChatGPT-4o (8.7 +- 1.1), compared to ChatGPT-4o-mini (8.2 +- 1.3), Copilot (8.2 +- 2.2), and Gemini (8.2 +- 1.6, p &lt; 0.001). Claude 3.5 Sonnet demonstrated superior consistency, changing responses in 5.4% of cases between the two iterations, compared to ChatGPT-4o (6.5%), ChatGPT-4o-mini (8.8%), Copilot (13.8%), and Gemini (18.5%). All chatbots outperformed human candidates from previous EDiR sessions, achieving a passing grade from this part of the examination. CONCLUSION: Claude 3.5 Sonnet exhibited superior accuracy, confidence, and consistency, with ChatGPT-4o performing nearly as well. The variation in performance among the evaluated models was substantial. RELEVANCE STATEMENT: Variation in performance, consistency, and confidence among chatbots in solving EDiR test-based questions highlights the need for cautious deployment, particularly in high-stakes clinical and educational settings. KEY POINTS: Claude 3.5 Sonnet outperformed other chatbots in accuracy and response consistency. ChatGPT-4o ranked second, showing strong but slightly less reliable performance. All chatbots surpassed EDiR candidates in text-based EDiR questions.

Klasifikace

  • Druh

    J<sub>imp</sub> - Článek v periodiku v databázi Web of Science

  • CEP obor

  • OECD FORD obor

    30224 - Radiology, nuclear medicine and medical imaging

Návaznosti výsledku

  • Projekt

  • Návaznosti

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    European Radiology Experimental

  • ISSN

    2509-9280

  • e-ISSN

    2509-9280

  • Svazek periodika

    9

  • Číslo periodika v rámci svazku

    1

  • Stát vydavatele periodika

    CH - Švýcarská konfederace

  • Počet stran výsledku

    8

  • Strana od-do

    79

  • Kód UT WoS článku

    001553684100001

  • EID výsledku v databázi Scopus

    2-s2.0-105013688714