Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00064203%3A_____%2F25%3A10500720" target="_blank" >RIV/00064203:_____/25:10500720 - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/00216208:11130/25:10500720
Výsledek na webu
<a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=CSQkTTEmVK" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=CSQkTTEmVK</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1186/s41747-025-00591-0" target="_blank" >10.1186/s41747-025-00591-0</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency
Popis výsledku v původním jazyce
BACKGROUND: We compared the performance, confidence, and response consistency of five chatbots powered by large language models in solving European Diploma in Radiology (EDiR) text-based multiple-response questions. METHODS: ChatGPT-4o, ChatGPT-4o-mini, Copilot, Gemini, and Claude 3.5 Sonnet were tested using 52 text-based multiple-response questions from two previous EDiR sessions in two iterations. Chatbots were prompted to evaluate each answer as correct or incorrect and grade its confidence level on a scale of 0 (not confident at all) to 10 (most confident). Scores per question were calculated using a weighted formula that accounted for correct and incorrect answers (range 0.0-1.0). RESULTS: Claude 3.5 Sonnet achieved the highest score per question (0.84 +- 0.26, mean +- standard deviation) compared to ChatGPT-4o (0.76 +- 0.31), ChatGPT-4o-mini (0.64 +- 0.35), Copilot (0.62 +- 0.37), and Gemini (0.54 +- 0.39) (p < 0.001). A self-reported confidence in answering the questions was 9.0 +- 0.9 for Claude 3.5 Sonnet followed by ChatGPT-4o (8.7 +- 1.1), compared to ChatGPT-4o-mini (8.2 +- 1.3), Copilot (8.2 +- 2.2), and Gemini (8.2 +- 1.6, p < 0.001). Claude 3.5 Sonnet demonstrated superior consistency, changing responses in 5.4% of cases between the two iterations, compared to ChatGPT-4o (6.5%), ChatGPT-4o-mini (8.8%), Copilot (13.8%), and Gemini (18.5%). All chatbots outperformed human candidates from previous EDiR sessions, achieving a passing grade from this part of the examination. CONCLUSION: Claude 3.5 Sonnet exhibited superior accuracy, confidence, and consistency, with ChatGPT-4o performing nearly as well. The variation in performance among the evaluated models was substantial. RELEVANCE STATEMENT: Variation in performance, consistency, and confidence among chatbots in solving EDiR test-based questions highlights the need for cautious deployment, particularly in high-stakes clinical and educational settings. KEY POINTS: Claude 3.5 Sonnet outperformed other chatbots in accuracy and response consistency. ChatGPT-4o ranked second, showing strong but slightly less reliable performance. All chatbots surpassed EDiR candidates in text-based EDiR questions.
Název v anglickém jazyce
Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance and consistency
Popis výsledku anglicky
BACKGROUND: We compared the performance, confidence, and response consistency of five chatbots powered by large language models in solving European Diploma in Radiology (EDiR) text-based multiple-response questions. METHODS: ChatGPT-4o, ChatGPT-4o-mini, Copilot, Gemini, and Claude 3.5 Sonnet were tested using 52 text-based multiple-response questions from two previous EDiR sessions in two iterations. Chatbots were prompted to evaluate each answer as correct or incorrect and grade its confidence level on a scale of 0 (not confident at all) to 10 (most confident). Scores per question were calculated using a weighted formula that accounted for correct and incorrect answers (range 0.0-1.0). RESULTS: Claude 3.5 Sonnet achieved the highest score per question (0.84 +- 0.26, mean +- standard deviation) compared to ChatGPT-4o (0.76 +- 0.31), ChatGPT-4o-mini (0.64 +- 0.35), Copilot (0.62 +- 0.37), and Gemini (0.54 +- 0.39) (p < 0.001). A self-reported confidence in answering the questions was 9.0 +- 0.9 for Claude 3.5 Sonnet followed by ChatGPT-4o (8.7 +- 1.1), compared to ChatGPT-4o-mini (8.2 +- 1.3), Copilot (8.2 +- 2.2), and Gemini (8.2 +- 1.6, p < 0.001). Claude 3.5 Sonnet demonstrated superior consistency, changing responses in 5.4% of cases between the two iterations, compared to ChatGPT-4o (6.5%), ChatGPT-4o-mini (8.8%), Copilot (13.8%), and Gemini (18.5%). All chatbots outperformed human candidates from previous EDiR sessions, achieving a passing grade from this part of the examination. CONCLUSION: Claude 3.5 Sonnet exhibited superior accuracy, confidence, and consistency, with ChatGPT-4o performing nearly as well. The variation in performance among the evaluated models was substantial. RELEVANCE STATEMENT: Variation in performance, consistency, and confidence among chatbots in solving EDiR test-based questions highlights the need for cautious deployment, particularly in high-stakes clinical and educational settings. KEY POINTS: Claude 3.5 Sonnet outperformed other chatbots in accuracy and response consistency. ChatGPT-4o ranked second, showing strong but slightly less reliable performance. All chatbots surpassed EDiR candidates in text-based EDiR questions.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
30224 - Radiology, nuclear medicine and medical imaging
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
European Radiology Experimental
ISSN
2509-9280
e-ISSN
2509-9280
Svazek periodika
9
Číslo periodika v rámci svazku
1
Stát vydavatele periodika
CH - Švýcarská konfederace
Počet stran výsledku
8
Strana od-do
79
Kód UT WoS článku
001553684100001
EID výsledku v databázi Scopus
2-s2.0-105013688714