Evaluation of Search Engines and AI Chatbots Using A/B Testing
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216275%3A25410%2F25%3A39922923" target="_blank" >RIV/00216275:25410/25:39922923 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1109/MIPRO65660.2025.11132074" target="_blank" >http://dx.doi.org/10.1109/MIPRO65660.2025.11132074</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/MIPRO65660.2025.11132074" target="_blank" >10.1109/MIPRO65660.2025.11132074</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Evaluation of Search Engines and AI Chatbots Using A/B Testing
Popis výsledku v původním jazyce
AI chatbots are today often used for learning and education without considering the quality of answers they provide. In this paper we evaluated the quality of responses to user queries from popular search engines and AI chatbots. We aimed to objectively determine using the A/B testing if the ChatGPT chatbot is better in answering various user queries than the Google Search. We used our own web interface for evaluating popular search engines and AI chatbots created using node.js and different APIs. In our experiment we devised a set of test queries in the form of factual data questions, mathematical problems and logic-based riddles. We rated the responses from various ChatGPT models and Google Search engine to all the test queries. Afterwards, the A/B testing method is run automatically through our web interface to find out if there is a statistically significant difference between the quality of their answers. We concluded that for our test query set there is no statistically significant difference between the earlier ChatGPT model 3.5 and the Google Search engine. However, we found that the ChatGPT model 4 is better in answering our test queries than the Google Search, and the difference is statistically significant.
Název v anglickém jazyce
Evaluation of Search Engines and AI Chatbots Using A/B Testing
Popis výsledku anglicky
AI chatbots are today often used for learning and education without considering the quality of answers they provide. In this paper we evaluated the quality of responses to user queries from popular search engines and AI chatbots. We aimed to objectively determine using the A/B testing if the ChatGPT chatbot is better in answering various user queries than the Google Search. We used our own web interface for evaluating popular search engines and AI chatbots created using node.js and different APIs. In our experiment we devised a set of test queries in the form of factual data questions, mathematical problems and logic-based riddles. We rated the responses from various ChatGPT models and Google Search engine to all the test queries. Afterwards, the A/B testing method is run automatically through our web interface to find out if there is a statistically significant difference between the quality of their answers. We concluded that for our test query set there is no statistically significant difference between the earlier ChatGPT model 3.5 and the Google Search engine. However, we found that the ChatGPT model 4 is better in answering our test queries than the Google Search, and the difference is statistically significant.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10200 - Computer and information sciences
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
2025 MIPRO 48th ICT and Electronics Convention : proceedings
ISBN
979-8-3315-3596-4
ISSN
2623-8764
e-ISSN
—
Počet stran výsledku
6
Strana od-do
373-378
Název nakladatele
Croatian Society for Information, Communication and Electronic Technology - MIPRO
Místo vydání
Rijeka
Místo konání akce
Opatija
Datum konání akce
2. 6. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—