Evaluation of Search Engines and AI Chatbots Using A/B Testing
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216275%3A25410%2F25%3A39922923" target="_blank" >RIV/00216275:25410/25:39922923 - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.1109/MIPRO65660.2025.11132074" target="_blank" >http://dx.doi.org/10.1109/MIPRO65660.2025.11132074</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/MIPRO65660.2025.11132074" target="_blank" >10.1109/MIPRO65660.2025.11132074</a>
Alternative languages
Result language
angličtina
Original language name
Evaluation of Search Engines and AI Chatbots Using A/B Testing
Original language description
AI chatbots are today often used for learning and education without considering the quality of answers they provide. In this paper we evaluated the quality of responses to user queries from popular search engines and AI chatbots. We aimed to objectively determine using the A/B testing if the ChatGPT chatbot is better in answering various user queries than the Google Search. We used our own web interface for evaluating popular search engines and AI chatbots created using node.js and different APIs. In our experiment we devised a set of test queries in the form of factual data questions, mathematical problems and logic-based riddles. We rated the responses from various ChatGPT models and Google Search engine to all the test queries. Afterwards, the A/B testing method is run automatically through our web interface to find out if there is a statistically significant difference between the quality of their answers. We concluded that for our test query set there is no statistically significant difference between the earlier ChatGPT model 3.5 and the Google Search engine. However, we found that the ChatGPT model 4 is better in answering our test queries than the Google Search, and the difference is statistically significant.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10200 - Computer and information sciences
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
2025 MIPRO 48th ICT and Electronics Convention : proceedings
ISBN
979-8-3315-3596-4
ISSN
2623-8764
e-ISSN
—
Number of pages
6
Pages from-to
373-378
Publisher name
Croatian Society for Information, Communication and Electronic Technology - MIPRO
Place of publication
Rijeka
Event location
Opatija
Event date
Jun 2, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—