All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Evaluation of Search Engines and AI Chatbots Using A/B Testing

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216275%3A25410%2F25%3A39922923" target="_blank" >RIV/00216275:25410/25:39922923 - isvavai.cz</a>

  • Result on the web

    <a href="http://dx.doi.org/10.1109/MIPRO65660.2025.11132074" target="_blank" >http://dx.doi.org/10.1109/MIPRO65660.2025.11132074</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1109/MIPRO65660.2025.11132074" target="_blank" >10.1109/MIPRO65660.2025.11132074</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Evaluation of Search Engines and AI Chatbots Using A/B Testing

  • Original language description

    AI chatbots are today often used for learning and education without considering the quality of answers they provide. In this paper we evaluated the quality of responses to user queries from popular search engines and AI chatbots. We aimed to objectively determine using the A/B testing if the ChatGPT chatbot is better in answering various user queries than the Google Search. We used our own web interface for evaluating popular search engines and AI chatbots created using node.js and different APIs. In our experiment we devised a set of test queries in the form of factual data questions, mathematical problems and logic-based riddles. We rated the responses from various ChatGPT models and Google Search engine to all the test queries. Afterwards, the A/B testing method is run automatically through our web interface to find out if there is a statistically significant difference between the quality of their answers. We concluded that for our test query set there is no statistically significant difference between the earlier ChatGPT model 3.5 and the Google Search engine. However, we found that the ChatGPT model 4 is better in answering our test queries than the Google Search, and the difference is statistically significant.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10200 - Computer and information sciences

Result continuities

  • Project

  • Continuities

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    2025 MIPRO 48th ICT and Electronics Convention : proceedings

  • ISBN

    979-8-3315-3596-4

  • ISSN

    2623-8764

  • e-ISSN

  • Number of pages

    6

  • Pages from-to

    373-378

  • Publisher name

    Croatian Society for Information, Communication and Electronic Technology - MIPRO

  • Place of publication

    Rijeka

  • Event location

    Opatija

  • Event date

    Jun 2, 2025

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article