Article

No Article found
CROSSREF PEER REVIEWED GOOGLE SCHOLAR PLAGIARISM CHECK
ISSN License

Evaluating the Accuracy and Limitations of AI Chatbots in Daily Tasks

Author: Dr. B. Anuja Beatrice and B. Perarasu

Published On: 2026-07-07

DOI: https://doi.org/10.70127/irjedt.vol.9.issue07.188

Abstract

Artificial intelligence (AI) chatbots such as ChatGPT, Gemini, and Copilot have moved from experimental novelties to everyday tools used for writing, scheduling, coding, tutoring, and information retrieval. Despite their rapid adoption, systematic understanding of how accurate and dependable these systems are across the range of tasks that ordinary users perform daily remains limited. Here we build and apply an evaluation framework that examines the accuracy, consistency, and failure modes of three widely used AI chatbots across six everyday task categories: factual question answering, mathematical reasoning, code generation, creative writing, scheduling and planning, and real-time information retrieval. Using a structured trial protocol of 600 prompts (100 per category) evaluated by human raters against verified ground truth, we quantify accuracy, hallucination rate, response relevance, and consistency under repeated querying. Results show that chatbots perform strongly (85–93%) on creative and general-knowledge tasks but degrade substantially (55–61%) on tasks requiring current, time-sensitive information, and that accuracy falls sharply as conversational complexity and dialogue length increase. We further identify a measurable gap between users' perceived trust in chatbot answers and the systems' actual measured accuracy, particularly in health, financial, and legal domains. We close by tracing these limitations to their architectural and training-data origins and by proposing concrete steps toward safer, more transparent everyday use of conversational AI.

Keywords— artificial intelligence, chatbots, conversational AI, large language models, accuracy evaluation, hallucination, natural language processing, human-AI interaction, reliability

Keywords
— artificial intelligence chatbots conversational AI large language models accuracy evaluation hallucination natural language processing human-AI interaction reliability
Article Information
Volume

9

Year

2026

Review Rounds

1

Article Type

Research Article

Indexed In
Publish your academic thesis as a book with ISBN Contact – connectirj@gmail.com
Get In Touch

2/11, SASTRI NAGAR, KOYEMBEDU, CHENNAI-600107

9488577176

editor@irjweb.com, connectirj@gmail.com

Follow Us

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License

Copyrights © IRJEdT. All Rights Reserved.

Visiters Count :