Misinformation propagation in AI chatbots: implications for scientific inquiry and biochemistry education

This study aims to assess the scientific accuracy of responses (outputs) from various Artificial Intelligence chatbots to biochemistry questions, with a particular focus on biochemical molecules. In this case study, five biochemistry questions were developed based on foundational disciplinary textbooks, with particular attention to concepts frequently associated with misconceptions or scientifically inaccurate explanations. Three widely used and publicly accessible AI chatbots – ChatGPT-4o, Gemini, and Co-pilot—were selected for evaluation. Each chatbot was presented with the same five questions, structured in a three-tier format to assess not only the accuracy of the selected answers but also the scientific validity of the accompanying explanations. The accuracy of their answers was assessed using a rubric designed to assess scientific rigour. The findings revealed differences in the scientific accuracy of the responses generated by the three chatbots. ChatGPT-4o demonstrated the highest overall accuracy, whereas Co-pilot and Gemini showed only a slight difference in their accuracy. Nevertheless, some chatbot responses contained scientifically inaccurate or misconception-consistent explanations. Although the findings are limited to the questions and chatbot versions examined, they highlight the importance of critically evaluating AI-generated content before using it in biochemistry education. Uncritical educational use of such responses may reinforce and disseminate scientific misconceptions.

Authors

Institutions

Publication Details

Journal
Journal of Biological Education
Published
2026-09-29
DOI
https://doi.org/10.1080/00219266.2026.2737444
Primary Topic
AI in Service Interactions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Misinformation propagation in AI chatbots: implications for scientific inquiry and biochemistry education

Rıdvan Elmas, Merve Adıgüzel Ulutaş, Mehmet Yalçın YILMAZ
Journal of Biological Education
AI in Service Interactions
article

Misinformation propagation in AI chatbots: implications for scientific inquiry and biochemistry education

Rıdvan Elmas, Merve Adıgüzel Ulutaş, Mehmet Yalçın YILMAZ
article en

Abstract

This study aims to assess the scientific accuracy of responses (outputs) from various Artificial Intelligence chatbots to biochemistry questions, with a particular focus on biochemical molecules. In this case study, five biochemistry questions were developed based on foundational disciplinary textbooks, with particular attention to concepts frequently associated with misconceptions or scientifically inaccurate explanations. Three widely used and publicly accessible AI chatbots – ChatGPT-4o, Gemini, and Co-pilot—were selected for evaluation. Each chatbot was presented with the same five questions, structured in a three-tier format to assess not only the accuracy of the selected answers but also the scientific validity of the accompanying explanations. The accuracy of their answers was assessed using a rubric designed to assess scientific rigour. The findings revealed differences in the scientific accuracy of the responses generated by the three chatbots. ChatGPT-4o demonstrated the highest overall accuracy, whereas Co-pilot and Gemini showed only a slight difference in their accuracy. Nevertheless, some chatbot responses contained scientifically inaccurate or misconception-consistent explanations. Although the findings are limited to the questions and chatbot versions examined, they highlight the importance of critically evaluating AI-generated content before using it in biochemistry education. Uncritical educational use of such responses may reinforce and disseminate scientific misconceptions.

Journal of Biological Education
Marmara University (TR), Gazi University (TR)
Quality Education
Openalex Percentile: Top 9%
AI in Service Interactions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.