Diagnostic performance of four AI tools in pharmacology MCQs: Accuracy, sensitivity, and specificity

A Ayah J. Al-Rahahleh M Mai Z. Rizik F Fahmi Y. Al-Ashwal R Rana K. Abu-Farha

Abstract

Background The rapid rise of AI in medical and pharmaceutical education has engendered much interest; however, a knowledge gap still exists in the evaluation of performances of these tools in critical academic contexts. Objectives The aim of this study was to assess and compare the performances of four openly accessible AI language tools, Microsoft Copilot, ChatGPT-3.5, Google Gemini, and DeepSeek AI, in responding to pharmacology-related MCQs with regard to diagnostic accuracy, sensitivity, specificity, and reproducibility. Methods A total of 80 MCQs were generated and validated, representing four therapeutic systems: cardiovascular, respiratory, gastrointestinal, and endocrine, including four pharmacological domains: mechanism of action, side effects, pharmacokinetics, and drug-drug interactions. Answers were classified into true/false positives and negatives in order to calculate accuracy, sensitivity, and specificity. After two weeks, a second round of testing was performed with the questions to assess answer reproducibility. Results The top overall performer was Microsoft Copilot: 87.5% accuracy, a sensitivity of 94.6%, and a specificity of 70.8%. It continued to perform strongly across all therapeutic systems, especially in the cardiovascular and respiratory domains, with the highest accuracy in identifying drug mechanisms and side effects. ChatGPT-3.5 performed similarly to Google Gemini (76.3% and 75.0% accuracy, respectively) but with higher sensitivity for ChatGPT-3.5 and higher specificity for Gemini. DeepSeek AI had the lowest accuracy overall (68.8%) and the lowest specificity (29.2%), but the highest consistency of reproducibility (97.5%). The performance of all tools decreased significantly with increasing level of question difficulty (p < 0.05). Conclusion All tools have some value in pharmacology education, but Microsoft Copilot was the most consistently accurate. Limitations in complexity and reproducibility suggest that caution should be exercised in academic and clinical use, particularly given the variability seen with ChatGPT-3.5.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 20, Issue 12
Published December 16, 2025
Pages e0337688
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (4)

A

Ayah J. Al-Rahahleh

M

Mai Z. Rizik

F

Fahmi Y. Al-Ashwal

R

Rana K. Abu-Farha