A comparative analysis of readability, quality, and reliability in large language model outputs pertaining to knee osteoarthritis queries
Abstract
This study aims to comparatively examine the readability, accuracy, and quality of responses provided by artificial intelligence (AI)-based chatbots such as Perplexity, ChatGPT-5, and Gemini to questions about knee osteoarthritis (KOA), which accounts for approximately four-fifths of the global osteoarthritis (OA) burden. In this study, 8 keywords were determined by excluding repetitive, irrelevant or synonymous ones from the 25 most frequently used English keywords associated with KOA based on Google Trends data, and these terms were asked as questions to three different artificial intelligence-based chatbots. The study measured readability using formulas like Coleman-Liau Index (CLI), Automated Readability Index (ARI), and Linsear Write (LW). Reliability of the information was assessed using the Journal of the American Medical Association (JAMA) benchmarks along with the modified DISCERN instrument. To determine overall content quality, the Global Quality Score (GQS) and the Ensuring Quality Information for Patients (EQIP) scale were applied. Together, these tools provided a comprehensive assessment of how understandable, reliable, and high-quality each chatbot’s responses were. The most frequently searched keywords related to OA were “osteoarthritis of knee,” “knee pain,” and “osteoarthritis knee pain.” A readability analysis of responses from three different AI-based chat systems revealed that all platforms had text levels above the Grade 6 threshold, and this difference was statistically significant (p < 0.05). Comparisons demonstrated that ChatGPT-5 produced the most readable content (FRES:45, GFOG:11.9, FKGL:9.24, CLI:14.03, SMOG:8.37, ARI:11.37, LW:7.2). However, Perplexity achieved significantly higher scores than ChatGPT-5 across all quality and reliability assessments, yielding superior median scores (DISCERN: 4, JAMA: 2, GQS: 4, EQIP: 92.8). Perplexity also outperformed Gemini in the mDISCERN reliability assessment (p = 0.001), while no significant difference in quality or reliability was found between Gemini and ChatGPT-5. No statistically significant difference was found between Gemini and ChatGPT in reliability and quality surveys. This analysis of KOA highlights significant challenges regarding the potential of popular AI chatbots for patient information. When examining readability levels, responses from these tools consistently exceed the recommended comprehensibility threshold, making it difficult for patients to absorb critical information. Furthermore, the relatively low scores recorded in reliability and content quality assessments raise significant concerns about the scientific validity and integrity of the medical information presented. Given these findings, the sufficient quality, robustness, and appropriate levels of understandability of future AI-based tools can only be ensured by the establishment and operation of an effective oversight mechanism.
Article Details
Authors (3)
Erdem Maraşlı
Erkan Ozduran
Volkan Hancı