Abstract 4370282: The Latest Large Language Model, Grok: Can It Provide Education about Atrial Fibrillation for Diverse Populations?

O Obaid Khan (California Health Sciences University, Clovis, California, United States) G Gloria Wu (UCSF School of Medicine, San Jose, California, United States) H Hrishi Paliath-Pathiyal (Nova Southeastern University, Fort Lauderdale, Florida, United States) P Paul Wang (Stanford University, Stanford, California, United States) I Ivan Chim (University of California, San Diego, San Diego, California, United States) B Brian Hoang (Department of Chemistry, Hunter College) E Emily Chung (Boston University, Boston, Massachusetts, United States) N Noemi Mendoza (San Francisco State University, San Francisco, California, United States) V Viki Toram (UCSF School of Medicine, San Jose, California, United States) R Riki Toram (UCSF School of Medicine, San Jose, California, United States) M Margaret Wang (Santa Clara University, Santa Clara, California, United States)

Abstract

Background: Large language models (LLMs) are used by atrial fibrillation patients. ChatGPT (OpenAI, San Francisco) and Grok (X.ai, San Francisco) have 450 M, 35 M monthly users, respectively. Grok is the newest LLM, open-sourced, uses Mixture of Experts algorithms, has 314 billion parameters, known for STEM answers and the use of X (formerly Twitter) as a data source. Grok was meant to be conversational in tone. LLMS are trained by data sets initially trained by software engineers and later by AI in part or exclusively. It is not known whether Grok responses about atrial fibrillation queries differ by patient gender and race/ethnicity. Methods: We used the query: “I am a 68-year-old [ethnic/racial group] [male/female] with atrial fibrillation. I had a heart attack 2 years ago with stents. What can I expect from my cardiologist?” Three ethnic groups (White, African American, and Latinx) and male/female gender. Response analysis: Word Count (WC) and Flesch-Kincaid Grade Level (FK). ChatGPT4.5 reviewed the LLM responses for cultural sensitivity. Results: Average WC: ChatGPT= 312.5±110.5, Grok= 830.7±104.7. Average FK: ChatGPT=10.7±0.9, Grok=10.3±1.0. Grok showed high cultural sensitivity, for African American female and Latinx users, e.g. diet, cardiovascular risk factors. Both male and female prompts were treated equitably in tone, depth, and scope. However, Grok did not incorporate culturally relevant content for White male or female users. For the Hispanic prompt, Grok mentioned the existence of “language services” but no website links or related organizations for further help. CHA2DS2-VASc is mentioned by both ChatGPT and Grok. Grok has a lower reading grade level for White males, Black females, Hispanic males than that of ChatGPT which may reflect their use of X (formerly Twitter) data. Grok had the longest response for Black females versus all other ethnic groups in this small study. Conclusion: Grok, the latest LLM, competes well with ChatGPT with its thoroughness and factual medical education answers. Reading level however varies by racial/ethnic group and gender.

Article Details

Journal Circulation
Volume / Issue Vol. 152, Issue Suppl_3
Published November 04, 2025
ISSN 0009-7322
Publisher Lippincott Williams & Wilkins

Journal Info

Circulation

Lippincott Williams & Wilkins

ISSN: 0009-7322 Health Sciences

Authors (11)

O

Obaid Khan

California Health Sciences University, Clovis, California, United States

G

Gloria Wu

UCSF School of Medicine, San Jose, California, United States

H

Hrishi Paliath-Pathiyal

Nova Southeastern University, Fort Lauderdale, Florida, United States

P

Paul Wang

Stanford University, Stanford, California, United States

I

Ivan Chim

University of California, San Diego, San Diego, California, United States

B

Brian Hoang

Department of Chemistry, Hunter College

E

Emily Chung

Boston University, Boston, Massachusetts, United States

N

Noemi Mendoza

San Francisco State University, San Francisco, California, United States

V

Viki Toram

UCSF School of Medicine, San Jose, California, United States

R

Riki Toram

UCSF School of Medicine, San Jose, California, United States

M

Margaret Wang

Santa Clara University, Santa Clara, California, United States