Abstract 4370603: Risk Stratification with AI-Predictive Models vs. Traditional Clinical Risk Scores in Patients Undergoing Ablation for Atrial Fibrillation: A Systematic Review and Meta-Analysis
Abstract
Background: Atrial fibrillation (AF) recurrence after catheter ablation remains difficult to predict. While traditional risk scores such as CHA2DS2-VASc and HATCH are widely used, their predictive accuracy is modest. Machine learning (ML) models have emerged as a potential alternative, integrating multimodal data to enhance individualized risk stratification. We conducted a systematic review and meta-analysis to evaluate their predictive performance, model design, and comparison with clinical risk scores. Methods: We searched PubMed, Embase, and Scopus for studies published between 2013 and 2024 using ML models to predict post-ablation AF recurrence. Eligible studies included adults undergoing catheter ablation and reported validation of ML model performance. Two reviewers independently extracted data on study design, sample size, input features, ML model type, validation method, AUROC, recurrence rates, and comparator clinical scores. Risk of bias was assessed using PROBAST. Results: Eleven studies comprising 2,994 patients were included. Most were retrospective and conducted between 2013 and 2023 across China, the United States, Portugal, and Europe. Sample sizes ranged from 90 to 1,606, with follow-up durations from 6 months to 5.8 years. AF recurrence rates ranged from 21% to 54%. ML model types included gradient boosting (n=4), convolutional neural networks (n=3), logistic regression (n=2), regularized linear models (n=1), and simulation-based models (n=1). Input data varied from clinical variables (age, LA diameter, comorbidities) to ECG morphology, cardiac CT-based LA wall thickness, and electrogram-derived features. In three head-to-head comparisons, ML models outperformed traditional scores. For example, the HAD-AF model achieved an AUROC of 0.938 versus 0.679 for CHA2DS2-VASc. Average patient age ranged from 56 to 66 years, with >60% male across cohorts. The pooled sensitivity and specificity of ML models for predicting AF recurrence were 80.2% (95% CI: 77.7%–82.7%) and 76.5% (95% CI: 73.9%–79.2%), respectively. The pooled AUROC from five studies was 0.89 (95% CI: 0.86–0.92), reflecting strong discriminative ability across diverse populations and input modalities. Conclusions: Machine learning models consistently outperformed traditional scores for predicting AF recurrence after ablation, with pooled AUROC nearing 0.90 and balanced sensitivity/specificity. Standardized external validation is essential for clinical implementation.
Article Details
Authors (11)
Snigdha Mandava
NRI Medical College and Hospital, NELLORE, India
Sri Lakshmi Ananya Bokka
Gandhi Medical College and Hospital, Secunderabad, India
Urja Sanghvi
Mayo clinic, rochester, minnesota, Rochester, Minnesota, India
Omer Farooq Mohammed
Osmania Medical College, Hyderabad, Hyderabad, India
Binay Panjiyar
Harvard Medical School, Jamaica, New York, United States
Simranjeet Nagoke
Government Medical College Jammu, Jammu, Kashmir, India
Tanzina Afroze
Texas tech university health services, Amarillo, Texas, United States
Hari Krishna Madamanchi
Siddhartha Medical College, Nellore, India
Bhavana Korlakunta
Osmania Medical College, Aswapuram, India
deekshith ameer shaik
Osmania medical college, Hyderabad, India
Venkata Ramana Katikala
KIMS, Amalapuram, Tadepalle, India