Abstract 10366: Comparative Evaluations of Six Risk Scores Proposed for Baseline Prediction of Atrial Fibrillation Recurrence Post Catheter Ablation
Bibliographic record
Abstract
Background: Being able to identify patients who are at risk of arrhythmia recurrence following catheter ablation is important for prognostication, shared-decision making, etc. While numerous predictive models for atrial fibrillation recurrence (AFR) have been proposed, few models have underwent external validation, and few studies compared multiple models using the same patient cohort. Aims: To evaluate six prediction models using an external dataset. Methods: Data from 632 patients was pooled from two independent clinical trials that enrolled patients with anti-arrhythmic drug refractory atrial fibrillation who underwent catheter ablation. Primary outcome for both trials was documented atrial tachyarrhythmia (atrial fibrillation/flutter/tachycardia as adjudicated by a clinical committee blinded to treatment strategy). We compared 6 models for predicting AFR recorded between days 91-365 post ablation using standard metrics and ranked them according to the positive and negative clinical utility indices. Results: As the figure shows, the top performing model was CHA2DS2-VASc, followed by DR-FLASH, albeit both achieved positive and negative predictive values lower than 65%. Many models achieved high specificity but low sensitivity, except for CAAP-AF, which achieved high sensitivity but low specificity. In summary, all achieved area under receiver operating characteristic curve lower than 56% and were deemed to have poor clinical utility (less than 0.5 out of 1) when evaluated on our dataset. Conclusion: All models performed poorly and had objectively limited clinical utility. There remains a need to develop generalizable tools for predicting post ablation AFR.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.039 |
| Meta-epidemiology (narrow) | 0.003 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".