Outcome prediction for treatment of brain arteriovenous malformations: performance of endovascular predictive scores in a single-center population
Bibliographic record
Abstract
BACKGROUND: Endovascular embolization is an accepted treatment modality for brain arteriovenous malformations (bAVM); however, treatment outcomes are highly variable, warranting accurate prediction for adequate patient selection. Several predictive scores have been proposed for this purpose. The objective of this study was to externally validate these scores for embolization of bAVM. METHODS: This study involved bAVM patients treated with transarterial embolization. Endovascular predictive scores were identified through literature search. Relevant data for scoring of included patients was extracted. Primary study outcomes were radiological cure and neurological complications. The performance of the scores was evaluated by analyzing calibration (z-scores from logistic regression), discrimination (area under the receiver operating characteristic curve, AUROC), and classification (Youden's index and corresponding sensitivity and specificity). Additionally, sensitivity analyses were performed restricting the study population by size, location, and embolization intent. RESULTS: A total of 198 bAVM (190 patients) were included. The rates of radiological cure and neurological complications were 18.2% and 14.1%, respectively. The literature search identified seven predictive scores. In the overall analysis, the Toronto score showed the best performance for radiological cure (AUROC 0.905). No significant difference was observed between the performance of the assessed scores for neurological complications. The sensitivity analysis showed improved performance of most scores. The Toronto score exhibited the highest performance for radiological cure (AUROC 0.857). The AVM Embolization Prognostic Risk Score (AVMEPRS) showed the highest performance for neurological complications (AUROC 0.751). The AVM Embocure Score (AVMES) showed fair to good performance for both efficacy and safety outcomes. CONCLUSION: Among the selected scores, the Toronto, AVMEPRS, and AVMES scores showed the best performances.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".