Reliability of the Diagnosis of Cerebral Vasospasm Using Catheter Cerebral Angiography: A Systematic Review and Inter- and Intraobserver Study
Bibliographic record
Abstract
BACKGROUND AND PURPOSE: Conventional angiography is the benchmark examination to diagnose cerebral vasospasm, but there is limited evidence regarding its reliability. Our goals were the following: 1) to systematically review the literature on the reliability of the diagnosis of cerebral vasospasm using conventional angiography, and 2) to perform an agreement study among clinicians who perform endovascular treatment. MATERIALS AND METHODS: Articles reporting a classification system on the degree of cerebral vasospasm on conventional angiography were systematically searched, and agreement studies were identified. We assembled a portfolio of 221 cases of patients with subarachnoid hemorrhage and asked 17 raters with different backgrounds (radiology, neurosurgery, or neurology) and experience (junior ≤10 and senior >10 years) to independently evaluate cerebral vasospasm in 7 vessel segments using a 3-point scale and to evaluate, for each case, whether findings would justify endovascular treatment. Nine raters took part in the intraobserver reliability study. RESULTS: The systematic review showed a very heterogeneous literature, with 140 studies using 60 different nomenclatures and 21 different thresholds to define cerebral vasospasm, and 5 interobserver studies reporting a wide range of reliability (κ = 0.14-0.87). In our study, only senior raters reached substantial agreement (κ ≥ 0.6) on vasospasm of the supraclinoid ICA, M1, and basilar segments and only when assessments were dichotomized (presence or absence of ≥50% narrowing). Agreement on whether to proceed with endovascular management of vasospasm was only fair (κ ≤ 0.4). CONCLUSIONS: Research on cerebral vasospasm would benefit from standardization of definitions and thresholds. Dichotomized decisions by experienced readers are required for the reliable angiographic diagnosis of cerebral vasospasm.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.052 | 0.216 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.008 | 0.008 |
| Bibliometrics | 0.019 | 0.014 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".