Observer Variability of an Angiographic Grading Scale Used for the Assessment of Intracranial Aneurysms Treated with Flow-Diverting Stents
Bibliographic record
Abstract
BACKGROUND AND PURPOSE: Novel angiographic grading scales for the assessment of intracranial aneurysms treated with flow-diverting stents have been recently developed because previous angiographic grading scales cannot be applied to these aneurysms. The purpose of this study was to evaluate the inter- and intraobserver variability of the novel O'Kelly Marotta grading scale, which was developed specifically for the angiographic assessment of aneurysms treated with flow-diverting stents. MATERIALS AND METHODS: Multiple raters (n = 31) from the disciplines of neuroradiology and neurosurgery were presented with pre- and posttreatment angiographic images of 14 aneurysms treated with intraluminal flow diverters. Raters were asked to classify pre- and posttreatment angiograms by using the OKM grading scale. Statistical analyses were subsequently performed with calculation of a generalized multirater κ statistic for assessment of inter- and intraobserver variability and by performing a Wilcoxon signed rank sum test for assessment of group differences. RESULTS: Variability analysis of the OKM grading scale yielded substantial (κ = 0.74) and almost perfect (κ = 0.99) inter- and intraobserver agreement, respectively, with no statistically significant differences between raters with a background of neuroradiology versus neurosurgery or attending physician versus trainee. CONCLUSIONS: The OKM grading scale for the assessment of intracranial aneurysms treated with flow-diverting stents is a reliable grading scale that can be used equally well by users of varying backgrounds and levels of training. Comparison with interobserver variability of pre-existing angiographic grading scales shows equal or better performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.086 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".