PHASES and the natural history of unruptured aneurysms: science or pseudoscience?
Bibliographic record
Abstract
Unruptured intracranial aneurysms (UIAs) are increasingly discovered and preventive interventions increasingly performed; from 2000 to 2010, rates of coiling have exploded (14-fold from 0.3 to 4.3 per 100 000 American Medicare beneficiaries).1 The problem is that no one knows whether preventive interventions do more harm than good. During the same period, rates of subarachnoid hemorrhage (SAH) also increased from 20 to 25 per 100 000.1 Questioning the merit of preventive interventions is long overdue, but once questions are raised, how to properly address them? On the surface, one reasonable answer is to try to identify patients for whom treatment would be indicated and separate them from patients for whom treatment would be futile or harmful. The idea makes sense, provided we proceed with the right methods. The core problem starts when a ‘natural history’ approach is employed. This method had noble beginnings, with the observation of animals and plants by Aristotle and Pliny the Elder,2 ,3 but it has now disappeared from all other domains in medicine, except ours. The most recent product of this approach is the PHASES scoring system, which proposes ‘a risk prediction chart to guide clinical decisions’.4 ,5 A fundamental question is whether this approach is scientific or pseudoscientific. Karl Popper attempted to demarcate science from pseudoscience: ‘A theory which is not refutable by any conceivable event is nonscientific. Irrefutability is not a virtue of a theory (as people often think) but a vice’.6 Let's see whether PHASES was designed to be refutable. Natural history studies have long looked for risk factors for rupture, such as aneurysm …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".