Hippocrates and prophecies: the unfulfilled promise of prediction rules
Bibliographic record
Abstract
Around 2500 BC (Before COVID-19), Hippocrates stated that it is unwise to prophesy either death or recovery in acute disease. 1 Despite realms of data and computing power available to generate predictions, this still seems valid today, as evidenced by the huge number of publications reporting new prediction models.An admittedly crude search on PubMed with ''prognosis OR prediction'' produces over 3.3 million papers and more than 20,000 when combined with ''COVID-19 OR SARS-CoV-2''.A systematic reviewed published in March 2020, just as the virus was beginning to spread globally, screened 2,696 titles to identify 27 studies that developed or validated a multivariable COVID-19 prediction model. 2 These authors conclude that the proposed models are ''poorly reported, at high risk of bias, and their reported performance is probably optimistic''.Statistical models that relate patient characteristics to outcomes can serve three general purposes.First, the least contentious purpose is consistent and standard reporting and risk adjustment of clinical trial results, which was purportedly the primary reason for the Sepsis-3 definition.3 The second purpose is controlling for heterogeneity when reporting quality metrics, although caution must be exercised when applying and interpreting these results, which are typically aggregated at an institutional level.4 The third purpose is using these results for individual prognostication.Often implemented as clinical prediction rules, these are frequently found wanting.5,6 In critical care, an early approach to determining futility was based on three or more organ failures for three or more days.7 Nevertheless, improvements in outcomes, for undetermined reasons, soon outdated that clinical prediction rule.8 ''Unreliable predictions could cause more harm than benefit in guiding clinical decisions'' 2 is a statement that we strongly endorse.We are particularly concerned with proposals to use scores for triage.For example, a framework to guide resource allocation for critically ill patients with COVID-19 included prediction of survival without specifying how that should be performed.9 Nevertheless, a retrospective evaluation of two triage scoring guidelines for the allocation of mechanical ventilators identified few patients as low priority and there was poor agreement between the two triage scoring guidelines.10 In fact, two well-established mortality prediction models were compared at the level of the individual patient using a threshold mortality prediction of 50% as proposed in triage models, and the resulting graph was a cloud.11 Prediction may be improved through longitudinal observations and modelling.8,12 In this issue of the Journal, Bartoszko et al. report such an approach using dynamic modelling based on three-day intervals.13 Their population was highly selected at a single quaternary centre where nearly one in three patients in the cohort received extracorporeal membrane oxygenation.They rightly concluded that external validation is required, but also suggest that the tool can be used to inform decision-making and resource allocation and allow population level comparisons across institutions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.056 | 0.314 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.003 | 0.014 |
| Scholarly communication | 0.007 | 0.016 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.014 | 0.036 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".