Hippocrates and prophecies: the unfulfilled promise of prediction rules
Notice bibliographique
Résumé
Around 2500 BC (Before COVID-19), Hippocrates stated that it is unwise to prophesy either death or recovery in acute disease. 1 Despite realms of data and computing power available to generate predictions, this still seems valid today, as evidenced by the huge number of publications reporting new prediction models.An admittedly crude search on PubMed with ''prognosis OR prediction'' produces over 3.3 million papers and more than 20,000 when combined with ''COVID-19 OR SARS-CoV-2''.A systematic reviewed published in March 2020, just as the virus was beginning to spread globally, screened 2,696 titles to identify 27 studies that developed or validated a multivariable COVID-19 prediction model. 2 These authors conclude that the proposed models are ''poorly reported, at high risk of bias, and their reported performance is probably optimistic''.Statistical models that relate patient characteristics to outcomes can serve three general purposes.First, the least contentious purpose is consistent and standard reporting and risk adjustment of clinical trial results, which was purportedly the primary reason for the Sepsis-3 definition.3 The second purpose is controlling for heterogeneity when reporting quality metrics, although caution must be exercised when applying and interpreting these results, which are typically aggregated at an institutional level.4 The third purpose is using these results for individual prognostication.Often implemented as clinical prediction rules, these are frequently found wanting.5,6 In critical care, an early approach to determining futility was based on three or more organ failures for three or more days.7 Nevertheless, improvements in outcomes, for undetermined reasons, soon outdated that clinical prediction rule.8 ''Unreliable predictions could cause more harm than benefit in guiding clinical decisions'' 2 is a statement that we strongly endorse.We are particularly concerned with proposals to use scores for triage.For example, a framework to guide resource allocation for critically ill patients with COVID-19 included prediction of survival without specifying how that should be performed.9 Nevertheless, a retrospective evaluation of two triage scoring guidelines for the allocation of mechanical ventilators identified few patients as low priority and there was poor agreement between the two triage scoring guidelines.10 In fact, two well-established mortality prediction models were compared at the level of the individual patient using a threshold mortality prediction of 50% as proposed in triage models, and the resulting graph was a cloud.11 Prediction may be improved through longitudinal observations and modelling.8,12 In this issue of the Journal, Bartoszko et al. report such an approach using dynamic modelling based on three-day intervals.13 Their population was highly selected at a single quaternary centre where nearly one in three patients in the cohort received extracorporeal membrane oxygenation.They rightly concluded that external validation is required, but also suggest that the tool can be used to inform decision-making and resource allocation and allow population level comparisons across institutions.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,056 | 0,314 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,004 | 0,003 |
| Études des sciences et des technologies | 0,003 | 0,014 |
| Communication savante | 0,007 | 0,016 |
| Science ouverte | 0,003 | 0,004 |
| Intégrité de la recherche | 0,014 | 0,036 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».