MétaCan
Menu
Retour à la cohorte
Enregistrement W2243523692 · doi:10.1093/ejcts/ezw263

Reply to Collins and Le Manach

2016· letter· en· W2243523692 sur OpenAlexaff
Teresa M. Kieser, Michael Rose, Stuart J. Head

Notice bibliographique

RevueEuropean Journal of Cardio-Thoracic Surgery · 2016
Typeletter
Langueen
DomaineMedicine
ThématiqueSepsis Diagnosis and Treatment
Établissements canadiensLibin Cardiovascular Institute of AlbertaUniversity of Calgary
Organismes subventionnairesnon disponible
Mots-clésSample size determinationSample (material)Logistic regressionStatisticsEconometricsComputer scienceMathematicsPhysics

Résumé

récupéré en direct d'OpenAlex

We thank Collins and Le Manach for their insightful comments on our paper [1, 2] and also for drawing our attention to their article [1] on sample sizes for external validation of prognostic models, which unfortunately was unavailable when our data were analysed. However, their article examined the effect of sample size for proportional hazards models and we used a logistic regression model (with previously published papers [3, 4] more applicable to the assessment of appropriate sample size). Our sample size, though small, is not as exaggerated as their examples (one study with 8 cases and one with 1 case). Although we reported c-statistics with confidence intervals and the P-value for the difference between them for consistency with previously published papers, we stated upfront that this has severe limitations and that huge sample sizes would be needed to detect clinically relevant differences between two c-statistics considered to be in the ‘excellent’ range. When calculating the sample size for external validation, it is necessary to choose one or two statistics believed most important. We chose calibration; since both calibration-in-the-large and the miscalibration-coefficient were statistically significant, the power of our study is not an issue, but we acknowledge that bias of these estimates may be. Even though it does not directly apply to our logistic regression model, we did use their simulation study for a sample size of ∼37. There was no difference in the coverage rates of confidence intervals between a sample of 37 and one of 100 (or even 200). Also bias in the calibration slope is huge when the number of events is ≤10, but <2.5% when the sample size is 37 and ∼1.6% when 100. Regarding our calibration plot, the scale of the axes was chosen to avoid uninformative white space in the figure; and to avoid confusion, we added a green diagonal line to illustrate the line of equality on which should lie perfect predictions. We believe that ‘the risk of x% of how many patients died’ is clearly presented in the legend of said table and the table itself. Our calibration plots differ only from others in that both are presented on the same graph to illustrate difference. We agree that a less-smoothed calibration plot with 95% CI would have been ideal but not appropriate in this study due to small sample sizes. Operative mortality rate for all coronary artery bypass graft procedures is low (<5%); therefore, any prognostic model for operative mortality will necessarily have zero deaths in the smallest risk groups. Risk score validation of low-risk groups is equally important as for high-risk groups. Our analysis showed that calibration was strongest in the low-risk groups but in the highest risk groups was underestimated by logistic EuroSCORE and overestimated by EuroSCORE II. Surgeons can therefore be confident in either score for low-risk patients. The missing ejection fraction data are unfortunate but are currently being updated by chart review. So far, most had an ejection fraction of >50% which would not alter the EuroSCORE values nor results of our study. Finally, validation of risk models is generally accepted best assessed in the settings in which they will be used [5, 6].

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,007
score de la tête « metaresearch » (Gemma)0,069
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,046
Score d'incertitude au seuil0,036

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0070,069
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0010,001
Études des sciences et des technologies0,0030,003
Communication savante0,0040,006
Science ouverte0,0030,002
Intégrité de la recherche0,0460,054
Charge utile insuffisante (le modèle a refusé de juger)0,0040,006

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,087
Tête enseignante GPT0,319
Écart entre enseignants0,232 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2016
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueEuropean Journal of Cardio-Thoracic SurgeryMême sujetSepsis Diagnosis and TreatmentTravaux en français237 207