MétaCan
Menu
Back to cohort
Record W2521059521 · doi:10.1002/cncr.30345

Reply to Nomograms need to be presented in full

2016· letter· en· W2521059521 on OpenAlexaffabout
Cheuk‐Wai Choi, Wei Xu, Anne W.M. Lee

Bibliographic record

VenueCancer · 2016
Typeletter
Languageen
FieldMathematics
TopicStatistical Methods and Inference
Canadian institutionsPrincess Margaret Cancer Centre
Fundersnot available
KeywordsBootstrapping (finance)NomogramResamplingStatisticsCalibrationSelection (genetic algorithm)EconometricsMedicineComputer scienceMathematicsMachine learningOncology

Abstract

fetched live from OpenAlex

We deeply appreciate the interest in our recent publication1 and the comments on the statistical methodology by Drs. Collins and Le Manach. We would like to take this opportunity to respond to their comments. First, we agree that bootstrapping could be another technique in model selection. However, there also are published reports that have demonstrated the potential drawbacks of the bootstrapping method (eg, retrieving overly complex models),2, 3 and a simulation study by Austin comparing the bootstrap method with the conventional backward elimination method demonstrated a similar performance by both methods in identifying variables.4 Furthermore, there are studies indicating that alternative methods by subsampling may have merits over bootstrapping and are worth considering in future investigations.2, 5 Nevertheless, we did repeat our analyses using the bootstrap resampling approach as suggested by Drs. Collins and Le Manach, and the final model returned was found to be the same as that obtained using the backward elimination method in our original article.1 Specifically, the 4 contributing factors (overall stage, age, gross primary tumor volume, and lactate dehydrogenase) were selected in >80% of 200 bootstrap replications whereas the 2 unselected variables (sex and performance status, which were chosen via univariable analysis) were excluded in >65% of the replications. We hereby confirm that our selection of prognostic factors for the nomogram were appropriate. Second, in the calibration plots, we compared observed versus predicted survival probability for the 5-year overall survival endpoint. We presented the intercept and slope of the joined lines based on the calibration plots, rather than the “calibration slope” as suggested by Drs. Collins and Le Manach. We agree that the number of groups may affect the estimation of the intercepts and slopes. We explored different numbers of groups for calibration, and found that the estimated values remained similar regardless of whether we used 4 (as reported in our article),1 5, or 10 groups. For example, with regard to the calibration plots based on the training cohort, the corresponding intercept (slope) for 4, 5, and 10 groups were −0.05 (1.06), −0.08 (1.10), and −0.02 (1.03), respectively. The 3 sets of results were insignificantly different from an intercept of 0 and a slope of 1. We have performed additional analysis using the methods proposed by Drs. Collins and Le Manach, and the results demonstrated similar findings as published in our original article.1 Given baseline hazard h0(t) as the intercept and lp(X) as the linear predictor, the “calibration slope” refers to the slope b in the linear fit of log(h(t|X)) = log(h0(t)) + blp(X).6 For the training cohort, the “calibration slopes” were 1.000, 0.997, and 0.978, respectively, for 4, 5, and 10 groups; all demonstrated no significant difference from 1. The smoothed regression line from flexible adaptive hazard regression (the blue line in Figure 1) also demonstrated good calibration over the training cohort, consistently supporting our conclusion. Calibration plots on 5-year overall survival based on the training cohort. Last, the baseline hazard simply refers to the hazard for the standard set of conditions that continuous variables equal 0 and categorical variables equal corresponding references.6 With this approach, investigators could directly obtain the baseline hazard from the nomogram and assess prognostication of individual patients with their data. We thank Drs. Collins and Le Manach for their comments, and agree that their suggested approaches are useful alternatives. However, repeating the analyses with the suggested methods and performing sensitivity analysis to obtain a comprehensive evaluation of the predictive model demonstrated results comparable to those obtained by the statistical methods used in our original article.1 We confirm that the clinical conclusions regarding the selection of prognostic factors and the nomogram calibration are robust and consistent, irrespective of the statistical approaches used. Our original article provides a valid predictive model for patients with nasopharyngeal cancer;1 the developed nomogram based on the newly proposed 8th edition of the American Joint Committee on Cancer/Union for International Cancer Control staging system, together with additional independent prognostic factors, provides a practical supplementary tool for refining the prediction of overall survival and tailoring treatment strategies for individual patients. No specific funding was disclosed. The authors made no disclosures. Horace C.W. Choi, PhD Department of Clinical Oncology University of Hong Kong Hong Kong, China Wei Xu, PhD Department of Biostatistics Princess Margaret Cancer Centre Toronto, Ontario, Canada Anne W.M. Lee, MD Department of Clinical Oncology University of Hong Kong; Department of Clinical Oncology University of Hong Kong-Shenzhen Hospital Hong Kong, China

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.049
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.139
GPT teacher head0.413
Teacher spread0.274 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2016
Admission routes2
Has abstractyes

Explore more

Same venueCancerSame topicStatistical Methods and InferenceFrench-language works237,207