Development of a Prediction Model for Survival Time in Esophageal Cancer Patients Treated with Resection.
Bibliographic record
Abstract
ObjectiveAccurate estimates of survival guide decision-making for patients and oncologists. Advances in the capacity to measure complex tumour biology and patient factors allow for concurrent consideration of clinical, pathological, molecular, and biological markers for prognostication. Clinical prediction tools are a mechanism to combine and personalize these increasingly large amounts of complex information for prognostication.
 ApproachWe describe the process of linking routinely collected health data, cancer registry, and pathology report data in two provinces to develop (Ontario, Canada) and validate (Manitoba, Canada) a clinical prediction tool in esophageal cancer. We compared the performance of a base model restricted to patient and disease characteristics available prior to surgical resection (e.g., age, sex, histology, comorbidities), and a more complex model including pathology specimen details (e.g., tumour stage). Cox proportional hazards models were fit to predict death at three years following resection. Internal and external validity was assessed using overall calibration and optimism corrected c-statistics. Equity was assessed through calibration in predefined patient subgroups.
 Results2124 patients who underwent surgical resection for esophageal cancer between May 1, 2004 and June 30, 2016 for whom a pathology record was available were included in the study cohort. Median age was 66, with 80% males and 85% adenocarcinomas. Survival data were available until March 31, 2020. The model with pathology data had superior discrimination and calibration (calibration slope of 1.02 and intercept -0.01, and optimism-corrected c-statistic 0.77), compared to the base model (calibration slope of 0.95, intercept 0.02, and c-statistic 0.60). External validation is ongoing.
 ConclusionOur study demonstrates that prediction models for cancer prognosis built solely on data from health administrative databases may be unreliable. The addition of high-quality pathology report data from electronic medical records or population-based cancer registries is necessary for accurate estimation. Our work provides a framework for combining administrative and clinical data which could be applied to the development of other clinical prediction models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".