SDPS-39 AN UPDATED VALIDATION OF PREDICTIVE ALGORITHMS FOR BRAIN METASTASES IN NON-SMALL CELL LUNG CANCER: SYSTEMATIC REVIEW AND INDEPENDENT COHORTVALIDATION ANALYSIS
Bibliographic record
Abstract
Abstract INTRODUCTION Many non-small cell lung cancer (NSCLC) patients eventually develop brain metastases (BM). Reliable risk stratification with predictive algorithms can lead to early intervention. Here, we evaluated the performance of published BM risk-stratification algorithms using an independent cohort of NSCLC patients. METHODS We evaluated statistical models predicting BM in NSCLC by systematically reviewing relevant studies and testing them on an independent cohort of NSCLC patients in electronic medical records (2011-2020) from Penn State Health. Patient data was randomly split into 70% training and 30% testing for modeling using L1-regularized logistic regression, and we assessed the models' performance using ROC analyses. RESULTS Out of 1,643 publications, 22 met our criteria, and 12 of those studies (527,258 patients) included variables consistently available in patient charts. Our validation cohort included 1,699 NSCLC patients, with a median age at diagnosis of 68 and 20.4% developing BM. Among feasible models, Zhang 2021 had the highest performance in our cohort (AUROC [95% CI]: 0.89 [0.85-0.93]) and was comprised of the following predictors: age at diagnosis; surgical, chemotherapy, and radiation status; T and N stage; histological grade; and number of organs with metastases. Within our independent cohort, logistic regression revealed that the most informative predictors were the number of organs with metastases (OR [95% CI]: 3.25 [2.70, 3.92]), age at NSCLC diagnosis (OR [95% CI]: 0.98 [0.96, 0.99]), N-stage 1 at diagnosis (OR [95% CI]: 1.83 [1.08, 3.11]), and non-carcinoid or -carcinoma histology (OR [95% CI]: 0.29 [0.10, 0.84]). CONCLUSION Our robust approach, including systematic review and one of the largest-to-date independent institutional datasets, proposes a clinically feasible, novel algorithm for stratifying NSCLC patients based on risk of developing BM. This work can inform pragmatic screening and surveillance guidelines to facilitate early detection of BM in NSCLC patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".