Development and Validation of a Multivariable Lung Cancer Risk Prediction Model That Includes Low-Dose Computed Tomography Screening Results
Bibliographic record
Abstract
Importance: Low-dose computed tomography lung cancer screening is most effective when applied to high-risk individuals. Objectives: To develop and validate a risk prediction model that incorporates low-dose computed tomography screening results. Design, Setting, and Participants: A logistic regression risk model was developed in National Lung Screening Trial (NLST) Lung Screening Study (LSS) data and was validated in NLST American College of Radiology Imaging Network (ACRIN) data. The NLST was a randomized clinical trial that recruited participants between August 2002 and April 2004, with follow-up to December 31, 2009. This secondary analysis of data from the NLST took place between August 10, 2013, and November 1, 2018. Included were LSS (n = 14 576) and ACRIN (n = 7653) participants who had 3 screens, adequate follow-up, and complete predictor information. Main Outcomes and Measures: Incident lung cancers occurring 1 to 4 years after the third screen (202 LSS and 96 ACRIN). Predictors included scores from the validated PLCOm2012 risk model and Lung CT Screening Reporting & Data System (Lung-RADS) screening results. Results: Overall, the mean (SD) age of 22 229 participants was 61.3 (5.0) years, 59.3% were male, and 90.9% were of non-Hispanic white race/ethnicity. During follow-up, 298 lung cancers were diagnosed in 22 229 individuals (1.3%). Eight result combinations were pooled into 4 groups based on similar associations. Adjusted for PLCOm2012 risks, compared with participants with 3 negative screens, participants with 1 positive screen and last negative had an odds ratio (OR) of 1.93 (95% CI, 1.34-2.76), and participants with 2 positive screens with last negative or 2 negative screens with last positive had an OR of 2.66 (95% CI, 1.60-4.43); when 2 or more screens were positive with last positive, the OR was 8.97 (95% CI, 5.76-13.97). In ACRIN validation data, the model that included PLCOm2012 scores and screening results (PLCO2012results) demonstrated significantly greater discrimination (area under the curve, 0.761; 95% CI, 0.716-0.799) than when screening results were excluded (PLCOm2012) (area under the curve, 0.687; 95% CI, 0.645-0.728) (P < .001). In ACRIN validation data, PLCO2012results demonstrated good calibration. Individuals who had initial negative scans but elevated PLCOm2012 six-year risks of at least 2.6% did not have risks decline below the 1.5% screening eligibility criterion when subsequent screens were negative. Conclusions and Relevance: According to this analysis, some individuals with elevated risk scores who have negative initial screens remain at elevated risks, warranting annual screening. Positive screens seem to increase baseline risk scores and may identify high-risk individuals for continued screening and enrollment into clinical trials. Trial Registration: ClinicalTrials.gov Identifier: NCT00047385.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".