The Impact of Model Assumptions on Personalized Lung Cancer Screening Recommendations
Bibliographic record
Abstract
BACKGROUND: Recommendations regarding personalized lung cancer screening are being informed by natural-history modeling. Therefore, understanding how differences in model assumptions affect model-based personalized screening recommendations is essential. DESIGN: Five Cancer Intervention and Surveillance Modeling Network (CISNET) models were evaluated. Lung cancer incidence, mortality, and stage distributions were compared across 4 theoretical scenarios to assess model assumptions regarding 1) sojourn times, 2) stage-specific sensitivities, and 3) screening-induced lung cancer mortality reductions. Analyses were stratified by sex and smoking behavior. RESULTS: Most cancers had sojourn times <5 y (model range [MR]; lowest to highest value across models: 83.5%-98.7% of cancers). However, cancer aggressiveness still varied across models, as demonstrated by differences in proportions of cancers with sojourn times <2 y (MR: 42.5%-64.6%) and 2 to 4 y (MR: 28.8%-43.6%). Stage-specific sensitivity varied, particularly for stage I (MR: 31.3%-91.5%). Screening reduced stage IV incidence in most models for 1 y postscreening; increased sensitivity prolonged this period to 2 to 5 y. Screening-induced lung cancer mortality reductions among lung cancers detected at screening ranged widely (MR: 14.6%-48.9%), demonstrating variations in modeled treatment effectiveness of screen-detected cases. All models assumed longer sojourn times and greater screening-induced lung cancer mortality reductions for women. Models assuming differences in cancer epidemiology by smoking behaviors assumed shorter sojourn times and lower screening-induced lung cancer mortality reductions for heavy smokers. CONCLUSIONS: Model-based personalized screening recommendations are primarily driven by assumptions regarding sojourn times (favoring longer intervals for groups more likely to develop less aggressive cancers), sensitivity (higher sensitivities favoring longer intervals), and screening-induced mortality reductions (greater reductions favoring shorter intervals). IMPLICATIONS: Models suggest longer screening intervals may be feasible and benefits may be greater for women and light smokers. HIGHLIGHTS: Natural-history models are increasingly used to inform lung cancer screening, but causes for variations between models are difficult to assess.This is the first evaluation of these causes and their impact on personalized screening recommendations through easily interpretable metrics.Models vary regarding sojourn times, stage-specific sensitivities, and screening-induced lung cancer mortality reductions.Model outcomes were similar in predicting greater screening benefits for women and potentially light smokers. Longer screening intervals may be feasible for women and light smokers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".