Improving Lung Cancer Screening Selection: The HUNT Lung Cancer Risk Model for Ever-Smokers Versus the NELSON and 2021 United States Preventive Services Task Force Criteria in the Cohort of Norway: A Population-Based Prospective Study
Bibliographic record
Abstract
Background Improving the method for selecting participants for lung cancer (LC) screening is an urgent need. Here, we compared the performance of the Helseundersøkelsen i Nord-Trøndelag (HUNT) Lung Cancer Model (HUNT LCM) versus the Dutch-Belgian lung cancer screening trial (Nederlands-Leuvens Longkanker Screenings Onderzoek (NELSON)) and 2021 United States Preventive Services Task Force (USPSTF) criteria regarding LC risk prediction and efficiency. Methods We used linked data from 10 Norwegian prospective population-based cohorts, Cohort of Norway. The study included 44,831 ever-smokers, of which 686 (1.5%) patients developed LC; the median follow-up time was 11.6 years (0.01–20.8 years). Results Within 6 years, 222 (0.5%) individuals developed LC. The NELSON and 2021 USPSTF criteria predicted 37.4% and 59.5% of the LC cases, respectively. By considering the same number of individuals as the NELSON and 2021 USPSTF criteria selected, the HUNT LCM increased the LC prediction rate by 41.0% and 12.1%, respectively. The HUNT LCM significantly increased sensitivity ( p < 0.001 and p = 0.028), and reduced the number needed to predict one LC case (29 versus 40, p < 0.001 and 36 versus 40, p = 0.02), respectively. Applying the HUNT LCM 6-year 0.98% risk score as a cutoff (14.0% of ever-smokers) predicted 70.7% of all LC, increasing LC prediction rate with 89.2% and 18.9% versus the NELSON and 2021 USPSTF, respectively (both p < 0.001). Conclusions The HUNT LCM was significantly more efficient than the NELSON and 2021 USPSTF criteria, improving the prediction of LC diagnosis, and may be used as a validated clinical tool for screening selection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".