Lessons on managing pulmonary nodules from NELSON: we have come a long way
Bibliographic record
Abstract
The results of the Dutch-Belgian low-dose CT (LDCT) screening trial (NELSON) have been eagerly awaited since the US National Lung Screening Trial (NLST), published in 2011, demonstrated annual LDCT screening of the chest led to a 20% decrease in lung cancer mortality compared with chest X-ray screening.1 While the US Preventative Service Task Force approved lung cancer screening in 2014,2 Europe and the rest of the globe have been paralysed by fear of implementation costs and the feasibility of introducing national LDCT screening programmes. The consensus from healthcare payers outside of the USA has been that we should wait for the results of the NELSON trial which, while smaller in size, would give us the confidence of a second randomised controlled trial and proof of effect in a population outside of the US healthcare system. Indeed, a recent Health Technology Assessment of LDCT screening in the UK specifically named the NELSON trial as an important source of future information and explicitly stated the results were required in order to make a decision about its efficacy and cost-effectiveness.3 It was on this background of hope and perhaps, dare we say it, healthcare payer fear that the NELSON trial preliminary results were released at the World Conference on Lung Cancer in Toronto in October. Despite NELSON being smaller in size than NLST and having a preponderance of male participants, the results were clear: LDCT screening compared with no screening leads to a statistically significant lung cancer mortality reduction of 26% for men and numbers hint that the benefit could be even greater in women (between 40% and 60%). With this new data the UK, and indeed the world outside of the USA, now needs to cast aside concerns over efficacy, as well as procrastination over implementation, and concentrate more …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".