Circulating Proteome for Pulmonary Nodule Malignancy
Bibliographic record
Abstract
ABSTRACT Background While lung cancer low-dose computed tomography (LDCT) screening is being rolled out in many regions around the world, differentiation of indeterminate pulmonary nodules between malignant and benign remains to a challenge for screening programs. We conducted one of the first systematic investigations of circulating protein markers for their ability to assess the risk of malignancy for screen-detected pulmonary nodules. Methods Based on four LDCT screening studies in the United States, Canada and Europe, we assayed 1078 unique protein markers in pre-diagnostic samples based on a nested case-control design with a total of 1253 participants. Protein markers were measured using proximity extension assays and the data were analyzed using multivariate logistic regression, random forest, and penalized regressions. Results We identified 36 potentially informative markers differentiating malignant nodules from benign nodules. Pathway analysis revealed a tightly connected network based on the 36 protein-coding genes. We observed a differential mRNA expression profile of the corresponding 36 mRNAs between lung tumors and adjacent normal tissues using data from The Cancer Genomic Atlas. We prioritized a panel of 9 protein markers through 10-fold nested cross-validations. We observed that circulating protein markers can increase sensitivity to 0.80 for nodule malignancy compared to the Brock model (p-value<0.001). Two additional markers were identified that were specific for lung tumors diagnosed within one year. All 11 protein markers showed general consistency in improving prediction across the four LDCT studies. Conclusions Circulating protein markers can help to differentiate between malignant and benign pulmonary nodules. Validating these results in an independent CT-screening study will be required prior to clinical implementation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".