Validation of Lung EpiCheck, a novel methylation-based blood assay, for the detection of lung cancer in European and Chinese high-risk individuals
Bibliographic record
Abstract
AIM: Lung cancer screening reduces mortality. We aim to validate the performance of Lung EpiCheck, a six-marker panel methylation-based plasma test, in the detection of lung cancer in European and Chinese samples. METHODS: A case-control European training set (n=102 lung cancer cases, n=265 controls) was used to define the panel and algorithm. Two cut-offs were selected, low cut-off (LCO) for high sensitivity and high cut-off (HCO) for high specificity. The performance was validated in case-control European and Chinese validation sets (cases/controls 179/137 and 30/15, respectively). RESULTS: The European and Chinese validation sets achieved AUCs of 0.882 and 0.899, respectively. The sensitivities/specificities with LCO were 87.2%/64.2% and 76.7%/93.3%, and with HCO they were 74.3%/90.5% and 56.7%/100.0%, respectively. Stage I nonsmall cell lung cancer (NSCLC) sensitivity in European and Chinese samples with LCO was 78.4% and 70.0% and with HCO was 62.2% and 30.0%, respectively. Small cell lung cancer (SCLC) was represented only in the European set and sensitivities with LCO and HCO were 100.0% and 93.3%, respectively. In multivariable analyses of the European validation set, the assay's ability to predict lung cancer was independent of established risk factors (age, smoking, COPD), and overall AUC was 0.942. CONCLUSIONS: Lung EpiCheck demonstrated strong performance in lung cancer prediction in case-control European and Chinese samples, detecting high proportions of early-stage NSCLC and SCLC and significantly improving predictive accuracy when added to established risk factors. Prospective studies are required to confirm these findings. Utilising such a simple and inexpensive blood test has the potential to improve compliance and broaden access to screening for at-risk populations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".