Test performance and clinical validity of circulating tumor DNA (ctDNA) in predicting relapse in solid tumors treated with curative intent therapy.
Bibliographic record
Abstract
3036 Background: Studies have explored the prognostic value of ctDNA in predicting relapse in solid tumors treated with curative intent. These studies have evaluated ctDNA at specific ‘landmark’ timepoint or over numerous ‘surveillance’ time points. However, variable results have led to uncertainty about the clinical validity of this tool. Here, we quantify the predictive and discriminatory accuracy of ctDNA and explore sources of heterogeneity at both landmark and surveillance time points across different tumor sites. Methods: A search of MEDLINE (host: PubMed) identified studies evaluating ctDNA after curative intent therapy in solid tumors. Odds ratios (OR) for disease recurrence at both landmark and surveillance time points for each study were calculated and pooled in a meta-analysis using the Peto method. Pooled sensitivity and specificity weighted by individual study inverse variance were estimated and meta-regression utilizing linear regression weighted by inverse variance was performed to explore associations between patient and tumor characteristics and the OR for disease recurrence. Results: Of 23 studies identified; 16 (750 patients) and 14 studies (853 patients) reported on landmark and surveillance time points respectively. The median time from completion of definitive therapy to landmark testing was 51.5 days (range 3-120). The pooled OR for recurrence at landmark was 22.22 (95% CI 14.82-33.30) and at surveillance was 27.51 (95% CI 19.1-39.63). The pooled sensitivity for ctDNA at landmark and surveillance time points were 59.9% and 73.2%. The corresponding specificities were 90.9% and 86.6%. Subgroup results are shown in the table. There was lower predictive accuracy with the use of tumor site specific panels, in patients receiving adjuvant chemotherapy and in lung cancer. Meta-regression showed that longer time to landmark and higher number of surveillance blood draws were associated with higher prognostic accuracy, as was a history of smoking. Conclusions: Although ctDNA at both landmark and surveillance time points shows high prognostic accuracy, it has low sensitivity, suboptimal specificity and therefore weak discriminatory accuracy to predict relapse in patients with solid tumors treated with curative intent. Testing methodology, time points and patient populations need to be optimized before it can be incorporated routinely in clinical practice.[Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.061 | 0.141 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.006 | 0.019 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".