Optimizing the number of variants tracked to follow disease burden with circulating tumor DNA assays in metastatic colorectal cancer
Bibliographic record
Abstract
Background: The number of somatic mutations detectable in circulating tumor DNA (ctDNA) is highly heterogeneous in metastatic colorectal cancer (mCRC). The optimal number of mutations required to assess disease kinetics is relevant and remains poorly understood. Objectives: To determine whether increasing panel breadth (the number of tracked variants in a ctDNA assay) would alter the sensitivity in detecting ctDNA in patients with mCRC. Design: We used archival tissue sequencing to perform an in silico assessment of the optimal number of tracked mutations to detect and monitor disease kinetics in mCRC using sequencing data from the Canadian Cancer Trials Group CO.26 trial. Methods: For each patient, 1, 2, 4, 8, 12, or 16 of the most clonal (highest variant allele frequency) somatic variants were selected from archival tissue-based whole-exome sequencing and assessed for the proportion of variants detected in matched ctDNA at baseline, week 8, and progression timepoints. Results: Data from 110 patients were analyzed. Genes most frequently encountered among the top four highest VAF variants in archival tissue were TP53 (51.9% of patients), APC (43.3%), KRAS (42.3%), and SMAD4 (9.6%). While the frequency of detecting at least one tracked variant increased when expanding beyond variant pool sizes of 1 and 2 in baseline ( p = 0.0030) and progression ( p = 0.0030) ctDNA samples, we observed no significant benefit to increases in variant pool size past four variants in any of the ctDNA timepoints ( p < 0.05). Conclusion: While increasing panel breadth beyond two tracked variants improved variant re-detection in ctDNA samples from patients with treatment refractory mCRC, increases beyond four tracked variants yielded no significant improvement in variant re-detection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".