Quantifying copy number variations in cell-free DNA for potential clinical utility from a large prostate cancer cohort.
Bibliographic record
Abstract
5072 Background: Prostate cancer (PrCa) is the most frequent non-dermatological malignancy in the male population. Genomic instability resulting in copy number variation (CNV) is a hallmark of malignant transformation. CNV traces from tumors in cell-free DNA (cfDNA) of prostate cancer patients may be identified through massive parallel sequencing (MPS) of serum DNA. These CNV traces may be biomarkers of cancer with clinical applications for screening and follow-up. Methods: DNA was extracted from serum of 205 PrCa patients (Gleason 2 to10), 207 age matched male controls (HC), 10 men with benign hyperplasia (BPH) and 10 with prostatitis (PiS). DNA was amplified using random primers, tagged with a unique molecular identifier per sample, sequenced on a SOLiD system and aligned to the human genome (Build 37). Hits were counted in sliding 100kbp intervals and normalized. Using a random-resampling procedure, genomic regions showing copy number variations in cfDNA that distinguish PrCa from HC were selected. A model using 20 cfDNA regions was cross-validated and used as cfDNA biomarker. Receiver operator characteristics (ROC) curves were calculated for assessment of diagnostic performance by means of area under the curve (AUC). Results: To assess whether CNVs in cfDNA are indicative of PrCa, the number of regions with significant CNV deviation was counted in a first subset of 82 PrCa. Using only the number of regions as measure resulted in an AUC of 0.81 (0.7 – 0.9, p<0.001). Therefore, all samples were used to select regions (n=80) in random resampling (50/50). These regions were used to define a highly significant 20-regions model using five rounds of 10-fold cross-validation (AUC: 0.85±0.7; p< 10-7). This final model discriminated between PrCa and HC with an AUC of 0.92 (0.87 – 0.95) reaching a calculated accuracy of 83%. Both BPH and PiS could be distinguished from PrCa using the cfDNA CNV biomarker with a predicted accuracy of 90%. Conclusions: MPS revealed that only a limited number of chromosomal regions showing CNVs are necessary to achieve statistical separation between prostate cancer and controls. This technique may prove to be clinically useful for screening and follow up of men with prostate cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".