Ovarian Cancer Biomarker Performance in Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial Specimens
Bibliographic record
Abstract
Establishing a cancer screening biomarker's intended performance requires "phase III" specimens obtained in asymptomatic individuals before clinical diagnosis rather than "phase II" specimens obtained from symptomatic individuals at diagnosis. We used specimens from the Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial to evaluate ovarian cancer biomarkers previously assessed in phase II sets. Phase II specimens from 180 ovarian cancer cases and 660 benign disease or general population controls were assembled from four Early Detection Research Network or Ovarian Cancer Specialized Program of Research Excellence sites and used to rank 49 biomarkers. Thirty-five markers, including 6 additional markers from a fifth site, were then evaluated in PLCO proximate specimens from 118 women with ovarian cancer and 474 matched controls. Top markers in phase II specimens included CA125, HE4, transthyretin, CA15.3, and CA72.4 with sensitivity at 95% specificity ranging from 0.73 to 0.40. Except for transthyretin, these markers had similar or better sensitivity when moving to phase III specimens that had been drawn within 6 months of the clinical diagnosis. Performance of all markers declined in phase III specimens more remote than 6 months from diagnosis. Despite many promising new markers for ovarian cancer, CA125 remains the single-best biomarker in the phase II and phase III specimens tested in this study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".