Use of the Total Cancer Care System to Enrich Screening for CD30-Positive Solid Tumors for Patient Enrollment Into a Brentuximab Vedotin Clinical Trial: A Pilot Study to Evaluate Feasibility
Bibliographic record
Abstract
BACKGROUND: One approach to identify patients who meet specific eligibility criteria for target-based clinical trials is to use patient and tumor registries to prescreen patient populations. OBJECTIVE: Here we demonstrate that the Total Cancer Care (TCC) Protocol, an ongoing, observational study, may provide a solution for rapidly identifying patients with CD30-positive tumors eligible for CD30-targeted therapies such as brentuximab vedotin. METHODS: The TCC patient gene expression profiling database was retrospectively screened for CD30 gene expression determined using HuRSTA-2a520709 Affymetrix arrays (GPL15048). Banked tumor tissue samples were used to determine CD30 protein expression by semiquantitative immunohistochemistry. Statistical comparisons of Z- and H-scores were performed using R statistical software (The R Foundation), and the predictive value, accuracy, sensitivity, and specificity of CD30 gene expression versus protein expression was estimated. RESULTS: As of March 2015, 120,887 patients have consented to the institutional review board-approved TCC Protocol. A total of 39,157 fresh frozen tumor specimens have been collected, from which over 14,000 samples have gene expression data available. CD30 RNA was expressed in a number of solid tumors; the highest median CD30 RNA expression was observed in primary tumors from lymph node, soft tissue (many sarcomas), lung, skin, and esophagus (median Z-scores 1.011, 0.399, 0.202, 0.152, and 1.011, respectively). High level CD30 gene expression significantly enriches for CD30-positive protein expression in breast, lung, skin, and ovarian cancer; accuracy ranged from 72% to 79%, sensitivity from 75% to 100%, specificity from 70% to 76%, positive predictive value from 20% to 40%, and negative predictive value from 95% to 100%. CONCLUSIONS: The TCC gene expression profiling database guided tissue selection that enriched for CD30 protein expression in a number of solid tumor types. Such an approach may improve screening efficiency for enrolling patients into biomarker-based clinical trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".