Abstract 7361: Scientific rationale and successful implementation of biospecimen collection in the NCI Connect Cohort for Cancer Prevention
Bibliographic record
Abstract
Abstract Introduction: The Connect for Cancer Prevention Study is a new prospective cohort with repeated exposure assessment and long-term follow-up to study the cancer continuum from initiation, multi-step carcinogenesis, diagnosis, to outcomes in a diverse US population. Over 50, 000 participants towards the goal of 200, 000 have been recruited so far at 10 integrated healthcare sites across the United States. Goals of Connect include studies of cancer etiology, risk prediction, and early detection. Biospecimens are a critical component for exposure assessment and biomarker measurement. We summarize the scientific rationale and successful implementation of biospecimen collections in Connect. Methods: A biospecimen collection protocol was informed by literature review, expert consultations, and pilot studies evaluating how different blood collection tubes and processing protocols influence DNA yield and quality, as well as nucleic acid and protein-based biomarkers. Baseline biospecimens include a 45ml blood draw, a urine collection, and a mouthwash sample. Biospecimens are collected in dedicated research labs or at clinical sites with home collection of mouthwash samples. A cell-free DNA collection for multi-cancer early detection is implemented at research collection sites. All biospecimens are shipped to an NCI laboratory for processing and long-term storage. Process metrics include sample completeness, sample deviations, temperature logging, and needle-to-freezer time, among others. Results: As of November 2024, 35, 790 participants of 53, 005 enrolled (68%, ranging from 57% to 75% across sites) donated blood and urine samples, with baseline collections still ongoing. Among collections, 58% were from clinical sites, and 42% from research labs. Over 80% of biospecimens collected at research labs were received at NCI within one day, while 70% of biospecimens collected at clinical sites were received within 2 days. Over 95% of all biospecimens were received within 4 days. The return of home-collected mouthwash samples was 77%. Among 276, 228 biospecimen tubes collected, 87% were complete with no deviations recorded. Over 90% of participants submitted a short survey relevant for sample collection after the biospecimen visit. Blood collections are planned every three years to study different exposure windows and biomarker changes within individuals. More frequent collections are considered among participants at high risk of cancer. Conclusions: Connect combines EHR data, state-of-the-art surveys, and biospecimen collections to address critical questions of cancer etiology and prevention. We successfully implemented a robust and efficient biospecimen collection approach at 10 recruitment sites across the U.S. The timeline and process for biospecimen access for the research community and initial biospecimen activities will be presented at AACR. Citation Format: Nicolas A. Wentzensen, Stephanie Weinstein, Amanda Black, Erin Schwartz, Hannah P. Yang, Michelle Brotzman, Paul Albert, Laura Beane-Freeman, Amy Berrington, Jonas De Almeida, Jonine D. Figueroa, Montserrat Garcia-Closas, Nicole Gerlanc, Gretchen Gierach, Rena Jones, Peter Kraft, Charles Matthews, Habib Ahsan, Brisa Aschebrook-Kilfoy, Chun-Hung Chan, Robert Greenlee, Stacey Honda, Ben Rybicki, Blythe Ryerson, Katherine Sanchez, Mark Schmidt, Kevin Sykes, Larissa White, Jeanette Ziegenfuss, Stephen Chanock, Christian Abnet, Mia Gaudet. Scientific rationale and successful implementation of biospecimen collection in the NCI Connect Cohort for Cancer Prevention [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7361.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.573 | 0.592 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.005 | 0.005 |
| Scholarly communication | 0.007 | 0.003 |
| Open science | 0.007 | 0.006 |
| Research integrity | 0.011 | 0.008 |
| Insufficient payload (model declined to judge) | 0.003 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".