Colon Cancer Family Registry: An International Resource for Studies of the Genetic Epidemiology of Colon Cancer
Bibliographic record
Abstract
BACKGROUND: Family studies have served as a cornerstone of genetic research on colorectal cancer. MATERIALS AND METHODS: The Colorectal Cancer Family Registry (Colon CFR) is an international consortium of six centers in North America and Australia formed as a resource to support studies on the etiology, prevention, and clinical management of colorectal cancer. Differences in design and sampling schemes ensures a resource that covers the continuum of disease risk. Two separate recruitment strategies identified colorectal cancer cases: population-based (incident case probands identified by cancer registries; all six centers) and clinic-based (families with multiple cases of colorectal cancer presenting at cancer family clinics; three centers). At this time, the Colon CFR is in year 10 with the second phase of enrollment nearly complete. In phase I recruitment (1998-2002), population-based sampling ranged from all incident cases of colorectal cancer to a subsample based on age at diagnosis and/or family cancer history. During phase II (2002-2007), population-based recruitment targeted cases diagnosed before the age of 50 years are more likely attributable to genetic factors. Standardized protocols were used to collect information regarding family cancer history and colorectal cancer risk factors, and biospecimens were obtained to assess microsatellite instability (MSI) status, expression of mismatch repair proteins, and other molecular and genetic processes. RESULTS: Of the 8,369 case probands enrolled to date, 2,602 reported having one or more colorectal cancer-affected relatives and 799 met the Amsterdam I criteria for Lynch syndrome. A large number of affected (1,324) and unaffected (19,816) relatives were enrolled, as were population-based (4,108) and spouse (983) controls. To date, 91% of case probands provided blood (or, for a few, buccal cell) samples and 75% provided tumor tissue. For a selected sample of high-risk subjects, lymphocytes have been immortalized. Nearly 600 case probands had more than two affected colorectal cancer relatives, and 800 meeting the Amsterdam I criteria and 128, the Amsterdam II criteria. MSI testing for 10 markers was attempted on all obtained tumors. Of the 4,011 tumors collected in phase I that were successfully tested, 16% were MSI-high, 12% were MSI-low, and 72% were microsatellite stable. Tumor tissues from clinic-based cases were twice as likely as population-based cases to be MSI-high (34% versus 17%). Seventeen percent of phase I proband tumors and 24% of phase II proband tumors had some loss of mismatch repair protein, with the prevalence depending on sampling. Active follow-up to update personal and family histories, new neoplasms, and deaths in probands and relatives is nearly complete. CONCLUSIONS: The Colon CFR supports an evolving research program that is broad and interdisciplinary. The greater scientific community has access to this large and well-characterized resource for studies of colorectal cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.036 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.016 | 0.021 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.031 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".