The Identification of Colon Cancer Susceptibility Genes by Using Genome-Wide Scans
Bibliographic record
Abstract
Recent studies have indicated that in approximately 35% of all colorectal cancer (CRC) cases, the CRC was inherited. Although a number of high-risk familial variants have been identified, these mutations explain <6% of CRC cases; therefore, further genome-wide scans will need to be conducted in the future. There are two popular approaches to genome-wide scans, namely linkage and association. The linkage approach utilizes several hundred markers (typically between 300 and 500 markers) throughout the genome and identifies candidate regions shared among affected family members. Candidate regions are then scrutinized for the presence of susceptibility loci. Linkage studies require no prior information and can provide new avenues for future research, but the regions identified are often large and include many candidate genes. The second and more recent approach is the genome-wide association study (GWAS) in which hundreds of thousands of markers called single nucleotide polymorphisms (SNPs) are used to identify the SNPs associated with traits of interest by employing family-based or case-control association methods. GWAS studies require no prior information and, because they use hundreds of thousands of SNPs, they can target specific candidate genes and/or narrow regions for investigation. Study design considerations, methodology, and the execution of linkage and genome-wide association studies that use both family and case-control designs are covered in this chapter.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".