Prediction of findings at screening colonoscopy using a machine learning algorithm based on complete blood counts (ColonFlag)
Bibliographic record
Abstract
Adenomatous polyps are a common precursor lesion for colorectal cancer. ColonFlag is a machine- learning-based algorithm that uses basic patient information and complete blood cell counts (CBC) to identify individuals at elevated risk of colorectal cancer for intensified screening. The purpose of this study was to determine whether ColonFlag is also able to predict the presence of high risk adenomatous polyps at colonoscopy. This study was conducted at a large colon cancer screening center in Calgary, Alberta. The study population included asymptomatic individuals between the ages of 50 and 75 who underwent a screening colonoscopy between January 2013 and June 2015. All subjects had at least one CBC result within the year prior to colonoscopy. Based on age, sex, red blood cell parameters, inflammatory cells and platelets, the ColonFlag algorithm generated a score from 0 to 100. We compared the ability of the ColonFlag test result to discriminate between individuals who were found to have a high risk polyp and those with a normal colonoscopy. Among the 17,676 individuals who underwent a screening colonoscopy there were 1,014 found to have a high risk precancerous lesion (5.7%) and 60 were found to have colorectal cancer (0.3%). At a specificity of 95%, the odds ratio for a positive ColonFlag was 2.0 for those with an advanced precancerous lesion compared with those with a normal colonoscopy. The odds ratio did not vary according to patient subgroup, colorectal cancer location or stage. ColonFlag is a passive test that can use routine blood test results to help identify individuals at elevated risk for high risk precancerous polyps as well as frank colorectal cancer. These individuals may be targeted in an effort to achieve greater compliance with conventional screening tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".