Measuring colorectal cancer incidence: the performance of an algorithm using administrative health data
Bibliographic record
Abstract
BACKGROUND: Certain cancer case ascertainment methods used in Quebec and elsewhere are known to underestimate the burden of cancer, particularly for some subgroups. Algorithms using claims data are a low-cost option to improve the quality of cancer surveillance, but have not frequently been implemented at the population-level. Our objectives were to 1) develop a colorectal cancer (CRC) case ascertainment algorithm using population-level hospitalization and physician billing data, 2) validate the algorithm, and 3) describe the characteristics of cases. METHODS: We linked physician billing, hospitalization, and tumor registry data for 2,013,430 Montreal residents age 20+ (2000-2010). We compared the performance of three algorithms based on diagnosis and treatment codes from different data sources. We described identified cases according to age, sex, socioeconomic status, treatment patterns, site distribution, and time trends. All statistical tests were two-sided. RESULTS: Our algorithm based on diagnosis and treatment codes identified 11,476 of the 12,933 incident CRC cases contained in the tumor registry as well as 2317 newly-captured cases. Our cases share similar overall time trends and site distributions to existing data, which increases our confidence in the algorithm. Our algorithm captured proportionally 35% more individuals age 50 and younger among CRC cases: 8.2% vs. 5.3%. The newly captured cases were also more likely to be living in socioeconomically advantaged areas. CONCLUSIONS: Our algorithm provides a more complete picture of population-wide CRC incidence than existing case ascertainment methods. It could be used to estimate long-term incidence trends, aid in timely surveillance, and to inform interventions, in both Quebec and other jurisdictions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.003 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".