A Population-Based Analysis of Breast Cancer Incidence and Survival by Subtype in Ontario Women
Bibliographic record
Abstract
Background: Breast cancer (bca) is the type of cancer most frequently diagnosed among women in Canada. Breast cancer is categorized into various molecular subtypes by the expression of estrogen receptor (er), progesterone receptor (pgr), and her2 (human epidermal growth factor receptor 2). Currently, Canada has no national cancer registry with epidemiology data by subtype. Thus, we conducted a study to determine incidence, survival, and clinicopathologic characteristics by bca subtype [triple negative breast cancer (tnbc); her2+; and hormone receptor-positive (hr+), her2-] in Canadian women newly diagnosed with bca. Methods: Female patients diagnosed between 1 April 2012 and 31 March 2016 (fiscal 2012-2015) were identified in the Ontario Cancer Registry, and individual patient data were linked to data in provincial health administrative databases. Descriptive statistics and Kaplan-Meier curves were generated. Results: In this cohort, 3277 women (9.5%) had tnbc, 4902 (14.3%) had her2+ bca, and 22,247 (64.8%) had hr+, her2-breast cancer. The annual incidence was 15 per 100,000 for the tnbc group, 21-23 per 100,000 for the her2+ group, and 97-105 per 100,000 for the hr+, her2- group. The lowest median overall survival (mos) of 8.9 months was observed in women with clinical stage iv tnbc. In comparison, the mos was 37.3 months in those with her2+ disease and 35.2 months in those with and hr+, her2- metastatic bca. Conclusions: In the present study, the most recent and largest administrative database analysis of a Canadian population to date, we observed a subtype distribution consistent with previously reported data, together with comparable annual incidence and overall survival patterns.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".