Breast cancer subtype variation by ethnicity in a population-based cohort in British Columbia.
Bibliographic record
Abstract
1585 Background: Reports suggest breast cancer subtypes (BCS) are different in developing countries, often attributed to ethnicity. We assessed BCS in our multiethnic population with access to screening mammography and universal health care. Methods: The Breast Cancer Outcomes Unit database was used to identify women diagnosed with invasive breast cancer in 2006 and referred to the BCCA. Ethnicity was abstracted by chart review from a patient questionnaire completed by all new referrals. Ethnicities were grouped into: Caucasian (CA), East Asian (EA), Aboriginal (A), South Asian (SA), South-East Asian (SEA) and Other (O). BCS were grouped as: ER/PR+ HER2-, ER/PR+ HER2+, ER/PR- HER2+, ER/PR- HER2-. Rates of ER and HER2 positivity and BCS were compared by ethnicity using the Chi-square test; age at diagnosis was compared by ethnicity using the Kruskal-Wallis test. Results: 2,222 women were eligible with 2072 having completed questionnaires. Median age was 59 years. The distribution of ethnicities: 73% CA, 8% EA, 4% A, 3% SA, 3% SEA and 9% O. T stage was T1 57%, T2 34%, T3 8% and Tx 1%. N stage was N0 55%, N1 25%, N2 9%, N3 3% and NX 9%. 4.5% were Stage IV at diagnosis. HER2 positivity was significant by ethnicity; 14% CA, 18% EA, 25% A, 29% SA, 26 % SEA, 17% O (p = 0.0008). Age varied significantly (p<0.0001); mean age in SEA was 53 compared to 60 years in CA. Conclusions: Although the subsets are small there appear to be differences in the rates of ER and HER2 positivity by ethnicities. ER negative disease was more frequent in Aboriginal and South Asians. These groups and SE Asians had more HER2 + disease. Validation on larger populations is pending. Genomics analysis may provide epidemiologic information. Differences in BCS may impact screening and prevention strategies. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".