Incidence of acute myeloid leukemia: A regional analysis of Canada
Bibliographic record
Abstract
National disease registries are used to study prevalence, incidence, and mortality trends in a given population. These registries have also been used to identify case clustering patterns, which aid greatly in finding external risk factors for a given ailment. These descriptive studies are often followed by other confirmatory investigations and are then further supported by biologic plausibility. The identification of case clustering patterns has been the root cause of the identification of external triggers for several diseases, such as cholera, mesothelioma, lung cancers, and select lymphoproliferative diseases.1 In our recent study,2 we investigated in detail the crude incidence and mortality trends and geographic distribution of acute myeloid leukemia (AML) across Canada. Before our report, there was a significant knowledge gap on the epidemiology, distribution, and populations at risk for AML in this country. The study was followed by constructive and valuable criticism, which is part of the normal scientific process, calling for additional analyses, including age standardization and adjustments for multiple comparisons.3, 4 The original article2 was prepared according to strict journal guidelines on word limits and the number of display figures and tables, and it should be viewed as the first step in this line of investigation. The initial manuscript underwent 2 rounds of peer review by 5 experts in the field from May 10 to September 8, 2018, at which point the article was endorsed for publication. For this initial study on AML, the research team followed the same approved protocol for data analysis that was used for a series of other published articles within approximately the last 2 years.5-16 In this study, we present additional analyses, which we hope will help to address the criticisms brought forward by experts in the research community. This study was conducted in accordance with the CISS-RDC-668035 and 13-SSH-MCG-3749 protocols, which were approved by the Social Sciences and Humanities Research Council of Canada and the Quebec Inter-University Centre for Social Statistics, respectively. In agreement with institutional policy, this study received an exemption from review by the McGill University research ethics board. We examined the data on AML incidence in 2 population-based cancer databases, the Canadian Cancer Registry (CCR) and Le Registre Québécois du Cancer (LRQC), for the period of 1992-2010, and we used International Classification of Diseases for Oncology, Third Edition (ICD-O-3) codes for 15 AML subtypes in a manner similar to that previously reported.6 The CCR is a dynamic database of Canadian residents from 12 Canadian provinces and territories (excluding Quebec) who were diagnosed with primary tumors from 1992 to 2013 (alive or dead). Data for Quebec patients were obtained from the LRQC. Data from the LRQC database were available only from 1992 to 2010. Data on new cases of AML were obtained from the CCR (2014 version), which spanned the period of 1992-2013. Because the data from the LRQC database for Quebec were available only up to 2010, we chose for this study to analyze the data from the period of 1992-2010 alone so that we could include all Canadian provinces and territories for the same period. The CCR and LRQC databases provide demographic, geographic, and clinical information, including the patient's sex, year of diagnosis, age at the time of diagnosis, and postal code of residence as well as the ICD-O-3 code of the tumor. The CCR and LRQC databases do not record data on several other demographic characteristics, such as the ethnic background of patients. For incidence calculations, data on population counts for the nation, per province, per city, and per forward sortation area (FSA) were obtained from the Canadian Census of Population for the years 1996, 2001, 2006, and 2011 from Statistics Canada. In Canada, postal codes consist of letters and numbers (eg, H3G 1A4), where the first 3 characters (eg, the FSA) define a region in the country; there are more than 1600 FSAs across Canada. For calculating rural and urban incidence rates by FSA, we used the previously described classification by Canada Post, which is based on the second character in the FSA, for rural delivery areas versus urban delivery areas.13 FSAs whose second character is 0 are considered rural, whereas FSAs whose second character is 1 to 9 are considered urban. Statistical analyses were performed with the SAS 9.3 statistical software package (SAS Institute, Cary, North Carolina). AML cases were defined on the basis of ICD-O-317 as follows: pure erythroid leukemia (ICD-O-3 code 9840); AML, not otherwise specified, with mutated NPM1 or with biallelic mutations of CEBPA (9861); acute myelomonocytic leukemia (9867); acute basophilic leukemia (9870); AML with inv(16)(p13.1q22) or t(16;16)(p13.1;q22);CBFB-MYH11, acute promyelocytic leukemia with PML-RARA (9871); AML with minimal differentiation (9872); AML with maturation (9874); acute monoblastic/monocytic leukemia (9891); AML with myelodysplasia-related changes (9895); AML with t(8;21)(q22;q22.1);RUNX1-RUNX1T1 (9896); AML with t(9;11)(p21.3;q23.3);MLLT3-KMT2A (9897); acute megakaryoblastic leukemia (9910); therapy-related myeloid neoplasms (9920); myeloid sarcoma (9930); and acute panmyelosis with myelofibrosis (9931). We followed a number of regulations established by the Social Sciences and Humanities Research Council of Canada to protect the identity, privacy, and personal information of individual patients. These measures included limitations on releasing stratified regional data by age or cancer subtypes, random rounding of counts to a multiple of 5, not releasing counts lower than 5 or full 6-character postal codes of patients, and requiring a given region or analyzed population to consist of more than 5000 individuals. In summary, data could be released by FSAs (representing the first 3 of 6 characters in a postal code) only when the population was larger than 5000 residents and 5 or more events were detected. This limited the release of stratified data by age for smaller FSAs when there were fewer than 5000 residents per age group and/or fewer than 5 events occurring during the study period. The data on AML cases per city/FSA per age group were rounded to the multiple of 5 (via a random rounding scheme as previously reported2) and subsequently were used to calculate age-standardized incidence rates (ASIRs). Confidence intervals (CIs) were based on the exact Poisson method. Statistical significance was defined by the 95% CI not overlapping with that of the Canadian national incidence rate. The CIs for the ASIRs were produced with the Spiegelman method.18 We present P values with false discovery rates (FDRs) as methods of controlling for multiple comparisons of the various cities and FSAs rates with the national rate.19,20 Age is a known risk factor for AML. In the previous study on the descriptive epidemiology of AML, we reported mostly crude rates.2 In this analysis, we extend our previous findings and report the ASIRs of AML in Canada based on the newly obtained data from the CCR and LRQC. The incidence rate of AML in Canada in 1992-2010 was 30.61 cases per million individuals per year (95% CI, 30.17-31.06). As detailed in Table 1 and Figure 1, several Canadian cities (notably Sarnia, Sault Ste Marie, Regina, and Ottawa) demonstrated elevated ASIRs in comparison with the national average; however, when we looked at the FDR P values, ASIRs for most of these cities were not statistically confirmed (Table 1). The lowest (2-fold less) age-standardized incidence was observed in cities such as Scarborough, York, and North Vancouver, and this was also confirmed by the FDR P values. The age-standardization analysis indicated that the previously observed high crude incidence rates in the cities of Sarnia, Hamilton, Sault Ste Marie, St. Catharines, and Thunder Bay were likely due at least in part to the overall older population residing in these cities (Table 1). In our comparisons, we used the FDR approach19 to control for multiple comparisons. This method calculates the expected proportion of true null hypotheses rejected out of the total number of rejections. Similarly, an ASIR analysis by FSA was conducted for populous FSAs for which data could be released without compromising patient privacy. As detailed in Table 2, several FSAs with high age-standardized AML incidences were located in Ontario (8 of 10 high-incidence FSA); they included N7V in Sarnia, Ontario, L8S and L8H in Hamilton, Ontario, and P6C in Sault Ste Marie, Ontario, which had more than 2-fold elevated ASIRs in comparison with the national average. Also, 2 areas in Vancouver (FSAs V6H and V6L) demonstrated similar trends of more than double the national incidence rate. While several FSAs had significant actual P value, FDR analysis did not demonstrate statistical significance, as shown in Table 2. Finally, to highlight the increase in AML incidence in urban regions, we compared the ASIRs of AML between rural and urban FSAs. On the basis of the ASIRs, the AML incidence was ~1.19-fold higher in urban FSAs versus rural FSAs (P < .001). On the basis of crude incidence rate calculations, the urban FSA incidence was 1.22-fold higher than the rural FSA incidence (P < .001). Collectively, the analyses presented here complement and, in important instances, modify/correct the conclusions in our original report.2 However, these combined results suggest the need for additional investigations into the incidence of AML in Sarnia and other cities/FSAs before definitive conclusions can be drawn. The crude rates presented previously help us to understand the burden of AML in specific regions and can inform resource allocation for AML for these regions, whereas the lack of a difference between crude and age-adjusted rates in cities indicates that overall the increase observed in the crude rates may be attributed to differences in age composition in the segments. However, the significant ASIRs observed in some FSAs within these cities (based on actual p value calculations) may indicate the presence of environmental risk factors in these segments. We believe that these 2 combined reports are valuable and serve as an essential foundation for subsequent investigations into the incidence of AML in Canada. No specific funding was disclosed. The authors made no disclosures. Cities data - for reviewer FSA data -for reviewer_tobesubmitted Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.014 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".