Data Resource Profile: The Canadian Hospitalization and Taxation Database (C-HAT)
Bibliographic record
Abstract
With the current emphasis that studies of health should evaluate patient-centred outcomes, the economic consequences of impaired health certainly qualify as ones ‘that people notice and care about’.1,2 For working age people, the abilities to work and earn income are essential aspects of life, affecting not only their own and their family’s economic well-being, but also their quality of life and sense of self.3,4 Most previous studies of how health is related to such labour market outcomes evaluated chronic health conditions.5–8 However, 47% of acute, non-obstetric hospitalizations occur among working age people,9,10 and little is known about the economic consequences of such acute health shocks, including work and earnings. Though it is likely that some health shocks result in significant and enduring earning and employment decrements (earnings shocks), quantifying such effects has been hindered by a lack of data that link individual health and economic data. Most existing studies have been limited by: short duration of follow-up,11 small samples,12,13 studying periods before modern treatment paradigms14,15 and lack of a comparison group.16 To enable analyses that obviate all these limitations, we linked administrative hospital records to individual tax returns for all Canadians (excluding residents of Quebec) for the 16 fiscal years 1999/2000–14/15. The final dataset included longitudinal tax records successfully linked to 16 142 064 individuals who experienced 48 203 982 hospitalizations. Ethics approval for this work was obtained from the Health Research Ethics Board of the University of Manitoba. The linkage was approved by Statistics Canada’s Executive Management Board,17 and use of the linked data is governed by the Directive on Record Linkage.18 Statistics Canada ensures respondent privacy during the linkage process and subsequent use of linked files. Only employees directly involved in the linkage process had access to the unique identifying information required for linkage, and these individuals do not access health- or tax-related information. After the data linkage was completed, an analytical file was created from which the identifying information was removed. This de-identified file can be accessed by analysts for research purposes. Statistics Canada linked national, population-based hospital and tax files. Hospital records were derived from the Discharge Abstract Database (DAD) supplied to Statistics Canada by the Canadian Institute for Health Information. DAD includes all discharges from acute care hospitals in Canada for fiscal years 1999/2000 to 2014/15, excluding Quebec for all years and Manitoba for fiscal years 1999/2000–03/04.19 With approximately 3 million yearly hospital discharges, it contains high quality administrative health data originating from clinical chart abstractions manually performed in every Canadian hospital, by uniformly trained personnel using standardized data definitions.20,21 The hospital data contains hundreds of data elements, including: (i) details of hospital admission, discharge, disposition and transfers; (ii) demographic and residence information; (iii) medical services and locations of care; (iv) up to 25 diagnoses coded in International Statistical Classification of Diseases (ICD)-9, 9-CM and 10-CA format, each accompanied by a diagnosis type which allows identification of whether the diagnosis occurred before or after hospital admission;22 and (v) up to 20 procedures coded using the Canadian Classification of Diagnostic, Therapeutic, and Surgical Procedures (before 2001/2002) and the Canadian Classification of Health Interventions (from 2001/02 onwards).23 Individuals are identified in each hospital file record by a health insurance number (HIN) that is unique within provinces and territories, being issued to access the system of universal, publicly funded health care. All 55 383 725 hospital discharges occurring in that interval were considered for linkage. Tax records, representing all filers in Canada from 1982 to 2014, were derived from the T1 Family File (T1FF) provided to Statistics Canada by the Canada Revenue Agency.24 Canadians must file returns if they have any taxable income, capital gains, self-employment or payments from or into any type of retirement account or if they wish to receive benefits.25 The tax file comprises over 25 million yearly records, containing information on demographics, residency and income by source. In an individual year, persons are identified in tax data by Social Insurance Numbers (SIN) and/or Dependent Identification Numbers (DIN) assigned by the federal government for employment and tax purposes. DINs identify children when filers claim child tax benefits. On a yearly basis, approximately 75% of all Canadians file individual tax returns. This rate is varies by age, being 98% for those ≥19 years of age, 64% for those precisely 18 years old, and <5% among those aged 0–14 years. However, approximately 95% of the total population is represented in the tax data since these include tax information filed on behalf of most children and some spouses who do not file individually. Tax data allow identification of census families, grouping parents and children living at the same address. A complicating factor for record linkage is that an individual may acquire multiple DINs and SINs over time; for example, temporary SINs are assigned to newcomers to Canada before they obtain a permanent SIN. To handle this complication, tax records were pre-processed, assigning a unique identifier in the tax file for each individual (TAXID) which associated all identifying information for that person (sex, date of birth, postal codes, SINs and DINs). There were 43 286 130 individuals represented in the processed tax file. Our goal was to link individuals’ health insurance and social insurance numbers in the hospital and tax files, respectively. In the absence of names on the hospital records, we used a Linkage Key composed of three identifiers common to both datasets: sex, date of birth and postal code. Hospital file re-abstraction studies have shown that these elements are highly reliable, with discrepancy rates of <5%.26 Since less than 40 individuals live in an average Canadian postal code,27 it is rare for multiple individuals to share all three of these. We aimed to maximize the overall linkage rate while avoiding false-positive links. We proceeded in sequential steps: data preparation, record linkage and quality/consistency assessment. Format check algorithms were used to identify invalid HINs. Overall 98.8% of records had a valid HIN, but the rate was lower for certain regions including the territories of Nunavut (76%) and the Northwest Territories (84%). Hospital records with invalid HINs or missing/incomplete Linkage Keys were excluded. We excluded records in the hospital file that were not expected to link: stillbirths, cadavers and non-residents of Canada. Since postal codes typically change when people move, an HIN could be associated with multiple Linkage Keys. For each TAXID in the tax file, we created one or more Linkage Keys. We excluded Linkage Keys associated with multiple TAXIDS and those associated TAXIDs (803 148; representing 1.9% of all TAXIDs). For example, same-sex multiple births (e.g. twins) residing at the same address (sharing the same postal code) represent the records dropped from the linked data because of their non-unique linkage key. Failure to do this would risk confusing multiple birth siblings, thereby creating false inks between the hospital and tax files. Of all hospital records, 98.5% (n = 54 553 113) were eligible for linkage, comprising 25 244 559 Linkage Keys and 18 747 756 HINs. Most HINs (75.6%) were associated with a single Linkage Key; another 17.7% were associated with two Linkage Keys. Of all TAXIDs, 98.1% (n = 42 482 982) were eligible for linkage, comprising 268 088 957 Linkage Keys. Approximately 12% of TAXIDs were associated with a single linkage key; on average each TAXID had 6.3 associated linkage keys. Initial HIN-TAXID links were created using an exact deterministic match on all three Linkage Key components. Unique (1: 1) links between HINs and TAXIDs were accepted, whereas non-unique links were rejected as likely to be erroneous. Finally, HIN-SIN links were created by extending the unique HIN-TAXID links to all SINs associated with a given TAXID (Figure 1B). (A) Record processing preparatory to linkage between the Hospital File and the Tax File. (B) Data linkage between Health Insurance Numbers in the Hospital File, and TAXIDs in the Tax File. Among the starting count of HINs, no matching Linkage Key(s) could be found in the Tax File for 11.1% (n = 2 085 613), and another 2.8% (n = 520 079) matched non-uniquely with TAXID(s) and were therefore but rejected. Thus 86.1% (n = 16 142 064) of HINs were successfully linked to a TAXID; the results were similar when considered at the level of individual hospital records, with 88.4% (n = 48 203 982) of individual hospital records being successfully linked to the Tax File. We first estimated the positive predictive value (PPV) of accepted HIN-TAXID links, in comparison with manual review of a random sample of such links, stratified by sex and four categories of age (0–14, 15–19, 20–29 and over 30) reflecting differences in the rate of representation in the tax data. Three independent reviewers conducted manual reviews, with any disagreements resolved by majority rule. To determine the true linkage status of a given pair, reviewers were provided with: (i) all Linkage Keys from the hospital and tax records; (ii) all identifying information associated with the TAXID; (iii) HINs from the hospital records; and (iv) Linkage Key matching information for the pair (1:1, 1:many, or many: many). We calculated that 500 manual reviews across all strata would be required for an expected false-positive rate of 0.05%; because of this very small proportion, we used a targeted overall coefficient of variation of 200% which allowed for the detection of a false-positive rate of up to 0.40% using a reasonable sample for manual review. Point estimates and 95% confidence intervals were calculated using weightings from the stratified sampling, and Clopper-Pearson exact binomial intervals combined across strata using Bonferroni’s method. Using this approach, the PPV for our linkage was estimated to be 97.8% [95% confidence interval (CI) 80.6–100.0). A high PPV ensures that the reported match rate is not achieved by allowing a large proportion of uncertain links.28 Next, we used a similar process to estimate the error rate in rejecting links that were initially identified (Figure 1B). A total sample size of 698 within the stratified sample was calculated to allow detection of a 5% rate of such errors. This analysis estimated that only 2.0% (95% CI 1.1–3.3) of rejected HINs-TAXID links were, in fact, accurate links. Finally, we assessed the consistency of linkage rates for hospital records across sub-groups of fiscal year of hospitalization, sex, province/territory and Most Responsible Hospital Diagnosis; this latter is the diagnosis accounting for the majority of hospital length of stay,19 categorized into ICD-10 chapters.29 Only linkage rates pertaining to hospital records coded using ICD-10 are presented here, as they represent the majority of hospital records included in this linkage (hospital records for fiscal years 2004/05–14/15 used ICD-10); linkage patterns in ICD-9-based hospital abstracts showed similar patterns. For comparison we evaluated linkage rates by birth year, for which the pattern was expected to reflect the likelihood of being a tax filer (i.e. working age). Because our very large sample sizes were expected to produce small P-values for even inconsequential differences in linkage rates, we assessed consistency numerically, not statistically. The linkage rate varied minimally (85.8–89.8%) over study years (Figure 2). Rates were similar for men (89.2%) and women (87.7%). Linkage rates were systematically lower in the territories than the provinces (Table 1). The main difference among provinces was that a minority had substantially higher rates of rejected links, as high as 16.8% in New Brunswick. Linkage rates were >85% across categories of Most Responsible Hospital Diagnosis, except for those related to pregnancy/childbirth/birth disorders and invalid diagnostic codes (Table 2) 0.29 Even for diagnostic categories showing lower linkage rates, these generally rose over the study years (data not shown). Our findings that linkage rates were generally consistently high across time, geography, diagnosis and sex indicates the fitness of this new linked database to support a broad range of research and policy studies related to the labour and income impacts of acute illnesses requiring hospitalization. Linkage rates exceeded 80% for all parts of Canada except the three territories and two small provinces, which altogether comprise 3% of the population.30 Rates exceeded 85% for most hospital diagnostic categories, being lowest related to pregnancy and children; and even this potential limitation was most marked in the earlier study years, improving in later years. Variation in linkage rates (%) of Heath Insurance Numbers (HIN) in Hospital File with TAXIDs in Tax File, fiscal years 1999/2000–14/15 records, by province/territory Variation in linkage rates (%) of Heath Insurance Numbers (HIN) in Hospital File with TAXIDs in Tax File, fiscal years 1999/2000–14/15 records, by province/territory Variation in linkage rates (%) of Heath Insurance Numbers (HIN) in Hospital File with TAXIDs in Tax File, by Most Responsible Hospital Diagnosis, Acute-care hospitals, fiscal years 2004/05–14/15 Variation in linkage rates (%) of Heath Insurance Numbers (HIN) in Hospital File with TAXIDs in Tax File, by Most Responsible Hospital Diagnosis, Acute-care hospitals, fiscal years 2004/05–14/15 Linkage rates (%) of Heath Insurance Numbers (HIN) in the Hospital File with TAXIDs in the Tax File, by fiscal year. Finally, there was substantial variation in linkage rate by birth year (Table 3). As expected from their predictably low rates of filing taxes during the study years, lower rates of linking hospital records to tax data were observed among the most elderly (born before 1910) and children (born after 1990). Lower rates among the elderly could also be attributable to: (i) incorrect birth dates; and (ii) difficulty matching postal codes due to tax forms containing postal codes representing long-term care institutions, or belonging to those who filed taxes on behalf of elderly persons. Lower linkage rates among children could also be due to: (i) sharing of HINs between mothers and babies in obstetric hospitalization records, a common practice in New Brunswick, Prince Edward Island, Newfoundland and the territories in earlier years;31 and (ii) the lack of requirement that Canadian children have an SIN. The sharing of HINs between mothers and newborns could also explain the lower observed linkage rates of pregnancy-related hospitalizations in the earlier years of the hospital data, which improved over time. Variation in linkage rates (%) of Discharge Abstract Database records, by date of birth Variation in linkage rates (%) of Discharge Abstract Database records, by date of birth This new data resource enables evaluation of the effect of any type of acute illness or injury requiring hospitalization on the ability to work and earn. Under the same funding umbrella that was used to create this database, we are currently performing an analysis of how acute myocardial infarction, cardiac arrest and stroke affect work and earnings in the 3 years after the event: for both those who had these health events and (using the identification of marriage in the tax data) their spouses. The method combines matching individuals with and without these acute health shocks, and performing difference-in-difference regression to assess how these health events affect the outcomes. Preliminary results demonstrate that compared with matched controls, individuals experiencing these health shocks are 5–20 percentage points less likely to work in the 3rd year after the event than the control group. The annual earnings of these health shock survivors are 8–30% lower. Among the three events, stroke has the largest effects on the labour market outcomes among survivors.32,33 This linked dataset could also be used in the reverse direction, to study the impact of income, or changes in income, on health as measured by hospital-related events. This new data resource has seven main strengths. First, it is a national population-based linked dataset (excluding Quebec). Second, it will be straightforward as time goes by to update to include more recent tax and hospital data—although this will require additional funding. Third, the range and breadth of information contained in the component datasets allows for numerous types of detailed investigations. Although the most common use may be among working people experiencing health shocks, inclusion in the tax data of sources of income other than wages and salaries means that it can be used across age groups and for those not working. Fourth, the longitudinal tax data allow for assessing before vs after differences in work and income related to the health shock. Fifth, this population-based database allows identification of matched control groups which did not have the health shock of interest. These last two factors makes it possible to combine matching with difference-in-difference analysis, a powerful combination that reduces bias due to both observed and unobserved characteristics.34,35 Sixth, the database is not restricted to specific health conditions, but includes the data for all causes of acute hospitalization. Therefore, it can be used to provide information on downstream wage and employment effects for any acute condition leading to hospitalization. Such information might be useful for augmenting cost-effectiveness analyses or estimates of burden of disease. Seventh, its longitudinal and unselected nature allows for modelling the dynamic and reciprocal interplay between income, employment and ill health requiring hospitalization, in the population at large, in social subgroups of particular interest, and even between members of the family. There are also four limitations. The first is that we took a conservative approach towards linkage, seeking to minimize false-positive links between HINs and TAXIDs. Strategies towards this goal included: (i) requiring exact matches on all three components of the Linkage Key; (ii) when a single Linkage Key was associated with multiple TAXIDs, excluding those TAXIDs; and (iii) the for rejecting links. These have for the use of the data to study specific For example, living at the same postal are excluded. links mothers and babies were not resolved but the use of these linked data for pregnancy-related and hospital events, at for earlier years. Our quality analysis that these a high PPV and that only of rejected links were Second, the linked data the of which contains of the national and from it with of the Third, there are some age in using the linked data. Although a large majority of Canadians tax this rate is lower for being only 64% for those 18 years assessing between earnings and acute health shocks in is limited by the that in that age do not have earnings. Finally, the linkage includes all hospitals to the national hospital database, only acute care hospitals are consistently and and of services and hospitalizations are The linked tax and hospital data are for use at Statistics Canada by approved For more information a linkage of population-based acute hospital abstracts and individual tax data in Canada. was created to support analyses of the economic effects on individuals of acute health events requiring hospitalization. contains longitudinal tax records for 16 142 064 individuals who experienced 48 203 982 hospitalizations excluding Quebec for the interval and Manitoba for fiscal years links the Canadian Discharge Abstract Database (DAD) for fiscal years and the T1 Family File (T1FF) containing tax data for There was no all individuals with data were includes all elements of the DAD and The hospital data (i) details of hospital admission, discharge, disposition and transfers; (ii) demographic information; (iii) medical services involved in care; (iv) use of care; (v) and procedures The tax data identify all sources of taxable income, including and other retirement income, and social including tax or child tax benefits. is for use at Statistics Canada by approved For more information This work was by the and of Canada and the Research Manitoba of
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.027 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.017 | 0.041 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.132 | 0.044 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".