Cohort Profile: The Care Trajectories—Enriched Data (TorSaDE) cohort
Bibliographic record
Abstract
The Care Trajectories—Enriched Data Cohort (TorSaDE) aims to identify the most effective and efficient care trajectories for people with chronic conditions. The TorSaDE cohort includes all Québec (Canada) residents aged ≥12 years old who participated in at least one of the 4 cycles (2007–08, 2009–10, 2011–12 and 2013–14) of the Canadian Community Health Survey (CCHS) and agreed to share the information collected in this survey with provincial and territorial ministries of health, the Institut de la Statistique du Québec (ISQ), Health Canada, and the Public Health Agency of Canada for research purposes (n = 81 093 distinct participants). The data contains all the information available in the CCHS questionnaire including sections on health status and reported health problems, lifestyle, prevention, use and access to health services, and sociodemographic characteristics. These responses to the CCHS questionnaires are linked with the participants’ medico-administrative data over a period of 21 years (1996–2016), and includes information on hospitalizations, emergency department visits, medical visits, interventions in local community service centres, prescription drugs, and date and cause of death. Researchers interested in using the cohort data must submit their project for approval to the TorSaDE Working Group and obtain the necessary authorizations. Please contact A.V. with any enquiries. In an aging population, chronic diseases are a growing burden on societies.1 For instance, cardiovascular disease, cancer, respiratory conditions and diabetes are responsible for >4 out of 5 premature deaths worldwide.2 In Canada, 67% of direct healthcare costs are attributable to chronic conditions whereas 44% of adults have at least one of the 10 most common chronic conditions.3,4 Patients diagnosed with chronic conditions have unique care trajectories, dependent upon the type of condition, the characteristics of the individuals and the organization of services. In this paper, a care trajectory is defined as a pattern of care and resource utilization over time (e.g. a pregnant woman who visits her physician at day 1, the imaging centre at day 3 and the mid-wife at day 7). Our definition is based on the ‘6 W’ multidimensional model of care trajectories. This model considers the patients' attributes (‘who’), the evolution of their chronic conditions (‘why’), their patterns of care use across care providers (‘which’), care units (‘where’), and treatments (‘what’), at specific periods (‘when’).5 Care trajectories are an important concept as they directly impact the appropriate use of health services, such as emergency department (ED) visits or hospitalizations, as well as patients’ health. In-depth understanding of those care trajectories, revealing their impact on patients’ health and the healthcare system, will faciliate determining optimal models of healthcare for the prevention and management of chronic diseases. However, describing care trajectories is a complex task as interactions with the healthcare system are numerous and varied, especially among people with chronic conditions. In recent years, care trajectories have been increasingly studied using medico-administrative data,6,7 although very few have proposed an operational and complete definition of care trajectories that takes into account all dimensions. This kind of study has been made possible through the pooling of medico-administrative data over a long period of time from different sources, such as outpatient and hospitalization data, combined with sociodemographic and geographic data. However, medico-administrative data may lack relevant sociodemographic characteristics (e.g. education level, income or gender) as well as other factors potentially related to care trajectories such as patients' lifestyle habits, risk factors, self-reported diseases and perceived health. On the other hand, that information is collected in the Canadian Community Health Surveys (CCHS), allowing it to be taken into account in the analysis of care trajectories by linking the information contained in the CCHS to medico-administrative data. This data linkage is also a way to create this type of cohort by starting with a representative sample of the population for which self-reported information is available and then linking it to medico-administrative data for each participant.8 Canada offers a universal healthcare system (i.e. it is paid for through taxes) and each province and territory has its own health insurance plan. In the province of Québec, the cohort Care Trajectories—Enriched Data (TorSaDE; derived from French: Trajectoire de Soins—Données Enrichies) was created to identify the most effective and efficient care trajectories for patients with chronic conditions, and to help decision-makers, stakeholders and clinicians in the planning and organization of healthcare services. The specific objectives of the TorSaDE cohort are 2-fold: (i) describing care trajectories of patients with chronic conditions; and (ii) measuring the impact of care trajectories on their health (e.g. by investigating which care trajectories are linked to better perceived health). The TorSaDE cohort includes all residents in the province of Québec, Canada, who participated in one of the four cycles of the general CCHS conducted between 2007 and 20149 and agreed to share the information collected in this survey with provincial and territorial ministries of health, the Institut de la Statistique du Québec (ISQ), Health Canada and the Public Health Agency of Canada for statistical purposes. Québec is a French-speaking province (about 79% of residents’ first language is French, as opposed to <5% in the rest of Canada) and is the second largest province in Canada with a population exceeding 8 400 000. The CCHS is a biennial cross-sectional survey conducted by Statistics Canada, which collects comprehensive information on Canadians, including social, economic, behavioural, health status, health services utilization and determinants of health from a representative sample of the population aged ≥12 years who are living in private dwellings in the ten provinces and three territories in Canada. Canadians living on Indian Reserves or Crown lands, residing in institutions, full-time members of the Canadian Forces and residents of certain remote regions were excluded from the survey.9 The CCHS covers >95% of the Canadian population aged ≥12 years.9 The objective of each CCHS biennial survey was to provide reliable estimates related to health status, healthcare utilization and health determinants at the national, provincial and health region (HR) level; the TorSaDE cohort is thus representative of the Québec population aged ≥12 years. To achieve this goal, a sample of 130 000 respondents across Canada was required biannually. A multi-stage sample allocation strategy gave relatively equal importance to the HRs and the provinces. The sample was allocated among the provinces according to the size of their respective populations and the number of HRs they contain.9 Depending on the cycle, between 90 and 95% of participants agreed to share the CCHS information. For these participants, the Québec components of the CCHS surveys were linked to the medico-administrative data using a probabilistic linkage, which was possible for >95% of them. The following variables were used for the linkage: Health Insurance Number provided by the Régie de l’Assurance Maladie du Québec (RAMQ), surname and first name, sex and date of birth. The numbers of Québec participants in each of the four CCHSs compared with the number of participants included in TorSaDE (∼90% of the CCHS sample for each cycle) are displayed in Table 1. In total, there are 81 744 participants in the TorSaDE cohort. However, some participants responded to more than one CCHS, resulting in 81 093 distinct participants. Number of participants within CCHS and TorSaDE This is the number of records in Statistics Canada's share files for Quebec participants. It includes only respondents who consented to have their data shared with another organization, namely ISQ. Sharing consent rates are ∼96% for the CCHS cycles. Includes participants who agreed for the use of their responses and the linkage with data from other sources for research purposes (according to the CCHS cycle, between 90 and 95% of respondents agreed) and whose linkage with health-administrative data was possible (>95% of cases). Number of participants within CCHS and TorSaDE This is the number of records in Statistics Canada's share files for Quebec participants. It includes only respondents who consented to have their data shared with another organization, namely ISQ. Sharing consent rates are ∼96% for the CCHS cycles. Includes participants who agreed for the use of their responses and the linkage with data from other sources for research purposes (according to the CCHS cycle, between 90 and 95% of respondents agreed) and whose linkage with health-administrative data was possible (>95% of cases). Participants included in the TorSaDE cohort are those who responded to one of the CCHS questionnaires (between 2007 and 2014). Although there was no active follow-up of participants, the cohort is a mixture of a cross-sectional survey combined with retrospective and prospective health information, from medico-administrative data that covers a 21-year period (between 1996 and 2016) for each of the participants (Figure 1). For example, for those who completed the 2007 CCHS, the TorSaDE cohort contains retrospective information on their healthcare utilization, from medico-administrative data over an 11-year period (1996–2006) prior to their participation in CCHS and a prospective period of 10 years (2007–16) after their participation. For participants of the 2014 CCHS, however, retrospective medico-administrative data were available for 19 years (1996–2014), whereas the prospective data covers only 2 years (2015 and 2016). Cross sectional survey dates and follow-up period with health-administrative data The databases included in the TorSaDE cohort are depicted in Figure 2. All the information in the CCHS questionnaires is available for each TorSaDE participant. The CCHS included sections on health status and reported health problems, lifestyle, prevention, use and access to health services, and sociodemographic characteristics (Table 2). Table 3 describes the medico-administrative data, which includes hospitalizations, ED visits (data from 2014 onward), medical visits, interventions from local community service centres (CLSCs) (data from 2012 onward), prescription drugs (for persons insured under the Québec public drug insurance plan), and date and cause of death. For each of these healthcare services, details are available regarding the date of the service, the diagnoses associated with the service, the interventions performed by a health professional and the cost of services. Furthermore, to ease the use of medico-administrative data for researchers not familiar with this type of data, the TorSaDE team has produced a number of predefined variables related to the yearly healthcare use [e.g. comorbidity index, number of hospitalizations, number of ED visits, number of general practitioner (GP) visits, number of specialist visits, number of prescription drugs]. Overall, the TorSaDE database contains >1000 variables from CCHS and >200 variables from the medico-administrative databases. Cohort TorSaDE databases. APR-DRG, All Patient Diagnosis Related Groups; BDCU, Joint Emergency Room Database; FIPA, Insured Persons Registration file; MED-ÉCHO, maintenance and exploitation of data for the study of hospital patients List of characteristics measured in the CCHS List of characteristics measured in the CCHS Description of health-administrative database included in TorSaDE Hospitalization (MED-ECHO) Hospitalization (APR-DRG) Medical visits (RAMQ billing) Interventions in CLSCs (I-CLSC) Hospitalization (MED-ECHO) Hospitalization (APR-DRG) Medical visits (RAMQ billing) Interventions in CLSCs (I-CLSC) AHFS, American Hospital Formulary Service Drug Information; APR-DRG, All Patient Diagnosis Related Groups; BDCU, Joint Emergency Room Database; CCI, Canadian Classification of Health Interventions;DIN, Drug Identification Number; ICD 9, International Classification of Diseases, 9th Revision; ICD 10, International Classification of Diseases, 10th Revision; MED-ÉCHO, Maintenance and exploitation of data for the study of hospital patients; NCHS, National Centre for Health Statistics. Description of health-administrative database included in TorSaDE Hospitalization (MED-ECHO) Hospitalization (APR-DRG) Medical visits (RAMQ billing) Interventions in CLSCs (I-CLSC) Hospitalization (MED-ECHO) Hospitalization (APR-DRG) Medical visits (RAMQ billing) Interventions in CLSCs (I-CLSC) AHFS, American Hospital Formulary Service Drug Information; APR-DRG, All Patient Diagnosis Related Groups; BDCU, Joint Emergency Room Database; CCI, Canadian Classification of Health Interventions;DIN, Drug Identification Number; ICD 9, International Classification of Diseases, 9th Revision; ICD 10, International Classification of Diseases, 10th Revision; MED-ÉCHO, Maintenance and exploitation of data for the study of hospital patients; NCHS, National Centre for Health Statistics. The following tables present descriptive statistics of the TorSaDE participants’ data derived from 4 CCHS cycles (2007–14) to summarize their socio-demographic characteristics (Table 4), their health status and risk factors as reported in the survey (Table 5) and their use of health services 1 year after CCHS participation from the medico-administrative health data (Table 6). The purpose of the tables is to provide information to researchers interested in working with these data on whether the sample is large enough (statistical power) to conduct analyses [e.g. the number of people ≥65 years of age or the number with self-reported chronic obstructive pulmonary disease (COPD)]. CCHS (2007–14)a derived sociodemographic characteristics of TorSaDE participants Information reported by the participant during the survey. Unweighted number of people in the cohort. Weighted percentages are used to obtain reliable estimates at the provincial level. 95% confidence intervals using bootstrap weights. Area from Statistics Canada census. CMA, census metropolitan area; CA, census agglomeration; MIZ, census metropolitan influenced zone. CCHS (2007–14)a derived sociodemographic characteristics of TorSaDE participants Information reported by the participant during the survey. Unweighted number of people in the cohort. Weighted percentages are used to obtain reliable estimates at the provincial level. 95% confidence intervals using bootstrap weights. Area from Statistics Canada census. CMA, census metropolitan area; CA, census agglomeration; MIZ, census metropolitan influenced zone. CCHS (2007–14)a derived health status and risk factors of TorSaDE participants Information reported by the participant during the survey. Unweighted number of people in the cohort. Weighted percentage are used to obtain reliable estimates at the provincial level. 95% Confidence intervals using bootstrap weights. CCHS (2007–14)a derived health status and risk factors of TorSaDE participants Information reported by the participant during the survey. Unweighted number of people in the cohort. Weighted percentage are used to obtain reliable estimates at the provincial level. 95% Confidence intervals using bootstrap weights. Health services utilization 1 year following CCHS (2007–14) participationa Information from health-administrative databases. Unweighted number of people in the cohort. Weighted percentages are used to obtain reliable estimates at the provincial level. 95% Confidence intervals using bootstrap weights. For participants admissible to the public drug insurance plan which includes people >65 years of age, people on social assistance and any resident not funded through private/employer insurance plans. Health services utilization 1 year following CCHS (2007–14) participationa Information from health-administrative databases. Unweighted number of people in the cohort. Weighted percentages are used to obtain reliable estimates at the provincial level. 95% Confidence intervals using bootstrap weights. For participants admissible to the public drug insurance plan which includes people >65 years of age, people on social assistance and any resident not funded through private/employer insurance plans. As the TorSaDE cohort was established in 2019, the research is not yet completed and very few papers have been published. Nevertheless, some methodological development has been completed on how to operationalize the measurement of care trajectories, considering the ‘6W’ multidimensional model of care trajectories.5 This was done using a state sequence analysis,10,11 a methodology developed mainly in social sciences for describing and classifying life courses such as professional careers or family pathways,12,13 and adapted to take into account the multidimensional structure of care trajectories. This adaptation is fully described in Vanasse et al.14 and, although it did not use the TorSaDE cohort, it will be one of the options to analyse care trajectories on TorSaDE. Another methodological paper that used the TorSaDE cohort developed a gender index15 (the GENDER Index). This composite index was built using several variables available in the CCHS that were deemed gender-related (e.g. occupation, receiving child support, number of working hours) and resulted in a multidimensional composite score (0–100). Higher score can be associated with being female/having more feminine characteristics. The GENDER Index may be useful to enhance the capacity of researchers using the TorSaDE cohort to conduct gender-based analysis among the working population,16 especially when self-reported gender data are unavailable. Apart from methodological developments, some additional work that used the TorSaDE cohort has been presented in conferences. Lunghi et al.17 aimed to estimate the prevalence of antidepressant drugs 12 months after the response to the CCHS and to identify factors associated with antidepressant prescriptions. King and Strumpf.18 aimed to develop a patient-centered measure of attachment to primary care physicians that can be used in administrative data. This was done using the response to the CCHS question ‘Do you have a regular medical doctor?’ In Henri et al.,19 the objective was to identify the typology of care trajectories in COPD patients using state sequence analysis, and determine the social, environmental and health characteristics of patients within each care trajectory. Chiu et al.20 aimed at comparing the performances of two prediction models for frequent ED users: (A) one using a comorbidity index derived from an administrative database and (B) another using self-perceived health variables. Letarte et al.21 explored the longitudinal exposure to neighbourhood deprivation, using the residential history of the TorSaDE cohort. Using a sequence analysis, respondents were classified according to their deprivation sequence up to 19 years before the CCHS survey. The TorSaDE cohort has many advantages for analyzing the care trajectories of patients with chronic conditions. It includes more than 80 000 participants and is a representative sample of the Quebec population aged ≥12 years. The TorSaDE cohort is a compilation of enriched data from two complementary sources, CCHS and medico-administrative data, facilitating cross-validation (i.e. comparing self-reported diseases from CCHS with medical diagnoses in medico-administrative data) and enabling analysis of care trajectories over a long period of time (up to 21 years). Furthermore, the linking of medico-administrative data with cross-sectional CCHSs will allow for analysis of care trajectories while taking into account the sociodemographic characteristics of individuals as well as other factors potentially related to care trajectories, such as patients' lifestyle habits and risk factors. In addition, information on the self-perceived health or diseases reported by patients in CCHSs could be used to evaluate care trajectories outcomes. Finally, other uses of the cohort are possible, such as ecological studies to investigate the relationships between environmental determinants (e.g. built environment, parks, food supply, etc.) and self-reported conditions and/or medical diagnosis (e.g. obesity, diabetes, etc.). The TorSaDE cohort refers only to the Québec population; its use is and will be restricted to this population. A Canadian cohort would give a more complete picture of care trajectories in Canada and allow comparisons between provinces. However, although Canada ensures universal access to publicly funded health services, healthcare delivery and planning occur mainly at the provincial level and different health legislation in each province does not allow for matching with medico-administrative data for Canada as a whole. One of the weaknesses is whether this cohort has sufficient statistical power to conduct the research of interest. One of the purposes of combining several CCHSs was indeed to increase the statistical power necessary to conduct research on chronic diseases that may be infrequent (e.g. COPD). Nevertheless, the TorSaDE cohort, which is representative of the population of Quebec (>8 000 000), has great potential for use by researchers. Other weaknesses of the TorSaDE cohort are the common challenges with medico-administrative data, e.g. it may be of lifestyle when compared with medical Furthermore, the Québec universal healthcare plan does not some specific (e.g. or medical (e.g. or and (e.g. or treatments necessary following an the CCHS is whereas the medico-administrative data are It can be to use these data in longitudinal However, as the CCHS information, it can be used as a starting or for care trajectories It is not possible to contact the CCHS participants to additional data, as the TorSaDE cohort in to the Statistics Canada Finally, prescription drug data are only available for the of participants by public prescription drug insurance plan adults and residents private allowing only analyses of this of healthcare management of the TorSaDE cohort data has been to the which access to data for researchers at a research centre in and another in Québec and has a of for data access which access to were out in including of more and remote access to the cohort. In addition, the of a research centre in is during the Researchers are to use the cohort data, and to they must submit their project for approval to the TorSaDE Working Group and obtain the necessary from and Québec Information through the Data The has the of the TorSaDE cohort for research purposes. that the approval for will be relatively The TorSaDE cohort data with is to researchers. It a of the databases by the definition of each medico-administrative database and the CCHS with their response A database that the cohort with data can be This researchers to available and to and their before using it in the research centre that the TorSaDE data. Finally, a of allowing the of using medico-administrative data according to health conditions of has been developed and information on and This it possible to in the to a of patients with a health of according to the and for TorSaDE is provided by Canadian of Health Quebec for and Institut de la du
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
| gpt | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".