MétaCan
Menu
Retour à la cohorte
Enregistrement W4390153996 · doi:10.1093/ije/dyad175

Cohort Profile: The Cardiovascular Research Data Catalogue

2023· article· en· W4390153996 sur OpenAlexafffundabout
Jaakko Reinikainen, Tarja Palosaari, Alejandro J Canosa-Valls, Carsten Oliver Schmidt, Rita Wissa, Sucharitha Chadalavada, Laia Codó, Josep Lluis Gelpí, Bijoy Joseph, Aad van der Lugt, Elsa Pacella, Steffen E. Petersen, Esmeralda Ruiz Pujadas, Liliána Szabó, Tanja Zeller, Teemu Niiranen, Karim Lekadir, Kari Kuulasmaa

Notice bibliographique

RevueInternational Journal of Epidemiology · 2023
Typearticle
Langueen
DomaineEnvironmental Science
ThématiqueHealth, Environment, Cognitive Aging
Établissements canadiensMcGill University Health Centre
Organismes subventionnairesCanadian Institutes of Health ResearchEuropean CommissionDeutsches Zentrum für Herz-Kreislaufforschung
Mots-clésMedicineCohortCohort studyInternal medicine

Résumé

récupéré en direct d'OpenAlex

The Cardiovascular Research Data Catalogue was established to provide a centralized web-based catalogue of cohort- and variable-level metadata with a focus on cardiovascular diseases to facilitate the discovery, sharing and co-analysis of different data types and sources needed in cardiovascular research. Currently, the catalogue contains information on 33 individual cohort studies and 2 harmonization initiatives. The documented cohort studies involve data from 16 countries and >1 million participants. The earliest measurements were carried out in 1972 and the most recent study started in 2016. The cohort studies have several data sources and data types, such as physical measures, biosamples, questionnaires, administrative databases, cognitive measures, imaging, omics and electrocardiogram data. All metadata entered in the catalogue as well as the access information for person-level data are freely available without need for authentication. The catalogue is openly accessible through the European Society of Cardiology website at https://www.escardio.org/Research/Cardiovascular-data. Data from a single study are rarely sufficient to comprehensively understand a phenomenon of interest or to build predictive models with statistical power. Meta-analyses, including aggregated data for several cohort studies, can be carried out to synthesize published results, but the use of harmonized individual-level data from different studies is preferable to reach the maximal comparability across the studies. A significant barrier to the co-analysis approach is finding relevant studies and accessing their data. Such barriers have, amongst others, led to designing the FAIR (Findable, Accessible, Interoperable and Reusable) principles, which have been recognized internationally.1 Currently, it is still difficult to obtain a comprehensive overview of available data from existing observational studies on cardiovascular diseases. Because there is no requirement to register studies other than clinical trials, information on existing studies and their characteristics is often lacking or highly fragmented. In addition, it is often unclear how to request access to the raw data or what kind of ethical, legal and societal issues (ELSI) may be related to the use of the data. All these factors are substantial hindrances to research in the cardiovascular domain and beyond. Hence, a first and critical step to improve the utilization of the data from existing studies is to facilitate their discovery and subsequent access requests. Data searches are burdensome if metadata are not publicly available or if they are available only on study-specific websites. Once suitable data sources have been identified, the next challenge for researchers is to find information on how to request access (including access conditions, contact points, access fees and procedures). Data catalogues have been created to overcome these problems. For instance, Maelstrom Research has developed a metadata cataloguing toolkit, which provides services to present information at the study and variable levels.2 The portal of Medical Data Models has been designed to foster sharing of medical data models by describing information systems at the variable level, such as data elements regarding previous diseases.3 The International Clinical Trials Registry Platform provides information about clinical trials,4 but information is limited to the study level. There are also disease-specific catalogues and analysis platforms. For example, Dementias Platform UK has a data portal focusing on cohorts for dementia research and allows remote data analyses.5 Brain-CODE is a Canadian neuroinformatics platform designed for data management and analysis in neuroscience.6 The Childhood Cancer Data Catalog provides study-level information on paediatric oncology data resources.7 An internationally centralized and structured data catalogue on existing cardiovascular disease research studies has still been lacking. Therefore, the euCanSHare (EU–Canada joint infrastructure for next-generation multi-study heart research, http://www.eucanshare.eu/) project has set up a highly rich and easy-to-use cardiovascular research data catalogue, which currently comprises detailed and easily accessible information on 33 individual studies and 2 harmonization initiatives. The Cardiovascular Research Data Catalogue is included in the euCanSHare platform (https://eucanshare.bsc.es) (Figure 1), which is hosted by the Barcelona Supercomputing Center (BSC). Furthermore, it is openly accessible through the European Society of Cardiology (ESC) website at https://www.escardio.org/Research/Cardiovascular-data. Its implementation is based on the OBiBa suite8 (www.obiba.org)—a software suite developed by Maelstrom Research (www.maelstrom-research.org). The OBiBa suite consists of several software applications, three of which were used in this project: Opal, Mica and Agate. Opal is a web-based database application that allows data managers to securely import and export data in multiple formats and data types, and to store and display these data. Browsing and data discovery in the catalogue use the Mica web portal. User authentication, user profile management and e-mail notifications to the Opal and Mica applications are managed using Agate. The Agate application in the euCanSHare project has been complemented with an OpenID Connect interface to allow authentication through providers such as the European Life Science Research Infrastructures Login and to offer a single sign-on for the available services. Schematic diagram of the main components within the euCanSHare platform. Components available for users are shown in the upper part of the diagram on a blue background and under them their technical implementation on a green background. In yellow cylinders, the external data repositories linked from the euCanSHare platform. EGA, European Genome-phenome Archive; EuBI, Euro-BioImaging; openID, open identifier standard; openVRE, open Virtual Research Environment The first version of the platform was released in April 2021. Currently, the platform provides a centralized documentation of cohort metadata, questionnaire data, physical measures, imaging data, omics, biosamples and ELSI information with a focus on cardiovascular diseases. Thus, it facilitates the discovery, sharing and co-analysis of different data types relevant for cardiovascular research in accordance with the FAIR principles. No authentication is needed for exploring and searching studies and metadata. To promote open science in cardiovascular research, we describe the data catalogue and its features in this article, as well as the studies and cohorts currently included. The Cardiovascular Research Data Catalogue provides a structured, searchable documentation of data originating from different target populations (such as general population, ethnic minority or patient group) and different data types and sources. Currently, the studies in the catalogue are observational cohort studies, but in principle there are no restrictions regarding the study design, geography, measurement period or any individual-level characteristics for studies to be included in the catalogue, apart from being relevant for cardiovascular research. The catalogue has detailed study descriptions including objectives, study design, contact and data access information, timeline of data collection events, and information on participants and data sources. A significant output of the euCanSHare project is a versatile search engine tailored specifically for cardiovascular data, which helps to find suitable studies. Definitions of variables, together with their categories or measurement units, have been documented, which facilitates evaluating the harmonization potential of studies. Numbers of missing data per variable are available for most of the studies. Some studies have also reported reasons for missingness, e.g. ‘not applicable’ for the year when stopped smoking among never smokers or ‘missing by design’ for a pre-defined subset of participants for whom the data were intentionally not collected. There are currently 33 individual studies in the catalogue and 2 harmonization initiatives that already have data harmonized to be as comparable as possible for joint analysis. Table 1 shows the countries, numbers of participants and start years of data collection in the individual studies described in the catalogue. Here, we will call the studies by their acronyms, which have been explained in Table 2. Both individual studies and harmonization initiatives have been listed in Table 2 including information on whether the study has been harmonized with respect to the existing MORGAM (MONICA Risk, Genetics, Archiving and Monograph Study) standard.9 Characteristics of the individual studies in the catalogue in September 2023a The catalogue has more detailed study-level information than this summary and many of the studies have locally more information than what has been entered into the catalogue. ECG, electrocardiogram. Study acronyms are listed in Table 2. Study acronyms and full names The study has been harmonized to the MORGAM Project. MONICA, Multinational monitoring of trends and determinants in cardiovascular disease. The included studies have been conducted in several European countries as well as in Canada, Australia and the Asian part of Russia. Detailed metadata are freely available in the catalogue at https://mica.eucanshare.bsc.es/. The catalogue does not contain the raw person-level data, but information on access procedures is described on each study description page. This information includes access conditions and procedures as well as the study representative’s contact information. All the studies currently in the catalogue are cohort studies. Most of them are based on the general population and some are patient cohorts. Two of the studies are population-based biobanks: UK Biobank (UKBB) and Estonia study. Caerphilly and PRIME studies involve only men and Alpha-Tocopherol, Beta-Carotene Lung Cancer Prevention (ATBC) study consists only of smoking men. CAHHM includes a cohort that is focused on the First Nations in Canada and another one on Chinese Canadians and South Asian Canadians. StenoCARDIA is composed of patients with suspected acute myocardial infarction, whereas AtheroGene includes patients with documented coronary artery disease with angiographical diagnosis. BHFC contains heart failure patients with an ejection fraction of <50% and LIFE Heart study has patients with suspected coronary heart disease. Duration and timing of data collection events vary considerably between the studies. The data collection periods of the individual studies are presented in Figure 2. The oldest study (FINRISK) started in 1972 and the most recent (HCHS) started in 2016. In most of the studies, continuous follow-up for death and onset of diseases is based on linkage with administrative databases whereas, in some studies, the disease incidence is identified through periodic follow-up contacts. Participants were followed up mainly through national or regional population or death registries and different disease-specific registries such as the MONICA coronary event and stroke registers. Disease events are obtained from hospital records or local health authorities in many studies conducted in Germany and Italy. Data collection periods of the individual studies, including all data sources, in the catalogue in September 2023. Open end points mean that the follow-up is still ongoing, but also other studies may be extended in the future. Study acronyms are listed in Table 2 Eighteen studies re-examined their cohorts (including studies e.g. from the Nordic countries, UK, Italy and Germany) and three studies (AtheroGene, ESTHER and PRIME) have re-contacted the participants only by questionnaires. Twelve studies (including e.g. the non-European studies) have not contacted the participants directly after the baseline examination but only used the above-mentioned passive follow-up using registry linkage or hospital records. When studies extend the follow-up of the participants, the information in the catalogue is updated accordingly by the representatives of the studies. In addition to morbidity and mortality follow-up, re-examinations have been carried out for the participants after the baseline measurements in many studies. Some studies encompass various cohorts, which have been sampled independently of each other. For example, FINRISK has collected data in 5-year intervals between 1972 and 2017 on new cohorts in which only some participants have been re-examined by chance. Table 1 shows the data sources and selected data types currently available in each individual study in the catalogue. All studies have data from physical measures and biosamples, and most also have data from questionnaires and administrative databases. Fifteen studies have reported having cognitive measures data available, 13 studies have imaging data, 16 studies have omics data and 21 studies have electrocardiogram (ECG) data. Common physical metrics include anthropometric and blood pressure measures. Cognitive measures include tests such as the Digit Symbol Substitution, MoCA (Montreal Cognitive Assessment) test, reaction time tests, hearing and vision tests, and various memory tests. Medical history at the baseline is usually obtained from self-reports or different administrative databases, such as healthcare registries and reimbursements for prescription medicines registries. Besides study-level data, the catalogue also allows users to search for information by applying variable-level criteria in their queries. The areas of information used to perform variable-level searches are based on the standardized Maelstrom Research variable classification taxonomies, which include a wide variety of categories, e.g. socio-demographic characteristics, lifestyle, diseases, and physical and laboratory measures. In addition, a refined classification of cardiovascular system-related diseases (such as hypertensive diseases, ischaemic heart diseases and cerebrovascular diseases) has been created for the needs of cardiovascular research. It is also possible to do word searches for variable names and labels. Figure 3 presents frequencies of areas of information for a subset of all the categories available in the catalogue. From cardiovascular disease variables, ischaemic heart diseases and variables combining multiple diseases are covered the most (30 studies). In addition, diabetes variables (e.g. self-reported diabetes diagnosis, treatment for diabetes and onset of different types of diabetes during follow-up) have been described in 30 studies. This condition is considered a cardiovascular-related disease due to its high relevance in cardiovascular research. A broad scope of information on relevant risk factors is also available, such as laboratory measures (e.g. blood lipids, glucose, and hormone and vitamin levels in serum), circulation and respiration measures (blood pressure, ECG) and anthropometric variables (e.g. measured height, weight, and waist and hip circumference) have been described in 31 studies. Number of studies that have reported to have variables belonging to different areas of information. In addition to these, there are many other variable categories in the catalogue Currently, the euCanSHare project contains metadata that describe 33 individual cardiovascular studies. Other studies are being added and the catalogue is expected to continue to grow over the years. To this end, a user interface with tutorials has been built to guide future researchers and data managers so they can add their studies (i.e. metadata) to the catalogue and hence increase visibility for their data. Several research projects have already used the catalogue. A study with UKBB data had previously demonstrated that the relative risk of heart failure (HF) is higher in women with type 1 diabetes than in men.10 To externally validate these findings, five studies with suitable data (DAN-MONICA, FINRISK, Moli-sani, NSW-MONICA and SHHEC) were identified from the Cardiovascular Research Data Catalogue. The validation study confirmed that diabetes was associated with increased risks of HF, but there was no evidence of a difference in the relative risk according to sex.11 The catalogue was also utilized in a case study of atrial fibrillation (AF) for machine-learning-based diagnosis and risk prediction. The identification of the cohorts with the needed data, including electrocardiogram (ECG), clinical data and cardiac magnetic resonance (CMR) imaging, was facilitated by the catalogue and its search engine. As a first step, machine-learning models for AF prediction were built based on the UK Biobank data.12 The combination of the ECG and CMR imaging resulted in a novel insight into AF-related electro-anatomic remodelling and its variation by sex. The conclusion was that a combination of ECG and radiomics models could improve the future diagnostic options of women at risk of AF. Subsequently, other cohorts identified in the Cardiovascular Research Data Catalogue, namely the SHIP and HCHS cohorts, are being used to externally validate the risk prediction models. Another AF study identified vascular risk factors and standard CMR indices through the catalogue. This study described that the combination of radiomics, standard CMR indices and risk factors provided the best performance for prediction of incident AF.13 However, the performance was similar to vascular risk factors with radiomics. Other cohorts will be searched to externally validate the results and harmonization will be performed using the tools provided in the catalogue to simplify the process. By providing a centralized and structured documentation together with a search engine, the Cardiovascular Research Data Catalogue facilitates discoverability, accessibility and the reuse of cardiovascular research data. Versatile options for search criteria both on the study and variable levels enable researchers to easily find relevant data for their purposes. We have focused on the user-friendliness of the catalogue by developing and optimizing the user interface and tutorials based on feedback from several researchers within and outside the consortium. An ESC short course and three webinars on the Cardiovascular Research Data Catalogue are also accessible as recordings on the ESC online platform (https://esc365.escardio.org/), which may be used to learn how to use the many features of the catalogue. A full list of the educational content is also available at http://www.eucanshare.eu/learn-more-on-the-use-of-the-cardiovascular-research-data-catalogue/. The contents of the catalogue are easily extendible. Representatives of the studies already documented can update the information when new data have been collected. Also, the catalogue is open for other studies that have data which can be used for multi-cohort cardiovascular research, including new emerging studies. The euCanSHare platform at Barcelona Supercomputing Center also provides a virtual research environment, which combines tools for accessing the granted data and their management, visualization and analysis. It is powered by a cloud-based computational infrastructure in which the consortium has made available a set of tools for data-quality assessment in epidemiological studies, radiomics applications for cardiac medical images and bioinformatics analysis. After authentication, the platform provides a private workspace in which a researcher can either upload data sets or import granted-access data sets with one click from well-known data repositories associated with the consortium, such as Euro-BioImaging14 or the European Genome-phenome Archive.15 Finally, the Cardiovascular Research Data Catalogue is compliant with data privacy requirements, as it only contains metadata and does not store any raw data. This enables data owners to promote their studies while maintaining full control over their data. For the time being, the catalogue does not describe every variable collected in the described studies. Rather, for many of the studies, the variables are described to the extent to which they have been harmonized to the MORGAM Project.9 Continuous efforts are needed from the data providers to keep the information up to date. Although the catalogue facilitates access to the data by providing information about the access options and contacts, the procedures for requesting and granting access to the data vary from study to study and researchers must contact the studies individually. All the metadata entered in the catalogue can be freely accessed through the ESC website at or directly at https://mica.eucanshare.bsc.es/. the study description of the catalogue, there are data access that the conditions and procedures for accessing the person-level study data. There are also to the which may have information (e.g. access conditions or and contact information of can researchers to access the individual-level data. about the catalogue can be to The studies described in the catalogue with the of all studies had been by local and was obtained from all included. For any new use of the data, possible need for needs to be from the data of the the All to the of the Cardiovascular Research Data Catalogue and to the and critical of this This was by the European research and under Canadian are by the Canadian of Research and the is by the for Cardiovascular Research The to all the studies for providing their metadata, the Maelstrom Research for technical and external researchers for providing feedback to improve the of the catalogue. provides for Cardiovascular is listed as a of an on the use of a to the of myocardial Number is a of the All other no of

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,042
score de la tête « metaresearch » (Gemma)0,013
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche, Charge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesMétarecherche
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,181
Score d'incertitude au seuil0,999

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0420,013
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,001
Communication savante0,0000,000
Science ouverte0,0020,001
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0010,002

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,286
Tête enseignante GPT0,462
Écart entre enseignants0,176 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations2
Publié2023
Routes d'admission3
Résumé présentoui

Explorer davantage

Même revueInternational Journal of EpidemiologyMême sujetHealth, Environment, Cognitive AgingTravaux en français237 207