MétaCan
Menu
Back to cohort
Record W3097986273 · doi:10.1136/jech-2020-214259

Overview of retrospective data harmonisation in the MINDMAP project: process and results

2020· review· en· W3097986273 on OpenAlexafffund
Tina W. Wey, Dany Doiron, Rita Wissa, Guillaume Fabre, Irina Motoc, J. Mark Noordzij, Milagros Ruiz, Erik J. Timmermans, Frank J. van Lenthe, Martin Bobák, Basile Chaix, Steinar Krokstad, Parminder Raina, Erik R. Sund, Mariëlle A. Beenackers, Isabel Fortier

Bibliographic record

VenueJournal of Epidemiology & Community Health · 2020
Typereview
Languageen
FieldEnvironmental Science
TopicHealth, Environment, Cognitive Aging
Canadian institutionsMcMaster UniversityImpactMcGill University Health Centre
FundersCanadian Institutes of Health ResearchHorizon 2020 Framework ProgrammeNederlandse Organisatie voor Wetenschappelijk OnderzoekCanada Foundation for Innovation
KeywordsSet (abstract data type)Process (computing)Data collectionMedicineData setMinimum Data SetData scienceVariable (mathematics)Computer scienceStatistics

Abstract

fetched live from OpenAlex

BACKGROUND: The MINDMAP project implemented a multinational data infrastructure to investigate the direct and interactive effects of urban environments and individual determinants of mental well-being and cognitive function in ageing populations. Using a rigorous process involving multiple teams of experts, longitudinal data from six cohort studies were harmonised to serve MINDMAP objectives. This article documents the retrospective data harmonisation process achieved based on the Maelstrom Research approach and provides a descriptive analysis of the harmonised data generated. METHODS: A list of core variables (the DataSchema) to be generated across cohorts was first defined, and the potential for cohort-specific data sets to generate the DataSchema variables was assessed. Where relevant, algorithms were developed to process cohort-specific data into DataSchema format, and information to be provided to data users was documented. Procedures and harmonisation decisions were thoroughly documented. RESULTS: The MINDMAP DataSchema (v2.0, April 2020) comprised a total of 2841 variables (993 on individual determinants and outcomes, 1848 on environmental exposures) distributed across up to seven data collection events. The harmonised data set included 220 621 participants from six cohorts (10 subpopulations). Harmonisation potential, participant distributions and missing values varied across data sets and variable domains. CONCLUSION: The MINDMAP project implemented a collaborative and transparent process to generate a rich integrated data set for research in ageing, mental well-being and the urban environment. The harmonised data set supports a range of research activities and will continue to be updated to serve ongoing and future MINDMAP research needs.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.394
metaresearch head score (Gemma)0.440
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: none
Teacher disagreement score0.394
Threshold uncertainty score0.748

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.3940.440
Meta-epidemiology (narrow)0.0030.003
Meta-epidemiology (broad)0.0020.005
Bibliometrics0.0110.015
Science and technology studies0.0040.004
Scholarly communication0.0110.006
Open science0.0070.021
Research integrity0.0030.005
Insufficient payload (model declined to judge)0.0150.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.474
GPT teacher head0.517
Teacher spread0.043 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations31
Published2020
Admission routes2
Has abstractyes

Explore more

Same venueJournal of Epidemiology & Community HealthSame topicHealth, Environment, Cognitive AgingFrench-language works237,207