MétaCan
Menu
Back to cohort
Record W4391320577 · doi:10.3389/fdgth.2024.1329630

INSPIRE datahub: a pan-African integrated suite of services for harmonising longitudinal population health data using OHDSI tools

2024· article· en· W4391320577 on OpenAlexfundno aff
Tathagata Bhattacharjee, Sylvia Kiwuwa-Muyingo, Chifundo Kanjala, Molulaqhooa L. Maoyi, David Amadi, Michael Ochola, Damazo T. Kadengye, Arofan Gregory, Agnes Kiragga, Amelia Taylor, Jay Greenfield, Emma Slaymaker, Jim Todd

Bibliographic record

VenueFrontiers in Digital Health · 2024
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsnot available
FundersInternational Development Research CentreWellcome TrustStyrelsen för Internationellt Utvecklingssamarbete
KeywordsCloud computingComputer scienceMetadataData integrationData scienceInteroperabilityPopulationWorld Wide WebDatabase

Abstract

fetched live from OpenAlex

Introduction: Population health data integration remains a critical challenge in low- and middle-income countries (LMIC), hindering the generation of actionable insights to inform policy and decision-making. This paper proposes a pan-African, Findable, Accessible, Interoperable, and Reusable (FAIR) research architecture and infrastructure named the INSPIRE datahub. This cloud-based Platform-as-a-Service (PaaS) and on-premises setup aims to enhance the discovery, integration, and analysis of clinical, population-based surveys, and other health data sources. Methods: The INSPIRE datahub, part of the Implementation Network for Sharing Population Information from Research Entities (INSPIRE), employs the Observational Health Data Sciences and Informatics (OHDSI) open-source stack of tools and the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) to harmonise data from African longitudinal population studies. Operating on Microsoft Azure and Amazon Web Services cloud platforms, and on on-premises servers, the architecture offers adaptability and scalability for other cloud providers and technology infrastructure. The OHDSI-based tools enable a comprehensive suite of services for data pipeline development, profiling, mapping, extraction, transformation, loading, documentation, anonymization, and analysis. Results: The INSPIRE datahub's "On-ramp" services facilitate the integration of data and metadata from diverse sources into the OMOP CDM. The datahub supports the implementation of OMOP CDM across data producers, harmonizing source data semantically with standard vocabularies and structurally conforming to OMOP table structures. Leveraging OHDSI tools, the datahub performs quality assessment and analysis of the transformed data. It ensures FAIR data by establishing metadata flows, capturing provenance throughout the ETL processes, and providing accessible metadata for potential users. The ETL provenance is documented in a machine- and human-readable Implementation Guide (IG), enhancing transparency and usability. Conclusion: The pan-African INSPIRE datahub presents a scalable and systematic solution for integrating health data in LMICs. By adhering to FAIR principles and leveraging established standards like OMOP CDM, this architecture addresses the current gap in generating evidence to support policy and decision-making for improving the well-being of LMIC populations. The federated research network provisions allow data producers to maintain control over their data, fostering collaboration while respecting data privacy and security concerns. A use-case demonstrated the pipeline using OHDSI and other open-source tools.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.030
metaresearch head score (Gemma)0.052
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesOpen science
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.995
Threshold uncertainty score0.159

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0300.052
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0070.008
Science and technology studies0.0020.002
Scholarly communication0.0060.009
Open science0.0050.017
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0150.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.194
GPT teacher head0.421
Teacher spread0.227 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designBench or experimental
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations10
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueFrontiers in Digital HealthSame topicResearch Data Management PracticesFrench-language works237,207