Leveraging artificial intelligence and data science techniques in harmonizing, sharing, accessing and analyzing SARS-COV-2/COVID-19 data in Rwanda (LAISDAR Project): study design and rationale
Bibliographic record
Abstract
BACKGROUND: Since the outbreak of COVID-19 pandemic in Rwanda, a vast amount of SARS-COV-2/COVID-19-related data have been collected including COVID-19 testing and hospital routine care data. Unfortunately, those data are fragmented in silos with different data structures or formats and cannot be used to improve understanding of the disease, monitor its progress, and generate evidence to guide prevention measures. The objective of this project is to leverage the artificial intelligence (AI) and data science techniques in harmonizing datasets to support Rwandan government needs in monitoring and predicting the COVID-19 burden, including the hospital admissions and overall infection rates. METHODS: The project will gather the existing data including hospital electronic health records (EHRs), the COVID-19 testing data and will link with longitudinal data from community surveys. The open-source tools from Observational Health Data Sciences and Informatics (OHDSI) will be used to harmonize hospital EHRs through the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM). The project will also leverage other OHDSI tools for data analytics and network integration, as well as R Studio and Python. The network will include up to 15 health facilities in Rwanda, whose EHR data will be harmonized to OMOP CDM. EXPECTED RESULTS: This study will yield a technical infrastructure where the 15 participating hospitals and health centres will have EHR data in OMOP CDM format on a local Mac Mini ("data node"), together with a set of OHDSI open-source tools. A central server, or portal, will contain a data catalogue of participating sites, as well as the OHDSI tools that are used to define and manage distributed studies. The central server will also integrate the information from the national Covid-19 registry, as well as the results of the community surveys. The ultimate project outcome is the dynamic prediction modelling for COVID-19 pandemic in Rwanda. DISCUSSION: The project is the first on the African continent leveraging AI and implementation of an OMOP CDM based federated data network for data harmonization. Such infrastructure is scalable for other pandemics monitoring, outcomes predictions, and tailored response planning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.006 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".