Digital healthcare tools for the world’s poorest people and places: A new framework for syndromic surveillance forged in the fight against COVID-19 in Somalia. (Preprint)
Bibliographic record
Abstract
UNSTRUCTURED Low- and middle-income (LMIC) countries often lack capacity for infectious disease surveillance. The COVID-19 pandemic has challenged many of them to implement non-pharmaceutical interventions while preparing to distribute vaccines. Because of the rapidly evolving situation in many LMICs, real-time data is crucial to understand the needs of the population, and to provide supportive evidence for healthcare interventions. In this paper, we highlight the role of international partners in a data collection effort (79,746 individuals surveyed over six weeks, beginning in April, 2020,) led by the municipal government in Mogadishu, Somalia, where health workers and officials are battling COVID-19 alongside numerous other communicable diseases, as well as violent extremism, and natural disasters. This effort united a collaboration between Canadian and Somalian digital health information engineers in developing a freely-accessible open source digital survey tool for syndromic surveillance. The project enabled extensive health data collection and sharing. It also spurred the rapid development of locally-customized digital solutions, and the training of local digital health engineers. The research insights gathered from this project have been instrumental in the implementation of non-pharmaceutical interventions and in health and sanitation resource allocation in Mogadishu, while providing an updated perspective of the socio-demographics of the population of a city that has not had a census since 1975. Nine collaborative principles are elucidated for developing digital survey tools for syndromic surveillance. The project offers a model for replication in other LMICs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.024 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.010 | 0.007 |
| Science and technology studies | 0.008 | 0.028 |
| Scholarly communication | 0.032 | 0.035 |
| Open science | 0.004 | 0.020 |
| Research integrity | 0.008 | 0.007 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".