Insights into COVID-19 data collection and management in Malawi: exploring processes, perceptions, and data discrepancies
Bibliographic record
Abstract
Background: The completion of case-based surveillance forms was vital for case identification during COVID-19 surveillance in Malawi. Despite significant efforts, the resulting national data suffered from gaps and inconsistencies which affected its optimal usability. The objectives of this study were to investigate the processes of collecting and reporting COVID-19 data, to explore health workers' perceptions and understanding of the collection tools and processes, and to identify factors contributing to data quality. Methods: A total of 75 healthcare professionals directly involved in COVID-19 data collection from the Malawi Ministry of Health in Lilongwe and Blantyre participated in Focus Group Discussions and In-Depth Interviews. We collected participants' views on the effectiveness of surveillance forms in collecting the intended data, as well as on the data collection processes and training needs. We used MAXQDA for thematic and document analysis. Results: Form design significantly influenced data quality and, together with challenges in applying case definitions, formed 44% of all issues raised. Concerns regarding processes used in data collection and training gaps comprised 49% of all the issues raised. Language issues (2%) and privacy, ethical, and cultural considerations (4%), although mentioned less frequently, offered compelling evidence for further review. Conclusions: Our study highlights the integral connection between data quality and the design and utilization of data collection forms. While the forms were deemed to contain the most relevant fields, deficiencies in format, order of fields, and the absence of an addendum with guidelines, resulted in large gaps and errors. Form design needs to be reviewed so that it appropriately fits into the overall processes and systems that capture surveillance data. This study is the first of its kind in Malawi, offering an in-depth view of the perceptions and experiences of health professionals involved in disease surveillance on the tools and processes they use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.009 | 0.004 |
| Open science | 0.011 | 0.192 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".