Tailoring Mobile Data Collection for Intervention Research in a Challenging Context: Development and Implementation in the Malakit Study
Bibliographic record
Abstract
BACKGROUND: An interventional study named Malakit was implemented between April 2018 and March 2020 to address malaria in gold mining areas in French Guiana, in collaboration with Suriname and Brazil. This innovative intervention relied on the distribution of kits for self-diagnosis and self-treatment to gold miners after training by health mediators, referred to in the project as facilitators. OBJECTIVE: This paper aims to describe the process by which the information system was designed, developed, and implemented to achieve the monitoring and evaluation of the Malakit intervention. METHODS: The intervention was implemented in challenging conditions at five cross-border distribution sites, which imposed strong logistical constraints for the design of the information system: isolation in the Amazon rainforest, tropical climate, and lack of reliable electricity supply and internet connection. Additional constraints originated from the interaction of the multicultural players involved in the study. The Malakit information system was developed as a patchwork of existing open-source software, commercial services, and tools developed in-house. Facilitators collected data from participants using Android tablets with ODK (Open Data Kit) Collect. A custom R package and a dashboard web app were developed to retrieve, decrypt, aggregate, monitor, and clean data according to feedback from facilitators and supervision visits on the field. RESULTS: Between April 2018 and March 2020, nine facilitators generated a total of 4863 form records, corresponding to an average of 202 records per month. Facilitators' feedback was essential for adapting and improving mobile data collection and monitoring. Few technical issues were reported. The median duration of data capture was 5 (IQR 3-7) minutes, suggesting that electronic data capture was not taking more time from participants, and it decreased over the course of the study as facilitators become more experienced. The quality of data collected by facilitators was satisfactory, with only 3.03% (147/4849) of form records requiring correction. CONCLUSIONS: The development of the information system for the Malakit project was a source of innovation that mirrored the inventiveness of the intervention itself. Our experience confirms that even in a challenging environment, it is possible to produce good-quality data and evaluate a complex health intervention by carefully adapting tools to field constraints and health mediators' experience. TRIAL REGISTRATION: ClinicalTrials.gov NCT03695770; https://clinicaltrials.gov/ct2/show/NCT03695770.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".