Testing the validity and feasibility of using a mobile phone-based method to assess the strength of implementation of family planning programs in Malawi
Bibliographic record
Abstract
BACKGROUND: To effectively deliver on proposed objectives, it is vital that practitioners, policymakers, and other stakeholders are able to clearly understand how strongly their large-scale program is being implemented. This study sought to test the feasibility, cost-effectiveness, and validity of a phone-based method as an innovative and cost-efficient approach to assessing program implementation strength (through an Implementation Strength Assessment - ISA), alternative to the traditional in-person field methods. METHODS: We conducted 701 mobile phone and 356 in-person interviews with facility in-Charges and two types of community health workers who provide family planning services in the Dowa and Ntcheu districts in Malawi. Responses received via the phone interview were validated through in-person review of records and inspections. Sensitivity and specificity were calculated to determine validity. RESULTS: Most indicators at the health facility and community health worker levels were above a 70% threshold for sensitivity. However, there were fewer indicators that met this threshold for specificity. The primary reason for lower specificity was due to poor recordkeeping. Collecting data via mobile phone was found to be feasible and twice as cost-efficient as collecting the same data via in-person inspections. CONCLUSIONS: The rapid increase in mobile phone ownership and network availability in lower income countries could offer an alternative, cost-effective avenue to collect data for a better understanding of program implementation. Through rigorous assessment, this study found that using mobile phones could be a low-cost alternative to collect data on health system delivery of services, especially in places where routine data quality is poor and traditional, in-person methods are costly.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.067 | 0.124 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".