Methodological issues of retrospective surveys for measuring mortality of highly clustered diseases: case study of the 2014–16 Ebola outbreak in Bo District, Sierra Leone
Bibliographic record
Abstract
BACKGROUND: There is a lack of empirical data on design effects (DEFF) for mortality rate for highly clustered data such as with Ebola virus disease (EVD), along with a lack of documentation of methodological limitations and operational utility of mortality estimated from cluster-sampled studies when the DEFF is high. OBJECTIVES: The objectives of this paper are to report EVD mortality rate and DEFF estimates, and discuss the methodological limitations of cluster surveys when data are highly clustered such as during an EVD outbreak. METHODS: We analysed the outputs of two independent population-based surveys conducted at the end of the 2014-2016 EVD outbreak in Bo District, Sierra Leone, in urban and rural areas. In each area, 35 clusters of 14 households were selected with probability proportional to population size. We collected information on morbidity, mortality and changes in household composition during the recall period (May 2014 to April 2015). Rates were calculated for all-cause, all-age, under-5 and EVD-specific mortality, respectively, by areas and overall. Crude and adjusted mortality rates were estimated using Poisson regression, accounting for the surveys sample weights and the clustered design. RESULTS: Overall 980 households and 6,522 individuals participated in both surveys. A total of 64 deaths were reported, of which 20 were attributed to EVD. The crude and EVD-specific mortality rates were 0.35/10,000 person-days (95%CI: 0.23-0.52) and 0.12/10,000 person-days (95%CI: 0.05-0.32), respectively. The DEFF for EVD mortality was 5.53, and for non-EVD mortality, it was 1.53. DEFF for EVD-specific mortality was 6.18 in the rural area and 0.58 in the urban area. DEFF for non-EVD-specific mortality was 1.87 in the rural area and 0.44 in the urban area. CONCLUSION: Our findings demonstrate a high degree of clustering; this contributed to imprecise mortality estimates, which have limited utility when assessing the impact of disease. We provide DEFF estimates that can inform future cluster surveys and discuss design improvements to mitigate the limitations of surveys for highly clustered data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".