Measuring the Masses: Domains Driving Data Collection and Analysis for the Health Outcomes of Mass Gatherings (Paper 3)
Bibliographic record
Abstract
INTRODUCTION: Without a robust evidence base to support recommendations for medical services at mass gatherings (MGs), levels of care will continue to vary and preventable morbidity and mortality will exist. Accordingly, researchers and clinicians publish case reports and case series to capture and explain some of the health interventions, health outcomes, and host community impacts of MGs. Streamlining and standardizing post-event reporting for MG medical services and associated health outcomes could improve inter-event comparability, thereby supporting and promoting growth of the evidence base for this discipline. The present paper is focused on theory building, proposing a set of domains for data that may support increasingly comprehensive, yet lean, reporting on the health outcomes of MGs. This paper is paired with another presenting a proposal for a post-event reporting template. METHODS: The conceptual categories of data presented are based on a textual analysis of 54 published post-event medical case reports and a comparison of the features of published data models for MG health outcomes. FINDINGS: A comparison of existing data models illustrates that none of the models are explicitly informed by a conceptual lens. Based on an analysis of the literature reviewed, four data domains emerged. These included: (i) the Event Domain, (ii) the Hazard and Risk Domain, (iii) the Capacity Domain, and (iv) the Clinical Domain. These domains mapped to 16 sub-domains. DISCUSSION: Data modelling for the health outcomes related to MGs is currently in its infancy. The proposed illustration is a set of operationally relevant data domains that apply equally to small, medium, and large-sized events. Further development of these domains could move the MG community forward and shift post-event health outcomes reporting in the direction of increasing consistency and comprehensiveness. CONCLUSION: Currently, data collection and analysis related to understanding health outcomes arising from MGs is not informed by robust conceptual models. This paper is part of a series of nested papers focused on the future state of post-event medical reporting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.294 | 0.470 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.012 | 0.016 |
| Science and technology studies | 0.004 | 0.012 |
| Scholarly communication | 0.012 | 0.013 |
| Open science | 0.004 | 0.011 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".