Assessing fitness-for-purpose and comparing the suitability of COVID-19 multi-country models for local contexts and users
Bibliographic record
Abstract
<ns3:p> <ns3:bold>Background:</ns3:bold> Mathematical models have been used throughout the COVID-19 pandemic to inform policymaking decisions. The COVID-19 Multi-Model Comparison Collaboration (CMCC) was established to provide country governments, particularly low- and middle-income countries (LMICs), and other model users with an overview of the aims, capabilities and limits of the main multi-country COVID-19 models to optimise their usefulness in the COVID-19 response. </ns3:p> <ns3:p> <ns3:bold>Methods:</ns3:bold> Seven models were identified that satisfied the inclusion criteria for the model comparison and had creators that were willing to participate in this analysis. A questionnaire, extraction tables and interview structure were developed to be used for each model, these tools had the aim of capturing the model characteristics deemed of greatest importance based on discussions with the Policy Group. The questionnaires were first completed by the CMCC Technical group using publicly available information, before further clarification and verification was obtained during interviews with the model developers. The fitness-for-purpose flow chart for assessing the appropriateness for use of different COVID-19 models was developed jointly by the CMCC Technical Group and Policy Group. </ns3:p> <ns3:p> <ns3:bold>Results:</ns3:bold> A flow chart of key questions to assess the fitness-for-purpose of commonly used COVID-19 epidemiological models was developed, with focus placed on their use in LMICs. Furthermore, each model was summarised with a description of the main characteristics, as well as the level of engagement and expertise required to use or adapt these models to LMIC settings. </ns3:p> <ns3:p> <ns3:bold>Conclusions:</ns3:bold> This work formalises a process for engagement with models, which is often done on an ad-hoc basis, with recommendations for both policymakers and model developers and should improve modelling use in policy decision making. </ns3:p>
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.043 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.008 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".