Constructing COVID-19 Pandemic Prediction Models Using Positivity Rate
Bibliographic record
Abstract
Data science-based techniques have been widely applied in studies related to COVID-19 spread prediction. In these studies, different modeling techniques have been deployed to estimate the current and future trajectories of the pandemic. The estimation models for these trajectories mainly consider the change in the daily number of confirmed cases as a key modeling input variable. The enforcement of precautionary measures related to COVID-19, in several countries, relied on the estimation results obtained using these prediction models. This paper demonstrates the use of the positivity rate, instead of the daily number of confirmed cases, to obtain more accurate estimates for future pandemic trajectories. We used COVID-19 datasets for eight countries to obtain the daily positivity rate. Susceptible–Infected–Recovered (SIR) modeling was applied for the two compared cases, the case of using the positivity rate and the case of using the daily number of confirmed cases. For each case, we obtained estimated dates for the end of the pandemic. Pairs of results are statistically compared. The results for the predicted pandemic end dates obtained using the positivity rate were found to be statistically different from those obtained using daily confirmed cases. Based on these results, health authorities are advised to consider the pandemic prediction results based on the positivity rates because these rates consider both the daily number of confirmed cases and the number of performed tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".