The World Health Organization COVID-19 surveillance database
Bibliographic record
Abstract
In January 2020, SARS-CoV-2 virus was identified as a cause of an outbreak in China. The disease quickly spread worldwide, and the World Health Organization (WHO) declared the pandemic in March 2020.From the first notifications of spread of the disease, the WHO's Emergency Programme implemented a global COVID-19 surveillance system in coordination with all WHO regional offices. The system aimed to monitor the spread of the epidemic over countries and across population groups, severity of the disease and risk factors, and the impact of control measures. COVID-19 surveillance data reported to WHO is a combination of case-based data and weekly aggregated data, focusing on a minimum global dataset for cases and deaths including disaggregation by age, sex, occupation as a Health Care Worker, as well as number of cases tested, and number of cases newly admitted for hospitalization. These disaggregations aim to monitor inequities in COVID-19 distribution and risk factors among population groups.SARS-CoV-2 epidemic waves continue to sweep the world; as of March 2022, over 445 million cases and 6 million deaths have been reported worldwide. Of these, over 327 million cases (74%) have been reported in the WHO surveillance database, of which 255 million cases (57%) are disaggregated by age and sex. A public dashboard has been made available to visualize trends, age distributions, sex ratios, along with testing and hospitalization rates. It includes a feature to download the underlying dataset.This paper will describe the data flows, database, and frontend public dashboard, as well as the challenges experienced in data acquisition, curation and compilation and the lessons learnt in overcoming these. Two years after the pandemic was declared, COVID-19 continues to spread and is still considered a Public Health Emergency of International Concern (PHEIC). While WHO regional and country offices have demonstrated tremendous adaptability and commitment to process COVID-19 surveillance data, lessons learnt from this major event will serve to enhance capacity and preparedness at every level, as well as institutional empowerment that may lead to greater sharing of public health evidence during a PHEIC, with a focus on equity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".