Detection of Spatiotemporal Clusters of COVID-19–Associated Symptoms and Prevention Using a Participatory Surveillance App: Protocol for the @choum Study
Bibliographic record
Abstract
BACKGROUND: The early detection of clusters of infectious diseases such as the SARS-CoV-2-related COVID-19 disease can promote timely testing recommendation compliance and help to prevent disease outbreaks. Prior research revealed the potential of COVID-19 participatory syndromic surveillance systems to complement traditional surveillance systems. However, most existing systems did not integrate geographic information at a local scale, which could improve the management of the SARS-CoV-2 pandemic. OBJECTIVE: The aim of this study is to detect active and emerging spatiotemporal clusters of COVID-19-associated symptoms, and to examine (a posteriori) the association between the clusters' characteristics and sociodemographic and environmental determinants. METHODS: This report presents the methodology and development of the @choum (English: "achoo") study, evaluating an epidemiological digital surveillance tool to detect and prevent clusters of individuals (target sample size, N=5000), aged 18 years or above, with COVID-19-associated symptoms living and/or working in the canton of Geneva, Switzerland. The tool is a 5-minute survey integrated into a free and secure mobile app (CoronApp-HUG). Participants are enrolled through a comprehensive communication campaign conducted throughout the 12-month data collection phase. Participants register to the tool by providing electronic informed consent and nonsensitive information (gender, age, geographically masked addresses). Symptomatic participants can then report COVID-19-associated symptoms at their onset (eg, symptoms type, test date) by tapping on the @choum button. Those who have not yet been tested are offered the possibility to be informed on their cluster status (information returned by daily automated clustering analysis). At each participation step, participants are redirected to the official COVID-19 recommendations websites. Geospatial clustering analyses are performed using the modified space-time density-based spatial clustering of applications with noise (MST-DBSCAN) algorithm. RESULTS: The study began on September 1, 2020, and will be completed on February 28, 2022. Multiple tests performed at various time points throughout the 5-month preparation phase have helped to improve the tool's user experience and the accuracy of the clustering analyses. A 1-month pilot study performed among 38 pharmacists working in 7 Geneva-based pharmacies confirmed the proper functioning of the tool. Since the tool's launch to the entire population of Geneva on February 11, 2021, data are being collected and clusters are being carefully monitored. The primary study outcomes are expected to be published in mid-2022. CONCLUSIONS: The @choum study evaluates an innovative participatory epidemiological digital surveillance tool to detect and prevent clusters of COVID-19-associated symptoms. @choum collects precise geographic information while protecting the user's privacy by using geomasking methods. By providing an evidence base to inform citizens and local authorities on areas potentially facing a high COVID-19 burden, the tool supports the targeted allocation of public health resources and promotes testing. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/30444.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.033 | 0.026 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.005 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.053 | 0.015 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".