Addressing bias in the definition of SARS-CoV-2 reinfection: implications for underestimation
Bibliographic record
Abstract
Introduction: Reinfections are increasingly becoming a feature in the epidemiology of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection. However, accurately defining reinfection poses methodological challenges. Conventionally, reinfection is defined as a positive test occurring at least 90 days after a previous infection diagnosis. Yet, this extended time window may lead to an underestimation of reinfection occurrences. This study investigated the prospect of adopting an alternative, shorter time window for defining reinfection. Methods: A longitudinal study was conducted to assess the incidence of reinfections in the total population of Qatar, from February 28, 2020 to November 20, 2023. The assessment considered a range of time windows for defining reinfection, spanning from 1 day to 180 days. Subgroup analyses comparing first versus repeat reinfections and a sensitivity analysis, focusing exclusively on individuals who underwent frequent testing, were performed. Results: The relationship between the number of reinfections in the population and the duration of the time window used to define reinfection revealed two distinct dynamical domains. Within the initial 15 days post-infection diagnosis, almost all positive tests for SARS-CoV-2 were attributed to the original infection. However, surpassing the 30-day post-infection threshold, nearly all positive tests were attributed to reinfections. A 40-day time window emerged as a sufficiently conservative definition for reinfection. By setting the time window at 40 days, the estimated number of reinfections in the population increased from 84,565 to 88,384, compared to the 90-day time window. The maximum observed reinfections were 6 and 4 for the 40-day and 90-day time windows, respectively. The 40-day time window was appropriate for defining reinfection, irrespective of whether it was the first, second, third, or fourth occurrence. The sensitivity analysis, confined to high testers exclusively, replicated similar patterns and results. Discussion: A 40-day time window is optimal for defining reinfection, providing an informed alternative to the conventional 90-day time window. Reinfections are prevalent, with some individuals experiencing multiple instances since the onset of the pandemic.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".