Understanding the Prevalence and Geographic Heterogeneity of SARS-CoV-2 Infection: Findings of the First Serosurvey in Uttar Pradesh, India
Bibliographic record
Abstract
Population-based serological antibody test for SARS-CoV-2 infection helps in estimating the exposure in the community. We present the findings of the first district representative seroepidemiological survey conducted between 4 and 10 September 2020 among the population aged 5 years and above in the state of Uttar Pradesh, India. Multi-stage cluster sampling was used to select participants from 495 primary sampling units (villages in rural areas and wards in urban areas) across 11 selected districts to provide district-level seroprevalence disaggregated by place of residence (rural/urban), age (5-17 years/aged 18 +) and gender. A venous blood sample was collected to determine seroprevalence. Of 16,012 individuals enrolled in the study, 22.2% [95% CI 21.5-22.9] equating to about 10.4 million population in 11 districts were already exposed to SARS-CoV-2 infection by mid-September 2020. The overall seroprevalence was significantly higher in urban areas (30.6%, 95% CI 29.4-31.7) compared to rural areas (14.7%, 95% CI 13.9-15.6), and among aged 18 + years (23.2%, 95% CI 22.4-24.0) compared to aged 5-17 years (18.4%, 95% CI 17.0-19.9). No differences were observed by gender. Individuals exposed to a COVID confirmed case or residing in a COVID containment zone had higher seroprevalence (34.5% and 26.0%, respectively). There was also a wide variation (10.7-33.0%) in seropositivity across 11 districts indicating that population exposed to COVID was not uniform at the time of the study. Since about 78% of the population (36.5 million) in these districts were still susceptible to infection, public health measures remain essential to reduce further spread.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | high |
| gpt | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".