Digital Twin COVID Tracker Using Wastewater Data: A Middle-School Led Study Within the U.S. NSF National Research Traineeship Program Framework
Bibliographic record
Abstract
Wastewater infrastructure exists in every municipality across the United States and many other nations, offering a universal, non-invasive platform for community-level disease surveillance. Because viruses such as SARS-CoV-2 shed into wastewater days before symptoms appear, wastewater-based epidemiology (WBE) can provide crucial early-warning signals for public health. This study presents an AI-enabled Digital Twin prototype that predicts short-term COVID-19 trends using Center for Disease Control (CDC) wastewater viral activity data. Uniquely, this project was conceived and executed by middleschool first authors, highlighting the importance of early STEM engagement and intentional mentoring of young professionals on societally relevant environmental and health challenges. This work was conducted as part of our ongoing National Science Foundation (NSF) and National Institutes of Health (NIH) projects led by senior authors, which focus on convergence research and workforce development in AI-enabled, omics-guided living-interface engineering. Computational modeling, Jupyter Notebook workflow, GitHub integration, and cloud deployment were supported by graduate mentors, while system design, experimental logic, and interpretation were led by the student authors. The resulting platform, accessible through an interactive web app and QR-code interface, illustrates how guided, ageappropriate research experiences can empower middle-school students to explore wastewater informatics, digital twin concepts, machine learning, and epidemiological modeling. This work was also recognized with a 3rd-place award in the Sixth Grade Engineering Category at the 2025 High Plains Regional Science & Engineering Fair, highlighting both scientific merit and the broader impact of engaging middle-school students in societally relevant STEM research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".