An air quality digital twin for real-time outdoor air quality monitoring and prediction
Bibliographic record
Abstract
Respiratory health is closely tied to air quality, making it essential to measure and model air quality. Digital twins provide a powerful approach for air quality monitoring and prediction due to their ability to integrate real-time data, generate air quality predictions, and provide actionable insights. In the City of Elizabeth, New Jersey, poor air quality is driven by heavy industrial activities, dense traffic on major highways, and emissions from the nearby marine terminals and Newark Liberty International Airport. To tackle these issues, this research aimed to develop a digital twin for air quality monitoring and management for the City of Elizabeth. Using LiDAR scans of Housing Authority buildings, Building Information Models (BIM) were created to digitally represent physical structures. A network of outdoor sensors was deployed to capture real-time data on pollutants, including particulate matter (PM₂.₅) and ozone (O 3 ). Unlike traditional physics-based air quality models that rely on complex mathematical equations and require significant computational resources, this study employed a data-driven approach. By analyzing spatial and temporal patterns in air quality data, this method efficiently generated real-time air quality predictions. Integrating these predictions into digital twins enhances our understanding of air quality dynamics and enables stakeholders to communicate complex information effectively to the public. Furthermore, residents can make informed choices to improve their living conditions, such as determining the best times to open windows, use air filtration systems, or spend time in outdoor environment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".