Satellites, Coding & Air Sensors: Air quality research using PurpleAir sensors, TEMPO and Python
Bibliographic record
Abstract
Air quality makes up a large portion of pollution. An important metric is particulate matter (PM), which can vary in sizes less than one micron and greater than ten microns and is measured in micrograms per cubic meter. Time was dedicated to measuring the concentration of PM using low-cost PurpleAir (PA) sensors in Northwest Indiana (NWI), locating the particles origins, reading articles and papers for appropriate conversion factors (CF), and running experiments on the PA sensors. The PA sensors take one data point every ten seconds. That equates to more than three million data points per sensor per year, while multiple PA sensors are operating in NWI. Previous work has relied on Excel for generating monthly and yearly plots and distributions of PM concentration. Utilizing Python for data processing has significantly reduced the time to get to the analyze step. Other issues surrounding the PA sensors is whether they are providing a correct and unbiased concentration to other commercial and scientific grade instruments. This has led to searching and optimizing for the best CF equation(s) and running high-grade sensors alongside PA sensors. Many questions surround the PA instruments for whether they are a high-quality tool for air quality research. Comparing PA data alongside the Indiana Department of Environmental Managements sensors is vital and has revealed issues in IDEMs lack of data points. Air quality is also being measured by TEMPO, a satellite currently measuring NO2, O3, and formaldehyde hourly across the US from Canada to Mexico.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".