Evaluation data for "Global, high-resolution, reduced-complexity air quality modeling for PM2.5 using InMAP (Intervention Model for Air Pollution)"
Bibliographic record
Abstract
This zip file contains data for performing Global InMAP model runs and evaluations. To the extent that any of the data is covered by third party licenses, it is the responsibility of the user to follow the terms of those licenses. A description of the contents of this directory is below: measurements.csv<br> Vetted global dataset of ground-level annual-average measurements of total PM2.5 and species (pNO3, pSO4, pNH4) compiled from monitoring networks, used for model performance evaluation. Data sources are: World Health Organization (Global), European Environment Agency (Europe), National Air Pollution Surveillance Program (Canada), Environmental Protection Agency (United States of America), Central Pollution Control Board (India), Australian Government State of the Environment (Australia), and Acid Deposition Monitoring Network In East Asia (EANET) (East Asia). population directory<br> Population count data is from the Gridded Population of The World (v4.10) projected to year 2020. The data is in 15x15 arcminute grids, except for in grid cells where the population is above 80,000, where the population data is 30x30 arcseconds. GlobalInMAPData_v1.ncf<br> Regular-grid Global InMAP input data for the year 2005 for use as the "InMAPData" variable in the InMAP configuration file. It was created from GEOS-Chem v.11-01 simulation outputs with the 'inmap preproc' command. global_inmap_004x003_v1.1.0.gob<br> Global InMAP variable grid resolution input data for coords for year 2016 for use as the "VariableGridData" variable in the InMAP configuration file. It was created with the 'inmap grid' command using GlobalInMAPData_v1.ncf and population.shp. 2016_emissions directory<br> Total PM2.5 and precursor emissions to arrive at total PM2.5 concentrations from Global InMAP. Units for polygonized emissions inputs (shapefiles) are short (US) tons/yr, and units for gridded emissions inputs (NetCDF files) are kg/yr. global_emission_changes directory<br> nh3.nc, nox.nc, and sox.nc are gridded emissions for changes in inorganic precursors for comparing Global InMAP and GEOS-Chem. Units are kg/yr. NH4-gc.nc, NIT-gc.nc, and SO4-gc.nc are results for changes in concentrations arising from these changes in emissions for 3 months, 1 month, and 2 months. usa_emission_changes directory<br> Emissions for comparing Global InMAP and US InMAP (described in Tessum et al., 2017).<br> Emissions are derived using the United States National Emissions Inventory (NEI) 2014v.1, processed exactly as in Thakrar et al., 2020.<br> Emissions are coal-powered electricity generation (NEI Source Classification Code: 10100212) and gasoline passenger vehicles (NEI Source Classification Code: 2201210080).<br> Units are ug/s. Tessum, C.W.; Hill, J.D.; Marshall, J.D. InMAP: A model for air pollution interventions. PloS One 2017, 12 (4) e0176131.<br> Thakrar, S.K.; Balasubramanian, S.; Adams, P.J.; Azevedo, I.M.; Muller, N.Z.; Pandis, S.N.; Polasky, S.; Pope III, C.A.; Robinson, A.L.; Apte, J.S.; Tessum, C.W.; Marshall, J.D.; Hill; J.D. Reducing mortality from air pollution in the United States by targeting specific emission sources. Environmental Science & Technology Letters 2020, 7(9), pp.639-645.<br> Gridded Population of the World, Version 4 (GPWv4): National Identifier Grid. Palisades, NY: NASA Socioeconomic Data and Applications Center (SEDAC). http://dx.doi.org/10.7927/H41V5BX1.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.004 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".