Code and data from: The risks to human health of air toxics, PM2.5, and ozone from the 2023 Canadian wildfires
Bibliographic record
Abstract
Supporting data for calculations of the risks to human health of air toxics, PM2.5, and ozone from the 2023 Canadian wildfires Data archive created by Havala Pye 0000-0002-2014-2140 For use of the data, users are encouraged to cite both this data archive (https://doi.org/10.5281/zenodo.16099070) for transparency and the following article for methods documentation: H. O. T. Pye, W. T. Hutzell, N. L. Fann, T. N. Skipper, M. Pye, J. Beidler, C. Allen, B. N. Murphy, E. L. D’Ambro, S. Lin, K. Talgo, L. Reynolds, D. Kang, J. Bash, K. M. Seltzer, S. L. Farrell, K. W. Appel, K. Brehme, R. C. Gilliam, B. H. Henderson, and A. W. H. Chan: The risks to human health of air toxics, PM2.5, and ozone from the 2023 Canadian wildfires, ChemRxiv, , 2025. This content is a preprint and has not been peer-reviewed. Please update to the final journal article when available. Additional model-ready data to run CMAQ (emissions, meteorology, and other input files) from this work are available at: US EPA, 2025, "CMAQ 2023 12US4 CRACMM2 Inputs and Meteorology", https://doi.org/10.15139/S3/GL41QC, UNC Dataverse, V1. Please see documentation on the CMAQ github repository for standard CMAQ conventions. Contents 12US4_files: netcdf and shape files with information for the 12US4 domain including shapefiles: shape files and other domain specification information population: population data for the U.S. and Canada on 12US4 grid as a netcdf file (acs_2023_5yr_bg_pop_12US4.nc, can_2021_cenus_subd_pop_12US4.nc) masks: mask for Canada used to remove fires for that region in CMAQ analysis/inputs: inputs used in analysis of CMAQ output cities.txt: Latitude and longitude of locations for data extraction in Figure 2. Formatted as a writesite (https://github.com/USEPA/CMAQ/tree/main/POST/writesite) file. 20250512Dose_Response_Library_Oct2024.xlsx: HEM toxicity values file from October 2024 and mapping to CRACMM species; CMAQCONCCOMBINE_UGM3 indicates combine species name (without units), CMAQHAP is the CMAQ-CRACMM species name, CMAQEMIS is the emitted species name (without T_ prepending); CMAQ names should not be cross-walked by line; See Pye et al. (2025) supplemental information Table S7 for more details and a readable version with crosswalk. degradelog.txt: excerpt from CMAQ simulation processor log file documenting active hazardous air pollutant (HAP) degradation reactions. TableS12_pye2024_erbymass.xlsx: Table S12 from Pye et al. 2024 (https://doi.org/10.1021/acs.est.4c06187) and supplemented with Gkatzelis et al. 2024 (https://doi.org/10.5194/acp-24-929-2024). 2023hc_CRACMM2_task4_final_sector_reports_withTSpecies_09apr2025.csv: Emissions totals in tons/yr by sector in CRACMM species names. Target_Organ_resp_Oct2024.xlsx: subset of HEM target organ file indicating all species with a respiratory target organ impact. Target_Organ_neuro_Oct2024.xlsx: subset of HEM target organ file indicating all species with a neurological target organ impact. analysis/scripts: scripts used to analyze concentrations and determine risk in the work of Pye et al. (2025) in python notebook (ipynb) and html. FigS19_NAPS_HAPS_Data_processing.ipynb/html: Creates Figure S17 evaluating CMAQ vs NAPS VOCs. Pye2025_allsrc_hapemiss_FigS1.ipnyb/html: Creates Figure S1 comparing HAP emissions across sectors and studies. Pye2025_FigS14S16S17_aqs_fire_analysis.ipnyb/html: Creates Figures S14,S16,S17 comparing CMAQ to AQS observations. Pye2025_Stackedbar_Fig2S6.ipynb/html: Creates Figures 2 and S6 with risk and photochemical age. Pye2025_SpeciesSIfigs_HAPs_HItotal_Fig1S5S20.ipynb/html: Creates Figure 1, S5, S20 including concentrations of HAPs. Pye2025_SpeciesSIfigs_PMox_S3S4S21S22.ipynmb/html: Creates PM and oxidant figures (S3, S4, S21, S22). Pye2025_TOC.ipynb/html: Creates Table of Contents (TOC) art. CMAQcode BLD: CMAQv5.5 fortran code with additional updates resulting in CRACMM3HAPs as described by Pye et al. (2025) base_config: build script, run scripts, control files, and example log files (text files) for the base simulation in Pye et al. (2025). Known issues: - while the control file indicates acrylonitrile emissions were scaled from CO for Canadian fires, the Canadian fire emissions of acrylonitrile were not implemented due to a bug in CMAQv5.5 DESID. - emissions of HAPs from biogenic sources were accidentally doubled in these scripts - lightning NOx was omitted nofire_config: build script, run script, control files, and example log files (text files) for the simulation without Canadian fires. post: species definitions files used to prepare ONLYHAPS and PMOXIDANTS netcdf output. CMAQoutput: netcdf gridded CMAQ output for base simulation and simulation without Canadian fires (nocanfire) Annual average HAP concentrations (and select other species such as CO): ANNUALAVG_20250502cracmm3hap_base_ONLYHAPS_2023_12US4.nc ANNUALAVG_20250502cracmm3hap_nocanfire_ONLYHAPS_2023_12US4.nc Annual average OH, O3, NO3, and PM2.5: ANNUALAVG_20250502cracmm3haps_base_PMOXIDANTS_2023_12US4.nc ANNUALAVG_20250502cracmm3haps_nocanfire_PMOXIDANTS_2023_12US4.nc Seasonal (Apr-Sept) average of max daily 8hr avg ozone: SEASAVG_APR2SEP_20250502cracmm3haps_base_2023_12US4.nc SEASAVG_APR2SEP_20250502cracmm3haps_nocanfire_2023_12US4.nc Seasonal (Apr-Sept) average of max daily 1 hour ozone: SEASAVG_1hrmaxo3_APR2SEP_20250502cracmm3haps_base_2023_12US4.nc SEASAVG_1hrmaxo3_APR2SEP_20250502cracmm3haps_nocanfire_2023_12US4.nc evaluation: output from CMAQ site compare utility by month baseaqs.tar: base simulation results by month nocanfireaqs.tar: no Canadian fire simulation results by month For each month: Ozone, PM, and major HAPS: AQS_Daily_2023_12US4_* files Select additional VOCs: AQS_Daily_VOC_2023_12US4_* files Canada_NAPS_VOC: intermediate files used in the evaluation of CMAQ predictions of select VOCs with NAPS observations boundaryconditions: sample scripts for how to download and map GEOSCF to CMAQ-CRACMM for boundary conditions python_batchfall.csh: cshell batch script that calls the following file to obtain GEOSCF files in original GEOSCF species (step 1) 2024fall_12US4_getgeoscfonly.py: gets GEOSCF files for fall (run in 2024) *expr: files used by aqmbc to map GEOSCF variables to CRACMM2 (also available in AQMBC on github) 2024_12US4_geoscfBCONbyhour.ipynb: creates BCON files along 12US4 boundary in CRACMM2 species on original GEOSCF time stamps (step 2) 2024_12US4_geoscfBCONplots.ipynb: QA check on BCON files (step 3) 2024_12US4_geoscf_modelready.ipynb: Concatenate and time shift the BCON files so they are CMAQ-ready (step 4) GRIDDESC: CMAQ input and grid description used by aqmbc benefits: BenMAP and AQBAT input files, outputs, and additional documentation. Individual files detailed in the Readme.docx. Public tools, data, and repositories related to this work: CMAQ: https://doi.org/10.5281/zenodo.1079878 and https://github.com/USEPA/CMAQ CRACMM: https://github.com/USEPA/CRACMM AMET: https://github.com/USEPA/AMET AQS data: https://www.epa.gov/aqs NAPS data: https://www.canada.ca/en/environment-climate-change/services/air-pollution/monitoring-networks-data/national-air-pollution-program.html BenMAP: https://www.epa.gov/benmap/benmap-downloads AQBAT: https://health-infobase.canada.ca/aqbat/ AQMBC: https://barronh.github.io/aqmbc/index.html
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.019 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.008 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.006 | 0.003 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.353 | 0.175 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".