Fossil Fuel CO₂ Emissions for the OCO2 Model Intercomparison Project (MIP)
Bibliographic record
Abstract
ATTENTION: Please only download the YYYY.tar files for the latest version. Due to a glitch in the Zenodo upload interface, older YYYYMM.tar files are showing up in this version as well. They should not be used for version 2023.0. These are fossil CO2 fluxes updated through July 2023 for atmospheric CO2 modeling. They were constructed primarily to be used for the OCO2 Model Intercomparison Project (MIP). For 2000-2021, they're based on ODIAC 2022, which in turn uses BP's energy use statistics for 2020 and 2021. ODIAC monthly emissions have been disaggregated to hourly using the TIMES emission factors for day of week and time of day (https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2012JD018196). For 2022 and 2023, ODIAC's 2021 emissions have been scaled by the ratio of 2022 (or 2023) to 2021 emissions reported by Carbon Monitor, downloaded on August 31, 2023 from https://carbonmonitor.org/. ODIAC does not have sectoral decomposition to the degree provided by Carbon Monitor, so total ODIAC emissions for each region have been scaled by the total emission change between 2021 and 2022 (or 2023) reported by Carbon Monitor, i.e., power, ground transport, etc. have not been separately scaled. Carbon Monitor data are daily, but ODIAC emissions are monthly. So Carbon Monitor data have been aggregated to monthly totals before deriving scaling factors between 2021 and 2022 (or 2023). Carbon Monitor reports international aviation emissions by country of origin, while ODIAC reports aviation emissions on a grid. Since there is no way to derive the points of emission for Carbon Monitor aviation emissions , all Carbon Monitor international aviation was aggregated to create a single number for each month, then that number was used to scale ODIAC's bunker fuel for each month in 2022 and 2023. Hourly global totals are given in the files as a check, in case you want to verify your units and file reading.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.032 | 0.013 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".