TROPOMI/WFMD v2.0: Improved retrievals of XCH <sub>4</sub> and XCO with XGBoost-based quality filtering
Bibliographic record
Abstract
Abstract. The TROPOspheric Monitoring Instrument (TROPOMI) on board the Sentinel-5 Precursor satellite provides daily global observations of atmospheric methane (CH4) and carbon monoxide (CO) at relatively high spatial resolution. The dense spatial and temporal coverage is achieved by the instrument's wide swath, which permits detailed mapping of the worldwide distribution of these important atmospheric constituents. The adaptation and optimisation of the Weighting Function Modified Differential Optical Absorption Spectroscopy (WFMD) algorithm for the simultaneous retrieval of the column-averaged dry-air mole fractions XCH4 and XCO from TROPOMI's shortwave infrared (SWIR) radiance measurements has proven to be a valuable complement and alternative to the operational TROPOMI products. The latest release of the TROPOMI/WFMD product (version 2.0) includes several improvements expanding its suitability for a wider range of scientific applications. Data yield at mid and high latitudes has increased, accompanied by improved accuracy and precision according to the validation with the ground-based Total Carbon Column Observing Network (TCCON). These advancements are primarily due to more refined quality filtering that has been accomplished by replacing the previous Random Forest Classifier with the more efficient and potentially higher performing Extreme Gradient Boosting (XGBoost) algorithm in conjunction with improved training data incorporating an updated cloud product from the Visible Infrared Imaging Radiometer Suite (VIIRS) and the TROPOMI Aerosol Index. This enhanced training data set enables more reliable identification of cloudy scenes and mitigates issues related to specific aerosol events over bright surfaces. Importantly, as with previous product versions, the actual quality classification does not depend on the real-time availability of these external data products, which are only required during the training phase.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".