Training datasets with manually labeled TROPOMI data for Machine Learning models [Schuit et al. 2023: Automated detection and monitoring of methane super-emitters using satellite data]
Notice bibliographique
Résumé
This repository contains the manually labeled training datasets of TROPOMI data used in Schuit et al. 2023 to train the Convolutional Neural Network (CNN) and Support Vector Classifier (SVC). The trainingdata is split into three files, SVC_trainingdata.nc, CNN_pos_trainingdata.nc and CNN_neg_trainingdata.nc. All training data originates from 2018, 2019 or 2020. This training dataset was generated using the SRON TROPOMI scientific xch4 data product version 18_17 (available at: https://ftp.sron.nl/open-access-data-2/TROPOMI/tropomi/ch4/18_17/, last access 21-06-2024) described by Lorente et al. (2021). Please note that this is an older version of the TROPOMI methane dataproduct. Users are recommended to use the latest operational TROPOMI methane product available on the Copernicus Dataspace (available at: https://documentation.dataspace.copernicus.eu/Data/SentinelMissions/Sentinel5P.html#sentinel-5p-level-2-methane), which includes important updates described by Lorente et al. (2023). CNN_pos_trainingdata.nc consists of 828 scenes of TROPOMI data with plume-like morphological structures in the methane channel. Every scene in this dataset is labeled as “plume_structures”. CNN_neg_trainingdata.nc consists of 2242 scenes of TROPOMI data without clear plumes. Every scene in this dataset is labeled as “no_plume”. SVC_trainingdata.nc consists of 843 scenes that were detected by the trained CNN, and were manually labeled as either “plume”, “artefact”, or “empty” by a human expert, taking into account information from the additional channels (e.g., surface albedo, windfield, aerosol optical depth) next to the methane channel. All three datasets contain all data fields/channels used at some point in the architecture, these are all taken from the TROPOMI Level 2 dataproduct. The CNN training dataset also contains these supporting channels, however only the xch4 channel was used to train the CNN in Schuit et al. (2023). For the SVC, this data was not used directly, but first feature engineering algorithms were applied to the channels present in this data, in order to generate a feature vector to represent the information in the scene. Further details on the computation of these ‘features’ are provided in Schuit et al. 2023, Section 2.1, 2.2, 2.3 and 2.4, and Table A1, and C1. Contents and data formats The three NetCDF files share the same structure, consisting of 13 channels with dimensions [N, 32, 32] where N is the number of scenes (either “CNN_pos_trainingdata_index”, “CNN_neg_trainingdata_index”, or “SVC_trainingdata_index”) and 32x32 are the spatial dimensions, represented as pixel indices in along-orbit and across-orbit direction. Next to these 13 channels with spatial data, there are 3 variables that correspond to the manual_label, the orbit_number and a unique_identifier with dimension [N]. The 13 spatial channels, including units and description are: xch4; [1e-9]; bias corrected column-averaged dry-air mole fraction of methane. The methane data was destriped as described in Section 2.1. latitude; [degrees]; latitude of the center of the TROPOMI pixel (not used for training). longitude; [degrees]; longitude of the center of the TROPOMI pixel (not used for training). albedo_SWIR; [-]; surface albedo in the SWIR channel. aerosol_optical_thickness_SWIR; [-]; aerosol optical thickness in the SWIR channel. surface_pressure; [hPa]; surface pressure. chi2; [-]; an indicator for retrieval fit quality. qa_value; [-]; quality flag windspeed_north_v10; [m/s]; northward component of the windspeed (south is negative) (originating from ERA5, but present in the TROPOMI Level 2 dataproduct). windspeed_east_u10; [m/s]; eastward component of the windspeed (west is negative) (originating from ERA5, but present in the TROPOMI Level 2 dataproduct). landflag_science; [-]; Land-water mask and surface classification based on a static database. 0=land, 1=water, 2=land+water, 3=coast. pixel_surface_area; [km2]; surface aera covered by the pixel. The coordinates of the four pixel corners are used to compute the surface area. pseudo_cloud_fraction; [-]; Because the VIIRS cloud fraction IFOV is not available in the science v18_17 dataproduct for pixels that do not contain a valid xch4 value, we have used the processing quality flags (PQF) instead to generate a pseudo-cloudfraction data channel. Most pixels that contain high cloud fractions are filtered out, and thus do not have a valid xch4 value, and thus no data on the regular cloud fraction. Each pixel in the pseudo-cloudfraction channel can have as value 0, 0.5 or 1. Pixels with the PQF “cloud_warning” are assigned a value of 0.5, pixels with PQF “cf_viirs_swir_ifov_filter” are assigned a value of 1. Pixels without either of the specified PQF are assigned a value of 0. Full citation of the paper: Schuit, B. J., Maasakkers, J. D., Bijl, P., Mahapatra, G., van den Berg, A.-W., Pandey, S., Lorente, A., Borsdorff, T., Houweling, S., Varon, D. J., McKeever, J., Jervis, D., Girard, M., Irakulis-Loitxate, I., Gorroño, J., Guanter, L., Cusworth, D. H., and Aben, I.: Automated detection and monitoring of methane super-emitters using satellite data, Atmos. Chem. Phys., 23, 9071–9098, https://doi.org/10.5194/acp-23-9071-2023, 2023.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,002 | 0,005 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».