The Ground-Based Musica Dataset: Tropospheric Water Vapour Isotopologues (H216O, H218O And Hd16O) As Obtained From Ndacc/Ftir Solar Absorption Spectra
Bibliographic record
Abstract
MUSICA (“MUlti-platform remote sensing of Isotopologues for investigating the Cycle of Atmospheric water”, http://www.imk-asf.kit.edu/english/musica.php) is a European Research Council (ERC) project. The project has developed tropospheric water vapour isotopologue retrievals (H2O and H2O-δD pairs) using ground-based FTIR spectra as well as thermal nadir spectra measured by the satellite sensor IASI. H2O-δD pairs allow studying tropospheric water transport pathways and in combination with models they can improve our understanding of important climate feedback mechanisms (see also WCRP Grand Challenges: http://www.wcrp-climate.org/grand-challenges).<br> <br> For MUSICA, the FTIR spectra have been analysed centrally at KIT using uniform and consistent retrieval settings, thereby guaranteeing ultimate consistency of the retrieval products generated for different FTIR stations. The FTIR products are H2O profiles for the lower, middle and upper troposphere as well as H2O-δD pairs for the lower and middle troposphere. The data have been produced for 12 FTIR stations and date back to 1996. The dataset has been extensively characterized and validated (theoretically and empirically). Furthermore, the spectra have been used to perform uniform retrievals of XCO<sub>2</sub>, which is then used for documenting the long-term stability of these kind of FTIR data. The data are provided in the form of two data types. The first type ("ftir.iso.h2o") is best-suited for tropospheric water vapour distribution studies that disregard the different isotopologues (comparison with radiosonde data, analyses of water vapour variability and trends, etc.). The second type ("ftir.iso.post.h2o") is needed for analysing moisture pathways by means of H<sub>2</sub>O-δD pair distribution. The data format is hdf4 and the files have been generated in compliance with GEOMS (Generic Earth Observation Metadata Standard). The complete MUSICA NDACC/FTIR dataset is also publicly available via the NDACC database (ftp://ftp.cpc.ncep.noaa.gov/ndacc/MUSICA). Details on the characteristics of the dataset are described in the paper "Tropospheric water vapour isotopoloque data (H\(_{2}^{16}\)O, H\(_{2}^{18}\)O and HD<sup>16</sup>O) as obtained from NDACC/FTIR solar absorption spectra" that has been prepared for ESSD in the context of the special issue “25th anniversary of NDACC”.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.004 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.007 | 0.014 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".