Testing the performance of cross-correlation techniques to search for molecular features in <i>JWST</i> NIRSpec G395H observations of transiting exoplanets
Bibliographic record
Abstract
ABSTRACT Cross-correlations techniques offer an alternative method to search for molecular species in James Webb Space Telescope (JWST) observations of exoplanet atmospheres. In a previous article, we applied cross-correlation functions for the first time to JWST NIRSpec/G395H observations of exoplanet atmospheres, resulting in a detection of CO in the transmission spectrum of WASP-39b and a tentative detection of CO isotopologues. Here, we present an improved version of our cross-correlation technique and an investigation into how efficient the technique is when searching for other molecules in JWST NIRSpec/G395H data. Our search results in the detection of more molecules via cross-correlations in the atmosphere of WASP-39b, including $\rm H_{2}O$ and $\rm CO_{2}$, and confirms the CO detection. This result proves that cross-correlations are a robust and computationally cheap alternative method to search for molecular species in transmission spectra observed with JWST. We also searched for other molecules ($\rm CH_{4}$, $\rm NH_{3}$, $\rm SO_{2}$, $\rm N_{2}O$, $\rm H_{2}S$, $\rm PH_{3}$, $\rm O_{3}$, and $\rm C_{2}H_{2}$) that were not detected, for which we provide the definition of their cross-correlation baselines for future searches of those molecules in other targets. We find that that the cross-correlation search of each molecule is more efficient over limited wavelength regions of the spectrum, where the signal for that molecule dominates over other molecules, than over broad wavelength ranges. In general, we also find that Gaussian normalization is the most efficient normalization mode for the generation of the molecular templates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".