Simultaneous calibration of spectro-photometric distances and the Gaia DR2 parallax zero-point offset with deep learning
Bibliographic record
Abstract
ABSTRACT Gaia measures the five astrometric parameters for stars in the Milky Way, but only four of them (positions and proper motion, but not distance) are well measured beyond a few kpc from the Sun. Modern spectroscopic surveys such as APOGEE cover a large area of the Milky Way disc and we can use the relation between spectra and luminosity to determine distances to stars beyond Gaia’s parallax reach. Here, we design a deep neural network trained on stars in common between Gaia and APOGEE that determines spectro-photometric distances to APOGEE stars, while including a flexible model to calibrate parallax zero-point biases in Gaia DR2. We determine the zero-point offset to be $-52.3 \pm 2.0\, \mu \mathrm{as}$ when modelling it as a global constant, but also train a multivariate zero-point offset model that depends on G, GBP − GRP colour, and Teff and that can be applied to all ≈58 million stars in Gaia DR2 within APOGEE’s colour–magnitude range and within APOGEE’s sky footprint. Our spectro-photometric distances are more precise than Gaia at distances ${\gtrsim} 2\, \mathrm{kpc}$ from the Sun. We release a catalogue of spectro-photometric distances for the entire APOGEE DR14 data set which covers Galactocentric radii $2\, \mathrm{kpc} \lesssim R \lesssim 19\, \mathrm{kpc}$; ${\approx} 150\, 000$ stars have ${\lt} 10{{\ \rm per\ cent}}$ uncertainty, making this a powerful sample to study the chemo-dynamical structure of the disc. We use this sample to map the mean [Fe/H] and 15 abundance ratios [X/Fe] from the Galactic Centre to the edge of the disc. Among many interesting trends, we find that the bulge and bar region at $R \lesssim 5\, \mathrm{kpc}$ clearly stands out in [Fe/H] and most abundance ratios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".