Hidden deep in the halo: selection of a reduced proper motion halo catalogue and mining retrograde streams in the velocity space
Bibliographic record
Abstract
ABSTRACT The Milky Way halo is one of the few galactic haloes that provides a unique insight into galaxy formation by resolved stellar populations. Here, we present a catalogue of ∼47 million halo stars selected independent of parallax and line-of-sight velocities, using a combination of Gaia DR3 proper motion and photometry by means of their reduced proper motion. We select high tangential velocity (halo) main sequence stars and fit distances to them using their simple colour-absolute-magnitude relation. This sample reaches out to ∼21 kpc with a median distance of 6.6 kpc thereby probing much further out than would be possible using reliable Gaia parallaxes. The typical uncertainty in their distances is $0.57_{-0.26}^{+0.56}$ kpc. Using the colour range 0.45 < (G0 − GRP, 0) < 0.715, where the main sequence is narrower, gives an even better accuracy down to $0.39_{-0.12}^{+0.18}$ kpc in distance. The median velocity uncertainty for stars within this colour range is 15.5 km s−1. The distribution of these sources in the sky, together with their tangential component velocities, are very well-suited to study retrograde substructures. We explore the selection of two complex retrograde streams: GD-1 and Jhelum. For these streams, we resolve the gaps, wiggles and density breaks reported in the literature more clearly. We also illustrate the effect of the kinematic selection bias towards high proper motion stars and incompleteness at larger distances due to Gaia’s scanning law. These examples showcase how the full RPM catalogue made available here can help us paint a more detailed picture of the build-up of the Milky Way halo.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".