DESI Early Data Release Milky Way Survey value-added catalogue
Bibliographic record
Abstract
ABSTRACT We present the stellar value-added catalogue based on the Dark Energy Spectroscopic Instrument (DESI) Early Data Release. The catalogue contains radial velocity and stellar parameter measurements for $\simeq$ 400 000 unique stars observed during commissioning and survey validation by DESI. These observations were made under conditions similar to the Milky Way Survey (MWS) currently carried out by DESI but also include multiple specially targeted fields, such as those containing well-studied dwarf galaxies and stellar streams. The majority of observed stars have $16\lt r\lt 20$ with a median signal-to-noise ratio in the spectra of $\sim$ 20. In the paper, we describe the structure of the catalogue, give an overview of different target classes observed, as well as provide recipes for selecting clean stellar samples. We validate the catalogue using external high-resolution measurements and show that radial velocities, surface gravities, and iron abundances determined by DESI are accurate to 1 km s−1, 0.3 dex, and $\sim$ 0.15 dex respectively. We also demonstrate possible uses of the catalogue for chemo-dynamical studies of the Milky Way stellar halo and Draco dwarf spheroidal. The value-added catalogue described in this paper is the very first DESI MWS catalogue. The next DESI data release, expected in less than a year, will add the data from the first year of DESI survey operations and will contain approximately 4 million stars, along with significant processing improvements.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.015 | 0.026 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".