The Pristine survey – I. Mining the Galaxy for the most metal-poor stars
Bibliographic record
Abstract
We present the Pristine survey, a new narrow-band photometric survey focused on the metallicity-sensitive Ca H&K lines and conducted in the Northern hemisphere with the wide-field imager MegaCam on the Canada-France-Hawaii Telescope. This paper reviews our overall survey strategy and discusses the data processing and metallicity calibration. Additionally we review the application of these data to the main aims of the survey, which are to gather a large sample of the most metal-poor stars in the Galaxy, to further characterize the faintest Milky Way satellites, and to map the (metal-poor) substructure in the Galactic halo. The current Pristine footprint comprises over 1000 deg2 in the Galactic halo ranging from b ˜ 30° to ˜78° and covers many known stellar substructures. We demonstrate that, for Sloan Digital Sky Survey (SDSS) stellar objects, we can calibrate the photometry at the 0.02-mag level. The comparison with existing spectroscopic metallicities from SDSS/Sloan Extension for Galactic Understanding and Exploration (SEGUE) and Large Sky Area Multi-Object Fiber Spectroscopic Telescope shows that, when combined with SDSS broad-band g and I photometry, we can use the CaHK photometry to infer photometric metallicities with an accuracy of ˜0.2 dex from [Fe/H] = -0.5 down to the extremely metal-poor regime ([Fe/H] < -3.0). After the removal of various contaminants, we can efficiently select metal-poor stars and build a very complete sample with high purity. The success rate of uncovering [Fe/H]SEGUE < -3.0 stars among [Fe/H]Pristine < -3.0 selected stars is 24 per cent, and 85 per cent of the remaining candidates are still very metal poor ([Fe/H]<-2.0). We further demonstrate that Pristine is well suited to identify the very rare and pristine Galactic stars with [Fe/H] < -4.0, which can teach us valuable lessons about the early Universe.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".