Spectroscopic classification of a complete sample of astrometrically-selected quasar candidates using <i>Gaia</i> DR2
Bibliographic record
Abstract
Here we explore the efficiency and fidelity of a purely astrometric selection of quasars as point sources with zero proper motions in the Gaia data release 2 (DR2). We have built a complete candidate sample including 104 Gaia -DR2 point sources, which are brighter than 20th magnitude in the Gaia G -band within one degree of the north Galactic pole (NGP); all of them have proper motions that are consistent with zero within 2 σ uncertainty. In addition to pre-existing spectra, we have secured long-slit spectroscopy of all the remaining candidates and find that all 104 stationary point sources in the field can be classified as either quasars (63) or stars (41). One of the new quasars that we discover is particularly interesting as the line-of-sight to it passes through the disc of a foreground ( z = 0.022) galaxy, which imprints both Na D absorption and dust extinction on the quasar spectrum. The selection efficiency of the zero-proper-motion criterion at high Galactic latitudes is thus ≈60%. Based on this complete quasar sample, we examine the basic properties of the underlying quasar population within the imposed limiting magnitude. We find that the surface density of quasars is 20 deg −2 (at G < 20 mag), the redshift distribution peaks at z ∼ 1.5, and only eight systems (13 -3 +5 %) show significant dust reddening. We then explore the selection efficiency of commonly used optical, near-, and mid-infrared quasar identification techniques and find that they are all complete at the 85−90% level compared to the astrometric selection. Finally, we discuss how the astrometric selection can be improved to an efficiency of ≈70% by including an additional cut requiring parallaxes of the candidates to be consistent with zero within 2 σ . The selection efficiency will further increase with the release of future, more sensitive astrometric measurements from the Gaia mission. This type of selection, which is purely based on the astrometry of the quasar candidates, is unbiased in terms of colours and intrinsic emission mechanisms of the quasars and thus provides the most complete census of the quasar population within the limiting magnitude of Gaia .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".