CatNorth: An Improved Gaia DR3 Quasar Candidate Catalog with Pan-STARRS1 and CatWISE
Bibliographic record
Abstract
Abstract A complete and pure sample of quasars with accurate redshifts is crucial for quasar studies and cosmology. In this paper, we present CatNorth, an improved Gaia Data Release 3 (Gaia DR3) quasar candidate catalog with more than 1.5 million sources in the 3π sky built with data from Gaia, Pan-STARRS1, and CatWISE2020. The XGBoost algorithm is used to reclassify the original Gaia DR3 quasar candidates as stars, galaxies, and quasars. To construct training/validation data sets for the classification, we carefully built two different master stellar samples in addition to the spectroscopic galaxy and quasar samples. An ensemble classification model is obtained by averaging two XGBoost classifiers trained with different master stellar samples. Using a probability threshold of p QSO_mean > 0.95 in our ensemble classification model and an additional cut on the logarithmic probability density of zero proper motion, we retrieved 1,545,514 reliable quasar candidates from the parent Gaia DR3 quasar candidate catalog. We provide photometric redshifts for all candidates with an ensemble regression model. For a subset of 89,100 candidates, accurate spectroscopic redshifts are estimated with the convolutional neural network from the Gaia BP/RP spectra. The CatNorth catalog has a high purity of ∼90%, while maintaining high completeness, which is an ideal sample to understand the quasar population and its statistical properties. The CatNorth catalog is used as the main source of input catalog for the Large Sky Area Multi-Object Fiber Spectroscopic Telescope phase III quasar survey, which is expected to build a highly complete sample of bright quasars with i < 19.5.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".