MétaCan
Menu
Back to cohort
Record W4393948564 · doi:10.3847/1538-4365/ad2ae6

CatNorth: An Improved Gaia DR3 Quasar Candidate Catalog with Pan-STARRS1 and CatWISE

2024· article· en· W4393948564 on OpenAlexfundno aff
Yuming Fu, Xue-Bing Wu, Yifan Li, Yuxuan Pang, Ravi Joshi, Shuo Zhang, Qiyue Wang, Jing Yang, FanLam Ng, Xingjian Liu, Yu Qiu, Rui Zhu, Huimei Wang, Christian Wolf, Yanxia Zhang, Zhi-Ying Huo, Y. L. Ai, Qinchun Ma, Xiaotong Feng, R. J. Bouwens

Bibliographic record

VenueThe Astrophysical Journal Supplement Series · 2024
Typearticle
Languageen
FieldPhysics and Astronomy
TopicGalaxies: Formation, Evolution, Phenomena
Canadian institutionsnot available
FundersCore Research for Evolutional Science and TechnologyPlanetary Science DivisionBasic and Applied Basic Research Foundation of Guangdong ProvinceLeibniz-GemeinschaftUniversity of Colorado BoulderUniversity of California, Los AngelesNational Astronomical Observatories, Chinese Academy of SciencesUniversity of Illinois at Urbana-ChampaignMax-Planck-Institut für AstronomieJet Propulsion LaboratoryLeibniz-Institut für Astrophysik PotsdamSmithsonian Astrophysical ObservatoryNational Development and Reform CommissionChina National Textile and Apparel CouncilEötvös Loránd TudományegyetemPeking UniversityNational Central UniversityChinese Academy of SciencesSmithsonian InstitutionNational Natural Science Foundation of ChinaÉcole Polytechnique Fédérale de LausanneShenzhen Technology UniversityGordon and Betty Moore FoundationQueen's University BelfastNational Key Research and Development Program of ChinaUniversidad Nacional Autónoma de MéxicoSpace Telescope Science InstituteLos Alamos National LaboratoryEuropean Space AgencyAlfred P. Sloan FoundationJohns Hopkins UniversityCarnegie Institution of WashingtonUniversity of UtahChina Postdoctoral Science FoundationHarvard UniversityQueen's UniversityOhio State UniversityDurham UniversityCalifornia Institute of TechnologyNational Aeronautics and Space AdministrationNanjing UniversityNew Mexico State UniversityScience Mission DirectorateUniversity of TorontoYale UniversityYunnan UniversityNational Science Foundation
KeywordsQuasarPhysicsAstrophysicsSkyGalaxyRedshiftOVV quasarPhotometric redshiftAstronomy

Abstract

fetched live from OpenAlex

Abstract A complete and pure sample of quasars with accurate redshifts is crucial for quasar studies and cosmology. In this paper, we present CatNorth, an improved Gaia Data Release 3 (Gaia DR3) quasar candidate catalog with more than 1.5 million sources in the 3π sky built with data from Gaia, Pan-STARRS1, and CatWISE2020. The XGBoost algorithm is used to reclassify the original Gaia DR3 quasar candidates as stars, galaxies, and quasars. To construct training/validation data sets for the classification, we carefully built two different master stellar samples in addition to the spectroscopic galaxy and quasar samples. An ensemble classification model is obtained by averaging two XGBoost classifiers trained with different master stellar samples. Using a probability threshold of p QSO_mean > 0.95 in our ensemble classification model and an additional cut on the logarithmic probability density of zero proper motion, we retrieved 1,545,514 reliable quasar candidates from the parent Gaia DR3 quasar candidate catalog. We provide photometric redshifts for all candidates with an ensemble regression model. For a subset of 89,100 candidates, accurate spectroscopic redshifts are estimated with the convolutional neural network from the Gaia BP/RP spectra. The CatNorth catalog has a high purity of ∼90%, while maintaining high completeness, which is an ideal sample to understand the quasar population and its statistical properties. The CatNorth catalog is used as the main source of input catalog for the Large Sky Area Multi-Object Fiber Spectroscopic Telescope phase III quasar survey, which is expected to build a highly complete sample of bright quasars with i < 19.5.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.016
Threshold uncertainty score0.031

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0050.003
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0060.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.005
GPT teacher head0.214
Teacher spread0.209 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations22
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueThe Astrophysical Journal Supplement SeriesSame topicGalaxies: Formation, Evolution, PhenomenaFrench-language works237,207