CURLING -- III. Identifying Candidates of Wide-separation Gravitationally Lensed Quasars from the CatNorth Catalogue
Bibliographic record
Abstract
Wide-separation lensed quasars (WSLQs) are a rare subclass of strongly lensed quasars produced by massive galaxy clusters. They provide valuable probes of dark-matter halos and quasar host galaxies. However, only about ten WSLQ systems are currently known, which limits further studies. To enlarge the sample from wide-area surveys, we developed a catalog-based pipeline and applied it to the CatNorth database, a catalog of quasar candidates constructed from Gaia DR3. CatNorth contains 1,545,514 quasar candidates with about 90% purity and a Gaia G-band limiting magnitude of roughly 21. The pipeline has three stages. First, we identify groups with separations between 10 and 72 arcsec using a HEALPix grid with 25.6 arcsec spacing and a friends-of-friends search. We then filter by intra-group color and spectral similarity, reducing the 1,545,514 sources to 14,244 groups while retaining all known, discoverable WSLQs. Finally, a visual check, guided by image geometry and the presence of likely foreground lenses, yields the candidate list with quality labels. We identify 333 new WSLQ candidates with separations from 10 to 56.8 arcsec. Using available SDSS DR16 and DESI DR1 spectroscopy, we uncover two new candidate systems; the remaining 331 candidates lack sufficient spectra and are labeled as 45 grade A, 98 grade B, and 188 grade C. We also compile 29 confirmed dual quasars as a by-product. When feasible, we plan follow-up spectroscopy and deeper imaging to confirm WSLQs among these candidates and enable the related science.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.021 | 0.013 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".