The second Herschel–ATLAS Data Release – III. Optical and near-infrared counterparts in the North Galactic Plane field
Bibliographic record
Abstract
Abstract This paper forms part of the second major public data release of the Herschel Astrophysical Terahertz Large Area Survey (H-ATLAS). In this work, we describe the identification of optical and near-infrared counterparts to the submillimetre detected sources in the 177 deg2 North Galactic Plane (NGP) field. We used the likelihood ratio method to identify counterparts in the Sloan Digital Sky Survey and in the United Kingdom InfraRed Telescope Imaging Deep Sky Survey within a search radius of 10 arcsec of the H-ATLAS sources with a 4σ detection at 250 μm. We obtained reliable (R ≥ 0.8) optical counterparts with r < 22.4 for 42 429 H-ATLAS sources (37.8 per cent), with an estimated completeness of 71.7 per cent and a false identification rate of 4.7 per cent. We also identified counterparts in the near-infrared using deeper K-band data which covers a smaller ∼25 deg2. We found reliable near-infrared counterparts to 61.8 per cent of the 250-μm-selected sources within that area. We assessed the performance of the likelihood ratio method to identify optical and near-infrared counterparts taking into account the depth and area of both input catalogues. Using catalogues with the same surface density of objects in the overlapping ∼25 deg2 area, we obtained that the reliable fraction in the near-infrared (54.8 per cent) is significantly higher than in the optical (36.4 per cent). Finally, using deep radio data which covers a small region of the NGP field, we found that 80–90 per cent of our reliable identifications are correct.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.009 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".