VEXAS: VISTA EXtension to Auxiliary Surveys
Bibliographic record
Abstract
Context.We present the second public data release of the VISTA EXtension to Auxiliary Surveys (VEXAS), where we classify objects into stars, galaxies, and quasars based on an ensemble of machine learning algorithms. Aims.The aim of VEXAS is to build the widest multi-wavelength catalogue, providing reference magnitudes, colours, and morphological information for a large number of scientific uses. Methods.We applied an ensemble of thirty-two different machine learning models, based on three different algorithms and on different magnitude sets, training samples, and classification problems (two or three classes) on the three VEXAS Data Release 1 (DR1) optical and infrared (IR) tables. The tables were created in DR1 cross-matching VISTA near-infrared data with Wide-field Infrared Survey Explorer far-infrared data and with optical magnitudes from the Dark Energy Survey (VEXAS-DESW), the Sky Mapper Survey (VEXAS-SMW), and the Panoramic Survey Telescope and Rapid Response System Survey (VEXAS-PSW). We assembled a large table of spectroscopically confirmed objects (VEXAS-SPEC-GOOD, 415 628 unique objects), based on the combination of six different spectroscopic surveys that we used for training. We developed feature imputation to also classify objects for which magnitudes in one or more bands are missing. Results.We classify in total ≈90 × 106objects in the Southern Hemisphere. Among these, ≈62.9 × 106(≈52.6 × 106) are classified as ‘high confidence’ (‘secure’) stars, ≈920 000 (≈750 000) as ‘high confidence’ (‘secure’) quasars, and ≈34.8 (≈34.1) million as ‘high confidence’ (‘secure’) galaxies, withpclass ≥ 0.7 (pclass ≥ 0.9). The DR2 tables update the DR1 with the addition of imputed magnitudes and membership probabilities to each of the three classes. Conclusions.The density of high-confidence extragalactic objects varies strongly with the survey depth: atpclass > 0.7, there are 11 deg−2quasars in the VEXAS-DESW footprint and 103 deg−2in the VEXAS-PSW footprint, while only 10.7 deg−2in the VEXAS-SM footprint. Improved depth in the mid-infrared and coverage in the optical and near-infrared are needed for the SM footprint that is not already covered by DESW and PSW.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.028 | 0.021 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".