Cataclysmic variables from Sloan Digital Sky Survey – V (2020–2023) identified using machine learning
Bibliographic record
Abstract
ABSTRACT SDSS-V is carrying out a dedicated survey for white dwarfs, single and in binaries, and we report the analysis of the spectroscopy of 504 cataclysmic variables (CVs) and CV candidates obtained during the first 34 months of observations of SDSS-V. We developed a convolutional neural network (CNN) to aid with the identification of CV candidates among the over 2 million SDSS-V spectra obtained with the BOSS spectrograph. The CNN reduced the number of spectra that required visual inspection to $\simeq 2$ per cent of the total. We identified 776 CV spectra among the CNN-selected candidates, plus an additional 27 CV spectra that the CNN misclassified, but that were found serendipitously by human inspection of the data. Analysing the SDSS-V spectroscopy and ancillary data of the 504 CVs in our sample, we report 61 new CVs, spectroscopically confirm 248 and refute 13 published CV candidates, and we report 82 new or improved orbital periods. We discuss the completeness and possible selection biases of the machine learning methodology, as well as the effectiveness of targeting CV candidates within SDSS-V. Finally, we re-assess the space density of CVs, and find $1.2\times 10^{-5}\, \mathrm{pc^{-3}}$.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".