CLASS Survey Description: Coronal-line Needles in the SDSS Haystack
Bibliographic record
Abstract
Abstract Coronal lines are a powerful, yet poorly understood, tool to identify and characterize active galactic nuclei. There have been few large-scale surveys of coronal lines in the general galaxy population in the literature so far. Using a novel preselection technique with a flux-to-rms ratio , followed by Markov Chain Monte Carlo fitting, we searched for the full suite of 20 coronal lines in the optical spectra of almost 1 million galaxies from the Sloan Digital Sky Survey Data Release 8. We present a catalog of the emission-line parameters for the resulting 258 galaxies with detections. The Coronal Line Activity Spectroscopic Survey includes line properties, host-galaxy properties, and selection criteria for all galaxies in which at least one line is detected. This comprehensive study reveals that a significant fraction of coronal-line activity is missed in past surveys based on a more limited set of coronal lines; ∼60% of our sample do not display the more widely surveyed [Fe x] λ6374. In addition, we discover a strong correlation between coronal-line and Wide-field Infrared Survey Explorer W2 luminosities, suggesting that the mid-infrared flux can be used to predict coronal-line fluxes. For each line we also provide a confidence level that the line is present, generated by a novel neural network, trained on fully simulated data. We find that after training the network to detect individual lines using 100,000 simulated spectra, we achieve an overall true-positive rate of 75.49% and a false-positive rate of only 3.96%.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.157 | 0.118 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".