Artificial neural network based calibrations for the prediction of galactic [N ii] λ6584 and Hα line luminosities
Bibliographic record
Abstract
The artificial neural network (ANN) is a well-established mathematical technique for data prediction, based on the identification of correlations and pattern recognition in input training sets. We present the application of ANNs to predict the emission line luminosities of Hα and [N ii] λ6584 in galaxies. These important spectral diagnostics are used for metallicities, active galactic nuclei (AGN) classification and star formation rates, yet are shifted into the infrared for galaxies above z ∼ 0.5, or may not be covered in spectra with limited wavelength coverage. The ANN is trained with a large sample of emission line galaxies selected from the Sloan Digital Sky Survey (SDSS) using various combinations of emission lines and stellar mass. The ANN is tested for galaxies dominated by both star formation and AGN; in both cases the Hα and [N ii] λ6584 line luminosities can be predicted with a scatter σ < 0.1 dex. We also show that the performance of the ANN does not depend significantly on the covering fraction, mass or metallicity of the data. Polynomial functions are derived that allow easy application of the ANN predictions to determine Hα and [N ii] λ6584 line luminosities. An ANN calibration for the Balmer decrement (Hα/Hβ) based on line equivalent widths and colours is also presented. The effectiveness of the ANN calibration is demonstrated with an independent data set (the Galaxy Mass and Assembly Survey). We demonstrate the application of our line luminosities to the determination of gas-phase metallicities and AGN classification. The ANN technique yields a significant improvement in the measurement of metallicities that require [N ii] and Hα when compared with the function-based conversions of Kewley & Ellison. The AGN classification is successful for 86 per cent of SDSS galaxies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".