Identifying Ly<i>α</i> emitter candidates with Random Forest: Learning from galaxies in the CANDELS survey
Bibliographic record
Abstract
The physical processes that make a galaxy a Lyman alpha emitter have been extensively studied over the past 25 yr. However, the correlations between physical and morphological properties of galaxies and the strength of the Lyα emission line are still highly debated. Here, we investigate the correlations between the rest-frame Lyα equivalent width and stellar mass, star formation rate, dust reddening, metallicity, age, half-light semi-major axis, Sérsic index, and projected axis ratio in a sample of 1578 galaxies in the redshift range of 2 ≤ z ≤ 7.9 from the GOODS-S, UDS, and COSMOS fields. From the large sample of Lyα emitters (LAEs) in the dataset, we find that LAEs are typically common main sequence (MS) star-forming galaxies that show a stellar mass ≤109 M⊙, star formation rate ≤ 100.5 M⊙ yr−1, E(B − V)≤0.2, and half-light semi-major axis ≤1 kpc. Building on these findings, we have developed a new method based on a random forest (RF) machine learning (ML) classifier to select galaxies with the highest probability of being Lyα emitters. When applied to a population in the redshift range z ∈ [2.5, 4.5], our classifier holds a (80 ± 2)% accuracy and (73 ± 4)% precision. At higher redshifts (z ∈ [4.5, 6]), we obtained an accuracy of 73% and precision of 80%. These results highlight the possibility of overcoming the current limitations in assembling large samples of LAEs by making informed predictions that can be used for planning future large-scale spectroscopic surveys.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".