The miniJPAS survey: A search for extreme emission-line galaxies
Bibliographic record
Abstract
Context. Galaxies with extreme emission lines (EELGs) may play a key role in the evolution of the Universe, as well as in our understanding of the star formation process itself. For this reason an accurate determination of their spatial density and fundamental properties in different epochs of the Universe will constitute a unique perspective towards a comprehensive picture of the interplay between star formation and mass assembly in galaxies. In addition to this, EELGs are also interesting in order to explain the reionization of the Universe, since their interstellar medium (ISM) could be leaking ionizing photons, and thus they could be low z, analogous of extreme galaxies at high z. Aims. This paper presents a method to obtain a census of EELGs over a large area of the sky by detecting galaxies with rest-frame equivalent widths ≥300 Å in the emission lines [O II]λλ3727,3729Å, [O III]λ5007Å, and Hα. For this, we aim to use the J-PAS survey, which will image an area of ≈8000 deg2 with 56 narrow band filters in the optical. As a pilot study, we present a methodology designed to select EELGs on the miniJPAS images, which use the same filter dataset as J-PAS, and thus will be exportable to this larger survey. Methods. We make use of the miniJPAS survey data, conceived as a proof of concept of J-PAS, and covering an area of ≈1 deg2. Objects were detected in the rSDSS images and selected by imposing a condition on the flux in a given narrow-band J-PAS filter with respect to the contiguous ones, which is analogous to requiring an observed equivalent width larger than 300 Å in a certain emission line within the filter bandwidth. The selected sources were then classified as galaxies or quasi-stellar objects (QSOs) after a comparison of their miniJPAS fluxes with those of a spectral database of objects known to present strong emission lines. This comparison also provided a redshift for each source, which turned out to be consistent with the spectroscopic redshifts when available (|Δz/(1 + zspec)| ≤ 0.01). Results. The selected candidates were found to show a compact appearance in the optical images, some of them even being classified as point-like sources according to their stellarity index. After discarding sources classified as QSOs, a total of 17 sources turned out to exhibit EW0 ≥ 300 Å in at least one emission line, thus constituting our final list of EELGs. Our counts are fairly consistent with those of other samples of EELGs in the literature, although there are some differences, which were expected due to biases resulting from different selection criteria.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".