Beyond UVJ: Color Selection of Galaxies in the JWST Era
Bibliographic record
Abstract
Abstract We present a new rest-frame color–color selection method using synthetic u s − g s and g s − i s , ( ugi ) s colors to identify star-forming and quiescent galaxies. Our method is similar to the widely used U − V versus V − J ( UVJ ) diagram. However, UVJ suffers known systematics. Spectroscopic campaigns have shown that UVJ -selected quiescent samples at z ≳ 3 include ∼10%–30% contamination from galaxies with dust-obscured star formation and strong emission lines. Moreover, at z > 3, UVJ colors are extrapolated because the rest-frame band shifts beyond the coverage of the deepest bandpasses at <5 μ m (typically Spitzer/IRAC 4.5 μ m or future JWST/NIRCam observations). We demonstrate that ( ugi ) s offers improvements to UVJ at z > 3, and can be applied to galaxies in the JWST era. We apply ( ugi ) s selection to galaxies at 0.5 < z < 6 from the (observed) 3D-HST and UltraVISTA catalogs, and to the (simulated) JAGUAR catalogs. We show that extrapolation can affect ( V − J ) 0 color by up to 1 mag, but changes ( g s − i s ) 0 color by ≤0.2 mag, even at z ≃ 6. While ( ugi ) s -selected quiescent samples are comparable to UVJ in completeness (both achieve ∼85%–90% at z = 3–3.5), ( ugi ) s reduces contamination in quiescent samples by nearly a factor of 2, from ≃35% to ≃17% at z = 3, and from ≃60% to ≃33% at z = 6. This leads to improvements in the true-to-false-positive ratio (TP/FP), where we find TP/FP ≳2.2 for ( ugi ) s at z ≃ 3.5 − 6, compared to TP/FP < 1 for UVJ -selected samples. This indicates that contaminants will outnumber true quiescent galaxies in UVJ at these redshifts, while ( ugi ) s will provide higher-fidelity samples.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".