Roman-Cosmic Noon: A Legacy Spectroscopic Survey of Massive Field and Protocluster Galaxies at $2
Bibliographic record
Abstract
Protoclusters are the densest regions in the distant universe ($z>2$) and are the progenitors of massive galaxy clusters ($M_{halo}>10^{14}{\rm M}_\odot$) in the local universe. They undoubtedly play a key role in early massive galaxy evolution and they may host the earliest sites of galaxy quenching or even induce extreme states of star formation. Studying protoclusters therefore not only gives us a window into distant galaxy formation but also provides an important link in our understanding of how dense structures grow over time and modify the galaxies within them. Current protocluster samples are completely unable to address these points because they are small and selected in a heterogeneous way. We propose the Roman-Cosmic Noon survey, whose centerpiece is an extremely deep (30ksec) and wide area (10 deg$^2$) prism slitless spectroscopy survey to identify the full range of galaxy structures at $210^{10.5} {\rm M}_\odot$ across the full range of star formation histories as well as many more lower mass star-forming galaxies. The survey will also contain field galaxies to much lower masses than in the High Latitude Wide Area Survey, but over an area dwarfing any current or planned deep spectroscopy probe at $z>2$. With the prism spectroscopy and some modest additional imaging this survey will measure precise stellar mass functions, quenched fractions, galaxy and protocluster morphologies, stellar ages, emission-line based SFRs, and metallicities. It will have extensive legacy value well beyond the key protocluster science goals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".