The importance of galaxy formation histories in models of reionization
Bibliographic record
Abstract
ABSTRACT Upcoming galaxy surveys and 21-cm experiments targeting high redshifts z ≳ 6 are highly complementary probes of galaxy formation and reionization. However, in order to expedite the large-volume simulations relevant for 21-cm observations, many models of galaxies within reionization codes are entirely subgrid and/or rely on halo abundances only. In this work, we explore the extent to which resolving and modelling individual galaxy formation histories affects predictions both for the galaxy populations detectable by upcoming surveys and the signatures of reionization accessible to upcoming 21-cm experiments. We find that a common approach, in which galaxy luminosity is assumed to be a function of halo mass only, is biased with respect to models in which galaxy properties are evolved through time via semi-analytic modelling and thus reflective of the diversity of assembly histories that naturally arise in N-body simulations. The diversity of galaxy formation histories also results in scenarios in which the brightest galaxies do not always reside in the centres of large-ionized regions, as there are often relatively low-mass haloes undergoing dramatic, but short-term, growth. This has clear implications for attempts to detect or validate the 21-cm background via cross-correlation. Finally, we show that a hybrid approach – in which only haloes hosting galaxies bright enough to be detected in surveys are modelled in detail, with the rest modelled as an unresolved field of haloes with abundance related to large-scale overdensity – is a viable way to generate large-volume ‘simulations‘ well suited to wide-area surveys and current-generation 21-cm experiments targeting relatively large k ≲ 1 h Mpc−1 scales.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".