Correcting for small-displacement interlopers in BAO analyses
Bibliographic record
Abstract
Abstract Due to the low resolution of slitless spectroscopy, future surveys including those made possible by the Roman and Euclid space telescopes will be prone to line mis-identification, leading to interloper galaxies at the wrong redshifts in the large-scale structure catalogues. The most pernicious of these have a small displacement between true and false redshift such that the interloper positions are correlated with the target galaxies. We consider how to correct for such contaminants, focusing on Hβ interlopers in [Oiii] catalogues as will be observed by Roman, which are misplaced by Δd = 97 h -1 Mpc at redshift z = 1. Because this displacement is close to the BAO scale, the peak in the interloper-target galaxy cross-correlation function at the displacement scale can change the shape of the BAO peak in the auto-correlation of the contaminated catalogue, and lead to incorrect cosmological measurements if not accounted for properly. We consider how to build a model for the monopole and quadrupole moments of the contaminated correlation function, including an additional free parameter for the fraction of interlopers. The key input to this model is the cross-correlation between the population of galaxies forming the interlopers and the main target sample. It will be important to either estimate this using calibration data or to use the contaminated small-scale auto-correlation function to model it, which may be possible if a number of requirements about the galaxy populations are met. We find that this method is successful in measuring the BAO dilation parameters without significant degradation in accuracy, provided the cross-correlation function is accurately known.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.023 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".