Detecting clusters and groups of galaxies populating the local Universe in large optical spectroscopic surveys
Bibliographic record
Abstract
Wide-field cosmological surveys provide hundreds of thousands of spectroscopically confirmed galaxy groups and clusters, valuable for tracing baryonic matter distribution. However, controlling systematics in identifying host dark matter halos and estimating their properties is crucial. We evaluate three group detection methods on a simulated dataset replicating the GAMA selection to understand systematics and selection effects. This is key for interpreting data from SDSS, GAMA, DESI, WAVES, and leveraging optical catalogues in the (X-ray) eROSITA era to quantify baryonic mass in galaxy groups. Using a lightcone from the Magneticum hydrodynamical simulation, we simulate a spectroscopic galaxy survey in the local Universe (down to $z<0.2$ and stellar mass completeness $M_{\star}\geq10^{9.8} M_{\odot}$). We assess completeness and contamination of reconstructed halo catalogues, evaluate membership accuracy, and analyse the halo mass recovery rate of group finders. All three group finders achieve high completeness ($>80\%$) at group and cluster scales, confirming optical selection's suitability for dense regions. Contamination at low masses ($M_{200}<10^{13} M_{\odot}$) arises from interlopers and fragmentation. Membership is at least 70\% accurate above the group mass scale, but inaccuracies bias halo mass estimates using galaxy velocity dispersion. Alternative proxies, like total stellar luminosity or mass, yield more accurate halo masses. The cumulative luminosity function of galaxy members matches predictions, showing the group finders' accuracy in identifying galaxy populations. These results confirm the reliability and completeness of spectroscopic catalogues produced by state-of-the-art group finders. This supports studies requiring large spectroscopic samples of galaxy groups and clusters, as well as investigations into galaxy evolution across diverse environments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".