Adaptive friend-of-friends algorithm for identifying gravitationally bound cosmological structures
Bibliographic record
Abstract
The Universe at the present epoch is found to be a network of matter overdense and underdense regions. Usually, the overdense regions are dominated by the dark-matter (DM) filaments where massive gravitationally bound structures such as groups and clusters of galaxies form at the nodes. At the cosmological timescales, the baryonic matter follows the flow of DM only, and together they form the cosmic web. To date, this picture of the Universe is best revealed through cosmological large-volume simulations and large-scale galaxy redshift surveys, in which, the most important step is the appropriate identification of structures. So far, these structures are identified using various group finding codes, mostly based on the friend of friends (FoF) or spherical overdensity (SO) algorithms. Although, the main purpose is to identify gravitationally bound structures, surprisingly, the mass information has hardly been used effectively by these codes. Moreover, while it is an established fact that the bound structures can best be formed at some particular mass overdense regions and practically the large-scale structures are barely spherical in shape, the methods used so far either constrain the overdensity or use the real unstructured geometry only. Even though these are key factors in the accurate determination of structures-mass information that can precisely constrain the cosmological models of the Universe, hardly any attempt has been made as yet to consider these important parameters together while formulating the grouping algorithms. In this paper, we present our proposed algorithm called the measure of increased tie with gravity order which takes care of all the above-mentioned relevant features and ensures the bound structures by means of physical quantities, mainly mass and the total energy information. Unlike the usual FoF method where a statistically chosen single linking length is used for all grouping elements, we introduced a novel concept of physically relevant arm length for each element depending on their individual gravity leading to a distinct linking length for each unique pair of elements. This proposed algorithm is thus fundamentally new such that, not only able to catch the gravitationally bound, real unstructured geometry very well, it does identify it roughly within a predefined physically motivated density threshold. Such a thing could not be simultaneously achieved before by any of the usual FoF or SO-based methods. We also demonstrate the unique ability of the code in the appropriate identification of structures, both from large volume cosmological simulations as well as from galaxy redshift surveys, highlighting the fact that it mitigates a few shortcomings of the basic FoF and SO algorithms and strengthens the foundation of clustering or halo-finding methods in general.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.005 | 0.002 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".