The impact from survey depth and resolution on the morphological classification of galaxies
Bibliographic record
Abstract
We consistently analyse for the first time the impact of survey depth and spatial resolution on the most used morphological parameters for classifying galaxies through non-parametric methods: Abraham and Conselice–Bershady concentration indices, Gini, M20 moment of light, asymmetry, and smoothness. Three different non-local data sets are used, Advanced Large Homogeneous Area Medium Band Redshift Astronomical (ALHAMBRA) and Subaru/XMM-Newton Deep Survey (SXDS, examples of deep ground-based surveys), and Cosmos Evolution Survey (COSMOS, deep space-based survey). We used a sample of 3000 local, visually classified galaxies, measuring their morphological parameters at their real redshifts (z ∼ 0). Then we simulated them to match the redshift and magnitude distributions of galaxies in the non-local surveys. The comparisons of the two sets allow us to put constraints on the use of each parameter for morphological classification and evaluate the effectiveness of the commonly used morphological diagnostic diagrams. All analysed parameters suffer from biases related to spatial resolution and depth, the impact of the former being much stronger. When including asymmetry and smoothness in classification diagrams, the noise effects must be taken into account carefully, especially for ground-based surveys. M20 is significantly affected, changing both the shape and range of its distribution at all brightness levels. We suggest that diagnostic diagrams based on 2–3 parameters should be avoided when classifying galaxies in ground-based surveys, independently of their brightness; for COSMOS they should be avoided for galaxies fainter than F814 = 23.0. These results can be applied directly to surveys similar to ALHAMBRA, SXDS and COSMOS, and also can serve as an upper/lower limit for shallower/deeper ones.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.042 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".