Automated morphological classification of Sloan Digital Sky Survey red sequence galaxies
Bibliographic record
Abstract
In the last decade, the advent of enormous galaxy surveys has motivated the development of automated morphological classification schemes to deal with large data volumes. Existing automated schemes can successfully distinguish between early- and late-type galaxies and identify merger candidates, but are inadequate for studying detailed morphologies of red sequence galaxies. To fill this need, we present a new automated classification scheme that focuses on making finer distinctions between early types roughly corresponding to Hubble types E, S0 and Sa. We visually classify a sample of 984 non-star-forming Sloan Digital Sky Survey galaxies with apparent sizes >14 arcsec. We then develop an automated method to closely reproduce the visual classifications, which both provides a check on the visual results and makes it possible to extend morphological analysis to much larger samples. We visually classify the galaxies into three bulge classes (BC) by the shape of the light profile in the outer regions: discs have sharp edges and bulges do not, while some galaxies are intermediate. We separately identify galaxies with features: spiral arms, bars, clumps, rings and dust. We find general agreement between BC and the bulge fraction B/T measured by the galaxy modelling package gim2d, but many visual discs have B/T > 0.5. Three additional automated parameters – smoothness, axial ratio and concentration – can identify many of these high-B/T discs to yield automated classifications that agree ∼70 per cent with the visual classifications (>90 per cent within one BC). Tests versus disc inclination indicate that both methods identify most face-on discs, but visually, features are lost in edge-on discs. 80 per cent of face-on visual discs have features while few visual bulges do, strongly validating the visual classifications. Given the good agreement between the visual and automated methods, we believe that the automated method can be applied to a much larger sample with confidence. Both methods are used to study the bulge versus disc frequency as a function of four measures of galaxy ‘size’: luminosity, stellar mass, velocity dispersion (σ) and radius (R). All size indicators show a fall in disc fraction and a rise in bulge fraction among larger galaxies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.006 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".