A robust morphological classification of high-redshift galaxies using support vector machines on seeing limited images
Bibliographic record
Abstract
Context. Morphology is the most accessible tracer of galaxies physical structure, but its interpretation in the framework of galaxy evolution still remains problematic. Its quantification at high redshift requires deep high-angular resolution imaging, which is why space data (HST) are usually employed. At , the HST visible cameras however probe the UV flux, which is dominated by the emission of young stars, which could bias the estimated morphologies towards late-type systems.Aims. In this paper we quantify the effects of this morphological k-correction at by comparing morphologies measured in the K and I-bands in the COSMOS area. The Ks-band data indeed have the advantage of probing old stellar populations in the rest frame for , enabling determination of galaxy morphological types unaffected by recent star formation.Methods. In Paper I we presented a new non-parametric method of quantifying morphologies of galaxies on seeing-limited images based on support vector machines. Here we use this method to classify ~50 000 Ks selected galaxies in the COSMOS area observed with WIRCam at CFHT. We use a 10-dimensional volume, including 5 morphological parameters, and other characteristics of galaxies such as luminosity and redshift. The obtained classification is used to investigate the redshift distributions and number counts per morphological type up to z ~ 2 and to compare them to the results obtained with HST/ACS in the I-band on the same objects. We associate to every galaxy with Ks < 21.5 and a probability between 0 and 1 of being late-type or early-type. We use this value to assess the accuracy of our classification as a function of physical parameters of the galaxy and to correct for classification errors.Results. The classification is found to be reliable up to z ~ 2. The mean probability is p ~ 0.8. It decreases with redshift and with size, especially for the early-type population, but remains above p ~ 0.7. The classification globally agrees with the one obtained using HST/ACS for . Above z ~ 1, the I-band classification tends to find less early-type galaxies than the Ks-band one by a factor ~1.5, which might be a consequence of morphological k-correction effects.Conclusions. We argue therefore that studies based on I-band HST/ACS classifications at could be underestimating the elliptical population. Using our method in a 21.5 magnitude-limited sample, we observe that the fraction of the early-type population is (21.9% ± 8%) at z ~ 1.5-2 and (32.0% ± 5%) at the present time. We will discuss the evolution of the fraction of galaxies in types from volume-limited samples in a forthcoming paper.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".