Young Stellar Objects in the Carina Nebula: Near-infrared Variability and Spectroscopy
Bibliographic record
Abstract
Abstract We present a catalog of 652 young stellar objects (YSOs) in the Carina star-forming region. The catalog was constructed by combining near-infrared K S -band variability from the VISTA Variables in the Vía Láctea eXtended survey and medium-resolution H-band spectroscopy from APOGEE-2, Sloan Digital Sky Survey IV (SDSS-IV). Variability analysis of 6.35 million sources identified 606 variable stars. The classification of the spectral lines by semisupervised K-means clustering of 704 stars, refined through comparison with known catalogs in literature and visual inspection of the spectra, was performed. Combined with K S variability, the final catalog contains three groups: Emission-line YSOs, Absorption-line YSOs, and Literature/Variable-identified YSOs. Cross validation with the Gaia DR3 proper motion and distance estimates supports Carina membership for 415 sources. The statistical characterization of YSO variability demonstrated that most Carina members (78%) exhibit variability patterns. Of these, 134 stars show emissions in their spectra, which is consistent with some accretion processes. Analysis of fundamental stellar parameters from StarHorse and Gaia DR3 reveals typical distributions of YSOs, dominated by low-mass (1–4M ⊙), solar-metallicity stars with temperatures between 4000 and 6000 K. Only a small fraction (4%) of the sources are more massive than 4M ⊙, suggesting limited ongoing massive star formation in Carina. This well-characterized catalog also offers a robust training data set for machine learning applications aimed at predicting YSO behavior.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".