NASCENT-stars large program: origin of molecular complexity towards emerging high-mass protostars
Bibliographic record
Abstract
During the star formation process, the interstellar medium gets enriched in a complex zoo of molecules. The extremely rich chemistry of the star forming gas holds the key to constrain the link between the chemistry of disks and their forming planetary systems. It is, however, unclear what are the key physical conditions that influence the emerging complex chemistry, and how the fundamental properties of the emerging star, or stellar clusters impact its evolution. In this context, wide-band receivers open a new era by allowing us to perform quasi-instantaneously spectral surveys that are required for a more robust estimation of molecular abundances. The NASCENT-stars project (PI: Csengeri) is a large observing program started in 2023 at the IRAM NOEMA observatory with 228 hours allocated at the telescope making use of the newly commissioned high spectral resolution (250 kHz) and wide-band observing mode. Overall, we observe a 46.5 GHz non-continuous bandwidth to characterize the physico-chemical conditions of the most active star forming regions of the Cygnus-X molecular complex at the scale of 1400 au probing statistically significant samples individual protostellar envelopes. I would like to highlight the first results of NASCENT-stars and our efforts to develop artificial intelligence and machine learning methods facilitating the exploitation of these rich datasets. Results of this project should pave the way to the spectacular science cases enabled by the ALMA WSU.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".