The integrated properties of the molecular clouds from the JCMT CO(3–2) High-Resolution Survey
Bibliographic record
Abstract
We define the molecular cloud properties of the Milky Way first quadrant using data from the JCMT CO(3–2) High-Resolution Survey. We apply the Spectral Clustering for Interstellar Molecular Emission Segmentation (SCIMES) algorithm to extract objects from the full-resolution data set, creating the first catalogue of molecular clouds with a large dynamic range in spatial scale. We identify more than 85000 clouds with two clear sub-samples: ∼35500 well-resolved objects and ∼540 clouds with well-defined distance estimations. Only 35 per cent of the catalogued clouds (as well as the total flux encompassed by them) appear enclosed within the Milky Way spiral arms. The scaling relationships between clouds with known distances are comparable to the characteristics of the clouds identified in previous surveys. However, these relations between integrated properties, especially from the full catalogue, show a large intrinsic scatter (∼0.5 dex), comparable to other cloud catalogues of the Milky Way and nearby galaxies. The mass distribution of molecular clouds follows a truncated-power-law relationship over three orders of magnitude in mass with a form dN/dM ∝ M−1.7 with a clearly defined truncation at an upper mass of |$M_0 \sim 3 \times 10^6\, \mathrm{ M}_\odot$|, consistent with theoretical models of cloud formation controlled by stellar feedback and shear. Similarly, the cloud population shows a power-law distribution of size with dN/dR ∝ R−2.8 with a truncation at R0 = 70 pc.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".