An Outer Galaxy Molecular Cloud Catalog
Bibliographic record
Abstract
We have generated a molecular cloud catalog from the Five College Radio Astronomy Observatory Outer Galaxy Survey of 12 CO ( J = 1-0) emission using a two-phase object identification procedure. The first phase consists of grouping pixels into contiguous structures above a radiation temperature threshold of 0.8 K. The second phase decomposes the first-phase objects by an enhanced version of the CLUMPFIND algorithm, using dynamic thresholding, and again with a threshold of 0.8 K used for discrimination. A detailed comparison of our method with the CLUMPFIND algorithm is given, highlighting the advantages of the use of dynamic (rather than quantized) thresholding. Basic attributes of the clouds—coordinates, bounding boxes, integrated intensities, peak observed temperatures—are tabulated in the catalog. A two-dimensional elliptical Gaussian is fitted to the velocity-integrated map of each cloud; the major and minor axis sizes and major axis position angles thus derived are included in the catalog. To the spatially integrated emission line of each cloud, a Gaussian profile is fitted to measure the global linewidth. Model Gaussian clouds, truncated at 0.8 K, are examined to determine the effects of biases on measured quantities, induced by truncation. Coupled with detailed analysis of the cataloged clouds, statistical corrections for the effects of truncation on measured sizes, linewidths, and integrated intensities are derived and applied, along with corrections for the effects of finite resolution on the measured attributes. The cataloged emission accounts for 76.4% of the total emission in the Outer Galaxy Survey. The deficit is shown to arise mainly from low-intensity emission on the periphery of larger objects, rather than from a large number of small and/or low-intensity features. From the measured parameters, Gaussian reconstructions of the emission are carried out, and these compare favorably to the raw data. A detailed analysis of the decomposition in complex regions is performed, showing that severe truncation at levels much in excess of 0.8 K is countered by the operation of a "thermostat," resulting in concatenation of emission into a larger object if severe blending is present, rather than the identification of a number of smaller, more heavily truncated objects. Two other tests are carried out: (1) an association test that examines the utility of using the decomposed 12 CO ( J = 1-0) emission, in comparison to CS emission, to identify possible sites of star formation as traced by IRAS point sources; (2) a test comparison of 12 CO and 13 CO decompositions to gauge the effects of emission saturation and blending. Overall, the results of these two tests show that emission enhancements in 12 CO ( J = 1-0) emission, induced by internal heating by embedded star formation, are in general usefully recorded in the decomposition, but that emission blending and saturation on scales of ~few arcminutes in complex regions can limit the precision to which associations with other tracers can be made. A statistical approach to source association that makes good use of the information contained in the catalog is developed and described. The new 12 CO cloud catalog, as a concise description of the OGS data, will facilitate intercomparisons of molecular clouds with other ISM components, available at comparable resolution within the Canadian Galactic Plane Survey. The catalog covers the Galactic longitude range 102 5 to 141 5, Galactic latitude range -3° to +5 4, and lsr velocity range +20.8 to -120.2 km s -1 and contains 14,592 objects. Due to its size, and for ease of access, the catalog is made available in electronic form.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.007 | 0.010 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.260 | 0.264 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".