The PHANGS-MUSE/HST-H <i>α</i> nebulae catalogue
Bibliographic record
Abstract
We present the PHANGS-MUSE/HST-H α nebulae catalogue, comprising 5177 spatially resolved nebulae across 19 nearby star-forming galaxies ( D < 20 Mpc), based on high-resolution H α imaging from HST, homogenised to a fixed (10 pc) physical resolution and sensitivity. Combined with MUSE integral field spectroscopy, this enables robust classification of 4882 H II regions and the separation of planetary nebulae and supernova remnants. We derive electron densities for 2544 H II regions using [S II ] diagnostics and adopt direct or representative electron temperatures for consistent physical characterisation. Nebular sizes are measured using circularised radii and intensity-weighted second moments, yielding a median radius of approximately 20 pc and extending down to (sub-)parsec (deconvolved) radii. A structural complexity score is introduced via hierarchical segmentation to trace substructure, highlighting that around a third of the regions are H II complexes containing several individual clusters and bubbles, with an increased fraction of these regions in galactic centres. A luminosity–size relation, calibrated using the resolved HST sample, is applied to 30 790 MUSE nebulae, allowing the recovery of nebular sizes down to ~1 pc and providing statistical completeness beyond the HST detection limit. Comparisons with classical Strömgren radii indicate that observed sizes are systematically larger, corresponding to typical volume filling factors with a median of ϵ ~ 0.22 (10th–90th percentile 0.06–0.78), with larger regions exhibiting progressively lower values. We associate 3349 H II regions with stellar populations from the PHANGS-HST association catalogue, finding median ages of ~3 Myr and typical stellar masses of around 10 4 –10 5 M ⊙ , supporting the link between ionised nebular and young stellar populations. We also assess the impact of diffuse ionised gas on emission-line diagnostics and after removing confirmed supernova remnants, find no strong variation in line ratios with nebular resolution, indicating minimal systematic bias in the MUSE catalogue. This dataset establishes a detailed, spatially resolved connection between nebular structure and ionising sources, and provides a benchmark for future studies of feedback, DIG contributions, and star formation regulation in the ISM, especially in combination with matched high-resolution observations. The full catalogue is made publicly available in machine-readable format.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.014 | 0.020 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".