WALLABY Pilot Survey: Public data release of ∼ 1800 H <scp>i</scp> sources and high-resolution cut-outs from Pilot Survey Phase 2
Bibliographic record
Abstract
Abstract We present the Pilot Survey Phase 2 data release for the Wide-field ASKAP L-band Legacy All-sky Blind surveY (WALLABY), carried-out using the Australian SKA Pathfinder (ASKAP). We present 1760 H i detections (with a default spatial resolution of 30′′) from three pilot fields including the NGC 5044 and NGC 4808 groups as well as the Vela field, covering a total of $\sim 180$ deg $^2$ of the sky and spanning a redshift up to $z \simeq 0.09$ . This release also includes kinematic models for over 126 spatially resolved galaxies. The observed median rms noise in the image cubes is 1.7 mJy per 30′′ beam and 18.5 kHz channel. This corresponds to a 5 $\sigma$ H i column density sensitivity of $\sim 9.1\times10^{19}(1 + z)^4$ cm $^{-2}$ per 30′′ beam and $\sim 20$ km s $^{-1}$ channel and a 5 $\sigma$ H i mass sensitivity of $\sim 5.5\times10^8 (D/100$ Mpc) $^{2}$ M $_{\odot}$ for point sources. Furthermore, we also present for the first time 12′′ high-resolution images (“cut-outs”) and catalogues for a sub-sample of 80 sources from the Pilot Survey Phase 2 fields. While we are able to recover sources with lower signal-to-noise ratio compared to sources in the Public Data Release 1, we do note that some data quality issues still persist, notably, flux discrepancies that are linked to the impact of side lobes associated with the dirty beams due to inadequate deconvolution. However, in spite of these limitations, the WALLABY Pilot Survey Phase 2 has already produced roughly a third of the number of HIPASS sources, making this the largest spatially resolved H i sample from a single survey to date.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.034 | 0.024 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".