THE ARECIBO LEGACY FAST ALFA SURVEY: THE α.40 H I SOURCE CATALOG, ITS CHARACTERISTICS AND THEIR IMPACT ON THE DERIVATION OF THE H I MASS FUNCTION
Bibliographic record
Abstract
We present a current catalog of 21 cm H i line sources extracted from the Arecibo Legacy Fast Arecibo L -band Feed Array (ALFALFA) survey over ∼2800 deg 2 of sky: the α.40 catalog. Covering 40% of the final survey area, the α.40 catalog contains 15,855 sources in the regions 07 h 30 m < R.A. < 16 h 30 m , +04° < decl. <+16°, and +24° < decl. <+28° and 22 h < R.A. < 03 h , +14° < decl. <+16°, and +24° < decl. < + 32°. Of those, 15,041 are certainly extragalactic, yielding a source density of 5.3 galaxies per deg 2 , a factor of 29 improvement over the catalog extracted from the H i Parkes All-Sky Survey. In addition to the source centroid positions, H i line flux densities, recessional velocities, and line widths, the catalog includes the coordinates of the most probable optical counterpart of each H i line detection, and a separate compilation provides a cross-match to identifications given in the photometric and spectroscopic catalogs associated with the Sloan Digital Sky Survey Data Release 7. Fewer than 2% of the extragalactic H i line sources cannot be identified with a feasible optical counterpart; some of those may be rare OH megamasers at 0.16 < z < 0.25. A detailed analysis is presented of the completeness, width-dependent sensitivity function and bias inherent of the α.40 catalog. The impact of survey selection, distance errors, current volume coverage, and local large-scale structure on the derivation of the H i mass function is assessed. While α.40 does not yet provide a completely representative sampling of cosmological volume, derivations of the H i mass function using future data releases from ALFALFA will further improve both statistical and systematic uncertainties.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.013 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".