CAVITY: Calar Alto Void Integral-field Treasury surveY
Bibliographic record
Abstract
The Calar Alto Void Integral-field Treasury surveY (CAVITY) is a legacy project aimed at characterising the population of galaxies inhabiting voids, which are the most under-dense regions of the cosmic web, located in the Local Universe. This paper describes the first public data release (DR1) of CAVITY, comprising science-grade optical data cubes for the initial 100 out of a total of ~300 galaxies in the Local Universe (0.005 < z < 0.050). These data were acquired using the integral-field spectrograph PMAS/PPak mounted on the 3.5m telescope at the Calar Alto observatory. The DR1 galaxy sample encompasses diverse characteristics in the color-magnitude space, morphological type, stellar mass, and gas ionisation conditions, providing a rich resource for addressing key questions in galaxy evolution through spatially resolved spectroscopy. The galaxies in this study were observed with the low-resolution V500 set-up, spanning the wavelength range 3745-7500 Å, with a spectral resolution of 6.0 Å (FWHM). Here, we describe the data reduction and characteristics and data structure of the CAVITY datasets essential for their scientific utilisation, highlighting such concerns as vignetting effects, as well as the identification of bad pixels and management of spatially correlated noise. We also provide instructions for accessing the CAVITY datasets and associated ancillary data through the project’s dedicated database.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.017 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".