The kinematic signature of damped Lyman alpha systems: using the<i>D</i>-index to screen for high column density H i absorbers<sup>★</sup>
Bibliographic record
Abstract
Using a sample of 21 damped Lyman alpha systems (DLAs) and 35 sub-DLAs, we evaluate the D-index from high-resolution spectra of the Mg iiλ 2796 profile. This sample represents an increase in the sub-DLA statistics by a factor of 4 over the original D-index sample. We investigate various techniques to define the velocity spread (Δv) of the Mg ii line to determine an optimal D-index for the identification of DLAs. The success rate of DLA identification is 50–55 per cent, depending on the velocity limits used, improving by a few per cent when the column density of Fe ii is included in the D-index calculation. We recommend the set of parameters that are judged to be most robust, have a combination of high DLA identification rate (57 per cent) and low DLA miss rate (6 per cent) and most cleanly separate the DLAs and sub-DLAs (Kolmogorov–Smirnov probability 0.5 per cent). These statistics demonstrate that the D-index is the most efficient technique for selecting low-redshift DLA candidates: 65 per cent more efficient than selecting DLAs based on the equivalent widths of Mg ii and Fe ii alone. We also investigate the effect of resolution on determining the N(H i) of sub-DLAs. We convolve echelle spectra of sub-DLA Lyα profiles with Gaussians typical of the spectral resolution of instruments on the Hubble Space Telescope and compare the best-fitting N(H i) values at both the resolutions. We find that the fitted H i column density is systematically overestimated by ∼0.1 dex in the moderate-resolution spectra compared to the best fits to the original echelle spectra. This offset is due to blending of nearby Lyα clouds that are included in the damping wing fit at low resolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".