A Catalog of Broad Absorption Line Quasars in Sloan Digital Sky Survey Data Release 5
Bibliographic record
Abstract
We present a catalog of 5039 broad absorption line (BAL) quasars (QSOs) in the Sloan Digital Sky Survey (SDSS) Data Release 5 (DR5) QSO catalog that have absorption troughs covering a continuous velocity range ⩾2000 km s−1. We have fitted ultraviolet (UV) continua and line emission in each case, enabling us to report common diagnostics of BAL strengths and velocities in the range −25, 000 to 0 km s−1 for Si iv λ1400, C iv λ1549, Al iii λ1857, and Mg ii λ2799. We calculate these diagnostics using the spectrum listed in the DR5 QSO catalog, and also for spectra from additional SDSS observing epochs when available. In cases where BAL QSOs have been observed with Chandra or XMM-Newton, we report the X-ray monochromatic luminosities of these sources. We confirm and extend previous findings that BAL QSOs are more strongly reddened in the rest-frame UV than non-BAL QSOs, and that BAL QSOs are relatively X-ray weak compared to non-BAL QSOs. The observed BAL fraction is dependent on the spectral signal-to-noise ratio (S/N); for higher S/N sources, we find an observed BAL fraction of ≈ 15%. BAL QSOs show a similar Baldwin effect as for non-BAL QSOs, in that their C iv emission equivalent widths decrease with increasing continuum luminosity. However, BAL QSOs have weaker C iv emission in general than do non-BAL QSOs. Sources with higher UV luminosities are more likely to have higher-velocity outflows, and the BAL outflow velocity and UV absorption strength are correlated with relative X-ray weakness. These results are in qualitative agreement with models that depend on strong X-ray absorption to shield the outflow from overionization and enable radiative acceleration. In a scenario in which BAL trough shapes are primarily determined by outflow geometry, observed differences in Si iv and C iv trough shapes would suggest that some outflows have ion-dependent structure.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.007 | 0.008 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.008 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".