Spectrum Classification of the Quasars from SDSS with the Redshift z~3 with PCA
Bibliographic record
Abstract
The spectrum lines of different quasars (QSOs) have been investigated, emphasis on the weak emission lines existent in the Lyα forest region using principal component analysis method (PCA) in the wavelength range from 1020 to 1600Å. The first, the second and the third principal component spectra (PCS) involve 63.4, 14.5 and 6.2% of the variance respectively, as the first seven PCS involve 96.1% of the total variance. The first PCS contain peak from high ionization emission lines namely the emission lines of Lyα and Lyβ (OVI, NV, SiIV, and CIV), these peaks are sharp and strong and the second PCS has peaks from low ionization emission lines (FeII, FeIII, SiII, and CII) that these emission lines are wide and almost rounded. By using the PCS, can be produced the QSO spectra artificially that are useful for investigation of how to discover QSOs and their continuum spectrum classification. By using the weights of the first two PCS can be defined five classes: class Zero and classes from ItoIV, these classifications will help to discover the continuum spectrum in the Lyα forest. 21 continuum (Listed in Table 1) have been used upon spectrum of bright QSOs from SDSS with z~3 and a signal to noise ratio more than 20 (S/N>20) and apparent magnitude less than 18 (mg<18). Initially, by investigating the spectrum of each one of these QSOs, the QSOs were classified and the peak of emission line of Lyα belonging to each one of classes was investigated. In the rest, the result mean spectrum of 21 QSOs were compared with mean spectrum of 50 QSOs with redshift 0.14
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".