A reliable activity proxy in SPIRou spectra of M dwarfs using machine learning
Bibliographic record
Abstract
Context. Recent instruments have extended radial velocity observations from the optical domain to the near-infrared range (NIR). In particular, this has allowed M dwarfs to be studied more extensively, which is notable because they are known to host rocky planets more frequently. However, these stars also have, on average, stronger magnetic activity compared to solar-type stars, and investigating this magnetic activity is key to uncovering any planets around such stars. Aims. This paper aims to extensively test a new reliable magnetic activity indicator named W1 and confirm it as a proxy for the small-scale magnetic field for M dwarf stars. Methods. The magnetic activity indicator W1 is derived from a principal component analysis (PCA) applied on the per-line differential line width (dLW). However, the PCA is highly sensitive to contamination from telluric residuals in the spectra. Therefore, we employed a filtering technique based on such machine learning tools as the unsupervised dimensional reduction (DR) algorithm and support vector machine (SVM) to remove faulty lines. We assessed the performance of this filtering method using a simulation of observations of the per-line dLW variations before applying it to NIR high-resolution spectroscopic observations from SPIRou (at the Canada-France-Hawaii Telescope) of five targets with various stellar magnetic activity levels, spectral types, and rotation periods contained in the SPIRou Legacy Survey, namely AU Mic, EV Lac, GJ1286, GJ1289, and G1 410. Results. The filtered W1 signal is modulated with a period consistent with the rotation period retrieved from activity indicators, corresponding to the magnetic activity for all the stars studied. It also correlates with the small-scale magnetic field of all five stars, with a direct Pearson correlation coefficient greater than 0.80. Additionally, we identified 201 stellar lines that are particularly sensitive to magnetic activity that could be valuable for the study of magnetic fields.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".