Searching for solar siblings in APOGEE and Gaia DR2 with N-body simulations
Bibliographic record
Abstract
ABSTRACT We make use of APOGEE and $Gaia\,$ data to identify stars that are consistent with being born in the same association or star cluster as the Sun. We limit our analysis to stars that match solar abundances within their uncertainties, as they could have formed from the same giant molecular cloud (GMC) as the Sun. We constrain the range of orbital actions that solar siblings can have with a suite of simulations of solar birth clusters evolved in static and time-dependent tidal fields. The static components of each galaxy model are the bulge, disc, and halo, while the various time-dependent components include a bar, spiral arms, and GMCs. In galaxy models without GMCs, simulated solar siblings all have JR < 122 km $\rm s^{-1}$ kpc, 990 < Lz < 1986 km $\rm s^{-1}$ kpc, and 0.15 < Jz < 0.58 km $\rm s^{-1}$ kpc. Given the actions of stars in APOGEE and $Gaia\,$, we find 104 stars that fall within this range. One candidate in particular, Solar Sibling 1, has both chemistry and actions similar enough to the solar values that strong interactions with the bar or spiral arms are not required for it to be dynamically associated with the Sun. Adding GMCs to the potential can eject solar siblings out of the plane of the disc and increase their Jz, resulting in a final candidate list of 296 stars. The entire suite of simulations indicate that solar siblings should have JR < 122 km $\rm s^{-1}$ kpc, 353 < Lz < 2110 km $\rm s^{-1}$ kpc, and Jz < 0.8 km $\rm s^{-1}$ kpc. Given these criteria, it is most likely that the association or cluster that the Sun was born in has reached dissolution and is not the commonly cited open cluster M67.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".