Exploiting Kepler’s Heritage: A Transfer Learning Approach for Identifying Exoplanets’ Transits in TESS Data
Bibliographic record
Abstract
Abstract In the last decade, exoplanets space missions started to collect a huge amount of photometric observations, with over ∼1,000,000 new light curves generated every month from the Transiting Exoplanet Survey Satellite (TESS) full-frame images alone. In order to analyze such an unprecedented volume of data, automated planet-candidate detection has become an appreciable replacement to human vetting. In this work, we present a Machine Learning approach, based on Deep Neural Networks, to perform a binary classification of TESS light curves in terms of planet candidate and not-planet. Since few TESS labeled data exist to date, we pre-train the network with Kepler DR24 data set, including ≳15,000 labeled light curves. Our pre-trained model is then tested on ExoFOP data, showing an appreciable gain in terms of reliability with respect to a randomly initialized model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".