Validation of Small <i>Kepler</i> Transiting Planet Candidates in or near the Habitable Zone
Bibliographic record
Abstract
Abstract A main goal of NASA’s Kepler Mission is to establish the frequency of potentially habitable Earth-size planets ( ). Relatively few such candidates identified by the mission can be confirmed to be rocky via dynamical measurement of their mass. Here we report an effort to validate 18 of them statistically using the BLENDER technique, by showing that the likelihood they are true planets is far greater than that of a false positive. Our analysis incorporates follow-up observations including high-resolution optical and near-infrared spectroscopy, high-resolution imaging, and information from the analysis of the flux centroids of the Kepler observations themselves. Although many of these candidates have been previously validated by others, the confidence levels reported typically ignore the possibility that the planet may transit a star different from the target along the same line of sight. If that were the case, a planet that appears small enough to be rocky may actually be considerably larger and therefore less interesting from the point of view of habitability. We take this into consideration here and are able to validate 15 of our candidates at a 99.73% (3 σ ) significance level or higher, and the other three at a slightly lower confidence. We characterize the GKM host stars using available ground-based observations and provide updated parameters for the planets, with sizes between 0.8 and 2.9 R ⊕ . Seven of them (KOI-0438.02, 0463.01, 2418.01, 2626.01, 3282.01, 4036.01, and 5856.01) have a better than 50% chance of being smaller than 2 R ⊕ and being in the habitable zone of their host stars.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".