Re-evaluating Small Long-period Confirmed Planets from Kepler
Bibliographic record
Abstract
Abstract We re-examine the statistical confirmation of small long-period Kepler planet candidates in light of recent improvements in our understanding of the occurrence of systematic false alarms in this regime. Using the final Data Release 25 (DR25) Kepler planet candidate catalog statistics, we find that the previously confirmed single-planet system Kepler-452b no longer achieves a 99% confidence in the planetary hypothesis and is not considered statistically validated in agreement with the finding of Mullally et al. For multiple planet systems, we find that the planet prior enhancement for belonging to a multiple-planet system is suppressed relative to previous Kepler catalogs, and we also find that the multiple-planet system member, Kepler-186f, no longer achieves a 99% confidence level in the planetary hypothesis. Because of the numerous confounding factors in the data analysis process that leads to the detection and characterization of a signal, it is difficult to determine whether any one planetary candidate achieves a strict criterion for confirmation relative to systematic false alarms. For instance, when taking into account a simplified model of processing variations, the additional single-planet systems Kepler-443b, Kepler-441b, Kepler-1633b, Kepler-1178b, and Kepler-1653b have a non-negligible probability of falling below 99% confidence in the planetary hypothesis. The systematic false alarm hypothesis must be taken into account when employing statistical validation techniques in order to confirm planet candidates that approach the detection threshold of a survey. We encourage those performing transit searches of K2, TESS, and other similar data sets to quantify their systematic false alarm rates. Alternatively, independent photometric detection of the transit signal or radial velocity measurements can eliminate the false alarm hypothesis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.087 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".