Re-evaluating Small Long-period Confirmed Planets from Kepler
Bibliographic record
Abstract
Abstract We re-examine the statistical confirmation of small long-period Kepler planet candidates in light of recent improvements in our understanding of the occurrence of systematic false alarms in this regime. Using the final Data Release 25 (DR25) Kepler planet candidate catalog statistics, we find that the previously confirmed single-planet system Kepler-452b no longer achieves a 99% confidence in the planetary hypothesis and is not considered statistically validated in agreement with the finding of Mullally et al. For multiple planet systems, we find that the planet prior enhancement for belonging to a multiple-planet system is suppressed relative to previous Kepler catalogs, and we also find that the multiple-planet system member, Kepler-186f, no longer achieves a 99% confidence level in the planetary hypothesis. Because of the numerous confounding factors in the data analysis process that leads to the detection and characterization of a signal, it is difficult to determine whether any one planetary candidate achieves a strict criterion for confirmation relative to systematic false alarms. For instance, when taking into account a simplified model of processing variations, the additional single-planet systems Kepler-443b, Kepler-441b, Kepler-1633b, Kepler-1178b, and Kepler-1653b have a non-negligible probability of falling below 99% confidence in the planetary hypothesis. The systematic false alarm hypothesis must be taken into account when employing statistical validation techniques in order to confirm planet candidates that approach the detection threshold of a survey. We encourage those performing transit searches of K2 , TESS , and other similar data sets to quantify their systematic false alarm rates. Alternatively, independent photometric detection of the transit signal or radial velocity measurements can eliminate the false alarm hypothesis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".