A STATISTICAL RECONSTRUCTION OF THE PLANET POPULATION AROUND<i>KEPLER</i>SOLAR-TYPE STARS
Bibliographic record
Abstract
Using the cumulative catalog of planets detected by the NASA Kepler mission, we reconstruct the intrinsic occurrence of Earth- to Neptune-size (1–4 R ⊕ ) planets and their distributions with radius and orbital period. We analyze 76,711 solar-type (0.8 < R * / R ☉ < 1.2) stars with 430 planets on 20–200 day orbits, excluding close-in planets that may have been affected by the proximity to the host star. Our analysis considers errors in planet radii and includes an "iterative simulation" technique that does not bin the data. We find a radius distribution that peaks at 2–2.8 Earth radii, with lower numbers of smaller and larger planets. These planets are uniformly distributed with logarithmic period, and the mean number of such planets per star is 0.46 ± 0.03. The occurrence is ∼0.66 if planets interior to 20 days are included. We estimate the occurrence of Earth-size planets in the "habitable zone" (defined as 1–2 R ⊕ , 0.99–1.7 AU for solar-twin stars) as . Our results largely agree with those of Petigura et al., although we find a higher occurrence of 2.8–4 Earth-radii planets. The reasons for this excess are the inclusion of errors in planet radius, updated Huber et al. stellar parameters, and also the exclusion of planets that may have been affected by proximity to the host star.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".