New Perspectives on the Exoplanet Radius Gap from a Mathematica Tool and Visualized Water Equation of State
Bibliographic record
Abstract
Abstract Recent astronomical observations obtained with the Kepler and TESS missions and their related ground-based follow-ups revealed an abundance of exoplanets with a size intermediate between Earth and Neptune (1 R ⊕ ≤ R ≤ 4 R ⊕). A low occurrence rate of planets has been identified at around twice the size of Earth (2 × R ⊕), known as the exoplanet radius gap or radius valley. We explore the geometry of this gap in the mass–radius diagram, with the help of a Mathematica plotting tool developed with the capability of manipulating exoplanet data in multidimensional parameter space, and with the help of visualized water equations of state in the temperature–density (T–ρ) graph and the entropy–pressure (s–P) graph. We show that the radius valley can be explained by a compositional difference between smaller, predominantly rocky planets (<2 × R ⊕) and larger planets (>2 × R ⊕) that exhibit greater compositional diversity including cosmic ices (water, ammonia, methane, etc.) and gaseous envelopes. In particular, among the larger planets (>2 × R ⊕), when viewed from the perspective of planet equilibrium temperature (T eq), the hot ones (T eq ≳ 900 K) are consistent with ice-dominated composition without significant gaseous envelopes, while the cold ones (T eq ≲ 900 K) have more diverse compositions, including various amounts of gaseous envelopes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.009 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".