A New Statistical Test of Fitness Set Data from Reciprocal Transplant Experiments Involving Intermediate Phenotypes
Bibliographic record
Abstract
Experimental biologists use reciprocal transplant experiments (RTEs) involving divergent forms to test hypotheses about fitness trade-offs across, and local adaptation to, native environments. Additional evolutionary hypotheses about diversifying selection, the evolution of specialization, and the coexistence of specialists and generalists are only testable when the RTE also includes intermediate (or alternatively generalist) forms. Environmental variation makes such RTEs challenging, and so strategies that increase their effectiveness are useful. Here, we focus on improvements to the efficiency of RTEs involving intermediate forms with respect to the experimental design and the analysis of the resulting data. We provide a likelihood ratio-based test that offers increased statistical power and robustness relative to another test involving nonlinear regression, when used both for simulated data sets and for data from a study of two divergent fish species and their hybrids transplanted between two lake habitats. The test can be used with unequal numbers of observations, unequal variances, and binomial-type survival data and other nonnormal data. Simulations suggest that having equal numbers of experimental units in each phenotype-environment combination is reasonable. The intentional pairing of observations between environmental conditions (by using clones, full sibs, or half-sibs) is beneficial when paired observations have fitnesses that are negatively related between conditions but is detrimental with positive relatedness. Our methods can be extended to study more than two divergent forms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".