Location–scale models in ecology and evolution: Heteroscedasticity in continuous, count and proportion data
Bibliographic record
Abstract
Abstract Biological data often violate the assumption of constant variance, yet such heteroscedasticity can reflect meaningful biological processes such as plasticity, canalization or stress responses. Despite this, most models treat variance as statistical noise. Here, we reintroduce location–scale regression as a general framework that jointly models the mean (location) and variance (scale) components of a response. We describe three hierarchical extensions: (1) fixed‐effects, (2) mixed‐effects and (3) double‐hierarchical models, which allow researchers to formally test variance structures alongside mean effects, enhancing biological interpretation. This framework is highly flexible and can extend beyond Gaussian assumptions to accommodate real‐world data. The framework accommodates over‐dispersed, under‐dispersed and zero‐inflated count data through the use of negative binomial and Conway–Maxwell–Poisson distributions, and bounded proportion data through beta‐binomial and beta regressions. Submodels can also be incorporated to account for structural zeros and ones when boundary outcomes are common. These extensions allow researchers to capture ecological processes such as presence–absence, success rates and bounded response rates. Using worked examples from published evolutionary and behavioural ecological studies, we illustrate how location–scale models can uncover biologically meaningful variance patterns that are overlooked in models focused solely on means. For instance, we show how food supplementation, hatching order and predation risk influence not only average trait values but also their variability. Each example corresponds to one of the model types and is implemented using widely used R packages such as glmmTMB and brms . All examples are accompanied by a freely accessible, step‐by‐step online tutorial, thereby lowering technical barriers and fostering broader adoption of location–scale modelling in ecological and evolutionary research. Finally, we propose a practical workflow for model selection and diagnostics and highlight recent extensions of the framework. These include multi‐response models, meta‐analytic models, phylogenetic comparative models and models including shape parameters such as skewness. Treating variance as a biologically informative response opens new avenues for us to explore the evolutionary, ecological and environmental processes that shape biological systems across diverse contexts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".