Understanding step selection analysis through numerical integration
Bibliographic record
Abstract
Abstract Step selection functions (SSFs) are flexible statistical models used to jointly describe animals' movement and habitat preferences. The popularity of SSFs has grown rapidly, and various extensions have been developed to increase their utility, including the ability to use multiple statistical distributions to describe movement constraints, interactions to allow movements to depend on local environmental features, and random effects and latent states to account for within‐ and among‐individual variability. Although the SSF is a relatively simple statistical model, its presentation has not been consistent in the literature, leading to confusion about model flexibility and interpretation. We believe that part of the confusion has arisen from the conflation of the SSF model with the methods used for statistical inference, and in particular, parameter estimation. Notably, conditional logistic regression (CLR) can be used to fit SSFs in exponential form, and this model fitting approach is often presented interchangeably with the actual model (the SSF itself). However, reliance on CLR reduces model flexibility, and suggests a misleading interpretation of step selection analysis as being equivalent to a case–control study. In this review, we explicitly distinguish between model formulation and inference technique, presenting a coherent framework to fit SSFs based on numerical integration and maximum likelihood estimation. We provide an overview of common numerical integration techniques (including Monte Carlo integration, importance sampling and quadrature), and explain how they relate to popular methods used in step selection analyses. This general framework unifies different model fitting techniques for SSFs, and opens the way for improved inferential methods. In this approach, it is straightforward to model movement with distributions outside the exponential family, and to apply different SSF model formulations to the same data set and compare them with AIC. By separating the model formulation from the inference technique, we hope to clarify many important concepts in step selection analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.069 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".