Contrasting beta diversity among regions: how do classical and multivariate approaches compare?
Bibliographic record
Abstract
Abstract Aim Approaches to calculating beta diversity (β) include classical measures based on alpha (α) and gamma (γ) diversity, and multivariate distance‐based measures. Species–area relationships cause measurements of γ to vary, making comparisons of classical β among regions contingent on sampling effort. A recent null‐modelling approach has attempted to account for variation in γ by calculating the degree to which β deviates from a random expectation. Here, we clarify the mathematical links between classical and multivariate approaches to measuring β, to derive predictions regarding the reliability of classical, null‐model and multivariate approaches. Next, we use four ecological datasets and simulated data to test the consistency of these approaches across sampling effort and γ. We focus on an issue that arises when making comparisons among regions, namely that even small changes to the area sampled can differentially increase measured γ in each region, potentially causing artefacts in β that are driven by methodology rather than biology. Innovation Comparisons among regions using classical and null‐model measures change dramatically as sampling effort and γ increase. This change is understood for classical β because of species–area relationships, but not for null‐model measures, making comparisons among regions impossible using the null‐model approach. Multiple‐site dissimilarity shows a similar sensitivity to γ as classical measures. In contrast, pairwise multivariate distances show no systematic effect of sampling effort and γ: increasing the number of sample plots decreases variability but does not alter mean β. Main conclusions Multivariate pairwise distances are independent of sample size, offering the most robust comparison among regions. The widespread influence of sampling effort and γ indicate that only scale‐dependent measures of classical and multiple‐site β are comparable, whereas null‐model β may not be comparable among regions. However, in cases where γ is well known, multiple‐site dissimilarity metrics offer several advantages, and should be strongly considered.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".